CkRdmaDeviceOnSender synchronises each buffer's stream before the metadata message goes out, on the MEMCPY (same-process) path:
if (transfer_mode == CkNcpyModeDevice::MEMCPY) {
// The receiver dereferences buffers[i]->ptr directly, so the producing
// kernel/copy has to have retired before the metadata message goes out.
for (int i = 0; i < numops; i++)
hapiCheck(hapiStreamSynchronize(buffers[i]->hapi_stream));
return;
}
It is correct — the receiver reads the pointer directly, so the producing work must have retired — but it blocks the PE until the GPU finishes, which is at odds with how the rest of the runtime treats device work. Every other completion in HAPI is asynchronous: an event is recorded and hapiPollEvents delivers a Charm++ callback from the scheduler loop. Here the PE simply waits.
The cost lands exactly where overdecomposition is supposed to help. With several chares per PE, one chare's send stalls the others' entry methods for the duration of a kernel, so the intra-process path — the one that should be cheapest — becomes the one that serialises a PE.
Why this is not a quick fix. Deferring it means the metadata message itself has to be sent asynchronously: record an event, and send from the completion callback rather than from CkRdmaDeviceOnSender. The send is issued by charmxi-generated code immediately after CkRdmaDeviceOnSender returns, so the generated path has to learn to defer, and ordering guarantees for messages issued after it need thought. That is a design change, not a local edit — which is why it was left alone in #3955.
Related: the same reasoning appears in the overdecomposition design notes: the runtime should hold work in software where it can be reordered, rather than blocking. A deferred-send mechanism here is a small instance of that.
Found during review of #3955 (GPU stage 9.2). Not a defect in that PR — the behaviour predates it — and deliberately postponed after discussion.
🤖 Generated with Claude Code
CkRdmaDeviceOnSendersynchronises each buffer's stream before the metadata message goes out, on the MEMCPY (same-process) path:It is correct — the receiver reads the pointer directly, so the producing work must have retired — but it blocks the PE until the GPU finishes, which is at odds with how the rest of the runtime treats device work. Every other completion in HAPI is asynchronous: an event is recorded and
hapiPollEventsdelivers a Charm++ callback from the scheduler loop. Here the PE simply waits.The cost lands exactly where overdecomposition is supposed to help. With several chares per PE, one chare's send stalls the others' entry methods for the duration of a kernel, so the intra-process path — the one that should be cheapest — becomes the one that serialises a PE.
Why this is not a quick fix. Deferring it means the metadata message itself has to be sent asynchronously: record an event, and send from the completion callback rather than from
CkRdmaDeviceOnSender. The send is issued by charmxi-generated code immediately afterCkRdmaDeviceOnSenderreturns, so the generated path has to learn to defer, and ordering guarantees for messages issued after it need thought. That is a design change, not a local edit — which is why it was left alone in #3955.Related: the same reasoning appears in the overdecomposition design notes: the runtime should hold work in software where it can be reordered, rather than blocking. A deferred-send mechanism here is a small instance of that.
Found during review of #3955 (GPU stage 9.2). Not a defect in that PR — the behaviour predates it — and deliberately postponed after discussion.
🤖 Generated with Claude Code