Repository navigation
Persistent device messaging (CkDevicePersistent) aborts for inter-node transfers #3970
Description
Activity
Using a pre-migration handle is undetected today, and fails silently
Raised by Kale: forgetting the teardown/rebuild after migration is likely to be a common mistake, so does the runtime say anything useful? Checked — it does not.
CkDevicePersistent::getvalidates exactly one thing:if (cnt < src.cnt) { CkAbort("CkDevicePersistent::get: Destination buffer is smaller than source buffer\n"); } CkNcpyModeDevice mode = findTransferModeDevice(src.pe, CkMyPe());
and
findTransferModeDeviceonly bounds-checks the index (CmiEnforce((srcPe >= 0) && (srcPe <= CmiNumPes()))). Neither can tell whether the chare that ownedsrc.peis still there.The pup method makes this materially worse, because a stale handle survives migration looking entirely valid:
void CkDevicePersistent::pup(PUP::er& p) { p((char*)&ptr, sizeof(ptr)); p|cnt; p|pe; p|cb; p((char*)&hapi_ipc_handle, sizeof(hapi_ipc_handle)); }
A raw device pointer, a stale
pe, and an IPC handle from the old process are all carried to the new PE, fully populated, with nothing marking them as needing recreation. The documented contract — "After the owner chare migrates,CkDevicePersistentneeds to be recreated and exchanged again" — is enforced only by the reader remembering it.Three distinct failure modes follow, none of which produces a diagnostic:
- Mode misclassification.
src.peis stale, sofindTransferModeDevicecan return MEMCPY when the source has actually moved to another node.hapiMemcpyAsyncthen dereferencessrc.ptr— a device pointer from a different process's address space — in this one. - Use-after-free. If the source chare migrated away and freed its buffer, the MEMCPY path reads freed device memory. Silent wrong answers.
- Stale IPC mapping.
ipc_openandipc_ptrare cached on the source descriptor, so a mapping into a process that has since freed the allocation keeps being reused.
A cheap guard for half of it
The pup method is exactly the moment a handle becomes stale, so one bool closes the case where the holder migrated:
if (p.isUnpacking()) needs_reopen = true; // in pup // and at the top of get/put: if (needs_reopen) CkAbort("CkDevicePersistent used after migration: it must be recreated and " "re-exchanged after the owning chare moves");
One flag, one branch, and every case above turns into a message that names the contract.
The harder half, worth naming
That guard does not cover the case where the source migrates. Chare B holds an exchanged copy of A's descriptor; if A moves, B's copy is stale but B was never pupped, so B has no local signal at all. Catching that needs either an epoch exchanged with the handle and validated on use, or A invalidating its outstanding copies when it migrates. That is a design question rather than a guard, and it should be settled alongside the inter-node work rather than after it — inter-node persistent plus load balancing is precisely the combination where this arises.
🤖 Generated with Claude Code
- Mode misclassification.
CkDevicePersistentimplements the persistent/Direct model for device buffers — one-time setup, then repeated transfers with no per-message control traffic — but only for transfers that stay inside a physical node. Inter-node aborts:So the model works for MEMCPY (same process) and IPC (same physical node, different process), and is unavailable exactly where it would pay most — iterative communication across nodes at scale.
Why this matters more after #3960
The persistent model is a third answer to the registration question that #3961 just addressed, and structurally the cheapest one.
open()establishes the handle once; subsequentget/putcarry no registration, no release, and no acknowledgement. Compare:CkDeviceMalloc)For communication whose shape repeats every iteration, and where program logic already knows when buffers are reusable (double buffering, or a collective between send and repack), persistent trades a one-time setup for the elimination of the remaining control messages. That is precisely the regime an iterative multi-node application lives in, and it is the one shape the API cannot express today.
What exists
CkNcpyBuffer::get/put, the Direct API (CMK_DIRECT_API,handleDirectApiCompletion).CkDevicePersistentwithopen/close/get/put, plusexamples/charm++/cuda/gpudirect/persistent/andbenchmarks/charm++/cuda/gpudirect/latency-persistent/. Intra-node only.Scope notes for whoever picks this up
CkDevicePersistentneeds to be recreated and exchanged again." An inter-node implementation should keep that contract rather than try to make handles migratable.open()time instead of with each transfer. The pool's arena registrations (Device zerocopy: release registrations on completion, register-once API, piggybacked acks (#3960) #3961) are the closest existing machinery.close()is where the registration is released, which is the same lifetime discipline the pool uses and the opposite of the per-message release the default path uses.🤖 Generated with Claude Code