Skip to content

Persistent device messaging (CkDevicePersistent) aborts for inter-node transfers #3970

Description

@lvkale

CkDevicePersistent implements the persistent/Direct model for device buffers — one-time setup, then repeated transfers with no per-message control traffic — but only for transfers that stay inside a physical node. Inter-node aborts:

// CkDevicePersistent::get, ckrdmadevice.C:1086
} else {
  CkAbort("Persistent GPU messaging is currently not supported for inter-node messages");
}

// CkDevicePersistent::put, ckrdmadevice.C:1121
} else {
  CkAbort("Persistent GPU messaging is not yet supported for inter-node messages");
}

So the model works for MEMCPY (same process) and IPC (same physical node, different process), and is unavailable exactly where it would pay most — iterative communication across nodes at scale.

Why this matters more after #3960

The persistent model is a third answer to the registration question that #3961 just addressed, and structurally the cheapest one. open() establishes the handle once; subsequent get/put carry no registration, no release, and no acknowledgement. Compare:

model per-message control traffic
per-message registration + ack (the #3961 default) register, release, ack id returned
device pool (CkDeviceMalloc) none for registration; ack only if a source callback is requested
persistent none — setup is one-time

For communication whose shape repeats every iteration, and where program logic already knows when buffers are reusable (double buffering, or a collective between send and repack), persistent trades a one-time setup for the elimination of the remaining control messages. That is precisely the regime an iterative multi-node application lives in, and it is the one shape the API cannot express today.

What exists

  • Host side: fully supported, inter-node included — CkNcpyBuffer::get/put, the Direct API (CMK_DIRECT_API, handleDirectApiCompletion).
  • Device side: CkDevicePersistent with open/close/get/put, plus examples/charm++/cuda/gpudirect/persistent/ and benchmarks/charm++/cuda/gpudirect/latency-persistent/. Intra-node only.

Scope notes for whoever picks this up

  • Migration is already out of scope by design and documented in the header: "Should only be used for exchanging between chares, not for migration. After the owner chare migrates, CkDevicePersistent needs to be recreated and exchanged again." An inter-node implementation should keep that contract rather than try to make handles migratable.
  • The inter-node path needs a registration that lives for the object's lifetime rather than per message, and the remote key must travel at open() time instead of with each transfer. The pool's arena registrations (Device zerocopy: release registrations on completion, register-once API, piggybacked acks (#3960) #3961) are the closest existing machinery.
  • Teardown at close() is where the registration is released, which is the same lifetime discipline the pool uses and the opposite of the per-message release the default path uses.

🤖 Generated with Claude Code

Activity

  1. lvkale commented on Sep 9, 2026

    @lvkale
    ContributorAuthor

    Using a pre-migration handle is undetected today, and fails silently

    Raised by Kale: forgetting the teardown/rebuild after migration is likely to be a common mistake, so does the runtime say anything useful? Checked — it does not.

    CkDevicePersistent::get validates exactly one thing:

    if (cnt < src.cnt) {
      CkAbort("CkDevicePersistent::get: Destination buffer is smaller than source buffer\n");
    }
    CkNcpyModeDevice mode = findTransferModeDevice(src.pe, CkMyPe());

    and findTransferModeDevice only bounds-checks the index (CmiEnforce((srcPe >= 0) && (srcPe <= CmiNumPes()))). Neither can tell whether the chare that owned src.pe is still there.

    The pup method makes this materially worse, because a stale handle survives migration looking entirely valid:

    void CkDevicePersistent::pup(PUP::er& p) {
      p((char*)&ptr, sizeof(ptr));
      p|cnt;
      p|pe;
      p|cb;
      p((char*)&hapi_ipc_handle, sizeof(hapi_ipc_handle));
    }

    A raw device pointer, a stale pe, and an IPC handle from the old process are all carried to the new PE, fully populated, with nothing marking them as needing recreation. The documented contract — "After the owner chare migrates, CkDevicePersistent needs to be recreated and exchanged again" — is enforced only by the reader remembering it.

    Three distinct failure modes follow, none of which produces a diagnostic:

    1. Mode misclassification. src.pe is stale, so findTransferModeDevice can return MEMCPY when the source has actually moved to another node. hapiMemcpyAsync then dereferences src.ptr — a device pointer from a different process's address space — in this one.
    2. Use-after-free. If the source chare migrated away and freed its buffer, the MEMCPY path reads freed device memory. Silent wrong answers.
    3. Stale IPC mapping. ipc_open and ipc_ptr are cached on the source descriptor, so a mapping into a process that has since freed the allocation keeps being reused.

    A cheap guard for half of it

    The pup method is exactly the moment a handle becomes stale, so one bool closes the case where the holder migrated:

    if (p.isUnpacking()) needs_reopen = true;   // in pup
    // and at the top of get/put:
    if (needs_reopen)
      CkAbort("CkDevicePersistent used after migration: it must be recreated and "
              "re-exchanged after the owning chare moves");

    One flag, one branch, and every case above turns into a message that names the contract.

    The harder half, worth naming

    That guard does not cover the case where the source migrates. Chare B holds an exchanged copy of A's descriptor; if A moves, B's copy is stale but B was never pupped, so B has no local signal at all. Catching that needs either an epoch exchanged with the handle and validated on use, or A invalidating its outstanding copies when it migrates. That is a design question rather than a guard, and it should be settled alongside the inter-node work rather than after it — inter-node persistent plus load balancing is precisely the combination where this arises.

    🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions