Skip to content

Device zerocopy: user documentation and diagnostics owed #3967

Description

@lvkale

Collecting the user-facing work owed around the device zerocopy API, so it is assignable rather than remembered. Standing rule (Kale): user-facing documentation changes with the change. #3961 does this for registration and buffer reuse; this is what remains.

1. The LCI-mediated path: warn, and fail informatively

Decision (Kale, 2026-09-03): the LCI-mediated path — the application allocates its own buffers, the runtime registers per send, and LCI's registration cache deduplicates — stays supported, on condition that users are warned in the documentation and that failures are informative.

It is the right default for the common case: a bounded set of long-lived buffers sent many times, where the first send registers and every later send is a cache lookup. Two properties need stating plainly, because both fail silently today:

  • A freed and reallocated device buffer can be served a stale registration. Nothing invalidates a cached device region on free: LCI asks for UCM_EVENT_MEM_TYPE_FREE, but the vendored UCX subset omits the ucm/cuda and ucm/rocm modules that would deliver it. Demonstrated by putting host memory in the same condition — freeing via a raw syscall so no UCM event fires — where the cache returned the registration created before the mapping was destroyed.
  • Registrations are never returned while the cache is on. max_regions, max_size and max_unreleased are all unlimited, so nothing evicts; a release only drops a refcount. An application with a bounded buffer set never notices. One that registers many distinct short-lived buffers accumulates memory regions without bound.

Documentation should therefore say: with application-allocated buffers, do not free and reallocate device buffers that have been sent; if the working set of distinct sent buffers is large or unbounded, use CkDeviceBufferRegister or the device pool.

Diagnostics worth adding, since neither condition is currently detectable by a user:

  • A better exhaustion message. Today the failure is LCI_Assert(false) ... err : No space left on device from register_memory_impl, which says nothing about what to do. It should name the cause (the NIC's memory-region table is full) and the remedies (register once, use the pool, enable eviction).
  • A registration-pressure warning. The runtime can count outstanding device registrations and warn once past a threshold, well before the table fills — the difference between a diagnosable warning and an abort thousands of iterations later.

2. hapiGetStream deprecation

Decided in the 2026-09-01 group meeting: users create their own streams; a return queue may come later. The documented model becomes user-created non-blocking streams — default-flagged streams measured 11–13 µs each to create at scale against a flat 1.6 µs, and they implicitly synchronise with the legacy null stream. Small surface: one implementation, one AMPI export, one test, one manual entry.

3. The hardware limits belong in the manual

A user choosing a stream count needs them: hardware queues 32 on NVIDIA (CUDA_DEVICE_MAX_CONNECTIONS) and about 8 on AMD (GPU_MAX_HW_QUEUES); concurrent kernels 128; copy engines a small fixed number.

4. +gpushm

Same-node IPC is opt-in and its absence is silent — without the flag, same-node transfers take the NIC instead. Measured slower than the RDMA-get path at 2–16 KB messages, so the manual should not present IPC as the obvious win. A one-time print when a same-node transfer falls back to the network would remove the surprise.

5. Pool scope

CkDeviceMalloc is for buffers intended to be sent, not a general device allocator: an arena is registered whole on the first send that lands in it, so scratch allocated from the pool is pinned along with everything else in its arena. See the separate note on #3961 about whether to add a will_send hint.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions