Collecting the user-facing work owed around the device zerocopy API, so it is assignable rather than remembered. Standing rule (Kale): user-facing documentation changes with the change. #3961 does this for registration and buffer reuse; this is what remains.
1. The LCI-mediated path: warn, and fail informatively
Decision (Kale, 2026-09-03): the LCI-mediated path — the application allocates its own buffers, the runtime registers per send, and LCI's registration cache deduplicates — stays supported, on condition that users are warned in the documentation and that failures are informative.
It is the right default for the common case: a bounded set of long-lived buffers sent many times, where the first send registers and every later send is a cache lookup. Two properties need stating plainly, because both fail silently today:
- A freed and reallocated device buffer can be served a stale registration. Nothing invalidates a cached device region on free: LCI asks for
UCM_EVENT_MEM_TYPE_FREE, but the vendored UCX subset omits the ucm/cuda and ucm/rocm modules that would deliver it. Demonstrated by putting host memory in the same condition — freeing via a raw syscall so no UCM event fires — where the cache returned the registration created before the mapping was destroyed.
- Registrations are never returned while the cache is on.
max_regions, max_size and max_unreleased are all unlimited, so nothing evicts; a release only drops a refcount. An application with a bounded buffer set never notices. One that registers many distinct short-lived buffers accumulates memory regions without bound.
Documentation should therefore say: with application-allocated buffers, do not free and reallocate device buffers that have been sent; if the working set of distinct sent buffers is large or unbounded, use CkDeviceBufferRegister or the device pool.
Diagnostics worth adding, since neither condition is currently detectable by a user:
- A better exhaustion message. Today the failure is
LCI_Assert(false) ... err : No space left on device from register_memory_impl, which says nothing about what to do. It should name the cause (the NIC's memory-region table is full) and the remedies (register once, use the pool, enable eviction).
- A registration-pressure warning. The runtime can count outstanding device registrations and warn once past a threshold, well before the table fills — the difference between a diagnosable warning and an abort thousands of iterations later.
2. hapiGetStream deprecation
Decided in the 2026-09-01 group meeting: users create their own streams; a return queue may come later. The documented model becomes user-created non-blocking streams — default-flagged streams measured 11–13 µs each to create at scale against a flat 1.6 µs, and they implicitly synchronise with the legacy null stream. Small surface: one implementation, one AMPI export, one test, one manual entry.
3. The hardware limits belong in the manual
A user choosing a stream count needs them: hardware queues 32 on NVIDIA (CUDA_DEVICE_MAX_CONNECTIONS) and about 8 on AMD (GPU_MAX_HW_QUEUES); concurrent kernels 128; copy engines a small fixed number.
4. +gpushm
Same-node IPC is opt-in and its absence is silent — without the flag, same-node transfers take the NIC instead. Measured slower than the RDMA-get path at 2–16 KB messages, so the manual should not present IPC as the obvious win. A one-time print when a same-node transfer falls back to the network would remove the surprise.
5. Pool scope
CkDeviceMalloc is for buffers intended to be sent, not a general device allocator: an arena is registered whole on the first send that lands in it, so scratch allocated from the pool is pinned along with everything else in its arena. See the separate note on #3961 about whether to add a will_send hint.
🤖 Generated with Claude Code
Collecting the user-facing work owed around the device zerocopy API, so it is assignable rather than remembered. Standing rule (Kale): user-facing documentation changes with the change. #3961 does this for registration and buffer reuse; this is what remains.
1. The LCI-mediated path: warn, and fail informatively
Decision (Kale, 2026-09-03): the LCI-mediated path — the application allocates its own buffers, the runtime registers per send, and LCI's registration cache deduplicates — stays supported, on condition that users are warned in the documentation and that failures are informative.
It is the right default for the common case: a bounded set of long-lived buffers sent many times, where the first send registers and every later send is a cache lookup. Two properties need stating plainly, because both fail silently today:
UCM_EVENT_MEM_TYPE_FREE, but the vendored UCX subset omits theucm/cudaanducm/rocmmodules that would deliver it. Demonstrated by putting host memory in the same condition — freeing via a raw syscall so no UCM event fires — where the cache returned the registration created before the mapping was destroyed.max_regions,max_sizeandmax_unreleasedare all unlimited, so nothing evicts; a release only drops a refcount. An application with a bounded buffer set never notices. One that registers many distinct short-lived buffers accumulates memory regions without bound.Documentation should therefore say: with application-allocated buffers, do not free and reallocate device buffers that have been sent; if the working set of distinct sent buffers is large or unbounded, use
CkDeviceBufferRegisteror the device pool.Diagnostics worth adding, since neither condition is currently detectable by a user:
LCI_Assert(false) ... err : No space left on devicefromregister_memory_impl, which says nothing about what to do. It should name the cause (the NIC's memory-region table is full) and the remedies (register once, use the pool, enable eviction).2.
hapiGetStreamdeprecationDecided in the 2026-09-01 group meeting: users create their own streams; a return queue may come later. The documented model becomes user-created non-blocking streams — default-flagged streams measured 11–13 µs each to create at scale against a flat 1.6 µs, and they implicitly synchronise with the legacy null stream. Small surface: one implementation, one AMPI export, one test, one manual entry.
3. The hardware limits belong in the manual
A user choosing a stream count needs them: hardware queues 32 on NVIDIA (
CUDA_DEVICE_MAX_CONNECTIONS) and about 8 on AMD (GPU_MAX_HW_QUEUES); concurrent kernels 128; copy engines a small fixed number.4.
+gpushmSame-node IPC is opt-in and its absence is silent — without the flag, same-node transfers take the NIC instead. Measured slower than the RDMA-get path at 2–16 KB messages, so the manual should not present IPC as the obvious win. A one-time print when a same-node transfer falls back to the network would remove the surprise.
5. Pool scope
CkDeviceMallocis for buffers intended to be sent, not a general device allocator: an arena is registered whole on the first send that lands in it, so scratch allocated from the pool is pinned along with everything else in its arena. See the separate note on #3961 about whether to add awill_sendhint.🤖 Generated with Claude Code