Device zerocopy examples repack send buffers with nothing guaranteeing the neighbour's get has finished reading them
Every .C under examples/charm++/cuda/gpudirect and benchmarks/charm++/cuda/gpudirect constructs its send buffers as CkDeviceBuffer(ptr, stream) and none passes a source callback, although CkDeviceBuffer accepts one. jacobi2d (and the copies derived from it) repacks the same d_send_* buffers every iteration.
On the inter-node path the receiver reads the sender's buffer with an RDMA get issued when the metadata message is processed. The sender proceeds to its next iteration as soon as it has received its own ghosts, which does not depend on whether every neighbour's get of its buffer has completed, so the repack can overwrite a buffer a neighbour is still reading. Under 32 chares per GPU with lagging chares the window is real; the examples check no values, so it shows as nothing.
Two fixes, either sufficient:
Reproduction and both fixes are in the jacobi-overdecomp benchmark on the jacobi-overdecomp-bench branch (-a requests callbacks with double-buffered sets).
A registered device pool (#3962) makes this more important: it removes the per-message release, so the source callback becomes the only signal a sender gets, and "the pool owns the lifetime" must not be read as "the buffer is safe to overwrite".
Device zerocopy examples repack send buffers with nothing guaranteeing the neighbour's get has finished reading them
Every
.Cunderexamples/charm++/cuda/gpudirectandbenchmarks/charm++/cuda/gpudirectconstructs its send buffers asCkDeviceBuffer(ptr, stream)and none passes a source callback, althoughCkDeviceBufferaccepts one.jacobi2d(and the copies derived from it) repacks the samed_send_*buffers every iteration.On the inter-node path the receiver reads the sender's buffer with an RDMA get issued when the metadata message is processed. The sender proceeds to its next iteration as soon as it has received its own ghosts, which does not depend on whether every neighbour's get of its buffer has completed, so the repack can overwrite a buffer a neighbour is still reading. Under 32 chares per GPU with lagging chares the window is real; the examples check no values, so it shows as nothing.
Two fixes, either sufficient:
Reproduction and both fixes are in the
jacobi-overdecompbenchmark on thejacobi-overdecomp-benchbranch (-arequests callbacks with double-buffered sets).A registered device pool (#3962) makes this more important: it removes the per-message release, so the source callback becomes the only signal a sender gets, and "the pool owns the lifetime" must not be read as "the buffer is safe to overwrite".