Skip to content

Device zerocopy examples repack send buffers with nothing guaranteeing the neighbour's get has finished reading them #3963

Description

@lvkale

Device zerocopy examples repack send buffers with nothing guaranteeing the neighbour's get has finished reading them

Every .C under examples/charm++/cuda/gpudirect and benchmarks/charm++/cuda/gpudirect constructs its send buffers as CkDeviceBuffer(ptr, stream) and none passes a source callback, although CkDeviceBuffer accepts one. jacobi2d (and the copies derived from it) repacks the same d_send_* buffers every iteration.

On the inter-node path the receiver reads the sender's buffer with an RDMA get issued when the metadata message is processed. The sender proceeds to its next iteration as soon as it has received its own ghosts, which does not depend on whether every neighbour's get of its buffer has completed, so the repack can overwrite a buffer a neighbour is still reading. Under 32 chares per GPU with lagging chares the window is real; the examples check no values, so it shows as nothing.

Two fixes, either sufficient:

Reproduction and both fixes are in the jacobi-overdecomp benchmark on the jacobi-overdecomp-bench branch (-a requests callbacks with double-buffered sets).

A registered device pool (#3962) makes this more important: it removes the per-message release, so the source callback becomes the only signal a sender gets, and "the pool owns the lifetime" must not be read as "the buffer is safe to overwrite".

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions