Repository navigation
Consolidate GPU Kernel launches - #186
Conversation
…n GPU tree walk enabled
…mmenceCalculateGraviyLocal is no longer called
|
We still need to decide what to do about the nodeGravityComputation and particleGravityComputation kernel launches. I don't think the remote gravity performance will benefit much from consolidating these, but this PR as-is probably breaks the local GPU gravity calculation if we aren't using the gpu-local-tree-walk option. |
|
This doesn't even compile if "--enable-gpu-local-tree-walk" is not specified. |
|
I just tested this out using the verbs comm layer and CUDA memory errors are back. I'm guessing the poor performance from MPI was actually preventing the remote gravity kernels from stepping on each other. We'll see if the CSA folks have any other suggestions when we talk to them next Monday, but I think we're going to need to move all of the nodeGravityComputation and particleGravityComputation kernel launches to the DataManager as well. |
nEwhLoop is checked for a valid value in DataManager before loading the Ewald h table onto the GPU. It needs to be set to an invalid value before the h-table is constructed.
…hWalk(). PELists only work with non-empty tree pieces. Also add CkAssert() in PEList::finishWalk() to catch too many checkings from TPs.
Using a vector of TreePiece pointers was overkill since we just needed a counter. Also added needed (for checkpoint/restart) code in Migration constructor and pup.
These two attributes kept track of the same thing.
…des. The traversal attempted to put links to NonLocal children of boundary nodes. Since they aren't in the list, this would occasionally create cycles in the tree and therefore infinite recursion.
The placement of the particles in GPU memory needs to be know before interaction lists can be serialized to send to the GPU. This commit adds code to be sure that happens.
Removed unused EwaldData attribute. Removed too-early set of data tranferred flag.
* Make EwaldInit() a DataManager method that gets run at tree construction time. This avoids a data race condition in the GPU code. It also saves some computation and memory since the h-loop table is only constructed once for an entire virtual node, instead of on every TreePiece. * Remove empty ewaldCPU() method. * DataManager::startEwaldGPU(): correct momcRoot copy.
The new PEList implementation moves a lot of these attributes to the PEList class. There were also some active particle bookkeeping on the GPU that were never used.
The accelerations get zeroed on the GPU, so there is no need of a tranfer from host to device.
Occasionally Ewald will complete after all the remote treewalks have finished. transferParticleVarsBack() also frees GPU memory so a use after free error would occur without this fix.
|
Also did minimal checks on the non-gpu-tree-walk code: it passes gravity tests. |
Having GPU kernel launches tied to the TreePieces degrades performance and is probably causing some data race conditions. The GPU versions of the local tree walk and ewald calculation are now handled by the data manager. Any kernel launches involving interaction list calculations are handled by node groups.
Note that another open PR #183 has already been merged into this branch.