Skip to content

Fix forked test-suite segfaults (OpenMP after fork) - #4

Closed
sbryngelson wants to merge 1 commit into
sbryngelson:mainfrom
comp-physics:fix/forked-suite-openmp-segfault
Closed

sbryngelson wants to merge 1 commit into
sbryngelson:mainfrom
comp-physics:fix/forked-suite-openmp-segfault

Conversation

@sbryngelson

Copy link
Copy Markdown
Owner

Problem

The two grad-vs-torch CIFAR tests (test_conv2d_pad_grad_matches_torch, test_group_norm_train_grad_matches_torch) crashed with SIGSEGV only late in the full pytest tests/ run, yet passed in isolation — so it looked environmental.

Root cause

The macOS crash report pinned the faulting thread entirely in libomp.dylib — not the ANE/Espresso/Metal stack:

libomp.dylib  __kmp_suspend_initialize_thread
libomp.dylib  __kmp_fork_barrier(int, int)
libomp.dylib  __kmp_launch_worker(void*)
EXC_BAD_ACCESS (SIGSEGV) KERN_INVALID_ADDRESS at 0x580

This is the classic OpenMP-thread-pool-after-fork() failure:

  1. numpy/torch start an OpenMP worker-thread pool in the parent pytest process.
  2. pytest-forked calls os.fork(); fork() does not copy the worker threads, so the child inherits a pool with dead thread descriptors.
  3. The child's next OpenMP op — torch CPU autograd in these tests — calls into libomp (__kmp_fork_barrier), dereferences a dead descriptor, and segfaults.

That explains both symptoms: it only triggers once enough preceding tests have initialized the pool in the parent, and it's about threads, not test size (the tests are tiny).

Fix

Force OpenMP/BLAS single-threaded in tests/conftest.py, before any test imports numpy/torch:

for _thread_var in ("OMP_NUM_THREADS", "MKL_NUM_THREADS",
                    "OPENBLAS_NUM_THREADS", "VECLIB_MAXIMUM_THREADS"):
    os.environ.setdefault(_thread_var, "1")

With one OpenMP thread there is no worker pool to break across fork(), so the child is fork-safe. setdefault lets a caller override. These tiny test ops gain nothing from OpenMP parallelism, so there's no meaningful slowdown.

Verification

  • The 9-file subset that reliably reproduced the crash: 224 passed, 0 segfaults.
  • Full forked suite: 527 passed in 497.78s, 0 crashes / 0 failures (was 2 SIGSEGV), with the env set only by conftest.py.

Independent of the cleanup in #3.

The two grad-vs-torch CIFAR tests segfaulted (SIGSEGV) only late in the full
pytest run, while passing in isolation. The crash report pinned the fault entirely
in libomp.dylib (__kmp_fork_barrier / __kmp_suspend_initialize_thread), not the ANE
stack: --forked calls os.fork() in a process where numpy/torch have started an
OpenMP worker-thread pool, fork() does not copy those threads, and the child's next
OpenMP op (torch CPU autograd) faults on the broken pool. Setting OMP/MKL/OPENBLAS/
VECLIB thread counts to 1 in conftest (before any import) removes the pool, so the
fork is safe. Full forked suite now 527 passed, 0 crashes (was 2 SIGSEGV).
@sbryngelson

Copy link
Copy Markdown
Owner Author

Folded into #3 (squashed into a single commit with the test/bench cleanup, per request).

@sbryngelson
sbryngelson deleted the fix/forked-suite-openmp-segfault branch June 14, 2026 02:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant