Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
84 commits
Select commit Hold shift + click to select a range
2e60a5d
feat: c library form with the per-point infer inline in the header
sbryngelson Sep 14, 2026
2ca5720
fix: prefix every external weight symbol with the model name
sbryngelson Sep 14, 2026
5ae8d5f
feat: device decoration for openmp, openacc, cuda and hip hosts; embe…
sbryngelson Sep 14, 2026
c26df4f
fix: explain the device macros and _dev pointers in the emitted header
sbryngelson Sep 14, 2026
05d2631
feat: native cuda/hip batched kernel with an openmp fallback; init co…
sbryngelson Sep 14, 2026
e2a0c4c
fix: bind per-TU tables from the header, check launches, tighten the …
sbryngelson Sep 14, 2026
fece60c
fix: use non-sticky GetLastError and clarify device_bind_here's per-i…
sbryngelson Sep 14, 2026
352c14a
feat: fortran device directives, embedded parameter weights, device c…
sbryngelson Sep 14, 2026
e078931
fix: give the fortran archive its own name so both recipes can build …
sbryngelson Sep 14, 2026
0fcd3bd
feat: gpu-gate command and the microfd closure worked example
sbryngelson Sep 14, 2026
8c7f0d3
fix: link cuda/hip harnesses with the device compiler, scope the nsys…
sbryngelson Sep 14, 2026
82aca44
fix: nsys parse reports inconclusive, forward host offload flags at l…
sbryngelson Sep 14, 2026
a4fd89e
docs: how to generate, build and call a model from C and Fortran on t…
sbryngelson Sep 14, 2026
df7741b
docs: state the constant-memory cutoff and verify's model-dtype toler…
sbryngelson Sep 14, 2026
0be7a1a
fix: device-pass guard needs the compiler macro, accept __HIP__, rele…
sbryngelson Sep 14, 2026
55d7cf4
test: library-form test asserts the storage-class-first infer signature
sbryngelson Sep 14, 2026
3cba88d
fix: gpu-gate builds the omp archive with the host compiler for the p…
sbryngelson Sep 14, 2026
4121d3f
fix: hoist g.nut into face()'s locals in the microfd patch; CI grep o…
sbryngelson Sep 14, 2026
b0f2a82
docs: state the cuda/hip archive limitation for per-point file-loaded…
sbryngelson Sep 14, 2026
375523f
fix: embedded fortran weights are initialized protected arrays so gfo…
sbryngelson Sep 14, 2026
0f8e553
fix: gpu-gate records an unrunnable harness as a failed step instead …
sbryngelson Sep 14, 2026
933909d
test: filter driver-level toolchain notices in every device test, as …
sbryngelson Sep 14, 2026
80b19f6
test: prove target regions by symbol, and skip rather than fail where…
sbryngelson Sep 14, 2026
17efb7a
Validate the GPU path on hardware, then cover every golden model and …
sbryngelson Sep 14, 2026
3d15162
Restore nnLSTM.py, which four golden generators import
sbryngelson Sep 14, 2026
7fe4121
Note that the nvc bailout is unreported, and why the reproducer is kept
sbryngelson Sep 14, 2026
c9457f9
Validate the HIP path on an MI210, and fix the three things that stop…
sbryngelson Sep 14, 2026
db06b47
Fix seven review findings, each pinned by a regression test
sbryngelson Sep 14, 2026
be3cbac
rocprof transfer count for HIP, MANDATORY tests skip where libgomp ig…
sbryngelson Sep 14, 2026
387627e
Prove no transfers across a time-step loop, for every harness, on bot…
sbryngelson Sep 14, 2026
583ea8e
Seed every golden generator once, in the runner, so the models are th…
sbryngelson Sep 15, 2026
1e63ae2
Several graph outputs, concatenated in y, and the Concat op
sbryngelson Sep 15, 2026
5905c26
examples/surrogates A: coarse-grid Burgers with a learned subgrid clo…
sbryngelson Sep 15, 2026
39b18c8
examples/surrogates B: reaction-diffusion with a batched learned step…
sbryngelson Sep 15, 2026
0ceae89
examples/surrogates C: bubbly acoustics with a recurrent per-cell sur…
sbryngelson Sep 15, 2026
766a8b1
examples/surrogates D: Poisson solves with a whole-field conv-net ini…
sbryngelson Sep 15, 2026
fc583d5
examples/surrogates: shared build rules, terser sources and READMEs, …
sbryngelson Sep 15, 2026
0e7ac3e
examples/surrogates: name the examples by what they are, not by letter
sbryngelson Sep 15, 2026
d5dba2b
Register-blocked dense layers, and infer_one: a launch per op for who…
sbryngelson Sep 15, 2026
3e09625
CI: compile the HIP backend under hipcc, as nvcc_compile does for CUDA
sbryngelson Sep 15, 2026
c044a1c
examples: poisson_guess takes NSTEPS; the CI-sized host run uses 3
sbryngelson Sep 15, 2026
802bb2a
examples: link the CUDA runtime and C++ runtime the way nvc asks for …
sbryngelson Sep 15, 2026
52ed1ab
Test the examples on the GPU toolchains, not only on the host
sbryngelson Sep 15, 2026
3c59802
Compile every golden model warning-free, not just the dense ones
sbryngelson Sep 15, 2026
f6ab040
CI: sanitizers over every golden model, and clang on the C backend
sbryngelson Sep 15, 2026
f9111e2
AveragePool: only count in-bounds cells when the divisor reads the count
sbryngelson Sep 15, 2026
b16a9b8
Fix the two macOS CI failures: sanitizer availability, and burgers' s…
sbryngelson Sep 15, 2026
ca25d45
verify takes ROSENNA_CC/ROSENNA_FC; CI adds flang and ifx Fortran jobs
sbryngelson Sep 15, 2026
799898f
Refresh the CUDA gate report: all six harness pairs, zero transfers
sbryngelson Sep 15, 2026
a63a4c6
infer_one orders itself across streams; verify tolerance follows the …
sbryngelson Sep 15, 2026
97bf5ad
Fuzz LSTM, several inputs and outputs, and broadcast Add
sbryngelson Sep 15, 2026
2551641
Measure coverage, with a floor; cover validate's refusals and fold's …
sbryngelson Sep 15, 2026
a9108c8
Drop the checked-in gate reports; consolidate four test files
sbryngelson Sep 15, 2026
81af694
Run the host examples on one thread: the macOS timeouts were oversubs…
sbryngelson Sep 15, 2026
057ce82
CI: cache pip, install CPU torch, and share one setup step
sbryngelson Sep 15, 2026
15d6d64
Cover every diagnostic the weights reader promises; floor to 97
sbryngelson Sep 15, 2026
c1be764
Make .gitignore allow-by-default instead of a whitelist
sbryngelson Sep 15, 2026
10b8143
Put the two python/examples/ trees where their kind lives, and test r…
sbryngelson Sep 15, 2026
4801481
Replace the microfd patch with our own solver: examples/cns_closure
sbryngelson Sep 15, 2026
58d82a5
Track cns_closure's model; give common.mk a TRAINER knob
sbryngelson Sep 15, 2026
38e12cb
Time the closure in cns.c, and add the sync its contract requires
sbryngelson Sep 15, 2026
aa11f24
Report the flow, not just the code: kinetic energy and mu_t in cns.c
sbryngelson Sep 15, 2026
317551d
Add NO_CLOSURE=1: the plain solver, as the baseline for the closure's…
sbryngelson Sep 15, 2026
922526a
cns_closure: use_device_ptr for the batched call; use_device_addr on …
sbryngelson Sep 15, 2026
692fcb5
verify.py: the embedded Fortran weights are protected, not parameter
sbryngelson Sep 15, 2026
e91d618
Softmax over the last axis, both backends (TODO item 3)
sbryngelson Sep 15, 2026
103c5b4
Fold an inference BatchNormalization into the Conv/Gemm feeding it (T…
sbryngelson Sep 15, 2026
f81acb1
Grouped and depthwise Conv (TODO item 3)
sbryngelson Sep 15, 2026
ed4acb8
Constant-mode Pad (TODO item 3)
sbryngelson Sep 15, 2026
055e407
Document the four new ops in the READMEs
sbryngelson Sep 15, 2026
20e1237
Pad: allow negative pads (crops)
sbryngelson Sep 15, 2026
2b7ebff
Pad: edge and reflect modes, the opset-18 axes operand; rename the op…
sbryngelson Sep 16, 2026
cd7e38d
One archive serves both call paths (TODO item 1)
sbryngelson Sep 16, 2026
ae253cb
Run the test suite in parallel: 264s -> 47s, same coverage
sbryngelson Sep 16, 2026
6620fba
Declare upload_device outside the CUDA guard; make the local test as …
sbryngelson Sep 16, 2026
657ac0a
Document the one-archive linking rule in the README
sbryngelson Sep 16, 2026
7949532
Fuse infer_one's small-op launches (TODO item 4); -fPIC in the recipe
sbryngelson Sep 16, 2026
4d37646
Item 5: per-device state; measured answers for async upload and omp l…
sbryngelson Sep 16, 2026
ebb24ce
Coverage badge in the README, live number in the CI job summary
sbryngelson Sep 16, 2026
7aa14c8
run_basic.sh: seed the generator the way everything else does
sbryngelson Sep 16, 2026
5882335
1-D Conv and pooling (rank-3 NCW)
sbryngelson Sep 16, 2026
b7c14fe
Fix a rejection message the 1-D change made stale
sbryngelson Sep 16, 2026
7d98746
GRU, both readings of linear_before_reset
sbryngelson Sep 16, 2026
caa39b4
Pre-merge: a doc claim that became false, and four dead imports
sbryngelson Sep 16, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 28 additions & 0 deletions .github/actions/setup/action.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
name: Set up roseNNa
description: Python with a cached pip, and the project's dependencies.

runs:
using: composite
steps:
- uses: actions/setup-python@v5
with:
python-version: '3.11'
# Every job installed the same ~500 MB of wheels from scratch:
# "Install dependencies" was 64-128s in each of the nine, the single
# largest repeated cost in the matrix.
cache: pip
cache-dependency-path: requirements.txt

- name: Install dependencies
shell: bash
run: |
set -euo pipefail
# CPU torch on Linux: the default wheel is 529 MB against 188 MB, and
# nothing in CI runs torch on a GPU -- it only exports the golden
# models. requirements.txt then finds torch already satisfied.
# macOS wheels are CPU-only anyway and are not on that index.
if [ "${{ runner.os }}" = "Linux" ]; then
pip install --index-url https://download.pytorch.org/whl/cpu torch
fi
pip install -r requirements.txt
pip install -e python
248 changes: 236 additions & 12 deletions .github/workflows/CI.yml
Original file line number Diff line number Diff line change
Expand Up @@ -28,21 +28,245 @@ jobs:
- name: Check gfortran version
run: gfortran --version

- name: Set up Python
uses: actions/setup-python@v5
- name: Set up Python and dependencies
uses: ./.github/actions/setup

- name: Python package tests
run: cd python && python3 -m pytest tests -v -n auto

# A dead-model skip (the golden models are regenerated unseeded each run
# and can come out all-zero) is legitimate; a skip for a missing compiler
# is not. pipefail keeps pytest's own exit status through the tee.
- name: Device-path tests ran on the host (not skipped for a missing compiler)
run: |
set -o pipefail
cd python && python3 -m pytest tests/test_device_c.py -v -rs 2>&1 | tee device.log
! grep -E -q "SKIPPED.*(no C compiler|-fopenmp|no gfortran)" device.log

# Compile-only check of the generated CUDA sources. There is no GPU here and
# no driver is installed: nvcc builds the kernel and the .c-as-C++ library for
# sm_80 and the test asserts the archive exists. Running it is the GPU gate's
# job. This job is the first time the generated CUDA code meets a compiler.
# ubuntu-22.04 on purpose: NVIDIA's ubuntu2404 apt repository starts at CUDA
# 12.5, and 12.4 is the toolkit version this job pins.
nvcc_compile:
runs-on: ubuntu-22.04

steps:
- name: Clone roseNNa
uses: actions/checkout@v4

- name: Install the CUDA 12.4 compiler and runtime headers (no driver)
run: |
set -euo pipefail
wget -q https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt-get update
sudo apt-get install -y --no-install-recommends cuda-nvcc-12-4 cuda-cudart-dev-12-4
echo "/usr/local/cuda-12.4/bin" >> "$GITHUB_PATH"

- name: Check nvcc
run: nvcc --version && gcc --version

- name: Set up Python and dependencies
uses: ./.github/actions/setup

# No -n here: this step uses -s (it prints the compiler's own output),
# and pytest-xdist cannot capture that.
- name: CUDA backend compiles under nvcc (not skipped)
run: |
set -o pipefail
cd python && python3 -m pytest tests/test_kernel.py -v -rs -s -k cuda 2>&1 | tee cuda.log
! grep -q "SKIPPED" cuda.log
grep -q "PASSED" cuda.log

# The HIP twin of nvcc_compile: hipcc and the HIP headers from AMD's apt
# repository, no GPU and no driver. It compiles the kernel and the .c-as-HIP
# library for gfx90a and links a driver against the archive, which is where
# hipcc-specific issues surfaced on the MI210 (an archive named after the .cu
# is compiled as HIP source; __HIP__ is not a HIP-compilation signal).
# Running is the GPU gate's job.
hipcc_compile:
runs-on: ubuntu-24.04

steps:
- name: Clone roseNNa
uses: actions/checkout@v4

- name: Install hipcc and the HIP headers (no driver)
run: |
set -euo pipefail
sudo mkdir -p --mode=0755 /etc/apt/keyrings
wget -qO- https://repo.radeon.com/rocm/rocm.gpg.key | gpg --dearmor | sudo tee /etc/apt/keyrings/rocm.gpg > /dev/null
echo "deb [arch=amd64 signed-by=/etc/apt/keyrings/rocm.gpg] https://repo.radeon.com/rocm/apt/6.4.1 noble main" \
| sudo tee /etc/apt/sources.list.d/rocm.list
printf 'Package: *\nPin: release o=repo.radeon.com\nPin-Priority: 600\n' | sudo tee /etc/apt/preferences.d/rocm-pin-600
sudo apt-get update
sudo apt-get install -y --no-install-recommends hipcc hip-dev rocm-device-libs
echo "/opt/rocm/bin" >> "$GITHUB_PATH"

- name: Check hipcc
run: hipcc --version

- name: Set up Python and dependencies
uses: ./.github/actions/setup

- name: HIP backend compiles under hipcc (not skipped)
run: |
set -o pipefail
cd python && python3 -m pytest tests/test_kernel.py -v -rs -s -k hip 2>&1 | tee hip.log
! grep -q "SKIPPED" hip.log
grep -q "PASSED" hip.log

# ASan/UBSan over the generated code for every golden model, both backends.
# A compiler warning cannot see the bug class this is for: the generated code
# is loop nests over fixed-size locals whose bounds come from the plan, so it
# goes wrong by computing an index from the wrong extent -- which writes past
# a stack array and returns plausible numbers. That shipped once already (an
# activation bounded by the previous op's output length), and it matched
# onnxruntime on every model that did not branch.
sanitizers:
runs-on: ubuntu-latest

steps:
- name: Clone roseNNa
uses: actions/checkout@v4

- name: Set up gfortran
uses: fortran-lang/setup-fortran@v1
with:
python-version: '3.11'
compiler: gcc
version: 13

- name: Install dependencies
run: pip install -r requirements.txt
- name: Set up Python and dependencies
uses: ./.github/actions/setup

- name: Python package tests
# -rs and the SKIPPED check together: a sanitizer job that skips every
# case because a compiler is missing would otherwise report success.
- name: Generated code is clean under ASan and UBSan
run: |
set -o pipefail
cd python && python3 -m pytest tests/test_sanitizers.py -v -rs -n auto 2>&1 | tee san.log
! grep -q "SKIPPED" san.log

# The C backend under a second compiler. clang is preinstalled on the runner,
# so this costs nothing and reads the generated code with a different set of
# warnings from gcc's.
clang_c:
runs-on: ubuntu-latest

steps:
- name: Clone roseNNa
uses: actions/checkout@v4

- name: Set up Python and dependencies
uses: ./.github/actions/setup

- name: Check clang
run: clang --version

- name: Every golden model compiles warning-free under clang
env:
ROSENNA_CC: clang
run: |
set -o pipefail
cd python && python3 -m pytest tests/test_regressions.py -v -rs -n auto -k warning 2>&1 | tee clang.log
! grep -q "SKIPPED" clang.log

# A second Fortran front end. Fortran is the backend with the least compiler
# diversity -- gfortran in CI, nvfortran only in the GPU gate -- and it is
# where the conformance risk sits: nvfortran is what found that
# `has_device_addr` is unimplemented there. flang and ifx each build every
# golden model and run it against onnxruntime, so this is a behaviour check
# and not only a compile.
#
# Both were verified on all 21 models before this job was written; what is
# unverified here is the install step, not the test.
flang_fortran:
runs-on: ubuntu-latest

steps:
- name: Clone roseNNa
uses: actions/checkout@v4

- name: Install flang
run: |
set -euo pipefail
wget -qO llvm.sh https://apt.llvm.org/llvm.sh
chmod +x llvm.sh
sudo ./llvm.sh 20
sudo apt-get install -y flang-20
flang-20 --version

- name: Set up Python and dependencies
uses: ./.github/actions/setup

- name: Every golden model builds with flang and matches onnxruntime
env:
ROSENNA_FC: flang-20
run: |
pip install -e python
cd python && python3 -m pytest tests -v
set -o pipefail
cd python && python3 -m pytest tests/test_golden_suite.py -v -rs -n auto 2>&1 | tee flang.log
! grep -q "SKIPPED" flang.log

ifx_fortran:
runs-on: ubuntu-latest

steps:
- name: Clone roseNNa
uses: actions/checkout@v4

- name: Install ifx
run: |
set -euo pipefail
wget -qO- https://apt.repos.intel.com/intel-gpg-keys/GPG-PUB-KEY-INTEL-SW-PRODUCTS.PUB \
| gpg --dearmor | sudo tee /usr/share/keyrings/oneapi-archive-keyring.gpg > /dev/null
echo "deb [signed-by=/usr/share/keyrings/oneapi-archive-keyring.gpg] https://apt.repos.intel.com/oneapi all main" \
| sudo tee /etc/apt/sources.list.d/oneAPI.list
sudo apt-get update
sudo apt-get install -y intel-oneapi-compiler-fortran

- name: Set up Python and dependencies
uses: ./.github/actions/setup

# setvars.sh is what puts ifx and its runtime on PATH/LD_LIBRARY_PATH;
# it has to be sourced in the same step that runs the tests.
- name: Every golden model builds with ifx and matches onnxruntime
run: |
set -o pipefail
source /opt/intel/oneapi/setvars.sh >/dev/null
cd python && ROSENNA_FC=ifx python3 -m pytest tests/test_golden_suite.py -v -rs -n auto 2>&1 | tee ifx.log
! grep -q "SKIPPED" ifx.log

# Coverage of the generator, with a floor. The number is not the point: the
# floor is, because it turns "this change quietly stopped testing something"
# into a failure. validate.py sat at 75% until every uncovered line turned out
# to be a `raise` -- a refusal nobody had ever run, in the file whose whole
# job is refusing what it cannot compile correctly.
#
# gate.py is omitted (see pyproject.toml): it drives real compilers and a real
# GPU, so measuring it here would report the runner, not the tests.
coverage:
runs-on: ubuntu-latest

steps:
- name: Clone roseNNa
uses: actions/checkout@v4

- name: Set up gfortran
uses: fortran-lang/setup-fortran@v1
with:
compiler: gcc
version: 13

- name: Set up Python and dependencies
uses: ./.github/actions/setup

- name: Run test cases
# The badge in the README states the floor, which is a guarantee CI
# enforces rather than a snapshot that rots. The exact number goes in
# this run's summary, where it costs no service and no write access.
- name: Coverage is above the floor
run: |
mkdir -p fLibrary/objFiles
chmod +x test/run.sh
cd test && ./run.sh
cd python && python3 -m pytest tests --cov -q -n auto
echo "## Coverage: $(python3 -m coverage report --format=total)% (floor: 97%)" \
>> "$GITHUB_STEP_SUMMARY"
77 changes: 45 additions & 32 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -1,33 +1,46 @@
*
!*/
!goldenFiles/*/
!openNP.fpp
!userTesting.fpp
!modelCreator.fpp
!variables.fpp
!*.f90
!*.py
!Makefile
!*.sh
!*.yml
!*.c
!*.toml
!goldenFiles/mnist/mnist.onnx
!instructions/*
reading.f90
userTesting.f90
linearV3copy.f90
test.txt
goldenFiles/gemm_huge/
goldenFiles/vgg16/
goldenFiles/turbulentShear/
graphs/
graph_scripts/
fLibrary/*.txt
# This file was a whitelist: `*` followed by `!*.py`, `!*.c` and so on, plus a
# one-off exception every time something new needed to ship (`!requirements.txt`,
# `!goldenFiles/mnist/mnist.onnx`). That inverts the failure mode. A file nobody
# remembered to whitelist is not merely untracked, it is invisible: absent from
# `git status`, skipped by `git add .`, and gone at the next clean checkout. The
# root `README.md` and `LICENSE` were both matched by it and survived only
# because they were already in the index.
#
# So: ignore what is generated, and let everything else be seen.

# Python
__pycache__/
*.py[cod]
*.egg-info/
.venv/
.pytest_cache/
.coverage

# Compiled objects, modules and archives
*.o
*.mod
*.smod
*.a
*.so

# Weights written beside a model, and the scratch file the golden generators
# drop next to their working directory.
*.rwt
*.fpp

# The golden models are exported by their own .py at test time and dump a .txt
# of expected outputs; mnist is the exception, a fixture whose .py reads it
# rather than generating it (deleting it once cost an afternoon).
goldenFiles/*/*.onnx
goldenFiles/*/*.txt
randomStuff/
fLibrary/modelCreator.f90
fLibrary/variables.fpp
test/modelCreator.f90
test/variables.fpp
!requirements.txt
!goldenFiles/mnist/mnist.onnx

# Example builds: generated sources go to gen/, and each surrogate's Makefile
# links a pair of drivers named <model>_c and <model>_f. The surrogates' own
# .onnx files ship (they are the example); microfd_closure's is exported by
# closure.py on demand.
examples/**/gen/
examples/surrogates/*/*_c
examples/surrogates/*/*_f
examples/cns_closure/cns_c
*.lock
Loading
Loading