Skip to content
Closed
Show file tree
Hide file tree
Changes from 43 commits
Commits
Show all changes
200 commits
Select commit Hold shift + click to select a range
064973e
feat: evidence-grade NIM discovery + all-modality cost-quality benchm…
seonghobae Aug 4, 2026
b48f396
merge: stack NIM benchmark on provider-egress security base
seonghobae Aug 4, 2026
0e19b3a
docs(changelog): record NIM benchmark evidence slice
seonghobae Aug 4, 2026
66c8d55
ci: apply reviewed NIM transport hardening
seonghobae Aug 4, 2026
9d546a7
fix(nim): harden egress and equalize policy budgets
seonghobae Aug 4, 2026
785fb15
fix(nim): install evidence and transport hardening
seonghobae Aug 4, 2026
e07a504
test(nim): prove secure transport and equal budgets
seonghobae Aug 4, 2026
d0312ea
docs(nim): document pinned egress and equal total budgets
seonghobae Aug 4, 2026
c7bcad2
chore(ci): remove temporary NIM repair workflow
seonghobae Aug 4, 2026
e1ae745
chore(ci): stage one-time NIM source repair
seonghobae Aug 4, 2026
1015608
chore(ci): remove privileged one-shot NIM repair workflow
seonghobae Aug 4, 2026
07128ac
chore(ci): stage exact-head NIM security repair
seonghobae Aug 4, 2026
a8393d7
chore(ci): harden bounded NIM repair job
seonghobae Aug 4, 2026
4cc0f55
chore(ci): stage bounded NIM source integration
seonghobae Aug 4, 2026
9ad742a
chore(ci): remove overlapping one-shot NIM workflow
seonghobae Aug 4, 2026
4fd96f1
fix(ci): replace write-capable PR repair with isolated publisher
seonghobae Aug 4, 2026
4bb7c85
docs(changelog): record reviewed NIM access evidence
seonghobae Aug 4, 2026
d9bedd7
chore(ci): stage equal-budget regression repair
seonghobae Aug 4, 2026
c4d47d2
ci: complete bounded NIM source integration
seonghobae Aug 4, 2026
db6ebb9
chore(ci): remove redundant NIM test-only repair workflow
seonghobae Aug 4, 2026
dba1d7e
ci: run bounded NIM source integration from PR checks
seonghobae Aug 4, 2026
a3741b5
fix(ci): execute bounded NIM source transformation deterministically
seonghobae Aug 4, 2026
18d3fe7
fix(ci): preserve embedded transformation literals
seonghobae Aug 4, 2026
27eb1b8
fix(security): remove PR-head OIDC source mutation
seonghobae Aug 4, 2026
ba6b34b
fix(security): delete privileged branch-trigger repair workflow
seonghobae Aug 4, 2026
e4578da
fix(security): delete branch-controlled OIDC repair source
seonghobae Aug 4, 2026
f80e503
test(security): pin PR workflow OIDC boundary
seonghobae Aug 4, 2026
bc77bfc
ci(review): export read-only exact-head workspace
seonghobae Aug 4, 2026
a71ebbc
ci(review): include inert transformation evidence
seonghobae Aug 4, 2026
6df72ee
test(ci): require secretless NIM dry-run job
seonghobae Aug 5, 2026
09fd519
fix(ci): isolate NIM dry runs from provider secrets
seonghobae Aug 5, 2026
2b622bd
docs(security): record NIM workflow secret isolation
seonghobae Aug 5, 2026
40b48ba
docs(changelog): record secretless NIM dry runs
seonghobae Aug 5, 2026
348e5e6
docs(benchmark): explain workflow credential isolation
seonghobae Aug 5, 2026
667b311
chore(review): stage verified NIM repair patch
seonghobae Aug 5, 2026
2498648
ci(nim): stage reviewed patch part 1 of 7
seonghobae Aug 5, 2026
4552084
ci(nim): stage reviewed patch part 2 of 7
seonghobae Aug 5, 2026
45c151e
ci(nim): stage reviewed patch part 3 of 7
seonghobae Aug 5, 2026
9dac6b2
ci(nim): stage reviewed patch part 4 of 7
seonghobae Aug 5, 2026
fecd25f
ci(nim): stage reviewed patch part 5 of 7
seonghobae Aug 5, 2026
19be572
ci(nim): stage reviewed patch part 6 of 7
seonghobae Aug 5, 2026
736893c
ci(nim): stage reviewed patch part 7 of 7
seonghobae Aug 5, 2026
07f28a9
ci(nim): apply exact reviewed benchmark integration
seonghobae Aug 5, 2026
7d5d6aa
ci: retrigger reviewed NIM integration with scoped permissions
seonghobae Aug 5, 2026
fd6b074
ci: verify and export reviewed NIM integration read-only
seonghobae Aug 5, 2026
cb7e9b3
ci: bind NIM verification to security ancestor
seonghobae Aug 5, 2026
3878e5b
fix(ci): bind NIM verification to pull-request head
seonghobae Aug 5, 2026
5bbedc2
fix(ci): reconcile reviewed NIM patch per file
seonghobae Aug 5, 2026
d98fe6c
fix(ci): reconstruct reviewed NIM tree from patch base
seonghobae Aug 5, 2026
428736c
fix(ci): bind reviewed NIM reconstruction to exact head
seonghobae Aug 5, 2026
e4e03c4
fix(ci): run reviewed NIM verification read-only on push
seonghobae Aug 5, 2026
6cd2012
fix(ci): use runner temp variable at execution time
seonghobae Aug 5, 2026
7595919
fix(ci): discover the exact reviewed patch base
seonghobae Aug 5, 2026
a93ec20
chore(ci): inspect the staged NIM integration payload
seonghobae Aug 5, 2026
f11193d
fix(ci): replay the reviewed NIM patch stack
seonghobae Aug 5, 2026
83b7906
fix(ci): rehydrate the reviewed patch index
seonghobae Aug 5, 2026
e9bfb85
fix(ci): execute the patch rehydration helper
seonghobae Aug 5, 2026
88274c0
fix(ci): apply the reviewed patch against the rehydrated index
seonghobae Aug 5, 2026
03b30e5
fix(ci): verify the reconstructed reviewed tree
seonghobae Aug 5, 2026
e256a9a
fix(ci): locate the exact reviewed patch context
seonghobae Aug 5, 2026
012f7c7
fix(ci): build the wheel in an isolated backend environment
seonghobae Aug 5, 2026
dde856f
ci(nim): publish the authenticated verified source tree
seonghobae Aug 5, 2026
8049ca8
feat(nim): integrate verified provider-neutral benchmark
seonghobae Aug 5, 2026
d1516a5
docs(nim): record immutable integration provenance
seonghobae Aug 5, 2026
f090ffc
test(ci): require stacked pull-request coverage
seonghobae Aug 5, 2026
d091b69
ci: validate stacked pull requests
seonghobae Aug 5, 2026
0d66e0a
ci: fuzz stacked pull requests
seonghobae Aug 5, 2026
928214c
ci: secure stacked pull requests
seonghobae Aug 5, 2026
c759cd1
test(ci): require isolated wheel builds
seonghobae Aug 5, 2026
8e19d64
fix(ci): isolate benchmark wheel builds
seonghobae Aug 5, 2026
e7f7238
ci(review): export read-only PR90 workspace
seonghobae Aug 5, 2026
c72f1cb
test(benchmark): stage complete-budget preflight red tests part 1
seonghobae Aug 5, 2026
cb20657
test(benchmark): stage complete-budget preflight red tests part 2
seonghobae Aug 5, 2026
636c3e0
fix(benchmark): stage complete request planner part 3
seonghobae Aug 5, 2026
1053f19
docs(benchmark): stage complete-budget evidence and verification part 4
seonghobae Aug 5, 2026
638a126
ci(benchmark): apply complete-budget preflight test-first
seonghobae Aug 5, 2026
1f6da53
ci(review): stage exact PR90 request-plan patch
seonghobae Aug 5, 2026
8a31de7
ci(review): refresh exact PR90 request-plan patch
seonghobae Aug 5, 2026
87496c8
ci(review): verify and publish exact PR90 request-plan fix
seonghobae Aug 5, 2026
b3c0471
ci(review): inspect exact PR90 patch hash
seonghobae Aug 5, 2026
2efcf08
ci(review): export decoded PR90 patch evidence
seonghobae Aug 5, 2026
a06f9c5
ci(review): verify exact staged PR90 patch
seonghobae Aug 5, 2026
77cc8cc
ci(review): publish verified PR90 request-plan fix
seonghobae Aug 5, 2026
796e6e9
test(benchmark): align scheduled complete-plan contract
seonghobae Aug 5, 2026
5956214
test(benchmark): require request-plan provenance schema
seonghobae Aug 5, 2026
9b72816
fix(nim): require a complete catalog benchmark plan
seonghobae Aug 5, 2026
a2af363
merge: integrate provider-egress security base into NIM benchmark
github-actions[bot] Aug 5, 2026
c5b3876
docs(adr): record exact NIM security integration evidence
seonghobae Aug 5, 2026
d43adaa
test(benchmark): expose undersized locked evaluation manifest
seonghobae Aug 5, 2026
16b5016
feat(benchmark): expand locked manifest to evidence floor
seonghobae Aug 5, 2026
ac13a56
test(benchmark): bind evidence floor to existing hard ceiling
seonghobae Aug 5, 2026
62e0e59
docs(benchmark): record thirty-task evidence floor
seonghobae Aug 5, 2026
43c629f
docs(benchmark): align request plan and evidence floor
seonghobae Aug 5, 2026
ff51584
docs(benchmark): doctor thirty-task evidence floor
seonghobae Aug 5, 2026
bec1129
test(benchmark): expose undersized default request caps
seonghobae Aug 5, 2026
cdfe08e
ci(benchmark): test and repair default request caps
seonghobae Aug 5, 2026
37c78b4
ci(benchmark): rearm default request-cap repair
seonghobae Aug 5, 2026
545cd64
ci(benchmark): add exact-head reopen trigger
seonghobae Aug 5, 2026
e56d1a3
ci(benchmark): use resolvable checkout action
seonghobae Aug 5, 2026
5216171
ci(benchmark): repair complete-plan default finalizer
seonghobae Aug 5, 2026
8cae8d4
test(benchmark): bind evidence floor to canonical request planner
seonghobae Aug 5, 2026
e31f901
ci(benchmark): cover all complete-plan regression anchors
seonghobae Aug 5, 2026
da8ec5f
ci(benchmark): repair invalid one-shot YAML
seonghobae Aug 5, 2026
1b97e48
ci(benchmark): split workflow-authorized publication
seonghobae Aug 5, 2026
0ebb1d5
fix(benchmark): align complete-plan source evidence
github-actions[bot] Aug 5, 2026
9e4b50b
fix(benchmark): align manual request cap with complete plan
seonghobae Aug 5, 2026
821b58e
ci(benchmark): remove completed one-shot workflow
seonghobae Aug 5, 2026
7b23bac
test(nim): cover buyer-facing complete request plan
seonghobae Aug 5, 2026
5c1a3f6
ci(nim): include request-plan regression in quality gate
seonghobae Aug 5, 2026
f5a494f
test: require NIM parser Atheris instrumentation
seonghobae Aug 5, 2026
48f007d
fix: instrument NIM parser imports for Atheris
seonghobae Aug 5, 2026
2736113
docs: distinguish historical NIM integration evidence
seonghobae Aug 5, 2026
b60ec5d
test: cover NIM review regressions
seonghobae Aug 5, 2026
bdc9673
chore: defer broader NIM regression batch
seonghobae Aug 5, 2026
646c417
docs: make NIM receipt head-stable
seonghobae Aug 5, 2026
318feb9
docs: record NIM fuzz instrumentation repair
seonghobae Aug 5, 2026
63438d0
test(nim): require completion-dependent README evidence status
seonghobae Aug 5, 2026
6fc53f1
docs(nim): clarify completion-dependent evidence status
seonghobae Aug 5, 2026
4a79964
test(nim): normalize README contract whitespace
seonghobae Aug 5, 2026
1217fc9
docs(changelog): record NIM evidence-status clarification
seonghobae Aug 5, 2026
11353d4
test(security): inspect every pull-request workflow structurally
seonghobae Aug 5, 2026
67c5c84
test(nim): enforce ordered non-empty secret-isolation steps
seonghobae Aug 5, 2026
989d7fe
test(security): parse pull-request branch filters structurally
seonghobae Aug 5, 2026
4329081
test(nim): lock review regressions before repair
seonghobae Aug 5, 2026
715ee76
fix(fuzz): require normalized NIM catalog depth failures
seonghobae Aug 5, 2026
8c4ee86
test: skip fixture-bound standalone NIM cases
seonghobae Aug 5, 2026
4ca0448
fix: close NIM evidence review gaps
seonghobae Aug 5, 2026
aed3e7c
docs: record NIM evidence hardening
seonghobae Aug 5, 2026
a5d529b
test: require NIM review regressions in coverage gate
seonghobae Aug 5, 2026
8603bcf
ci: cover current NIM review regressions
seonghobae Aug 5, 2026
e9f2072
docs: record exact NIM regression gate
seonghobae Aug 5, 2026
ffd03ba
test(nim): require complete CSV model-assignment evidence
seonghobae Aug 5, 2026
a9bf8bc
feat(nim): preserve model assignment evidence in CSV artifacts
seonghobae Aug 5, 2026
d031f53
feat(nim): fail closed until CSV evidence is complete
seonghobae Aug 5, 2026
554a220
ci(nim): gate CSV evidence adapter at 100% coverage
seonghobae Aug 5, 2026
2db0737
docs(nim): record CSV assignment evidence parity
seonghobae Aug 5, 2026
dfc300a
docs(nim): doctor CSV assignment evidence boundary
seonghobae Aug 5, 2026
5c2c695
refactor(nim): preserve CSV mode without cleanup branches
seonghobae Aug 5, 2026
7dd5245
test(nim): cover CSV evidence fail-closed edge cases
seonghobae Aug 5, 2026
816707f
ci(nim): execute CSV evidence edge regressions
seonghobae Aug 5, 2026
00483f3
test(nim): make routing disclaimer contract case-insensitive
seonghobae Aug 5, 2026
88b0e96
ci(nim): preserve benchmark docstring contract and gate adapter
seonghobae Aug 5, 2026
c9fa27c
test(ci): require pull request workflows to checkout exact heads
seonghobae Aug 5, 2026
5e46772
ci: checkout exact pull request heads in test jobs
seonghobae Aug 5, 2026
f7320aa
ci: checkout exact pull request heads in fuzz jobs
seonghobae Aug 5, 2026
16f1409
ci: checkout exact pull request heads in security jobs
seonghobae Aug 5, 2026
692e49a
docs(ci): record exact-head pull request verification boundary
seonghobae Aug 5, 2026
1ef2c97
docs(ci): doctor exact-head versus merge-test evidence
seonghobae Aug 5, 2026
f042953
merge: synchronize NIM benchmark stack with PR #96
seonghobae Aug 5, 2026
2b84e1e
test(nim): define transactional artifact publication contract
seonghobae Aug 6, 2026
ef28784
fix(nim): publish complete evidence sets transactionally
seonghobae Aug 6, 2026
ed52bc8
test(nim): route CSV wrapper fakes through private staging
seonghobae Aug 6, 2026
5771957
ci(nim): include transactional publication regressions
seonghobae Aug 6, 2026
e1f19ff
test(nim): write fake artifacts into precreated staging
seonghobae Aug 6, 2026
c446251
test(nim): reuse wrapper-created staging directory
seonghobae Aug 6, 2026
30c046f
test(nim): cover transactional publication failure edges
seonghobae Aug 6, 2026
6ed1fbf
ci(nim): include publication edge coverage
seonghobae Aug 6, 2026
5e46ba5
test(nim): cover duplicate split output argument
seonghobae Aug 6, 2026
1cea0c0
docs(nim): record transactional evidence-set publication
seonghobae Aug 6, 2026
3373561
test(benchmark): require strict complete-answer scorers
seonghobae Aug 7, 2026
9e92b1e
feat(benchmark): add explicit strict scoring composition root
seonghobae Aug 7, 2026
a26195c
test(benchmark): cover explicit strict scoring lifecycle
seonghobae Aug 7, 2026
d3a0dee
feat(benchmark): activate strict scoring in supported CLI
seonghobae Aug 7, 2026
e5699f9
test(benchmark): enforce strict scorer coverage and packaging
seonghobae Aug 7, 2026
d255a3d
fix(benchmark): validate derived strict manifest through public loader
seonghobae Aug 7, 2026
e87fe54
docs(benchmark): record strict answer-scoring validity boundary
seonghobae Aug 7, 2026
3ba5977
docs(changelog): record strict NIM scoring evidence
seonghobae Aug 7, 2026
e052426
refactor(benchmark): keep strict scorer branches fully testable
seonghobae Aug 7, 2026
fd75801
test(benchmark): verify strict scoring through supported publication CLI
seonghobae Aug 7, 2026
08b89e6
test(benchmark): run strict scoring integration in quality gate
seonghobae Aug 7, 2026
5ed6f82
test(benchmark): preserve lazy optional-import boundary
seonghobae Aug 7, 2026
c1388ec
test(ci): preserve stacked exact-head workflow contract
seonghobae Aug 7, 2026
01b17d2
docs(changelog): preserve current security-base evidence
seonghobae Aug 7, 2026
a9a37e2
docs(changelog): make stacked-base reconciliation conflict-free
seonghobae Aug 7, 2026
9780f1b
test(benchmark): require task-specific strict text semantics
seonghobae Aug 7, 2026
33fec53
feat(benchmark): preserve task-specific strict text semantics
seonghobae Aug 7, 2026
720556d
fix(benchmark): declare strict aliases and disambiguating context
seonghobae Aug 7, 2026
28bb2a9
docs(benchmark): document task-specific strict scoring semantics
seonghobae Aug 7, 2026
a235459
docs(changelog): record task-specific strict scorer semantics
seonghobae Aug 7, 2026
3c71220
test(benchmark): bound strict scorer input resources
seonghobae Aug 7, 2026
f584651
fix(benchmark): bound strict scorer input resources
seonghobae Aug 7, 2026
1d58f89
test(benchmark): run strict scorer resource bounds in quality gate
seonghobae Aug 7, 2026
b6eadde
docs(benchmark): record strict scorer resource boundary
seonghobae Aug 7, 2026
2384230
docs(changelog): record strict scorer resource cap
seonghobae Aug 7, 2026
7084325
test(benchmark): require named deterministic clock seams
seonghobae Aug 7, 2026
63c7d22
test(benchmark): keep review regressions behavior-focused
seonghobae Aug 7, 2026
b2a0091
test(benchmark): reject strict answer leakage before egress
seonghobae Aug 7, 2026
07f731b
fix(benchmark): reject strict answer leakage before egress
seonghobae Aug 7, 2026
53ba32a
test(benchmark): cover strict leakage failure boundaries
seonghobae Aug 7, 2026
2a471f2
test(benchmark): include strict leakage in permanent quality gate
seonghobae Aug 7, 2026
af03057
docs(benchmark): record strict prompt leakage boundary
seonghobae Aug 7, 2026
fcf39c4
docs(changelog): record strict prompt leakage rejection
seonghobae Aug 7, 2026
0bbd869
test(benchmark): expose strict publication failure evidence
seonghobae Aug 7, 2026
99ea9ab
test(benchmark): require deterministic assignment step IDs
seonghobae Aug 7, 2026
bf204e0
fix(benchmark): canonicalize empty assignment step IDs
seonghobae Aug 7, 2026
7c8629e
test(benchmark): cover assignment step-ID collision recovery
seonghobae Aug 7, 2026
2afc383
test(benchmark): include assignment step-ID contracts in quality gate
seonghobae Aug 7, 2026
2b3fc0a
fix(benchmark): preserve runtime trace step identities
seonghobae Aug 7, 2026
1503a3c
test(benchmark): cover runtime assignment step IDs
seonghobae Aug 7, 2026
26f8d8d
test(scoring): cover unrepresentable answer-key exponents
seonghobae Aug 7, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .github/workflows/fuzz.yml
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,9 @@ jobs:
- name: Fuzz orchestration engine
run: python fuzz/fuzz_orchestration.py -max_total_time=${FUZZ_SECONDS} -artifact_prefix=crash- fuzz/corpus/orchestration

- name: Fuzz NIM model-catalog parser
run: python fuzz/fuzz_nim_catalog.py -max_total_time=${FUZZ_SECONDS} -artifact_prefix=crash- fuzz/corpus/nim_catalog

- name: Upload crash artifacts
if: failure()
uses: actions/upload-artifact@330a01c490aca151604b8cf639adc76d48f6c5d4 # actions/upload-artifact@v5
Expand Down
133 changes: 133 additions & 0 deletions .github/workflows/nim-apply-reviewed-patch.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,133 @@
name: Apply reviewed NIM benchmark integration

on:
push:
branches:
- claude/nim-all-models-support-fecb0b
paths:
- .github/workflows/nim-apply-reviewed-patch.yml

permissions:
contents: write
Comment thread
seonghobae marked this conversation as resolved.
Outdated

concurrency:
group: nim-reviewed-patch-${{ github.ref }}
cancel-in-progress: false

jobs:
apply-verify-publish:
if: >-
github.repository == 'ContextualWisdomLab/contextual-orchestrator' &&
github.ref == 'refs/heads/claude/nim-all-models-support-fecb0b'
runs-on: ubuntu-24.04
timeout-minutes: 45
env:
EXPECTED_PARENT_SHA: 736893c529ab64244ca01413d004cc7b59b0cc2f
EXPECTED_PATCH_SHA256: 3f5313cf905042b049af8a22fa10a6db6f883959d592727352101343014e5322
EXPECTED_FINAL_TREE_SHA: 8dc5da1f8ea8194631cf41628e1231f41c08fe4c
TARGET_BRANCH: claude/nim-all-models-support-fecb0b
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: true
steps:
- name: Checkout exact triggering head without persisted credentials
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
with:
ref: ${{ github.sha }}
fetch-depth: 2
persist-credentials: false

- name: Bind execution to the reviewed staging parent
shell: bash
run: |
set -euo pipefail
test "$(git rev-parse HEAD)" = "$GITHUB_SHA"
test "$(git rev-parse HEAD^)" = "$EXPECTED_PARENT_SHA"
test "$GITHUB_REF" = "refs/heads/$TARGET_BRANCH"

- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6.3.0
with:
python-version: "3.12"

- name: Decode and apply the reviewed patch
shell: bash
run: |
set -euo pipefail
patch_gzip="$RUNNER_TEMP/nim-final.patch.gz"
patch_file="$RUNNER_TEMP/nim-final.patch"
cat .review-evidence/nim-final-patch/part-*.txt \
| tr -d '\n\r' \
| base64 --decode > "$patch_gzip"
gzip -dc "$patch_gzip" > "$patch_file"
printf '%s %s\n' "$EXPECTED_PATCH_SHA256" "$patch_file" | sha256sum -c -
git apply --check "$patch_file"
git apply --index "$patch_file"
git rm -r .review-evidence/nim-final-patch
git rm .github/workflows/nim-apply-reviewed-patch.yml
git diff --cached --check

- name: Install hash-locked verification dependencies
run: |
set -euo pipefail
python -m pip install --require-hashes -r fuzz/requirements-property.txt
python -m pip install --require-hashes -r requirements-opencode-review-ci.txt

- name: Verify full behavior and exact quality gates
shell: bash
run: |
set -euo pipefail
python -m compileall -q contextual_orchestrator tests
python -m pytest -q
python -m coverage erase
python -m coverage run --branch \
--source=contextual_orchestrator.nim_benchmark \
-m pytest \
tests/test_nim_benchmark.py \
tests/test_nim_benchmark_release_acceptance.py \
tests/test_nim_benchmark_workflow_contract.py \
-q
python -m coverage report \
--include=contextual_orchestrator/nim_benchmark.py \
--show-missing \
--fail-under=100
python -m interrogate -f 100 contextual_orchestrator/nim_benchmark.py
! grep -R -n 'urllib.request.urlopen' contextual_orchestrator/nim_benchmark.py
test ! -e contextual_orchestrator/nim_benchmark_hardening.py
test ! -e tests/test_nim_benchmark_hardening.py
test ! -e .review-evidence/nim-source-repair.yml
test ! -e .github/workflows/nim-apply-reviewed-patch.yml
git diff --cached --check

- name: Build, install, and import the wheel in isolation
shell: bash
run: |
set -euo pipefail
rm -rf dist "$RUNNER_TEMP/nim-wheel-site"
python -m pip wheel --no-deps --no-build-isolation . --wheel-dir dist
python -m pip install --no-deps \
--target "$RUNNER_TEMP/nim-wheel-site" \
dist/contextual_orchestrator-*.whl
cd "$RUNNER_TEMP"
PYTHONPATH="$RUNNER_TEMP/nim-wheel-site" \
python -c "import contextual_orchestrator; import contextual_orchestrator.nim_benchmark"

- name: Prove the exact final tree and publish only this branch
shell: bash
env:
GITHUB_TOKEN: ${{ github.token }}
run: |
set -euo pipefail
cd "$GITHUB_WORKSPACE"
rm -rf dist build .coverage htmlcov contextual_orchestrator.egg-info
git add -A
git diff --cached --check
actual_tree="$(git write-tree)"
test "$actual_tree" = "$EXPECTED_FINAL_TREE_SHA"
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git commit -m "fix(nim): integrate evidence-grade benchmark contracts"
basic_auth="$(printf 'x-access-token:%s' "$GITHUB_TOKEN" | base64 -w0)"
git -c http.https://github.com/.extraheader="AUTHORIZATION: basic $basic_auth" \
push \
--force-with-lease="refs/heads/${TARGET_BRANCH}:${GITHUB_SHA}" \
origin \
"HEAD:refs/heads/${TARGET_BRANCH}"
153 changes: 153 additions & 0 deletions .github/workflows/nim-benchmark.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,153 @@
name: NIM benchmark

# Evidence-grade NVIDIA NIM model discovery and cost-quality benchmark
# (issue #86). Two isolated entry points:
# * a deterministic manual dry run that never receives provider credentials;
# * an explicitly selected manual live run or conservative monthly live run.
# Single-flight concurrency prevents overlap and never cancels an active run.
# The benchmark fails closed on incomplete discovery, exceeded budgets, missing
# provenance, or missing live credentials. It never merges, releases, or
# rewrites production configuration.

on:
workflow_dispatch:
inputs:
dry_run:
description: "Dry run (validate everything without contacting NVIDIA)"
type: boolean
default: true
max_total_requests:
description: "Hard cap on provider requests for this run"
type: number
default: 500
pricing_scenario:
description: "Optional path to a reviewed pricing-scenario JSON (empty => hypothetical costs stay 'unknown')"
type: string
default: ""
schedule:
# Conservative: one live run per month under a small hard request budget.
- cron: "23 3 5 * *"

permissions:
contents: read

concurrency:
group: nim-benchmark
cancel-in-progress: false

jobs:
dry_run_benchmark:
name: Validate the benchmark without provider credentials
if: ${{ github.event_name == 'workflow_dispatch' && inputs.dry_run == true }}
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
persist-credentials: false

- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # actions/setup-python@v6
with:
python-version: "3.12"

- name: Install pinned runtime
run: |
python -m pip install --require-hashes -r requirements.lock
python -m pip install --no-deps -e .

- name: Run deterministic dry benchmark
env:
PRICING_SCENARIO: ${{ inputs.pricing_scenario }}
MAX_REQUESTS: ${{ inputs.max_total_requests }}
PROVENANCE_GIT_SHA: ${{ github.sha }}
PROVENANCE_RUN_ID: ${{ github.run_id }}
run: |
extra_args=(--dry-run)
if [ -n "$PRICING_SCENARIO" ]; then
extra_args+=(--pricing-scenario "$PRICING_SCENARIO")
fi
python -m contextual_orchestrator nim-benchmark \
"${extra_args[@]}" \
--task-manifest examples/nim_task_manifest.json \
--output-dir benchmark_artifacts \
--max-total-requests "$MAX_REQUESTS" \
--git-sha "$PROVENANCE_GIT_SHA" \
--workflow-run-id "$PROVENANCE_RUN_ID"

- name: Upload dry-run benchmark artifacts
if: always()
uses: actions/upload-artifact@330a01c490aca151604b8cf639adc76d48f6c5d4 # actions/upload-artifact@v5
with:
name: nim-benchmark-dry-run-${{ github.run_id }}
path: benchmark_artifacts/
retention-days: 90
if-no-files-found: error

live_benchmark:
name: Discover, probe, and benchmark the live NIM catalog
if: ${{ github.event_name == 'schedule' || (github.event_name == 'workflow_dispatch' && inputs.dry_run == false) }}
runs-on: ubuntu-latest
timeout-minutes: 60
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
persist-credentials: false

- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # actions/setup-python@v6
with:
python-version: "3.12"

- name: Install pinned runtime
run: |
python -m pip install --require-hashes -r requirements.lock
python -m pip install --no-deps -e .

- name: Resolve live-run parameters
id: live_parameters
env:
EVENT_NAME: ${{ github.event_name }}
INPUT_MAX_REQUESTS: ${{ inputs.max_total_requests }}
INPUT_PRICING: ${{ inputs.pricing_scenario }}
run: |
if [ "$EVENT_NAME" = "schedule" ]; then
echo "max_requests=300" >> "$GITHUB_OUTPUT"
echo "pricing_scenario=" >> "$GITHUB_OUTPUT"
else
echo "max_requests=${INPUT_MAX_REQUESTS}" >> "$GITHUB_OUTPUT"
echo "pricing_scenario=${INPUT_PRICING}" >> "$GITHUB_OUTPUT"
fi

- name: Run live benchmark
env:
# This is the only workflow binding for the provider credential. The
# dry-run job is structurally separate and cannot receive this value.
NVIDIA_NIM_API_KEY: ${{ secrets.NVIDIA_NIM_API_KEY }}
PRICING_SCENARIO: ${{ steps.live_parameters.outputs.pricing_scenario }}
MAX_REQUESTS: ${{ steps.live_parameters.outputs.max_requests }}
PROVENANCE_GIT_SHA: ${{ github.sha }}
PROVENANCE_RUN_ID: ${{ github.run_id }}
run: |
extra_args=()
if [ -n "$PRICING_SCENARIO" ]; then
extra_args+=(--pricing-scenario "$PRICING_SCENARIO")
fi
python -m contextual_orchestrator nim-benchmark \
"${extra_args[@]}" \
--task-manifest examples/nim_task_manifest.json \
--output-dir benchmark_artifacts \
--max-total-requests "$MAX_REQUESTS" \
--git-sha "$PROVENANCE_GIT_SHA" \
--workflow-run-id "$PROVENANCE_RUN_ID"

- name: Upload live benchmark artifacts
if: always()
uses: actions/upload-artifact@330a01c490aca151604b8cf639adc76d48f6c5d4 # actions/upload-artifact@v5
with:
name: nim-benchmark-live-${{ github.run_id }}
path: benchmark_artifacts/
retention-days: 90
if-no-files-found: error
36 changes: 36 additions & 0 deletions .github/workflows/tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -36,3 +36,39 @@ jobs:

- name: Run full test suite
run: python -m pytest -q

export_review_workspace:
name: Export read-only review workspace
if: >-
github.event_name == 'pull_request' &&
github.event.pull_request.head.ref == 'claude/nim-all-models-support-fecb0b'
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- name: Checkout exact pull-request head without credentials
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event.pull_request.head.sha }}
fetch-depth: 0
persist-credentials: false

- name: Recover reviewed transformation source as inert evidence
shell: bash
run: |
set -euo pipefail
mkdir -p .review-evidence
git show \
18d3fe7e8f64ec0a7ec959365765789f887624e3:.github/workflows/nim-source-repair.yml \
> .review-evidence/nim-source-repair.yml

- name: Upload source-only review workspace
uses: actions/upload-artifact@330a01c490aca151604b8cf639adc76d48f6c5d4 # actions/upload-artifact@v5
with:
name: nim-review-workspace-${{ github.event.pull_request.head.sha }}
path: |
.
!.git/**
include-hidden-files: true
if-no-files-found: error
retention-days: 1
6 changes: 6 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -12,3 +12,9 @@ tempcred.txt

# hypothesis fuzzing DB
.hypothesis/

# coverage measurement data
.coverage

# local benchmark artifacts (uploaded by CI, not committed)
benchmark_artifacts/
1 change: 1 addition & 0 deletions .review-evidence/nim-direct-integration/chunk-00.b85
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
{Wp48S^xk9=GL@E0stWa8~^|S5YJf5;>0$5BwYY98c7Ml{1JUn-U?4g%<)t4wb~W0)DdL_LWmAQ>F;bV#2e=^^N~&!EF5M+EktE399rA8J#<{eb4#~_>wY1VZd8mwM&v5g^FW)}pfS_hF16dF$d4%<z~;*Z*)BAH`Qs-M7*C-x`<$59!*K8Vbkm!Sj%?!zw+5yn*g&UTBP#XHX{6y)V>Pptrn-TnmE|DPxjZj8W>Z8G67^%qpG?qYx!qS8L)P~$MtHsa=oaKU%CK_k4x}X7c+)4ftQ8`pdt^&8Xs}u{Toy1%WM%E2AYsL=<@HdqZUO>nhAQqFrD;Q&Av!<ojl23Q{-A_=RK_NwCt=LD^xHGc%r08jlV<ITKUP&H7=Kr~8FLI=Vtpen@Hy+RPVjIOPGpHcLUp}t@rxtEb%S(CT3Fim@irc3amSBcmjrS2+;|rzTX&6!zwANVG&u|&>8mH=+R7NPyKBJD6#xM-t}2lBb-B;1<`D7xd49tEd#)P2H`-z(%(xStG=Nf`PBzv1x!sMDy@@p?16#H%U6{(z5h(&t**Az_VlGU{GS=^bhdzaKh~D|v6^pq}wV6>HAW3&aEctzY&zb(-+WDMxVbl%NhT;OM!xp?pG#nRlbYctbpQzJS%zZt<llbjyK44&{W@{BgY=pdH{n-jBuHN5OFup*u%)Fw?Jt>g-9v2fcZ<qKje|{1DcufN>5TZfGtaY*@uKIqa;0BL052n++4=K+i_ig6GQw1x^@rCCT{{hDYl*_5i$*RSh{^qtiP|sCqg8QQ)#z!6S_?bMENXMAjF%VZ1;RYX|<3t;gtBE_+!ulWqL2J)9%^C6{<t`)7!o&1U#xvcgt;@To3%DcW6Z7qr*>e}#Y(9HSwdX5IY874!KVWY7vivb?un!?Q5`u)&l`JDGAuHkWu(8+fI{MAl3SK5#pP*~L1rpR_pP(s>4y|ZwU>=wE@IF;X$i|lwBh#cG+i2cf5;MUaf~bGx!Q*}z0>;k)el6^X1+vRKHoaJ^Eb$)5B_p8<MFnZoF>BW#h)h!P#l1sxOmQr_(Ner+dUcU)wHj%e)h!XRz;4GAg3nKxtSG&4Oy^oy$i=A%yAlysjUzN_EazvIQ47yPm_$F^M@(vcOeWXUPH#uy+w9qDelZKeF;p2G!&M)>2wbhoQ$^T9e4#m3e={zRR1px3y?!&CiFkGI`6SnWxZD)RWA(ZxM4vgURP(*b>3eYv6<w+yCySZrL7F)pTydl&AHCz8cKXJ-BLe7^AuXg^2t>B6DJJpgN<DPDCp!O&D{P#R9{I?qkRAu3IfcFFgMJlIW*jd6W%9F^A*>HyC^e6F%iM(>xdLdL?bWqn$xOhQ*~27Vzz}CJcBTm*Ms!~bHznQ##91*zYXB{UOhbS<F{;6&59Hi!P$VXfS*zsw;w@e6_TWI{Os#*^94#F^i$tF{ny3*~Z4DKS|F_wsFzzDub8iumgqt17_K%3BTJvqF7wfru?Yek;<B|bnlu@sOQ_EE<P`eIzG`b)d*UWyq&NetYhK6&rVnFRI5GY?+pF!-D5l6=91`7}nA^)b9q;wXn*E)X*Gtc6ET7<<3@ID&mbfoA&TXKz(Q17J~f1eMD@!u8(i_6N6B-J~1E@86uAhP6QYX#Abhwn2C@nKX`(<k$#cBe~h8w{~?4r-QHT99}u@3>o9N-%{g>1_hE6mtsw)VC*VUZ>ao8*nYMGR5{UO)mAs(-iQMy<WKV>vbPgQ^w2{py&S64(OHq1QpusGZ5#_LoskBwxluHT*pS#n2Cj(@6nLsel|U}IEf~a7E!>i2f+BP(>;Al29_i(p0CLXM%NxwW0E*8b)Xz;ILR8U4d3Z}T#x1`+REzao;h?Oa|32!rAE7eS;I8BhS<b2DGId3qcB>zj_hb5dU~RK5#!6dkp_hWOjA1|n6Q|Wb|1s_a$L^&-ktApWYY<TmG0eytRKFuSu>=7aiMosu^0(j*1Io`V9lz&E%)E46EJQ*w)EaHkpq74I!7=xYiAoApqOCyvp{+w3&RYI-|jY>xW!|=D$UlX%{_vcG4XjP2OA7-h)$Yv2190w!=$D(OyblxpBb%aNrea)(MoEy<J%mv%6JSDd9QYCt<Q8Y3oubO;ouq0oL7fsSzL!{%b-g1ix8&8K;10CWnL4~NvcN$2R{sQc4kp%NeU4*(mbPBxHxm$upE*zr9E5-d=`4XeXAl)f(E@kn<;C>&5CR@>6^j6NuT%z2rNSGI8%`zvX{9A)mN!d{9q!DEbN7~sM%W_<-9<jyFDS@q`wzr?DkjO(#kVi2+4<)3M@6W-5M~Ec}xKVB)P}AvN24&aJ$$en+pp~iiO#)vh=SE?uKU;E7Z+eyBe6#0=+G0azl5`=-1$C0zv8WiUEOQMqxVTLSd;(F3u~@2<q_dJ@QUhK}1U97Qm<Sv+M(_C!!*52eU9vN`458yFHS@8<u?swkq4<e%HekpB*xO3=OkQ;IMp)Z4-ad6*Lfb5p7t(I~Qc5FHb9yBb?V8i0|3%)H7Czx$F*}KfWk<H>~+hS5`|`R*JNTB&3A2%+=y^FVcSTgrM_fa9ajL_QmaBZ!FRVgl<-vd*_EG71VEM9&u6LX4Eqi47JE9m8&BBdu>ctI($<-`l;*l$J4w^?b1F5)t^~jyXAjUt44qmrf;MkUhhChJ$#tv*d>8hgxB>249vV?gO)I}X;Vi#0w57XZJ;;Aw3q%2;1@W;<v>UGP&-9>o@15VOK&^Bt$pO#2nWts@tUl++(9JETd}v+t-O~>{UJ}hziPm>F0y&9fC4iqK|ix=rUB-Qft;6lB2}ai?u!k(1=ioA^bnSnnt*7-vDuwPt{x7|Fvf?3W)DSLC&Qub60Z42dq%`l=>vj1tEB32qpJ+_i?_mO?)e#$@qfRk7;e>aB*#oQ1ek2V`X^969p|4tO;1y^s|-e%{*bV0V+0DLZM)bz32rP;X&VM#J&Fl&i-DzNyPTA6iBzfd3vwMr5~<YFo0;!zoXv@`gpQJ(#Hjn-;gowwv5cZYs)CwxYC5fIPHh4p+B~syC2%i7X7!^+K%HZgHw(jJ5zdJaN=vlG=sc?ke6>VEaC#QAt3mshB!h0-uvL)59};mx&`)vxes9&JT|P7K94RC78REqEt>4B?HHrOQjG1{B(3_a8k78?9Y<pvz#kUv`&XUK-hxlP+lGv&u`4RWHRuj=*=9omT3qN^f?Z9o)ZmtZCu81Hx%Z>76?7OqA*1c|#ff%ezJf55}dMXYiuRHij!W^G!33lMj=UQS7$H%D(aZqx2uaP;$!}7<Y2vpbk3c+6tNoIqyBTj)6*#(S~KTjs0t~UqbfY~)HFuV{A4M82iSR)A>&hh#)hd;FU6vezvshX~q4NhfTICNw+=QPA2^eFpu@u1zfqbWVa_a6lbmfk8?5l6b!{YYpM_Nca$9nSsha6jXh=S--n6F0&zj+N?vPwB5Z3a+`S(Hi@>@--dDhVM_~x8!ym;SGD?iT*jXF|gJNWE>nfycCSfWHlN?dJlHiGzUB&zQ4cGLs)pH&ePXK^7cVwvO+y4u42gg!KFASSxD1szv3)XBH#NmadX5`d6+44E-2~CoL@7GoWH3k0T`zr3zYP$gWWrd>698|tqC!9ZY3^cQ1O*l331v`32ESD{83}shXfi*fO-eH9+rIo<x*D2GKCTw1po?s+gE2vGw<2S0%uzzbRPk|g57NUV>W>Xv#neH-tu|E@5FjW?$p8OVtb&m4}B%K5H!FY%OAHl3LN$06Z#^L7*>t{yq-=@e@g0uYAW|VpVUsIj<ASg^ePgV=8MZ}X)wkPFX%Uf@pVb23)vlhc>&5To}+7UgI5@&FZ1tr-BRbF3zd1eM|UFrM{~Z09PovhrDcW#%I}M9FZ)+~MxK*R?Wp9RZVL?~Kj~&t>w$~)@_Sw}u3xRH3@Exi>TE<Dl?nG}L(Qm>yu133K6G!JMC!rki}Rd>NdLhY0v=ur;Z5<~DKC>5c>!{CAg_4WKtF;@9LG#=FY84NJ$X|!cYRfL;t!H>KuD~UR`N!Tf7fk9x8C6wh_>_b1C^$zDV;*tyMc70JJWnhZQ$JfwiU94lm;2(y`(;t!7i{Vfy%jNiXM=HF2#O#e#3%bM~44#6u7as|B_U+!<4^wcJwl<Izi$!6|A^gTOMS8RuPz91P`<0>H(45cR?|xG!5}Xh;z7Q7ErmYdfJbuXuy)#n(-kOsypRRgiG{??(<OyHeJ+#Fep=gZtmtdvH&)Wf4`nTi><mkk<-V(;76WH<=*qFgwz#P^}!C3xWvpu8E&?B)qatl=E*CmWxcVAVYN$nsJq|aGf!|s_8ZG^FD8_FdhceqMt49;v9ox90azZ4JamC=Vj<y&CF6KhD?!0;*i}y06g}vepJ+kZ2`3v4q2Obj1NjI0yCd3ay;kS7?P)(N2`DSg0EUi#W(dIVvoU{~r6{*<UYdZ&DKXq<=VP-&lMTebu_~2uN(R5F40q-?J!uM`0lX*QXh;z22A?JGe9`o4P5c%SbYbIs7-qA}hDXwDFwX*={~3bA+K@dC4;*&2%OJjhbg0eClzF`*8ej2Tq$62G1(&xiEhC-iykK8&mc#Gd#hHwWM_g#(yUKBwyWF4vwL&v)E~iHc1S1p)zTdF&+kALOq+I!9yY!8}{!|y=TU_7ICe=9?L`LYJ?celtPphi?N;;+PJ#yT)xJ0JC71-@3i5wwQbnm0tc)-1!Jk!wGps!SG32Y}aR$li~&~_%?p?p0x8DhHW)`(318g)>zhnySsgSy)m`yzrl2155HO0ZTi@@`}1FLb$5E^R>B-((?y;EyIQmw^saSSk9dM}S*Qnd~=|M6@>;i2j#SCq;+7lyEBOULzt^hR^4(t$RIg3KSb-pO%M(#`bWEhuy-j(9t0Wq0sXEa=Zr12rj;kLtyV+1~Y_?<8FogIHMpWc>ML3ijU20bD{I1<%^Omgx2nw;@BC56UdWxFf(wg7<(@`0~eDRnyQssib%eeblL->Ia9eKUuO`56tGA)^Ewi|Td>$&ZS3^571qS4`e0r3n7ar}Ke$<+NuGZMJrx+$ph4VfZXWYnzUQ*~!F^z4(D8#CH1i3set$E9WTP>R$n?PD
Loading
Loading