-
Notifications
You must be signed in to change notification settings - Fork 1
feat: evidence-grade NIM discovery + all-modality cost-quality benchmark #90
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Closed
seonghobae
wants to merge
200
commits into
fix/atheris-interpreter-lock
from
claude/nim-all-models-support-fecb0b
Closed
Changes from 43 commits
Commits
Show all changes
200 commits
Select commit
Hold shift + click to select a range
064973e
feat: evidence-grade NIM discovery + all-modality cost-quality benchm…
seonghobae b48f396
merge: stack NIM benchmark on provider-egress security base
seonghobae 0e19b3a
docs(changelog): record NIM benchmark evidence slice
seonghobae 66c8d55
ci: apply reviewed NIM transport hardening
seonghobae 9d546a7
fix(nim): harden egress and equalize policy budgets
seonghobae 785fb15
fix(nim): install evidence and transport hardening
seonghobae e07a504
test(nim): prove secure transport and equal budgets
seonghobae d0312ea
docs(nim): document pinned egress and equal total budgets
seonghobae c7bcad2
chore(ci): remove temporary NIM repair workflow
seonghobae e1ae745
chore(ci): stage one-time NIM source repair
seonghobae 1015608
chore(ci): remove privileged one-shot NIM repair workflow
seonghobae 07128ac
chore(ci): stage exact-head NIM security repair
seonghobae a8393d7
chore(ci): harden bounded NIM repair job
seonghobae 4cc0f55
chore(ci): stage bounded NIM source integration
seonghobae 9ad742a
chore(ci): remove overlapping one-shot NIM workflow
seonghobae 4fd96f1
fix(ci): replace write-capable PR repair with isolated publisher
seonghobae 4bb7c85
docs(changelog): record reviewed NIM access evidence
seonghobae d9bedd7
chore(ci): stage equal-budget regression repair
seonghobae c4d47d2
ci: complete bounded NIM source integration
seonghobae db6ebb9
chore(ci): remove redundant NIM test-only repair workflow
seonghobae dba1d7e
ci: run bounded NIM source integration from PR checks
seonghobae a3741b5
fix(ci): execute bounded NIM source transformation deterministically
seonghobae 18d3fe7
fix(ci): preserve embedded transformation literals
seonghobae 27eb1b8
fix(security): remove PR-head OIDC source mutation
seonghobae ba6b34b
fix(security): delete privileged branch-trigger repair workflow
seonghobae e4578da
fix(security): delete branch-controlled OIDC repair source
seonghobae f80e503
test(security): pin PR workflow OIDC boundary
seonghobae bc77bfc
ci(review): export read-only exact-head workspace
seonghobae a71ebbc
ci(review): include inert transformation evidence
seonghobae 6df72ee
test(ci): require secretless NIM dry-run job
seonghobae 09fd519
fix(ci): isolate NIM dry runs from provider secrets
seonghobae 2b622bd
docs(security): record NIM workflow secret isolation
seonghobae 40b48ba
docs(changelog): record secretless NIM dry runs
seonghobae 348e5e6
docs(benchmark): explain workflow credential isolation
seonghobae 667b311
chore(review): stage verified NIM repair patch
seonghobae 2498648
ci(nim): stage reviewed patch part 1 of 7
seonghobae 4552084
ci(nim): stage reviewed patch part 2 of 7
seonghobae 45c151e
ci(nim): stage reviewed patch part 3 of 7
seonghobae 9dac6b2
ci(nim): stage reviewed patch part 4 of 7
seonghobae fecd25f
ci(nim): stage reviewed patch part 5 of 7
seonghobae 19be572
ci(nim): stage reviewed patch part 6 of 7
seonghobae 736893c
ci(nim): stage reviewed patch part 7 of 7
seonghobae 07f28a9
ci(nim): apply exact reviewed benchmark integration
seonghobae 7d5d6aa
ci: retrigger reviewed NIM integration with scoped permissions
seonghobae fd6b074
ci: verify and export reviewed NIM integration read-only
seonghobae cb7e9b3
ci: bind NIM verification to security ancestor
seonghobae 3878e5b
fix(ci): bind NIM verification to pull-request head
seonghobae 5bbedc2
fix(ci): reconcile reviewed NIM patch per file
seonghobae d98fe6c
fix(ci): reconstruct reviewed NIM tree from patch base
seonghobae 428736c
fix(ci): bind reviewed NIM reconstruction to exact head
seonghobae e4e03c4
fix(ci): run reviewed NIM verification read-only on push
seonghobae 6cd2012
fix(ci): use runner temp variable at execution time
seonghobae 7595919
fix(ci): discover the exact reviewed patch base
seonghobae a93ec20
chore(ci): inspect the staged NIM integration payload
seonghobae f11193d
fix(ci): replay the reviewed NIM patch stack
seonghobae 83b7906
fix(ci): rehydrate the reviewed patch index
seonghobae e9bfb85
fix(ci): execute the patch rehydration helper
seonghobae 88274c0
fix(ci): apply the reviewed patch against the rehydrated index
seonghobae 03b30e5
fix(ci): verify the reconstructed reviewed tree
seonghobae e256a9a
fix(ci): locate the exact reviewed patch context
seonghobae 012f7c7
fix(ci): build the wheel in an isolated backend environment
seonghobae dde856f
ci(nim): publish the authenticated verified source tree
seonghobae 8049ca8
feat(nim): integrate verified provider-neutral benchmark
seonghobae d1516a5
docs(nim): record immutable integration provenance
seonghobae f090ffc
test(ci): require stacked pull-request coverage
seonghobae d091b69
ci: validate stacked pull requests
seonghobae 0d66e0a
ci: fuzz stacked pull requests
seonghobae 928214c
ci: secure stacked pull requests
seonghobae c759cd1
test(ci): require isolated wheel builds
seonghobae 8e19d64
fix(ci): isolate benchmark wheel builds
seonghobae e7f7238
ci(review): export read-only PR90 workspace
seonghobae c72f1cb
test(benchmark): stage complete-budget preflight red tests part 1
seonghobae cb20657
test(benchmark): stage complete-budget preflight red tests part 2
seonghobae 636c3e0
fix(benchmark): stage complete request planner part 3
seonghobae 1053f19
docs(benchmark): stage complete-budget evidence and verification part 4
seonghobae 638a126
ci(benchmark): apply complete-budget preflight test-first
seonghobae 1f6da53
ci(review): stage exact PR90 request-plan patch
seonghobae 8a31de7
ci(review): refresh exact PR90 request-plan patch
seonghobae 87496c8
ci(review): verify and publish exact PR90 request-plan fix
seonghobae b3c0471
ci(review): inspect exact PR90 patch hash
seonghobae 2efcf08
ci(review): export decoded PR90 patch evidence
seonghobae a06f9c5
ci(review): verify exact staged PR90 patch
seonghobae 77cc8cc
ci(review): publish verified PR90 request-plan fix
seonghobae 796e6e9
test(benchmark): align scheduled complete-plan contract
seonghobae 5956214
test(benchmark): require request-plan provenance schema
seonghobae 9b72816
fix(nim): require a complete catalog benchmark plan
seonghobae a2af363
merge: integrate provider-egress security base into NIM benchmark
github-actions[bot] c5b3876
docs(adr): record exact NIM security integration evidence
seonghobae d43adaa
test(benchmark): expose undersized locked evaluation manifest
seonghobae 16b5016
feat(benchmark): expand locked manifest to evidence floor
seonghobae ac13a56
test(benchmark): bind evidence floor to existing hard ceiling
seonghobae 62e0e59
docs(benchmark): record thirty-task evidence floor
seonghobae 43c629f
docs(benchmark): align request plan and evidence floor
seonghobae ff51584
docs(benchmark): doctor thirty-task evidence floor
seonghobae bec1129
test(benchmark): expose undersized default request caps
seonghobae cdfe08e
ci(benchmark): test and repair default request caps
seonghobae 37c78b4
ci(benchmark): rearm default request-cap repair
seonghobae 545cd64
ci(benchmark): add exact-head reopen trigger
seonghobae e56d1a3
ci(benchmark): use resolvable checkout action
seonghobae 5216171
ci(benchmark): repair complete-plan default finalizer
seonghobae 8cae8d4
test(benchmark): bind evidence floor to canonical request planner
seonghobae e31f901
ci(benchmark): cover all complete-plan regression anchors
seonghobae da8ec5f
ci(benchmark): repair invalid one-shot YAML
seonghobae 1b97e48
ci(benchmark): split workflow-authorized publication
seonghobae 0ebb1d5
fix(benchmark): align complete-plan source evidence
github-actions[bot] 9e4b50b
fix(benchmark): align manual request cap with complete plan
seonghobae 821b58e
ci(benchmark): remove completed one-shot workflow
seonghobae 7b23bac
test(nim): cover buyer-facing complete request plan
seonghobae 5c1a3f6
ci(nim): include request-plan regression in quality gate
seonghobae f5a494f
test: require NIM parser Atheris instrumentation
seonghobae 48f007d
fix: instrument NIM parser imports for Atheris
seonghobae 2736113
docs: distinguish historical NIM integration evidence
seonghobae b60ec5d
test: cover NIM review regressions
seonghobae bdc9673
chore: defer broader NIM regression batch
seonghobae 646c417
docs: make NIM receipt head-stable
seonghobae 318feb9
docs: record NIM fuzz instrumentation repair
seonghobae 63438d0
test(nim): require completion-dependent README evidence status
seonghobae 6fc53f1
docs(nim): clarify completion-dependent evidence status
seonghobae 4a79964
test(nim): normalize README contract whitespace
seonghobae 1217fc9
docs(changelog): record NIM evidence-status clarification
seonghobae 11353d4
test(security): inspect every pull-request workflow structurally
seonghobae 67c5c84
test(nim): enforce ordered non-empty secret-isolation steps
seonghobae 989d7fe
test(security): parse pull-request branch filters structurally
seonghobae 4329081
test(nim): lock review regressions before repair
seonghobae 715ee76
fix(fuzz): require normalized NIM catalog depth failures
seonghobae 8c4ee86
test: skip fixture-bound standalone NIM cases
seonghobae 4ca0448
fix: close NIM evidence review gaps
seonghobae aed3e7c
docs: record NIM evidence hardening
seonghobae a5d529b
test: require NIM review regressions in coverage gate
seonghobae 8603bcf
ci: cover current NIM review regressions
seonghobae e9f2072
docs: record exact NIM regression gate
seonghobae ffd03ba
test(nim): require complete CSV model-assignment evidence
seonghobae a9bf8bc
feat(nim): preserve model assignment evidence in CSV artifacts
seonghobae d031f53
feat(nim): fail closed until CSV evidence is complete
seonghobae 554a220
ci(nim): gate CSV evidence adapter at 100% coverage
seonghobae 2db0737
docs(nim): record CSV assignment evidence parity
seonghobae dfc300a
docs(nim): doctor CSV assignment evidence boundary
seonghobae 5c2c695
refactor(nim): preserve CSV mode without cleanup branches
seonghobae 7dd5245
test(nim): cover CSV evidence fail-closed edge cases
seonghobae 816707f
ci(nim): execute CSV evidence edge regressions
seonghobae 00483f3
test(nim): make routing disclaimer contract case-insensitive
seonghobae 88b0e96
ci(nim): preserve benchmark docstring contract and gate adapter
seonghobae c9fa27c
test(ci): require pull request workflows to checkout exact heads
seonghobae 5e46772
ci: checkout exact pull request heads in test jobs
seonghobae f7320aa
ci: checkout exact pull request heads in fuzz jobs
seonghobae 16f1409
ci: checkout exact pull request heads in security jobs
seonghobae 692e49a
docs(ci): record exact-head pull request verification boundary
seonghobae 1ef2c97
docs(ci): doctor exact-head versus merge-test evidence
seonghobae f042953
merge: synchronize NIM benchmark stack with PR #96
seonghobae 2b84e1e
test(nim): define transactional artifact publication contract
seonghobae ef28784
fix(nim): publish complete evidence sets transactionally
seonghobae ed52bc8
test(nim): route CSV wrapper fakes through private staging
seonghobae 5771957
ci(nim): include transactional publication regressions
seonghobae e1f19ff
test(nim): write fake artifacts into precreated staging
seonghobae c446251
test(nim): reuse wrapper-created staging directory
seonghobae 30c046f
test(nim): cover transactional publication failure edges
seonghobae 6ed1fbf
ci(nim): include publication edge coverage
seonghobae 5e46ba5
test(nim): cover duplicate split output argument
seonghobae 1cea0c0
docs(nim): record transactional evidence-set publication
seonghobae 3373561
test(benchmark): require strict complete-answer scorers
seonghobae 9e92b1e
feat(benchmark): add explicit strict scoring composition root
seonghobae a26195c
test(benchmark): cover explicit strict scoring lifecycle
seonghobae d3a0dee
feat(benchmark): activate strict scoring in supported CLI
seonghobae e5699f9
test(benchmark): enforce strict scorer coverage and packaging
seonghobae d255a3d
fix(benchmark): validate derived strict manifest through public loader
seonghobae e87fe54
docs(benchmark): record strict answer-scoring validity boundary
seonghobae 3ba5977
docs(changelog): record strict NIM scoring evidence
seonghobae e052426
refactor(benchmark): keep strict scorer branches fully testable
seonghobae fd75801
test(benchmark): verify strict scoring through supported publication CLI
seonghobae 08b89e6
test(benchmark): run strict scoring integration in quality gate
seonghobae 5ed6f82
test(benchmark): preserve lazy optional-import boundary
seonghobae c1388ec
test(ci): preserve stacked exact-head workflow contract
seonghobae 01b17d2
docs(changelog): preserve current security-base evidence
seonghobae a9a37e2
docs(changelog): make stacked-base reconciliation conflict-free
seonghobae 9780f1b
test(benchmark): require task-specific strict text semantics
seonghobae 33fec53
feat(benchmark): preserve task-specific strict text semantics
seonghobae 720556d
fix(benchmark): declare strict aliases and disambiguating context
seonghobae 28bb2a9
docs(benchmark): document task-specific strict scoring semantics
seonghobae a235459
docs(changelog): record task-specific strict scorer semantics
seonghobae 3c71220
test(benchmark): bound strict scorer input resources
seonghobae f584651
fix(benchmark): bound strict scorer input resources
seonghobae 1d58f89
test(benchmark): run strict scorer resource bounds in quality gate
seonghobae b6eadde
docs(benchmark): record strict scorer resource boundary
seonghobae 2384230
docs(changelog): record strict scorer resource cap
seonghobae 7084325
test(benchmark): require named deterministic clock seams
seonghobae 63c7d22
test(benchmark): keep review regressions behavior-focused
seonghobae b2a0091
test(benchmark): reject strict answer leakage before egress
seonghobae 07f731b
fix(benchmark): reject strict answer leakage before egress
seonghobae 53ba32a
test(benchmark): cover strict leakage failure boundaries
seonghobae 2a471f2
test(benchmark): include strict leakage in permanent quality gate
seonghobae af03057
docs(benchmark): record strict prompt leakage boundary
seonghobae fcf39c4
docs(changelog): record strict prompt leakage rejection
seonghobae 0bbd869
test(benchmark): expose strict publication failure evidence
seonghobae 99ea9ab
test(benchmark): require deterministic assignment step IDs
seonghobae bf204e0
fix(benchmark): canonicalize empty assignment step IDs
seonghobae 7c8629e
test(benchmark): cover assignment step-ID collision recovery
seonghobae 2afc383
test(benchmark): include assignment step-ID contracts in quality gate
seonghobae 2b3fc0a
fix(benchmark): preserve runtime trace step identities
seonghobae 1503a3c
test(benchmark): cover runtime assignment step IDs
seonghobae 26f8d8d
test(scoring): cover unrepresentable answer-key exponents
seonghobae File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,133 @@ | ||
| name: Apply reviewed NIM benchmark integration | ||
|
|
||
| on: | ||
| push: | ||
| branches: | ||
| - claude/nim-all-models-support-fecb0b | ||
| paths: | ||
| - .github/workflows/nim-apply-reviewed-patch.yml | ||
|
|
||
| permissions: | ||
| contents: write | ||
|
|
||
| concurrency: | ||
| group: nim-reviewed-patch-${{ github.ref }} | ||
| cancel-in-progress: false | ||
|
|
||
| jobs: | ||
| apply-verify-publish: | ||
| if: >- | ||
| github.repository == 'ContextualWisdomLab/contextual-orchestrator' && | ||
| github.ref == 'refs/heads/claude/nim-all-models-support-fecb0b' | ||
| runs-on: ubuntu-24.04 | ||
| timeout-minutes: 45 | ||
| env: | ||
| EXPECTED_PARENT_SHA: 736893c529ab64244ca01413d004cc7b59b0cc2f | ||
| EXPECTED_PATCH_SHA256: 3f5313cf905042b049af8a22fa10a6db6f883959d592727352101343014e5322 | ||
| EXPECTED_FINAL_TREE_SHA: 8dc5da1f8ea8194631cf41628e1231f41c08fe4c | ||
| TARGET_BRANCH: claude/nim-all-models-support-fecb0b | ||
| FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: true | ||
| steps: | ||
| - name: Checkout exact triggering head without persisted credentials | ||
| uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0 | ||
| with: | ||
| ref: ${{ github.sha }} | ||
| fetch-depth: 2 | ||
| persist-credentials: false | ||
|
|
||
| - name: Bind execution to the reviewed staging parent | ||
| shell: bash | ||
| run: | | ||
| set -euo pipefail | ||
| test "$(git rev-parse HEAD)" = "$GITHUB_SHA" | ||
| test "$(git rev-parse HEAD^)" = "$EXPECTED_PARENT_SHA" | ||
| test "$GITHUB_REF" = "refs/heads/$TARGET_BRANCH" | ||
|
|
||
| - name: Set up Python | ||
| uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6.3.0 | ||
| with: | ||
| python-version: "3.12" | ||
|
|
||
| - name: Decode and apply the reviewed patch | ||
| shell: bash | ||
| run: | | ||
| set -euo pipefail | ||
| patch_gzip="$RUNNER_TEMP/nim-final.patch.gz" | ||
| patch_file="$RUNNER_TEMP/nim-final.patch" | ||
| cat .review-evidence/nim-final-patch/part-*.txt \ | ||
| | tr -d '\n\r' \ | ||
| | base64 --decode > "$patch_gzip" | ||
| gzip -dc "$patch_gzip" > "$patch_file" | ||
| printf '%s %s\n' "$EXPECTED_PATCH_SHA256" "$patch_file" | sha256sum -c - | ||
| git apply --check "$patch_file" | ||
| git apply --index "$patch_file" | ||
| git rm -r .review-evidence/nim-final-patch | ||
| git rm .github/workflows/nim-apply-reviewed-patch.yml | ||
| git diff --cached --check | ||
|
|
||
| - name: Install hash-locked verification dependencies | ||
| run: | | ||
| set -euo pipefail | ||
| python -m pip install --require-hashes -r fuzz/requirements-property.txt | ||
| python -m pip install --require-hashes -r requirements-opencode-review-ci.txt | ||
|
|
||
| - name: Verify full behavior and exact quality gates | ||
| shell: bash | ||
| run: | | ||
| set -euo pipefail | ||
| python -m compileall -q contextual_orchestrator tests | ||
| python -m pytest -q | ||
| python -m coverage erase | ||
| python -m coverage run --branch \ | ||
| --source=contextual_orchestrator.nim_benchmark \ | ||
| -m pytest \ | ||
| tests/test_nim_benchmark.py \ | ||
| tests/test_nim_benchmark_release_acceptance.py \ | ||
| tests/test_nim_benchmark_workflow_contract.py \ | ||
| -q | ||
| python -m coverage report \ | ||
| --include=contextual_orchestrator/nim_benchmark.py \ | ||
| --show-missing \ | ||
| --fail-under=100 | ||
| python -m interrogate -f 100 contextual_orchestrator/nim_benchmark.py | ||
| ! grep -R -n 'urllib.request.urlopen' contextual_orchestrator/nim_benchmark.py | ||
| test ! -e contextual_orchestrator/nim_benchmark_hardening.py | ||
| test ! -e tests/test_nim_benchmark_hardening.py | ||
| test ! -e .review-evidence/nim-source-repair.yml | ||
| test ! -e .github/workflows/nim-apply-reviewed-patch.yml | ||
| git diff --cached --check | ||
|
|
||
| - name: Build, install, and import the wheel in isolation | ||
| shell: bash | ||
| run: | | ||
| set -euo pipefail | ||
| rm -rf dist "$RUNNER_TEMP/nim-wheel-site" | ||
| python -m pip wheel --no-deps --no-build-isolation . --wheel-dir dist | ||
| python -m pip install --no-deps \ | ||
| --target "$RUNNER_TEMP/nim-wheel-site" \ | ||
| dist/contextual_orchestrator-*.whl | ||
| cd "$RUNNER_TEMP" | ||
| PYTHONPATH="$RUNNER_TEMP/nim-wheel-site" \ | ||
| python -c "import contextual_orchestrator; import contextual_orchestrator.nim_benchmark" | ||
|
|
||
| - name: Prove the exact final tree and publish only this branch | ||
| shell: bash | ||
| env: | ||
| GITHUB_TOKEN: ${{ github.token }} | ||
| run: | | ||
| set -euo pipefail | ||
| cd "$GITHUB_WORKSPACE" | ||
| rm -rf dist build .coverage htmlcov contextual_orchestrator.egg-info | ||
| git add -A | ||
| git diff --cached --check | ||
| actual_tree="$(git write-tree)" | ||
| test "$actual_tree" = "$EXPECTED_FINAL_TREE_SHA" | ||
| git config user.name "github-actions[bot]" | ||
| git config user.email "41898282+github-actions[bot]@users.noreply.github.com" | ||
| git commit -m "fix(nim): integrate evidence-grade benchmark contracts" | ||
| basic_auth="$(printf 'x-access-token:%s' "$GITHUB_TOKEN" | base64 -w0)" | ||
| git -c http.https://github.com/.extraheader="AUTHORIZATION: basic $basic_auth" \ | ||
| push \ | ||
| --force-with-lease="refs/heads/${TARGET_BRANCH}:${GITHUB_SHA}" \ | ||
| origin \ | ||
| "HEAD:refs/heads/${TARGET_BRANCH}" | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,153 @@ | ||
| name: NIM benchmark | ||
|
|
||
| # Evidence-grade NVIDIA NIM model discovery and cost-quality benchmark | ||
| # (issue #86). Two isolated entry points: | ||
| # * a deterministic manual dry run that never receives provider credentials; | ||
| # * an explicitly selected manual live run or conservative monthly live run. | ||
| # Single-flight concurrency prevents overlap and never cancels an active run. | ||
| # The benchmark fails closed on incomplete discovery, exceeded budgets, missing | ||
| # provenance, or missing live credentials. It never merges, releases, or | ||
| # rewrites production configuration. | ||
|
|
||
| on: | ||
| workflow_dispatch: | ||
| inputs: | ||
| dry_run: | ||
| description: "Dry run (validate everything without contacting NVIDIA)" | ||
| type: boolean | ||
| default: true | ||
| max_total_requests: | ||
| description: "Hard cap on provider requests for this run" | ||
| type: number | ||
| default: 500 | ||
| pricing_scenario: | ||
| description: "Optional path to a reviewed pricing-scenario JSON (empty => hypothetical costs stay 'unknown')" | ||
| type: string | ||
| default: "" | ||
| schedule: | ||
| # Conservative: one live run per month under a small hard request budget. | ||
| - cron: "23 3 5 * *" | ||
|
|
||
| permissions: | ||
| contents: read | ||
|
|
||
| concurrency: | ||
| group: nim-benchmark | ||
| cancel-in-progress: false | ||
|
|
||
| jobs: | ||
| dry_run_benchmark: | ||
| name: Validate the benchmark without provider credentials | ||
| if: ${{ github.event_name == 'workflow_dispatch' && inputs.dry_run == true }} | ||
| runs-on: ubuntu-latest | ||
| timeout-minutes: 30 | ||
| steps: | ||
| - name: Checkout repository | ||
| uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7 | ||
| with: | ||
| persist-credentials: false | ||
|
|
||
| - name: Set up Python | ||
| uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # actions/setup-python@v6 | ||
| with: | ||
| python-version: "3.12" | ||
|
|
||
| - name: Install pinned runtime | ||
| run: | | ||
| python -m pip install --require-hashes -r requirements.lock | ||
| python -m pip install --no-deps -e . | ||
|
|
||
| - name: Run deterministic dry benchmark | ||
| env: | ||
| PRICING_SCENARIO: ${{ inputs.pricing_scenario }} | ||
| MAX_REQUESTS: ${{ inputs.max_total_requests }} | ||
| PROVENANCE_GIT_SHA: ${{ github.sha }} | ||
| PROVENANCE_RUN_ID: ${{ github.run_id }} | ||
| run: | | ||
| extra_args=(--dry-run) | ||
| if [ -n "$PRICING_SCENARIO" ]; then | ||
| extra_args+=(--pricing-scenario "$PRICING_SCENARIO") | ||
| fi | ||
| python -m contextual_orchestrator nim-benchmark \ | ||
| "${extra_args[@]}" \ | ||
| --task-manifest examples/nim_task_manifest.json \ | ||
| --output-dir benchmark_artifacts \ | ||
| --max-total-requests "$MAX_REQUESTS" \ | ||
| --git-sha "$PROVENANCE_GIT_SHA" \ | ||
| --workflow-run-id "$PROVENANCE_RUN_ID" | ||
|
|
||
| - name: Upload dry-run benchmark artifacts | ||
| if: always() | ||
| uses: actions/upload-artifact@330a01c490aca151604b8cf639adc76d48f6c5d4 # actions/upload-artifact@v5 | ||
| with: | ||
| name: nim-benchmark-dry-run-${{ github.run_id }} | ||
| path: benchmark_artifacts/ | ||
| retention-days: 90 | ||
| if-no-files-found: error | ||
|
|
||
| live_benchmark: | ||
| name: Discover, probe, and benchmark the live NIM catalog | ||
| if: ${{ github.event_name == 'schedule' || (github.event_name == 'workflow_dispatch' && inputs.dry_run == false) }} | ||
| runs-on: ubuntu-latest | ||
| timeout-minutes: 60 | ||
| steps: | ||
| - name: Checkout repository | ||
| uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7 | ||
| with: | ||
| persist-credentials: false | ||
|
|
||
| - name: Set up Python | ||
| uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # actions/setup-python@v6 | ||
| with: | ||
| python-version: "3.12" | ||
|
|
||
| - name: Install pinned runtime | ||
| run: | | ||
| python -m pip install --require-hashes -r requirements.lock | ||
| python -m pip install --no-deps -e . | ||
|
|
||
| - name: Resolve live-run parameters | ||
| id: live_parameters | ||
| env: | ||
| EVENT_NAME: ${{ github.event_name }} | ||
| INPUT_MAX_REQUESTS: ${{ inputs.max_total_requests }} | ||
| INPUT_PRICING: ${{ inputs.pricing_scenario }} | ||
| run: | | ||
| if [ "$EVENT_NAME" = "schedule" ]; then | ||
| echo "max_requests=300" >> "$GITHUB_OUTPUT" | ||
| echo "pricing_scenario=" >> "$GITHUB_OUTPUT" | ||
| else | ||
| echo "max_requests=${INPUT_MAX_REQUESTS}" >> "$GITHUB_OUTPUT" | ||
| echo "pricing_scenario=${INPUT_PRICING}" >> "$GITHUB_OUTPUT" | ||
| fi | ||
|
|
||
| - name: Run live benchmark | ||
| env: | ||
| # This is the only workflow binding for the provider credential. The | ||
| # dry-run job is structurally separate and cannot receive this value. | ||
| NVIDIA_NIM_API_KEY: ${{ secrets.NVIDIA_NIM_API_KEY }} | ||
| PRICING_SCENARIO: ${{ steps.live_parameters.outputs.pricing_scenario }} | ||
| MAX_REQUESTS: ${{ steps.live_parameters.outputs.max_requests }} | ||
| PROVENANCE_GIT_SHA: ${{ github.sha }} | ||
| PROVENANCE_RUN_ID: ${{ github.run_id }} | ||
| run: | | ||
| extra_args=() | ||
| if [ -n "$PRICING_SCENARIO" ]; then | ||
| extra_args+=(--pricing-scenario "$PRICING_SCENARIO") | ||
| fi | ||
| python -m contextual_orchestrator nim-benchmark \ | ||
| "${extra_args[@]}" \ | ||
| --task-manifest examples/nim_task_manifest.json \ | ||
| --output-dir benchmark_artifacts \ | ||
| --max-total-requests "$MAX_REQUESTS" \ | ||
| --git-sha "$PROVENANCE_GIT_SHA" \ | ||
| --workflow-run-id "$PROVENANCE_RUN_ID" | ||
|
|
||
| - name: Upload live benchmark artifacts | ||
| if: always() | ||
| uses: actions/upload-artifact@330a01c490aca151604b8cf639adc76d48f6c5d4 # actions/upload-artifact@v5 | ||
| with: | ||
| name: nim-benchmark-live-${{ github.run_id }} | ||
| path: benchmark_artifacts/ | ||
| retention-days: 90 | ||
| if-no-files-found: error |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1 @@ | ||
| {Wp48S^xk9=GL@E0stWa8~^|S5YJf5;>0$5BwYY98c7Ml{1JUn-U?4g%<)t4wb~W0)DdL_LWmAQ>F;bV#2e=^^N~&!EF5M+EktE399rA8J#<{eb4#~_>wY1VZd8mwM&v5g^FW)}pfS_hF16dF$d4%<z~;*Z*)BAH`Qs-M7*C-x`<$59!*K8Vbkm!Sj%?!zw+5yn*g&UTBP#XHX{6y)V>Pptrn-TnmE|DPxjZj8W>Z8G67^%qpG?qYx!qS8L)P~$MtHsa=oaKU%CK_k4x}X7c+)4ftQ8`pdt^&8Xs}u{Toy1%WM%E2AYsL=<@HdqZUO>nhAQqFrD;Q&Av!<ojl23Q{-A_=RK_NwCt=LD^xHGc%r08jlV<ITKUP&H7=Kr~8FLI=Vtpen@Hy+RPVjIOPGpHcLUp}t@rxtEb%S(CT3Fim@irc3amSBcmjrS2+;|rzTX&6!zwANVG&u|&>8mH=+R7NPyKBJD6#xM-t}2lBb-B;1<`D7xd49tEd#)P2H`-z(%(xStG=Nf`PBzv1x!sMDy@@p?16#H%U6{(z5h(&t**Az_VlGU{GS=^bhdzaKh~D|v6^pq}wV6>HAW3&aEctzY&zb(-+WDMxVbl%NhT;OM!xp?pG#nRlbYctbpQzJS%zZt<llbjyK44&{W@{BgY=pdH{n-jBuHN5OFup*u%)Fw?Jt>g-9v2fcZ<qKje|{1DcufN>5TZfGtaY*@uKIqa;0BL052n++4=K+i_ig6GQw1x^@rCCT{{hDYl*_5i$*RSh{^qtiP|sCqg8QQ)#z!6S_?bMENXMAjF%VZ1;RYX|<3t;gtBE_+!ulWqL2J)9%^C6{<t`)7!o&1U#xvcgt;@To3%DcW6Z7qr*>e}#Y(9HSwdX5IY874!KVWY7vivb?un!?Q5`u)&l`JDGAuHkWu(8+fI{MAl3SK5#pP*~L1rpR_pP(s>4y|ZwU>=wE@IF;X$i|lwBh#cG+i2cf5;MUaf~bGx!Q*}z0>;k)el6^X1+vRKHoaJ^Eb$)5B_p8<MFnZoF>BW#h)h!P#l1sxOmQr_(Ner+dUcU)wHj%e)h!XRz;4GAg3nKxtSG&4Oy^oy$i=A%yAlysjUzN_EazvIQ47yPm_$F^M@(vcOeWXUPH#uy+w9qDelZKeF;p2G!&M)>2wbhoQ$^T9e4#m3e={zRR1px3y?!&CiFkGI`6SnWxZD)RWA(ZxM4vgURP(*b>3eYv6<w+yCySZrL7F)pTydl&AHCz8cKXJ-BLe7^AuXg^2t>B6DJJpgN<DPDCp!O&D{P#R9{I?qkRAu3IfcFFgMJlIW*jd6W%9F^A*>HyC^e6F%iM(>xdLdL?bWqn$xOhQ*~27Vzz}CJcBTm*Ms!~bHznQ##91*zYXB{UOhbS<F{;6&59Hi!P$VXfS*zsw;w@e6_TWI{Os#*^94#F^i$tF{ny3*~Z4DKS|F_wsFzzDub8iumgqt17_K%3BTJvqF7wfru?Yek;<B|bnlu@sOQ_EE<P`eIzG`b)d*UWyq&NetYhK6&rVnFRI5GY?+pF!-D5l6=91`7}nA^)b9q;wXn*E)X*Gtc6ET7<<3@ID&mbfoA&TXKz(Q17J~f1eMD@!u8(i_6N6B-J~1E@86uAhP6QYX#Abhwn2C@nKX`(<k$#cBe~h8w{~?4r-QHT99}u@3>o9N-%{g>1_hE6mtsw)VC*VUZ>ao8*nYMGR5{UO)mAs(-iQMy<WKV>vbPgQ^w2{py&S64(OHq1QpusGZ5#_LoskBwxluHT*pS#n2Cj(@6nLsel|U}IEf~a7E!>i2f+BP(>;Al29_i(p0CLXM%NxwW0E*8b)Xz;ILR8U4d3Z}T#x1`+REzao;h?Oa|32!rAE7eS;I8BhS<b2DGId3qcB>zj_hb5dU~RK5#!6dkp_hWOjA1|n6Q|Wb|1s_a$L^&-ktApWYY<TmG0eytRKFuSu>=7aiMosu^0(j*1Io`V9lz&E%)E46EJQ*w)EaHkpq74I!7=xYiAoApqOCyvp{+w3&RYI-|jY>xW!|=D$UlX%{_vcG4XjP2OA7-h)$Yv2190w!=$D(OyblxpBb%aNrea)(MoEy<J%mv%6JSDd9QYCt<Q8Y3oubO;ouq0oL7fsSzL!{%b-g1ix8&8K;10CWnL4~NvcN$2R{sQc4kp%NeU4*(mbPBxHxm$upE*zr9E5-d=`4XeXAl)f(E@kn<;C>&5CR@>6^j6NuT%z2rNSGI8%`zvX{9A)mN!d{9q!DEbN7~sM%W_<-9<jyFDS@q`wzr?DkjO(#kVi2+4<)3M@6W-5M~Ec}xKVB)P}AvN24&aJ$$en+pp~iiO#)vh=SE?uKU;E7Z+eyBe6#0=+G0azl5`=-1$C0zv8WiUEOQMqxVTLSd;(F3u~@2<q_dJ@QUhK}1U97Qm<Sv+M(_C!!*52eU9vN`458yFHS@8<u?swkq4<e%HekpB*xO3=OkQ;IMp)Z4-ad6*Lfb5p7t(I~Qc5FHb9yBb?V8i0|3%)H7Czx$F*}KfWk<H>~+hS5`|`R*JNTB&3A2%+=y^FVcSTgrM_fa9ajL_QmaBZ!FRVgl<-vd*_EG71VEM9&u6LX4Eqi47JE9m8&BBdu>ctI($<-`l;*l$J4w^?b1F5)t^~jyXAjUt44qmrf;MkUhhChJ$#tv*d>8hgxB>249vV?gO)I}X;Vi#0w57XZJ;;Aw3q%2;1@W;<v>UGP&-9>o@15VOK&^Bt$pO#2nWts@tUl++(9JETd}v+t-O~>{UJ}hziPm>F0y&9fC4iqK|ix=rUB-Qft;6lB2}ai?u!k(1=ioA^bnSnnt*7-vDuwPt{x7|Fvf?3W)DSLC&Qub60Z42dq%`l=>vj1tEB32qpJ+_i?_mO?)e#$@qfRk7;e>aB*#oQ1ek2V`X^969p|4tO;1y^s|-e%{*bV0V+0DLZM)bz32rP;X&VM#J&Fl&i-DzNyPTA6iBzfd3vwMr5~<YFo0;!zoXv@`gpQJ(#Hjn-;gowwv5cZYs)CwxYC5fIPHh4p+B~syC2%i7X7!^+K%HZgHw(jJ5zdJaN=vlG=sc?ke6>VEaC#QAt3mshB!h0-uvL)59};mx&`)vxes9&JT|P7K94RC78REqEt>4B?HHrOQjG1{B(3_a8k78?9Y<pvz#kUv`&XUK-hxlP+lGv&u`4RWHRuj=*=9omT3qN^f?Z9o)ZmtZCu81Hx%Z>76?7OqA*1c|#ff%ezJf55}dMXYiuRHij!W^G!33lMj=UQS7$H%D(aZqx2uaP;$!}7<Y2vpbk3c+6tNoIqyBTj)6*#(S~KTjs0t~UqbfY~)HFuV{A4M82iSR)A>&hh#)hd;FU6vezvshX~q4NhfTICNw+=QPA2^eFpu@u1zfqbWVa_a6lbmfk8?5l6b!{YYpM_Nca$9nSsha6jXh=S--n6F0&zj+N?vPwB5Z3a+`S(Hi@>@--dDhVM_~x8!ym;SGD?iT*jXF|gJNWE>nfycCSfWHlN?dJlHiGzUB&zQ4cGLs)pH&ePXK^7cVwvO+y4u42gg!KFASSxD1szv3)XBH#NmadX5`d6+44E-2~CoL@7GoWH3k0T`zr3zYP$gWWrd>698|tqC!9ZY3^cQ1O*l331v`32ESD{83}shXfi*fO-eH9+rIo<x*D2GKCTw1po?s+gE2vGw<2S0%uzzbRPk|g57NUV>W>Xv#neH-tu|E@5FjW?$p8OVtb&m4}B%K5H!FY%OAHl3LN$06Z#^L7*>t{yq-=@e@g0uYAW|VpVUsIj<ASg^ePgV=8MZ}X)wkPFX%Uf@pVb23)vlhc>&5To}+7UgI5@&FZ1tr-BRbF3zd1eM|UFrM{~Z09PovhrDcW#%I}M9FZ)+~MxK*R?Wp9RZVL?~Kj~&t>w$~)@_Sw}u3xRH3@Exi>TE<Dl?nG}L(Qm>yu133K6G!JMC!rki}Rd>NdLhY0v=ur;Z5<~DKC>5c>!{CAg_4WKtF;@9LG#=FY84NJ$X|!cYRfL;t!H>KuD~UR`N!Tf7fk9x8C6wh_>_b1C^$zDV;*tyMc70JJWnhZQ$JfwiU94lm;2(y`(;t!7i{Vfy%jNiXM=HF2#O#e#3%bM~44#6u7as|B_U+!<4^wcJwl<Izi$!6|A^gTOMS8RuPz91P`<0>H(45cR?|xG!5}Xh;z7Q7ErmYdfJbuXuy)#n(-kOsypRRgiG{??(<OyHeJ+#Fep=gZtmtdvH&)Wf4`nTi><mkk<-V(;76WH<=*qFgwz#P^}!C3xWvpu8E&?B)qatl=E*CmWxcVAVYN$nsJq|aGf!|s_8ZG^FD8_FdhceqMt49;v9ox90azZ4JamC=Vj<y&CF6KhD?!0;*i}y06g}vepJ+kZ2`3v4q2Obj1NjI0yCd3ay;kS7?P)(N2`DSg0EUi#W(dIVvoU{~r6{*<UYdZ&DKXq<=VP-&lMTebu_~2uN(R5F40q-?J!uM`0lX*QXh;z22A?JGe9`o4P5c%SbYbIs7-qA}hDXwDFwX*={~3bA+K@dC4;*&2%OJjhbg0eClzF`*8ej2Tq$62G1(&xiEhC-iykK8&mc#Gd#hHwWM_g#(yUKBwyWF4vwL&v)E~iHc1S1p)zTdF&+kALOq+I!9yY!8}{!|y=TU_7ICe=9?L`LYJ?celtPphi?N;;+PJ#yT)xJ0JC71-@3i5wwQbnm0tc)-1!Jk!wGps!SG32Y}aR$li~&~_%?p?p0x8DhHW)`(318g)>zhnySsgSy)m`yzrl2155HO0ZTi@@`}1FLb$5E^R>B-((?y;EyIQmw^saSSk9dM}S*Qnd~=|M6@>;i2j#SCq;+7lyEBOULzt^hR^4(t$RIg3KSb-pO%M(#`bWEhuy-j(9t0Wq0sXEa=Zr12rj;kLtyV+1~Y_?<8FogIHMpWc>ML3ijU20bD{I1<%^Omgx2nw;@BC56UdWxFf(wg7<(@`0~eDRnyQssib%eeblL->Ia9eKUuO`56tGA)^Ewi|Td>$&ZS3^571qS4`e0r3n7ar}Ke$<+NuGZMJrx+$ph4VfZXWYnzUQ*~!F^z4(D8#CH1i3set$E9WTP>R$n?PD |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.