Skip to content

fix(nim): retain failed tasks in paired quality and latency evidence - #1074

Draft
seonghobae wants to merge 53 commits into
codex/psychometric-kpi-successorfrom
codex/paired-policy-outcomes-20260905
Draft

seonghobae wants to merge 53 commits into
codex/psychometric-kpi-successorfrom
codex/paired-policy-outcomes-20260905

Conversation

@seonghobae

@seonghobae seonghobae commented Sep 5, 2026

Copy link
Copy Markdown
Contributor

Current integrated checkpoint — d7e9d9a (2026-09-08)

Head d7e9d9a240854f9b06c891290d5d249bec8833ce normally integrates parent 84a6052369a7bf8b6faae5db475bb68a5ad54a91; Draft / Proposed. Child changes were retained without force push.

Fresh local CI-equivalent benchmark gate on this unchanged head: 169 passed in 55.47 seconds, 1271 statements / 480 branches, zero missing, 100% coverage, public docstring check 100%, terminal exit 0. Command: coverage run with branch coverage and source contextual_orchestrator.nim_benchmark over test_nim_benchmark.py, test_nim_benchmark_release_acceptance.py and test_nim_benchmark_workflow_contract.py; then coverage report --fail-under=100 and interrogate -f 100. This validates the earlier coverage repair on the integrated source, not all repository code or all edge cases.

Separate integrated focused suite: 159 passed in 19.56 seconds across uptime, psychometric boundary, paper contract and reasoning-effort tests. No full current-head suite, hosted acceptance, independent approval, production performance or release is claimed.

All sections below are historical checkpoints. Their original “current” labels and check outcomes do not describe this integrated head.

Current exact-head checkpoint — 9f5d0ce (2026-09-07)

  • Exact head: 9f5d0ce2f7e95eb3bfcaf5cbd4e2a165dab7c86c; base: ec1c4e66512615ea1f00fa044f3bfba787aab567; lifecycle: Draft / Proposed.
  • Hosted quality job 101628312201 failed after a green full suite (3618 passed, 2 skipped) because nim_benchmark.py coverage was 99% (2791, 2798, 2809, 2814, 2819, 2827 in validate_report_schema).
  • Source 9f5d0ce2 adds test_report_schema_rejects_invalid_evaluation_contract. Local three-file coverage command: 169 passed, 1271/480 statements/branches, 0 missed, interrogate 100%.
  • Fresh source, fuzz, and supply-chain Checks must run on this SHA. No predecessor coverage result is transferred. No production policy or release is authorized.

Current exact-head checkpoint — 4351af9 (2026-09-07)

  • Exact head: 4351af9f1b6dca2f97526a46ac6dca28048ff623; base: ec1c4e66512615ea1f00fa044f3bfba787aab567; lifecycle: Draft / Proposed.
  • This ordinary non-force integration preserves aa4ff156f3001927: tied observed worker maxima yield no hindsight winner while measurement-only artifacts remain available. Concurrent commit e94dd035 adds prospective observation-coverage validation; the merge retains both deltas and their regression fixtures.
  • Fresh source, fuzz, and supply-chain Checks are running on this exact SHA. No predecessor result is transferred.
  • Fixed 2,000-resample bootstrap, fixed 95% interval, hard-coded policy comparison subset, and repository-authored inference/token/workflow budgets remain substantive no-heuristics blockers. No production policy or release is authorized.

Current exact-head checkpoint — f300192 (2026-09-07)

  • Exact head: f3001927ad28852c5cca874eb3bf61920107e42f; base: ec1c4e66512615ea1f00fa044f3bfba787aab567; lifecycle: Draft / Proposed.
  • Regression-first commits f1662dcdd217c053, followed by integration repair aa4ff156f3001927, remove the lexicographic model-id tie-break from best_single_worker_hindsight. A unique observed maximum is required; equal maxima now produce no hindsight winner while the measurement-only report remains available.
  • The existing fixed 2,000-resample percentile interval, hard-coded policy comparison subset, and repository-authored inference/token/workflow budgets remain unresolved. This checkpoint does not transfer predecessor GREEN or claim statistical release readiness.
  • The predecessor exact-head test run failed because tied dry-run fixtures aborted the whole report; f3001927 repairs that propagation. Fresh exact-head hosted Checks have not materialized yet. Do not merge until source, fuzz, security, and independent model-review gates are terminal on this SHA.

Current verified checkpoint — 621cf4d (2026-09-07)

Current pushed head: 621cf4d9cd47d166bd568cfb07a25692c47df210, base ec1c4e66512615ea1f00fa044f3bfba787aab567, Draft. This checkpoint supersedes old current-head wording below without discarding historical evidence.

  • Count/completion-only promotion is removed, and the separate reasoning-effort 55% declaration-only authority is now removed too. Its retained Boolean function returns false for every report; the removed threshold export and formerly true outcomes are explicitly documented breaking changes. Profiles, provider propagation and route/conduct defaults are unchanged.
  • Exact-head local full regression: 3612 passed, 2 skipped / 1110.73s, exit 0, matching clean start/end HEAD and JUnit3614/0failures/0errors/2skips. Focused96 passed; profile module168statements/62branches and public docstrings100%. These are software checks, not buyer KPI evidence.
  • Wheel built; separate-cwd Python-I archive import and45Python-file source-byte parity verified. This is not an independent environment installation or release.
  • New hosted run34077430336 is pending. Skipped bot reviews are not independent approval. No deployment is established.
  • Real held-out task/policy evidence, calibrated scoring, failure denominators, dependence-aware uncertainty, a released Rust statistical-owner contract and independent operational acceptance remain incomplete. No synthetic fixture or arbitrary replacement cutoff substitutes for these.

See the exact verification receipt and prospective design. Actual Edge visual inspection confirms the rendered receipt, Draft, pending checks and no-deployment status; this is not product-admin/mobile E2E.

Historical endpoint collection repair — 93bd66d

Historical head: 93bd66d494f8f44a15b9989fe21b075412103da4, historical base a4f693c41f642f960d59bb5a9849eadedd412bd5. All results in this section retain that revision scope and do not describe the current head.

Whole-model-ID URL encoding returned HTTP 404 for a public model while the documented author/slug path returned 200. Source 98cdc3ec validates exactly two nonempty non-dot segments and encodes each separately using the existing stdlib. Malformed IDs cause no request. Fixed origin, no-update behavior, authentication, routing weights and statistical-owner boundaries remain unchanged.

  • Committed RED aee1e497: 16 failed / 61 passed / exit 1 / 1.28s.
  • Source 98cdc3ec: 147 related tests passed in 4.80s, exit 0. Separate 77-test coverage run: changed fetch method 18/18 statements and 6/6 branches; default Ruff passes. Not exhaustive input or whole-repository coverage.
  • Current child focused: 270 passed in 63.75s, exit 0. The initial parent command named a nonexistent test file and ran no tests (exit 4); it is excluded from passing evidence and was corrected.
  • One actual isolated current-source collector poll at 2026-09-06 12:05:50 UTC successfully added transport window mass for openai/gpt-4o, with observed attempts still zero. No model inference, background service, or supplied provider credential. This does not establish fleet access, calibrated delivery probability, deployed behavior or buyer accuracy/latency.
  • Exact-head full suite completed: 3597 passed / 2 skipped in 1762.13s, exit 0. Matching clean start/end 93bd66d494f8f44a15b9989fe21b075412103da4, parsed JUnit 3599 cases, zero failures/errors. Session 28306 is terminal, not running. Artifacts: /tmp/co-uptime-path.T7v9Rj/child-full-*. The verification JSON's seconds field measures the outer command (1782.32s); the pytest duration above is from the terminal test log. Neither duration is a buyer latency KPI.
  • Current local Trivy filesystem scans both exited 0 with zero HIGH/CRITICAL vulnerability/secret findings in four lock targets under ignore-unfixed and default dev/test exclusions. These are not hosted Security acceptance.
  • Prior input-admission fulls are terminal: parent db4d21da 3566 passed/two skipped/739.03s and child 5eebac47 3581 passed/two skipped/754.61s, both exit 0, matching clean heads and parsed JUnit. Those and prior Trivy results predate this path repair.
  • Visual Inspection: the exact parent a4f693c4 rendered doctoring section was inspected in a real Edge desktop screenshot; the complete new section, counts and limitations are legible without observed clipping or overlap. Current parent and child PR headers/checkpoint bodies were also inspected in real Edge screenshots: Draft status, main/parent target, matching head/base, focused counts and running-full wording are readable without observed overlap. This is not product-admin/mobile E2E.

The normal child merge retains parent source/tests/documentation and the child's NIM comparison, observed audit and licensed paper delta. Current collection evidence records alternatives, live scope, prior fulls and remaining gates. RED/GREEN logs, coverage and live collector output are in /tmp/co-uptime-path.T7v9Rj; original public HTTP baseline is retained in /tmp/co-uptime-input.rjxaB9.

Rolling-window dependence, endpoint-route mix, benchmark-prior calibration, judge validation and released-owner/buyer evidence remain open. ADR0034 stays Proposed; ADR0021 acceptance and production defaults are unchanged. Both PRs stay Draft. Independent exact-head review, required hosted checks, parent-first protected merge and immutable release remain gates.

Problem and result

The policy summary counted failures in its denominator, but paired confidence intervals excluded any task where either policy failed. A policy delivering one of two answers could therefore appear tied with a policy delivering both.

The comparison now retains every shared locked task, treats unsuccessful delivery as zero reward, and reports mean delivered-score and terminal-outcome-time differences with paired bootstrap intervals. Raw failed answer scores remain null. Success counts and unmatched task counts expose the comparison denominator; duplicates and invalid measurements fail closed.

In the hand-checked unit fixture, the old apparent tie becomes a delivered-score difference of -0.5. Time differences of -50 and 1950 ms produce mean 950 ms with interval [-50, 1950]. This proves a calculation repair, not improved model accuracy or live latency.

Contract and scope

  • Report schema 2.0.0 identifies the new estimand. Version 1 comparisons must be regenerated from their original observations rather than pooled with version 2.
  • The existing paired mean-bootstrap implementation is reused. No new statistical dependency or provider call is added.
  • The former count/completion promotion thresholds have been removed. Current output is measurement-only with null recommendation/thresholds; retained paired statistics do not authorize deployment.
  • Mean intervals are conditional on the shared task set and selected policies. They do not establish p95 improvement, dependent-task validity, or uncertainty from hindsight model selection.
  • Product and technical requirements, UML sequence, research reference, migration, and remaining owner prerequisites are recorded in the NIM guide, doctoring record, research inventory, and gap baseline.

Historical identity-batch integration at fa8eef9

Historical head: fa8eef97cd7368e8985a367dc5a7a8e0147fe4d5.
Parent #1067: d740602fdd8c0e4f7d55e4d3ad37b9f560c09e01.

This normal two-parent merge retains previous child 9207412e and the new parent identity-batch change. The baseline conflict preserves both complete additions and records the new lineage. Explicit diffs show child NIM source/tests, comparison guide, observed-data audit, and papers are unchanged from 9207412e; the parent's orchestration source, psychometric tests, doctoring, and complete profile artifact are byte-identical to d740602f.

  • 243 focused tests passed in 35.24 seconds, terminal exit 0 on clean fa8eef97: NIM comparisons, release/workflow contracts, profiles, paper/planning checks, and psychometric routing. Evidence: /tmp/co-1074-catalog-batch.MIcuwO/focused-pytest.log and focused-junit.xml.
  • Exact current fa8eef9 full suite completed: 3,481 passed, two skipped, terminal exit 0 in 664.63 seconds. Start/end SHA both equal fa8eef97cd7368e8985a367dc5a7a8e0147fe4d5, and the tracked tree stayed clean. JUnit contains 3,483 cases, zero failures, zero errors, and two skips. Evidence: /tmp/co-1074-catalog-batch.MIcuwO/full-pytest.log and full-junit.xml. Session 79929 is terminal; do not restart it. Parent d740602f completed 3,466 passed, two skipped, exit 0 in 733.63 seconds on its separate clean tree. Hosted acceptance remains separate.
  • Exact fa8eef97 Trivy 0.74.0 filesystem scan passed (exit 0), default vulnerability/secret scanners, HIGH/CRITICAL and --ignore-unfixed: zero findings in four lock targets. Evidence: /tmp/co-1074-catalog-batch.MIcuwO/trivy.json and trivy.log. DB UpdatedAt 2026-09-06 07:00:11 UTC, DownloadedAt 07:08:45 UTC, NextUpdate 2026-09-07 07:00:11 UTC; this run reused that fresh DB. Default dev/test exclusions apply, and complete hosted Security acceptance remains separate.
  • The parent eliminates repeated catalog validation per batch, preserves real repeated attempts, and observes later configuration changes without a persistent cache. Its controlled profile is internal unit-fixture computation only, not a buyer latency/accuracy result. Both favorable and earlier regressing samples remain in the artifact.

Draft status, parent-first protected delivery, and independent approval remain unchanged.

Previous integration at 9207412

Historical head: 9207412ea8ac529d7d2622ab298989ef7899befb.
Historical parent #1067: bfeb73a6c58add7a23df052110593cdf43c0b0db.

This normal two-parent merge preserves prior child 402fd63a and the updated parent. The one baseline-document conflict retains both complete additions and adds the current lineage note. There were no source conflicts. Explicit diffs show the child's NIM source changes only the inherited reviewed-cost dates; its test fixture gains the parent's wall-time isolation. The observed-data audit, paper inventory, and separately licensed LLMRouter paper are byte-identical to the preceding child. Parent transport, psychometric source, both harnesses, transport tests, and psychometric doctoring are byte-identical to bfeb73a6.

  • 211 focused tests passed in 2.33 seconds, terminal exit 0, on clean 9207412e: NIM comparisons, release acceptance, workflow contracts, reasoning profiles, paper inventory, and planning contracts. Log/JUnit: /tmp/co-1074-current-integration.PBuxSQ.
  • Exact historical 9207412 full suite completed: 3,476 passed, two skipped, terminal exit 0 in 790.27 seconds. Start/end heads match and the tracked tree stayed clean. JUnit has 3,478 cases, zero failures, zero errors, and two skips. Evidence: /tmp/co-1074-current-integration.PBuxSQ. Session 90398 is terminal and must not be restarted.
  • Same-head coverage run: 211 tests passed in 7.83 seconds, with NIM 1,223 statements / 446 branches and reasoning profiles 185 statements / 68 branches, all 100%. Coverage JSON/log are in the same evidence directory.
  • Same-head Trivy 0.74.0 default vulnerability/secret filesystem scan: exit 0, zero HIGH/CRITICAL findings under --ignore-unfixed in four reported lock targets. A fresh DB download completed at 2026-09-06 07:08:45 UTC (UpdatedAt 07:00:11 UTC, NextUpdate 2026-09-07 07:00:11 UTC). JSON, scan log, and download log are in the same evidence directory. Default development/test dependency exclusions apply; this does not replace the complete hosted Security workflow.
  • Parent bfeb73a6 completed 3,461 passed, two skipped, terminal exit 0 in 675.10 seconds, with matching start/end head and clean tracked tree. Its JUnit has 3,463 cases and no failures/errors. Session 72900 is terminal. This parent result is not a full-suite result for the current child.
  • Both PRs remain Draft. Parent-first protected delivery, exact-head hosted acceptance, and independent approval remain required. No paper removed by the rights audit was restored, and no production policy was promoted.

Historical verification at 402fd63

Historical head: 402fd63a4f6a16558e8cb4e4cd8a3378da3827f6.
Historical parent #1067: ae704491fc24cd4618c709b036cc9353ea9923e5.

Ordinary merge 4fc184a020f1d63816f9e847d98a306ff27797f8 preserves both preceding child e2f3e2af8e6b723c6870fb6c4dcee2b4f7bc31af and parent histories. Both changelog and baseline additions were retained. Follow-up 402fd63a corrects only a document-location word. The child's NIM source, tests, observed-data audit, and licensed LLMRouter paper retain their preceding bytes; parent profile guards, tests, benchmark harnesses, and IRT interpretation document retain their parent bytes.

  • Exact historical 402fd63 focused verification: 210 tests passed in 2.66 seconds, covering NIM comparisons, release acceptance, workflow contracts, reasoning profiles, paper inventory, and product planning.
  • Exact historical 402fd63 full suite: 3,475 passed, two skipped, terminal exit 0 in 677.72 seconds. Start/end SHA both equal 402fd63a4f6a16558e8cb4e4cd8a3378da3827f6, and the tracked tree stayed clean. JUnit has 3,477 cases, zero failures, zero errors, and two skips. Evidence: /tmp/co-1074-measurement-integration.pCjFWW. Session 50435 completed; do not treat it as an active run.
  • Command: .venv/bin/python -m pytest -q --junitxml=<evidence-directory>/junit.xml.
  • Exact-head coverage recheck: 210 passed in 2.93 seconds, with NIM 1,223 statements / 446 branches and reasoning profiles 185 statements / 68 branches, all 100%. Configured public docstring coverage is 100%; that scope excludes private/nested functions and tests and is not the wider CodeRabbit changed-function scope.
  • git diff --check passes. Explicit source comparisons verify preservation of the separate parent and child changes.
  • Exact-head Trivy 0.74.0 filesystem scan: exit 0, no HIGH/CRITICAL findings under --ignore-unfixed, with default vulnerability and secret scanners. All four reported lock targets have zero findings. trivy fs --download-db-only checked the existing DB, updated 2026-09-05 07:05:47 UTC (cached download 10:07:59 UTC; next update 2026-09-06). This does not claim a new DB download. Default development/test dependency exclusions apply. JSON: /tmp/co-1074-measurement-integration.pCjFWW/trivy.json.
  • Hosted snapshot still has only CodeRabbit/Devin status contexts; no independent approval or complete repository Security/Quality acceptance was established. Draft status and parent-first protected delivery remain unchanged.
  • Parent executable/document head 47ae9d65 completed 3,460 passed / 2 skipped, exit 0, in 675.71 seconds on an unchanged clean tree; later parent documentation-only ae704491 passed 61 checks in 0.62 seconds. Neither result is a current-child full-suite result.

Historical child e2f3e2af completed 3,447 passed / 2 skipped, exit 0, in 677.14 seconds, with matching start/end head and clean tree; JUnit contains 3,449 cases, zero failures, zero errors, and two skips. That head also passed 149 NIM checks in 3.55 seconds with 1,223 statements / 446 branches at 100% and configured public docstrings at 100%. Its refreshed Trivy 0.74.0 scan exited 0 under HIGH/CRITICAL + ignore-unfixed with vulnerability and secret scanners; default development/test dependency exclusions apply. Those source/security results remain historical, not evidence of a scan or full run on the new integration.

Historical evidence: /tmp/co-1074-integration.naZc1Z. Earlier 1600b1d5 had 3,432 passed / 2 skipped in 664.58 seconds. None substitutes for exact-head hosted dependency review, pip-audit, CodeQL, the complete Security workflow, independent review, or protected delivery.

Inherited paper distribution correction

The parent audited its six bundled papers. FrugalGPT, RouteLLM, Hybrid LLM, and the fuzzing survey are removed from this current tree too; citations, source links, summaries, and historical hashes remain. HELM and IRT-Router retain explicit CC BY 4.0 sources and attribution. The independently licensed child LLMRouter PDF remains with its own record. No deleted copy was restored by merging, no source-code license was repurposed as paper permission, and Git history or prior release artifacts were not purged.

Public response-time evidence audit

  • Distinguish Feng et al. LLMRouter / xRouteBench from Li et al. LLMRouterBench; the latter estimates latency from tokens and serving statistics, not observed request-level p95.
  • Record immutable dataset and collector revisions. Three isolated source-function checks preserve seven-second successful and returned-error records, while an escaped exception records zero. This is a controlled source-contract finding, not an audited error count in the published dataset.
  • Dataset lineage, outcome completeness, sampling validity, and dataset redistribution rights remain unverified. No new estimator, dependency, provider call, production default, or performance claim is introduced.
  • Attach the unmodified CC BY 4.0 LLMRouter v1 PDF with SHA-256 and APA 7 attribution. A repeat download is byte-identical; Poppler's original dictionary warnings and the HTML reading alternative are documented. No raw dataset rows are redistributed.

Observed generic-matrix follow-up

Read and hash-verified all 147,924 rows from the immutable generic train/test files, using the existing PyArrow 25.0.1 reader with network denied during decoding. The committed aggregate JSON was checked again against every row. No dependency was installed and no raw prompt or response is redistributed.

  • Zero, negative, and non-finite duration counts are 0 in both inspected files; the source-code exception path does not prove zero-time contamination in this dataset.
  • 2,068 empty response strings have zero output tokens and scores, positive input/total tokens, and positive recorded time. Their terminal cause is unknown.
  • Train contains 18 repeated scored-task/model rows across one ARC-Challenge task and all 18 models. They are not exact duplicate records: 6 groups differ in output and 17 in elapsed time. No automatic deletion or independent-item interpretation is authorized.
  • No exact stimulus fingerprint overlaps train/test, and the observed task-by-model rectangles are complete. Neither result proves semantic independence or the full attempted-cell denominator.
  • Explicit terminal state, attempt identity, timing provenance, and dataset redistribution rights remain unresolved. This is an observation-contract audit, not a p95, ability, or accuracy result.

Dependency and review boundary

Draft, stacked on #1067. The parent remains open. Local Security and Quality now runs on this stacked PR, including current run 34077430336. This does not establish complete central required-workflow admission, independent review, or protected release; those owner boundaries remain separately tracked. Local validation is not protected-main or hosted acceptance. Require parent integration, complete validation on the then-current commit, and independent review before protected merge.

The coverage audit also repaired two pricing tests that previously accepted an unrelated hosted-access-expiry error. That independent test repair is already carried by the parent; this PR retains its history without duplicating its effective delta.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
@coderabbitai

coderabbitai Bot commented Sep 5, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Record pinned source and dataset provenance, distinguish returned API errors from zero-duration escaped exceptions, and reject estimated latency as observed p95 proof. Attach the CC BY 4.0 paper without redistributing dataset rows.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Read every row in the pinned generic train and test files using the existing Parquet reader. Record verified aggregate counts and hashes without redistributing prompts or responses. Distinguish blank outputs and repeated attempts from inferred provider failures.

Signed-off-by: Seongho Bae <me@seonghobae.me>
…ixes

Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>

# Conflicts:
#	docs/product-technical-gap-baseline.md
Signed-off-by: Seongho Bae <me@seonghobae.me>

# Conflicts:
#	docs/product-technical-gap-baseline.md
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>

# Conflicts:
#	CHANGELOG.md
#	docs/product-technical-gap-baseline.md
Signed-off-by: Seongho Bae <me@seonghobae.me>

# Conflicts:
#	CHANGELOG.md
#	docs/product-technical-gap-baseline.md
Signed-off-by: Seongho Bae <me@seonghobae.me>

# Conflicts:
#	CHANGELOG.md
#	docs/product-technical-gap-baseline.md
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
BREAKING CHANGE: benchmark reports require schema 3 with a prospective evaluation identity plan.

Signed-off-by: Seongho Bae <me@seonghobae.me>
…60905' into codex/paired-policy-outcomes-20260905

Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
…60905' into codex/paired-policy-outcomes-20260905

Signed-off-by: Seongho Bae <me@seonghobae.me>

# Conflicts:
#	tests/test_nim_benchmark.py
@seonghobae

Copy link
Copy Markdown
Contributor Author

독립 Visual Inspection: 원격 head 4351af9f1b6dca2f97526a46ac6dca28048ff623을 API로 확인한 뒤, 동일 SHA의 docs/nim_benchmark.md#comparison-report-version-3를 실제 Edge Preview 화면과 스크린샷으로 검사했습니다. 버전 3의 계획/관측 정합성, 내부 검증과 독립 사전등록의 구분, 동점일 때 관측·다른 정책 비교를 보존하는 설명, 구버전 보고서 처리 문단이 데스크톱 화면에서 잘림·겹침 없이 표시됐습니다. 이번 검증은 문서 렌더링에 한정하며 모바일·실제 공급자 실행·hosted CI·배포를 입증하지 않습니다.

Hosted quality coverage failed at 99% because validate_report_schema's
empty-identity, invalid-identity, count, worker-mismatch, and unknown
skip-reason branches were untested. Keep those incomplete plans from
publishing as complete paired evidence.
@seonghobae

Copy link
Copy Markdown
Contributor Author

Hosted Tests and package quality failed after a green full suite because validate_report_schema coverage was 99% (lines 2791, 2798, 2809, 2814, 2819, 2827).

Exact head is now 9f5d0ce2. The added contract tests reject empty evaluation identities, invalid identity fields, non-positive counts, catalog worker mismatches, and unknown cheapest-worker skip reasons. Local three-file coverage: 169 passed, 100% statements/branches, interrogate 100%. Fresh hosted Checks on this SHA are the remaining quality evidence; this does not authorize a production default change.

@seonghobae

Copy link
Copy Markdown
Contributor Author

Focused verification of commit 9f5d0ce2f7e95eb3bfcaf5cbd4e2a165dab7c86c: tests/test_reasoning_effort_profile.py completed with 58 passed in 87.09 seconds, exit 0. The worktree still matched this commit and had no tracked or untracked changes after the run.

The command ran from the existing PR #1074 worktree using the main checkout's existing project virtual-environment Python because this worktree had no local interpreter. This is focused source regression evidence, not an isolated dependency installation, full-suite result, hosted acceptance, measured buyer accuracy, or release. It covers the retained promotion function rejecting caller-declared report authority. Parent integration and independent protected review remain required.

@seonghobae

Copy link
Copy Markdown
Contributor Author

Integrated parent 84a6052 by normal merge, preserving the child delta. Integrated head d7e9d9a passed 159 tests across test_openrouter_uptime.py, test_psychometric_benchmark_boundaries.py, test_paper_contracts.py, and test_reasoning_effort_profile.py in 19.56 seconds; git diff --check passed. This is focused local integration evidence, not the full suite, hosted acceptance, production performance, or protected merge. Parent citation-only rights handling and per-experiment baseline identities are retained.

@seonghobae

Copy link
Copy Markdown
Contributor Author

Full local validation completed for exact head 2157d70. The worktree remained clean and HEAD unchanged. Result: 3628 passed, 2 skipped in 2633.80 seconds (exit 0). JUnit independently reports 3630 tests, 0 failures, 0 errors, 2 skipped. Local artifact: /tmp/co1074-full-validation.SbOIqj/results.xml; SHA-256: a0b1120c3e21ff9d07283a0160b436c7357600428d0cb835244813b53d7ebf9a. This is local regression evidence only, not protected-branch approval, release publication, production validation, or permission to change production routing defaults.

seonghobae commented Sep 8, 2026

Copy link
Copy Markdown
Contributor Author

Exact-head no-heuristics review for 2157d702c4609cc51829d7423911cadf26cf83d0 (current contextual_orchestrator/nim_benchmark.py blob 4e6094215892c9fc85ef3eb58f749b1b58f680d4).

The earlier 30/0.9 admission finding is now fail-closed in the report (minimum_paired_task_count=None, required_completion_fraction=None, routing_recommendation=None), and the lexicographic tie-break is gone. Those valid successor deltas must be preserved.

A separate merge blocker remains live in the current source:

  • paired_bootstrap_mean_difference(..., iterations=2000) still derives fixed percentile indices 0.025/0.975 and publishes method="paired_bootstrap_percentile_95"; this is inferential quality/latency evidence without an identified error target, sampling design, coverage validation, or fast-mlsirm contract.
  • The benchmark still changes the evaluated candidate set and test-time compute through caller defaults: first-ID max_eval_models=7, MAX_WORKFLOW_DEPTH=5, max_total_requests=2000, DEFAULT_MAX_OUTPUT_TOKENS=264, timeout_seconds=60.0, and probe_concurrency=4.
  • The token budget is enforced using estimate_tokens, explicitly documented as the ~4 chars/token heuristic, so the equal-budget claim is not measurement-valid.
  • The hard-coded comparison set (conduct_bounded / route_once / cheapest_eligible_worker plus hindsight leader) and Pareto output still select which contrasts are reported without a preregistered released measurement contract.

The canonical owner boundary is already consumable rather than hypothetical: contextual-orchestrator's exact branch pins immutable fast-mlsirm v0.9.1, and the protected consumer already imports that release for PsychometricRoutingEvidence. v0.9.1 exposes the validated judge-construct/IRT projection contract; #1074 currently bypasses it with a local scalar-score bootstrap. This is therefore a consumer integration gap, not permission to copy owner source or hand-roll another estimator.

Acceptance boundary: route applicable response-quality/calibration/uncertainty through the pinned immutable fast-mlsirm contract with executable provenance and a declared estimand/error target. Until the released contract can represent this benchmark design, fail closed: do not publish inferential CI, hindsight/Pareto selection, or an "equal-budget" result based on approximate token counts; retain complete raw per-cell evidence and descriptive fields only. Compute/model/time/token limits must come from the released allocation contract or be required explicit provenance that cannot authorize routing/promotion.

This review does not invalidate the child’s valid failed-task denominator and locked-cohort repairs. It keeps #1074 Draft and stacked on #1067; no force update, threshold substitution, or source ownership bypass is proposed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working priority: medium Normal-priority or P2 work status: draft type: bug Defect or incorrect behavior

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant