Skip to content

Add bounded benchmark modes and reliable resumable evidence - #387

Merged
acgxv merged 4 commits into
mainfrom
xv/benchmark-modes
Sep 21, 2026
Merged

acgxv merged 4 commits into
mainfrom
xv/benchmark-modes

Conversation

@acgxv

@acgxv acgxv commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Replace the default full competitor run with an offline quick check and explicit measure/compare modes. Quick runs repository tests and two scoring-blind file baselines; measure selects ELF plus those baselines over 18 named scenarios; compare retains the full frozen suites and current products. The report separates execution, quality metrics, reused results, and unmeasured capabilities.

Fix sampled-run false greens: missing targets, incomplete phases, execution failures, cleanup failures, and replay failures now fail acceptance. Add atomic identity-bound receipts, reevaluation before reuse, unit/wall-time budgets with process-group termination and a bounded cleanup reserve, stage timings, dry planning, and actionable value-free provider prerequisite reports. Split ELF into a dedicated Docker build stage to avoid qmd and mem0 setup for ELF-only measurement.

Validation: 55 Python benchmark tests pass; documentation, format, and whitespace checks pass. An actual quick run with repository tests completed in about 25 seconds on the existing local cache. CLI resume/budget checks are covered. The live baseline probe correctly failed before model calls because LITELLM_API_KEY is unavailable. No current product-quality or dollar-savings result is claimed; native host memory, agent task outcomes, Chinese inputs, restart/ACL runtime coverage, and large-corpus runs remain explicitly unmeasured by this sample.

Additional validation: the dedicated elf-runtime image built successfully within a 300-second cap, without the qmd builder or mem0 package layer. Both its Python dispatcher and Rust adapter started successfully. A real quick CLI run passed 455 Rust tests (92 integration cases skipped as declared), 54 benchmark tests, and 3 tooling tests in roughly 20 seconds on the local cache. A real resume run reused all eight baseline receipts while rerunning source tests. Chained resume and stale ELF-image rejection have regression coverage.

Final review added v2 bundle/report separation, deterministic report regeneration, chained resume receipts, and interruption cleanup with a regression test. Live measurement remains blocked by the missing benchmark LiteLLM credential.

@acgxv
acgxv merged commit 9138a4a into main Sep 21, 2026
9 checks passed
@acgxv
acgxv deleted the xv/benchmark-modes branch September 21, 2026 12:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant