Repository navigation
Add bounded benchmark modes and reliable resumable evidence - #387
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replace the default full competitor run with an offline quick check and explicit measure/compare modes. Quick runs repository tests and two scoring-blind file baselines; measure selects ELF plus those baselines over 18 named scenarios; compare retains the full frozen suites and current products. The report separates execution, quality metrics, reused results, and unmeasured capabilities.
Fix sampled-run false greens: missing targets, incomplete phases, execution failures, cleanup failures, and replay failures now fail acceptance. Add atomic identity-bound receipts, reevaluation before reuse, unit/wall-time budgets with process-group termination and a bounded cleanup reserve, stage timings, dry planning, and actionable value-free provider prerequisite reports. Split ELF into a dedicated Docker build stage to avoid qmd and mem0 setup for ELF-only measurement.
Validation: 55 Python benchmark tests pass; documentation, format, and whitespace checks pass. An actual quick run with repository tests completed in about 25 seconds on the existing local cache. CLI resume/budget checks are covered. The live baseline probe correctly failed before model calls because LITELLM_API_KEY is unavailable. No current product-quality or dollar-savings result is claimed; native host memory, agent task outcomes, Chinese inputs, restart/ACL runtime coverage, and large-corpus runs remain explicitly unmeasured by this sample.
Additional validation: the dedicated
elf-runtimeimage built successfully within a 300-second cap, without the qmd builder or mem0 package layer. Both its Python dispatcher and Rust adapter started successfully. A real quick CLI run passed 455 Rust tests (92 integration cases skipped as declared), 54 benchmark tests, and 3 tooling tests in roughly 20 seconds on the local cache. A real resume run reused all eight baseline receipts while rerunning source tests. Chained resume and stale ELF-image rejection have regression coverage.Final review added v2 bundle/report separation, deterministic report regeneration, chained resume receipts, and interruption cleanup with a regression test. Live measurement remains blocked by the missing benchmark LiteLLM credential.