Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions config/benchmark/benchmark-v3.json
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,9 @@
"targets": [
{
"id": "elf",
"image": "elf-benchmark-elf:v5",
"dockerfile": "docker/benchmark/Dockerfile",
"build_target": "elf-runtime",
"adapter": "elf",
"role": "source_linked_agent_memory_and_knowledge",
"score_eligible": true,
Expand Down
13 changes: 13 additions & 0 deletions docker/benchmark/Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,19 @@ WORKDIR /opt/qmd
RUN pnpm install --frozen-lockfile \
&& pnpm run build

FROM python:3.11-bookworm AS elf-runtime

ARG ELF_SOURCE_COMMIT=unknown
COPY --from=elf-builder /out/real_world_live_adapter /usr/local/bin/real_world_live_adapter
COPY config/local/elf.docker.toml /opt/elf/
COPY config/local/tokenizer.wordlevel.json /config/local/tokenizer.wordlevel.json
COPY scripts/benchmark-unit.py /opt/benchmark/benchmark-unit.py
COPY scripts/benchmark_targets /opt/benchmark/benchmark_targets
LABEL org.opencontainers.image.revision="${ELF_SOURCE_COMMIT}"
ENV PYTHONUNBUFFERED=1
WORKDIR /
CMD ["python3", "/opt/benchmark/benchmark-unit.py", "--target", "elf"]

FROM node:22-bookworm

ARG ELF_SOURCE_COMMIT=unknown
Expand Down
1 change: 1 addition & 0 deletions docker/benchmark/compose.yml
Original file line number Diff line number Diff line change
Expand Up @@ -149,6 +149,7 @@ services:

elf-unit:
<<: *unit
image: ${BENCHMARK_ELF_IMAGE:-elf-benchmark-elf:v5}
depends_on:
postgres:
condition: service_healthy
Expand Down
5 changes: 5 additions & 0 deletions docs/log.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,11 @@ logs.

## 2026-09-21

- Added bounded quick, measure, and compare benchmark modes; explicit scenario
selection, scoring-blind local baselines, atomic resumable receipts, phase
timings, and failure-aware acceptance. Offline evidence is separated from
model-backed measurements and unimplemented agent-task coverage.

- Removed the retired SAG v1 and Letta Python server implementations from the
default competitor matrix, dispatcher, container setup, and dependency locks.
Preserved dated evidence and documented the exact historical reproduction ref.
Expand Down
10 changes: 8 additions & 2 deletions docs/reference/benchmark-target-lifecycle.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ Not this document: A fresh measurement or a claim that retained target pins are

## Current set

The default manifest retains ten targets: ELF, mem0, qmd, LightRAG, OpenViking,
The competitor manifest retains ten targets: ELF, mem0, qmd, LightRAG, OpenViking,
Graphiti, GraphRAG, PageIndex, OpenKB, and Honcho. The four suite definitions,
scoring rules, and PageIndex eligibility boundary are unchanged. A reduction in
the matrix is not an improvement in measured performance. Do not compare an
Expand All @@ -42,7 +42,7 @@ Neither project is declared inactive. OpenKB and GraphRAG remain available:
slower development alone does not prove retirement. Other retained pins still
need separate freshness reviews; this change does not silently upgrade them.

The default manifest, target dispatcher, build list, Compose services, Dockerfiles,
The competitor manifest, target dispatcher, build list, Compose services, Dockerfiles,
and retired adapter dependencies no longer expose these two old implementations.
Explicit selection of a retired target fails before provider access or artifact
creation. Retired implementation-specific tests are removed; shared provenance,
Expand All @@ -61,3 +61,9 @@ Preserve original bundle contents and matrix identities when rendering old
reports. Do not remove rows from a historical result or relabel an old SAG/Letta
measurement as a measurement of its replacement. This cleanup performs no paid
provider calls and produces no new competitor-quality results.

## Execution modes

The command now defaults to offline quick checks. Use `--mode compare` for this
competitor set. The measure and compare modes also include two local baselines.
See [bounded benchmark modes](../runbook/benchmarking/benchmark_modes.md).
133 changes: 133 additions & 0 deletions docs/runbook/benchmarking/benchmark_modes.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,133 @@
---
type: Runbook
title: "Bounded Benchmark Modes"
description: "Run offline checks, ELF baseline measurements, and explicit competitor comparisons."
resource: docs/runbook/benchmarking/benchmark_modes.md
status: active
authority: procedural
owner: benchmarking
last_verified: 2026-09-21
code_refs:
- scripts/benchmark_runner/cli.py
- scripts/benchmark_runner/profiles.py
- scripts/benchmark_runner/checkpoints.py
- scripts/benchmark_runner/baselines.py
- scripts/benchmark_runner/runtime.py
- scripts/tests/test_benchmark_modes.py
related:
- docs/reference/benchmark-target-lifecycle.md
- docs/research/2026-09-20-elf-value-review.md
---
# Bounded Benchmark Modes

Purpose: Obtain bounded, reproducible feedback without running the entire competitor matrix.
Read this when: Checking a change, measuring ELF against simple baselines, or resuming a run.
Not this document: Proof of superiority, production acceptance, or a full agent-task benchmark.

## One entrypoint

Run all commands from the repository root. The existing Python benchmark owner is
retained under the repository engineering conventions; no new runtime is required.

```sh
cargo make benchmark-competitors --plan
cargo make benchmark-competitors --mode quick --max-seconds 240
cargo make benchmark-competitors --mode measure --max-seconds 1800 --max-units 12
cargo make benchmark-competitors --mode compare --only-target qmd --max-seconds 900
```

The default is now `quick`. `compare` explicitly selects the former full competitor
scope, plus the two baselines. Use `--plan` to inspect all scheduled jobs and pins
without reading provider configuration, building images, or creating artifacts.
`--only-target` selects one maintained product or baseline; it does not bypass
suite eligibility. `--job-limit` is an explicitly diagnostic deterministic subset,
not a representative quality claim. Zero and out-of-range limits fail.

| Mode | Execution | Default sample | Result authority |
| --- | --- | --- | --- |
| quick | `cargo make test`, then no-memory and persisted-file search | 18 named scenarios across four suites; 8 baseline units | Source tests and offline pipeline integrity. No ELF runtime or model answers. |
| measure | ELF, no-memory, files-search; one shared answer model and fixed suite inputs | Same 18 scenarios; 12 units | A small product retrieval/lifecycle comparison, not end-to-end agent task utility. |
| compare | Ten retained products plus both baselines | All 48 scenarios with per-product suite eligibility; 32 units | Frozen-version comparative evidence; execution failures remain visible. |

The selected sample includes direct lookup, synthesis, correction, work resumption,
scope traps, abstention, update, delete, combined mutations, document structure,
repository changes, and citation boundaries. `report.md` lists every selected job.
A passing sampled run cannot claim coverage of omitted jobs.

## Baseline boundaries

`no-memory` supplies no retrieved context. `files-search` writes each task's input
records to disk, applies explicit update/delete operations, reads them back, and
ranks by case-folded query-token overlap. It has no learned memory, ACL engine,
semantic index, or inferred updates. It is a transparent lexical baseline, not a
replacement for native host memory plus maintained project files.

Only opaque product fixtures enter either baseline. Expected answers, qrels,
privacy labels, and scoring rules remain in the central evaluator. Offline runs
leave answers absent; they do not manufacture answers from gold data. Live runs
use the same answer stage as ELF. Source preparation/curation cost and multi-session
agent behavior are not measured by these baselines.

## Budgets and timing

`--max-seconds` bounds external subprocess and shared-answer request waits.
Command timeouts terminate their process group. Cleanup has a separate 180-second
reserve so a time limit does not intentionally leave containers behind. Cleanup
failure fails acceptance. Small in-process JSON operations are not preempted.

`--max-units` bounds newly attempted units. Reused results do not consume a unit.
Unstarted scheduled units are explicitly recorded as timeout failures; they never
vanish from the denominator. No unbounded automatic retry is performed.

The report bundle records preflight, build, total, and unit duration. Container
units also record runtime, log/cleanup, shared-answer, and evaluation duration;
product payloads retain ingest/query measurements where supplied. Shared-answer
usage is retained when the provider supplies it. This is not a dollar cap: native
product adapters may make multiple model calls within a unit. Do not claim a cost
or speed improvement without matched repeated measurements. Image preparation can
consume the full budget on a cold host. ELF uses the `elf-runtime` build stage
and does not prepare qmd or mem0 dependencies for an ELF-only measurement.
Products remain serial to preserve the existing provider contention boundary.

## Resume safely

```sh
cargo make benchmark-competitors --mode measure --resume tmp/benchmark-v5/PRIOR_RUN --max-seconds 1800
```

New runs use `elf.benchmark_bundle/v2`. The report command supports v1 and v2
without relabeling quick evidence as a complete measured run.

Each run writes to a new directory. Receipts are atomic and bind source state,
manifest, exact selected suites, mode, provider routes/models, and built image IDs.
Only complete units with successful cleanup, replay, and coverage can be reused.
The raw result is evaluated again and must match the recorded evaluation. Changed
identity, malformed receipts, and failed units cause execution rather than reuse.
The hashes detect accidental modification; these are trusted local artifacts, not
signed evidence from an untrusted contributor.

Reused rows name the original receipt and remain historical timing samples. The
current implementation still prepares providers/images before reuse, and it does
not reuse live mutable product databases across scenarios. `--skip-build` is an
operator option for existing images; their image IDs are recorded. The shared ELF image must also declare the current Git HEAD in its source label;
an older or missing label is rejected. Dirty-source runs remain diagnostic only.

## Acceptance and limits

Execution acceptance checks exact scheduled target coverage, both phases, typed
completion, cleanup, and deterministic replay. Empty/partial/failed runs fail.
Quality is separate: recall, nDCG, answer correctness, and seeded invariant failures
remain in the bundle/report. ELF seeded forbidden-content and mutation violations
return a nonzero exit code even when execution completed. Baseline quality failures
remain comparison evidence; they do not make an otherwise valid run disappear.

Unmeasured in this sample: native host memory, autonomous task completion, Chinese
inputs, actual tenant ACL enforcement, restart recovery, and large-corpus behavior.
Use the existing integration and E2E tasks for their supported contracts. Adding
these benchmark dimensions requires new scenarios and independent outcome checks;
renaming the old fixtures does not supply that evidence.

Live execution requires the existing embedding and LiteLLM configuration. Missing
configuration produces a failed report with value-free prerequisite status. Never
put secret values in CLI arguments or published artifacts. Private run artifacts
remain local. A provider-backed run is not implied by successful offline tests.
2 changes: 2 additions & 0 deletions docs/runbook/benchmarking/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,3 +13,5 @@ Routes to: Benchmarking runbooks under `docs/runbook/benchmarking/`.
- `real_world_agent_memory_benchmark.md`: operator map for creating, extending, and
interpreting real-world agent memory benchmark jobs.
- `real_world_memory_evolution.md`: memory-evolution fixture runbook.

- [Bounded Benchmark Modes](benchmark_modes.md): offline checks, baseline measurements, budgets, and resumable evidence.
2 changes: 1 addition & 1 deletion scripts/benchmark-report.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ def parse_args() -> argparse.Namespace:
def main() -> int:
args = parse_args()
bundle = json.loads(args.bundle.read_text(encoding="utf-8"))
if bundle.get("schema") != "elf.benchmark_bundle/v1":
if bundle.get("schema") not in {"elf.benchmark_bundle/v1", "elf.benchmark_bundle/v2"}:
raise ValueError("input is not an ELF benchmark bundle")
args.out.parent.mkdir(parents=True, exist_ok=True)
args.out.write_text(publish(bundle), encoding="utf-8")
Expand Down
40 changes: 40 additions & 0 deletions scripts/benchmark_report/modes.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
"""Render v2 mode evidence without claiming a complete measured competitor run."""
from __future__ import annotations


def publish_modes(bundle):
lines = ["# Benchmark result", "", f"Mode: {bundle['mode']}",
f"Execution passed: {bundle['acceptance']['passed']}",
"", "Execution success does not establish product quality or superiority.",
"Quick mode does not execute ELF, external products, or model answers.", "",
"| Suite | Target | Status | Reused | Seconds |", "| --- | --- | --- | --- | --- |"]
for name, suite in sorted(bundle["suite_results"].items()):
for row in suite["results"]:
lines.append(f"| {name} | {row['target']} | {row['evaluation']['classification']} | "
f"{row.get('reuse', {}).get('reused', False)} | {row.get('duration_seconds', 0)} |")
lines += ["", "## Warm retrieval metrics", "",
"| Suite | Target | Recall@5 | nDCG@5 | Answer correctness |", "| --- | --- | --- | --- | --- |"]
for name, suite in sorted(bundle["suite_results"].items()):
for row in suite["results"]:
metrics = row["evaluation"].get("phases", {}).get("warm", {}).get("metrics") or {}
def value(key):
measured = metrics.get(key)
return "not measured" if measured is None else str(measured)
lines.append(f"| {name} | {row['target']} | {value('mean_recall_at_5')} | "
f"{value('mean_ndcg_at_5')} | {value('programmatic_answer_correctness')} |")
lines += ["", "## Findings", ""] + [f"- {f}" for f in bundle["acceptance"]["findings"]]
lines += ["", "## Seeded invariant findings", ""]
for finding in bundle.get("quality", {}).get("seeded_violations", []):
lines.append(f"- {finding['target']}/{finding['job']}: {finding['reason']}")
missing = [key for key, present in bundle.get("provider_configuration", {}).items() if not present]
if missing:
lines += ["", "Missing provider configuration: " + ", ".join(missing)]
lines += ["", "## Selected scenarios", ""]
for suite, jobs in sorted(bundle["coverage"].items()):
lines.append(f"- {suite}: " + ", ".join(jobs))
lines += ["", "## Coverage limits", "", "Unmeasured: native host memory, end-to-end agent actions, Chinese inputs, "
"ACL enforcement, restart recovery, and large-corpus behavior. Use the repository integration/E2E gates "
"for their existing guarantees. File search is a lexical baseline, not a simulated competitor.", "",
"All raw unit results and evaluator metrics are in bundle.json. Reused timings are historical samples."]
return "\n".join(lines) + "\n"

4 changes: 4 additions & 0 deletions scripts/benchmark_report/publish.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,8 @@
from typing import Any
import json

from .modes import publish_modes

from .coverage import (
common_ci,
coverage_table,
Expand All @@ -20,6 +22,8 @@


def publish(bundle: dict[str, Any]) -> str:
if bundle.get("schema") == "elf.benchmark_bundle/v2":
return publish_modes(bundle)
source = bundle.get("source") or {}
routes = bundle.get("provider_routes") or {}
preflight = bundle.get("provider_preflight") or {}
Expand Down
4 changes: 3 additions & 1 deletion scripts/benchmark_runner/answers.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,8 @@
from benchmark_contract import NativeContextError, answer_cases


from .runtime import remaining_seconds

def failure_unit(
target: dict[str, Any], suite: dict[str, Any], classification: str, message: str
) -> dict[str, Any]:
Expand Down Expand Up @@ -101,7 +103,7 @@ def attach_shared_answers(
)
native: dict[str, Any] | None = None
try:
with urllib.request.urlopen(request, timeout=300) as response:
with urllib.request.urlopen(request, timeout=remaining_seconds(300)) as response:
native = json.loads(response.read().decode())
unit["provider_raw"] = {"shared_answer": native}
content = native["choices"][0]["message"]["content"]
Expand Down
75 changes: 75 additions & 0 deletions scripts/benchmark_runner/baselines.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,75 @@
"""Scoring-blind local baselines over persisted files, without external services."""
from __future__ import annotations

import json
import re
import time
from pathlib import Path
from typing import Any

from benchmark_contract import assert_product_payload_is_blind, load_json
from .runtime import write_json

BASELINES = ("no-memory", "files-search")


def target_contract(name: str, suites: list[str]) -> dict[str, Any]:
return {"id": name, "adapter": name, "suites": suites, "score_eligible": True,
"pin": {"kind": "benchmark_source_commit"}}


def terms(text: str) -> set[str]:
return set(re.findall(r"[\w]+", text.casefold()))


def retrieve(items: list[dict[str, Any]], query: str, name: str) -> list[dict[str, Any]]:
if name == "no-memory":
return []
query_terms = terms(query)
ranked = [(len(query_terms & terms(item["text"])), item) for item in items]
ranked.sort(key=lambda pair: (-pair[0], pair[1]["evidence_id"]))
return [item for score, item in ranked[:5] if score > 0]


def run_baseline(name: str, inputs: Path, state: Path) -> dict[str, Any]:
if name not in BASELINES:
raise ValueError("unknown baseline")
ingest_ms = 0.0
phases: dict[str, Any] = {p: {"status": "completed", "jobs": []} for p in ("cold", "warm")}
for path in sorted(inputs.glob("*.json")):
job = load_json(path)
assert_product_payload_is_blind(job)
store = state / (job["job_id"] + ".json")
items = {item["evidence_id"]: item for item in job["corpus"]["items"]}
ingest_started = time.monotonic()
write_json(store, items)
ingest_ms += (time.monotonic() - ingest_started) * 1000
for phase in ("cold", "warm"):
operations = []
items = json.loads(store.read_text())
if phase == "warm":
for operation in job.get("operations", []):
key = operation["evidence_id"]
if operation["type"] == "delete":
items.pop(key, None)
elif operation["type"] == "update":
items[key] = {"evidence_id": key, "text": operation["text"]}
else:
raise ValueError("unsupported file operation")
operations.append({"requested_type": operation["type"],
"native_type": operation["type"], "classification": "completed",
"native_success": True})
write_json(store, items)
items = json.loads(store.read_text())
query_started = time.monotonic()
contexts = retrieve(list(items.values()), job["prompt"]["content"], name)
phases[phase]["jobs"].append({"job_id": job["job_id"], "classification": "completed",
"evidence_ids": [item["evidence_id"] for item in contexts],
"contexts": contexts, "returned_count": len(contexts),
"latency_ms": (time.monotonic() - query_started) * 1000,
"operations": operations})
return {"schema": "elf.benchmark_unit_result/v4", "target": name,
"native_mode": name, "score_eligible": True, "result_class": "completed",
"ingest_count": 1, "warm_reused_state": True,
"ingest_duration_ms": ingest_ms, "phases": phases,
"baseline_boundary": "Lexical token-overlap file search; no native host memory, agent actions, or ACL enforcement."}
Loading
Loading