Skip to content

research: verifier head-to-head — re-feed vs true-KV replay (H100, vLLM 0.23) - #109

Merged
matteso1 merged 1 commit into
mainfrom
research/verifier-headtohead
Jun 25, 2026
Merged

matteso1 merged 1 commit into
mainfrom
research/verifier-headtohead

Conversation

@matteso1

@matteso1 matteso1 commented Jun 25, 2026 •

Copy link
Copy Markdown
Member

The GPU run behind the inference-receipt thesis (the $0.65 H100 experiment). A replay verifier compares a recomputed token signature to a commitment; the commitment is the decode-time logprob, but DiFR/TOPLOC/SVIP/VeriLLM all reconstruct by re-feeding the transcript — and re-feed (bulk prefill) != decode. A thaw verifier replays from the true decode-time KV (.thawkv) and matches by construction.

Fresh generations, Qwen2.5-7B, 1× H100, vLLM 0.23.0, 4,652 tokens, prefix caching off (re-feed = clean fresh prefill):

  • false-reject: at τ=0.01 nats a re-feed verifier wrongly rejects 6.9% of genuine tokens; true-KV verifier 0%.
  • forgery masking: at low-margin decision tokens, re-feed drift exceeds the top1–top2 gap 19.8% of the time (1 in 5) — a single-token swap hides inside the verifier's own noise; a true-KV verifier (drift ~0) catches it.

So re-feed verification pays a false-reject cost on genuine output and has a forgery-acceptance hole at the tokens that matter; thaw's true-KV inference receipt removes both.

benchmarks/verifier_headtohead.py + site/receipts/2026-06-25_h100_verifier_headtohead.json. Honest scope: true-KV arm sound by construction (decode-time logprob = commitment, .thawkv bit-identical per separate receipts); a deployed-verifier head-to-head + cross-model generality are the follow-ups.

Summary by CodeRabbit

  • New Features
    • Added a new benchmark script to compare two verifier replay approaches on math problem prompts.
    • Now reports drift statistics, false-reject rates across tolerance settings, and token-masking rates in JSON output.
    • Added a new experiment receipt with hardware, version, and configuration details plus reproducible command information.

…LM 0.23)

The GPU run behind the inference-receipt thesis. A replay verifier reconstructs
a token's signature and compares to a commitment; the commitment is the
decode-time logprob, but every replay verifier (DiFR, TOPLOC, SVIP, VeriLLM)
reconstructs by RE-FEEDING the transcript. Re-feed (bulk prefill) != decode, so
the verifier drifts at low-margin tokens. A thaw verifier replays from the true
decode-time KV (.thawkv) and matches the commitment by construction.

Fresh generations, Qwen2.5-7B, 1x H100, vLLM 0.23.0, 4652 tokens, prefix
caching OFF so re-feed is a clean fresh prefill:
- false-reject: at tau=0.01 nats a re-feed verifier wrongly rejects 6.9% of
  GENUINE tokens; the true-KV verifier rejects 0%.
- forgery masking: at low-margin decision tokens, re-feed drift exceeds the
  top1-top2 gap 19.8% of the time (1 in 5) — a single-token swap hides inside
  the verifier's own re-feed noise; a true-KV verifier (drift ~0) catches it.

So replay-from-re-feed verification pays a false-reject cost on genuine output
AND has a forgery-acceptance hole at the tokens that matter; replay-from-true-KV
(thaw's inference receipt) removes both. Harness + receipt; honest scope noted
(true-KV arm sound by construction; deployed-verifier head-to-head is next).
@vercel

vercel Bot commented Jun 25, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
site Ready Ready Preview, Comment Jun 25, 2026 3:27am

Request Review

@coderabbitai

coderabbitai Bot commented Jun 25, 2026 •

Copy link
Copy Markdown

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 3c2da30f-b167-4c53-aa95-6ac3cfea7063

📥 Commits

Reviewing files that changed from the base of the PR and between 627597d and 893f962.

📒 Files selected for processing (2)
  • benchmarks/verifier_headtohead.py
  • site/receipts/2026-06-25_h100_verifier_headtohead.json

📝 Walkthrough

Walkthrough

Adds a benchmark script that compares decode-time and re-feed verifier logprobs on fixed prompts, computes drift and false-reject metrics across tolerances, and writes a JSON receipt with the measured results.

Changes

Verifier head-to-head benchmark

Layer / File(s) Summary
Benchmark scaffold and prompt set
benchmarks/verifier_headtohead.py
Defines the embedded prompt set and logprob helper functions for token lookup and margin calculation.
Decode-time and re-feed capture
benchmarks/verifier_headtohead.py
Parses CLI args, runs generation with prefix caching disabled, and re-feeds each transcript to collect prompt logprobs.
Drift and masking analysis
benchmarks/verifier_headtohead.py
Aligns decode-time and re-feed logprobs, computes drift, margin, and masking metrics, and emits JSON results.
Experiment receipt
site/receipts/2026-06-25_h100_verifier_headtohead.json
Adds the experiment metadata, drift statistics, false-reject sweep, forgery-masking metrics, scope text, and reproducer command.

Sequence Diagram(s)

sequenceDiagram
  participant main as main()
  participant llm as LLM
  participant tokenizer as tokenizer
  participant stdout as stdout
  participant jsonout as json-out

  main->>tokenizer: format PROMPTS with chat template
  main->>llm: generate completions with decode-time logprobs
  llm-->>main: committed token logprobs and top-5 alternatives
  main->>llm: re-feed each transcript as a single prefill
  llm-->>main: prompt-logprobs for each token
  main->>stdout: print results as JSON
  main->>jsonout: write JSON when requested
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Poem

🐇 I hop through prompts and logprob dew,
Re-fed the transcript, line by line anew.
Drift stats twinkle, margins glow,
JSON receipts make the burrow know.
Thump-thump—truth and replay both show.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch research/verifier-headtohead

Comment @coderabbitai help to get the list of available commands.

@matteso1
matteso1 merged commit de9a828 into main Jun 25, 2026
8 of 9 checks passed
@matteso1
matteso1 deleted the research/verifier-headtohead branch June 25, 2026 03:28

This branch was successfully deployed

1 active deployment
Preview — 893f962b Deployed Jun 25, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant