Repository navigation
Add GrammarAwareRanker for best-of-N self-consistency - #3
Merged
Merged
Conversation
When the LLM emits N candidates and none fully validate, "first-valid
wins" silently degenerates to "first candidate wins" — wasting the
self-consistency signal entirely. This change makes the rerank loop
aware of how far each candidate progressed through the grammar so that:
- if any candidate validates, it wins;
- otherwise, the candidate that parsed furthest before failing is
chosen — its error feedback is the most specific, which the retry
loop then folds back into the next prompt.
Changes:
- types.ts: ValidationResult now carries an optional errorOffset; the
Ranker interface gets a `context?: { validations? }` parameter so
rankers can score off the validator's progress without re-parsing.
- src/rankers/grammar.ts: GrammarAwareRanker — score is 0 for valid,
else max(1, len - errorOffset). Exported as a public ranker.
- src/noroshi.ts:
- pick() is now async and threads validations through to the ranker;
- when out of retries the final return surfaces the picked output,
not lastCandidates[0] (was dropping the ranker's choice).
- examples/creative-coding-p5js/validator.ts: lex + parser now report
errorOffset on every failure path (uses a lastPos cursor for parser
errors that don't have an obvious token-level position).
- examples/creative-coding-p5js/harness.mjs: mirrors the async pick,
the lastOutput fix, and ships a GrammarAwareRanker for in-browser
use.
- examples/creative-coding-p5js/test-ranker.ts: 6 cases — direct rank()
ordering (valid < deep-fail < shallow-fail) and end-to-end through
generate() (all-invalid → deepest wins; mixed → valid wins; no-ranker
baseline → first-valid still wins).
- npm test runs the new suite alongside the existing three; total is
now 38 assertions (9 + 8 + 15 + 6) per run.
Background: addresses the "small-LLM 1-shot miss" weakness identified
in the next-step research spike; closes the gap between Wang et al.'s
first-valid heuristic and what CRANE / GAD-style validity-aware
sampling buys at this scale, while staying logit-free.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
piroz
added a commit
that referenced
this pull request
May 22, 2026
Five-row ablation ladder (baseline / +grammar / +few-shot / +retry / +rerank) over 20 creative-coding tasks distinct from the few-shot bank, every output evaluated by the same CreativeCodingValidator. Frozen run on Qwen2.5-1.5B-Instruct (Ollama, RTX 2060 Mobile): baseline 0/20 ( 0%) 1.0 attempts 492 ms grammar-only 1/20 ( 5%) 1.0 attempts 645 ms +few-shot 11/20 (55%) 1.0 attempts 443 ms +retry 15/20 (75%) 1.8 attempts 727 ms +rerank 17/20 (85%) 4.7 attempts 1539 ms Files: - examples/creative-coding-p5js/bench/tasks.ts: 20 novel tasks across the grammar surface (static / repeat / tick / mouse / time). - examples/creative-coding-p5js/bench/run.ts: configurable runner — endpoint, model, API key via env vars; prints per-cell ✓/✗ trace and a summary table; writes results-<model>.json. - examples/creative-coding-p5js/bench/results-qwen2.5_1.5b.json: the frozen run, checked in so the README numbers are reproducible. - examples/creative-coding-p5js/bench/README.md: methodology, full table, and per-row commentary (where each component buys what). - README.md (root): headline table — "1.5B on-device model goes from 0% to 85% valid without fine-tuning." Background: completes step 2 from the post-PR-#3 plan (small-LLM × novel-DSL public bench). The +rerank row is the GrammarAwareRanker shipped in PR #3 doing its intended job. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements the next-step "B" from the research spike: grammar-aware best-of-N reranking to close the small-LLM 1-shot miss gap, while staying logit-free.
Problem
When
selfConsistency.n > 1and none of the N candidates fully validates, the old "first-valid wins" picker silently degraded to "first candidate wins" — the rest of the self-consistency signal was thrown away. The error feedback fed into the retry loop was therefore whichever sample happened to be sampled first, not whichever sample got closest to a valid program.Approach
ValidationResultgains an optionalerrorOffsetso a validator can report how far it parsed before giving up.Ranker.rankgains acontext?: { validations }parameter so a ranker can score off that progress without re-parsing.GrammarAwareRanker(new) scoresvalid → 0, elsemax(1, len - errorOffset)— lower wins, so:src/noroshi.ts::pick) is now async, threads validations through, and the out-of-retries return surfaces the picked output (was dropping the ranker's choice intolastCandidates[0]).Why this is worth doing now
FetchAdapter,WebLLMAdapter, futureChromePromptAdapter— gets the benefit equally.retry.includeErrorInPrompt— basically validity-aware sampling in the spirit of CRANE / GAD without needing constrained decoding.Tests
New suite
test-ranker.ts(6 cases):rank()ordering: valid (0) < deep-fail (1) < shallow-fail (24).generate():Combined with the existing suites (
test.ts9 /test-fetch.ts8 /test-safety.ts15),npm testnow runs 38 assertions, all green locally and on Node 20 + 22 (waiting on CI).Files
src/types.ts—ValidationResult.errorOffset?,Ranker.rank(..., context?)src/rankers/grammar.ts(new) —GrammarAwareRankersrc/index.ts— exportsGrammarAwareRankersrc/noroshi.ts— async pick,lastOutputfallback fixexamples/creative-coding-p5js/validator.ts— lex + parser emiterrorOffsetexamples/creative-coding-p5js/harness.mjs— mirrors the above for the browser pathexamples/creative-coding-p5js/test-ranker.ts(new) — 6 casespackage.json—npm testruns the new suite alongside the others