Skip to content

Add GrammarAwareRanker for best-of-N self-consistency - #3

Merged
piroz merged 1 commit into
mainfrom
feat/grammar-aware-ranker
May 22, 2026
Merged

piroz merged 1 commit into
mainfrom
feat/grammar-aware-ranker

Conversation

@piroz

@piroz piroz commented May 21, 2026

Copy link
Copy Markdown
Contributor

Implements the next-step "B" from the research spike: grammar-aware best-of-N reranking to close the small-LLM 1-shot miss gap, while staying logit-free.

Problem

When selfConsistency.n > 1 and none of the N candidates fully validates, the old "first-valid wins" picker silently degraded to "first candidate wins" — the rest of the self-consistency signal was thrown away. The error feedback fed into the retry loop was therefore whichever sample happened to be sampled first, not whichever sample got closest to a valid program.

Approach

  • ValidationResult gains an optional errorOffset so a validator can report how far it parsed before giving up.
  • Ranker.rank gains a context?: { validations } parameter so a ranker can score off that progress without re-parsing.
  • GrammarAwareRanker (new) scores valid → 0, else max(1, len - errorOffset) — lower wins, so:
    • any fully-valid candidate beats every invalid one;
    • among invalid candidates, the deepest parse beats a shallow one.
  • The picker (src/noroshi.ts::pick) is now async, threads validations through, and the out-of-retries return surfaces the picked output (was dropping the ranker's choice into lastCandidates[0]).

Why this is worth doing now

  • Directly attacks noroshi's biggest weakness: small models (1B-3B) frequently miss on the first try.
  • Stays inside the prompt-side contract (no logit access required) so every adapter — FetchAdapter, WebLLMAdapter, future ChromePromptAdapter — gets the benefit equally.
  • The deepest-failed-parse heuristic feeds the retry loop the most specific error message, which compounds with retry.includeErrorInPrompt — basically validity-aware sampling in the spirit of CRANE / GAD without needing constrained decoding.

Tests

New suite test-ranker.ts (6 cases):

  • Direct rank() ordering: valid (0) < deep-fail (1) < shallow-fail (24).
  • End-to-end through generate():
    • all 3 candidates invalid → deepest parse wins
    • mixed set → valid candidate wins regardless of position
    • no-ranker baseline → first-valid still wins (regression)

Combined with the existing suites (test.ts 9 / test-fetch.ts 8 / test-safety.ts 15), npm test now runs 38 assertions, all green locally and on Node 20 + 22 (waiting on CI).

Files

  • src/types.ts — ValidationResult.errorOffset?, Ranker.rank(..., context?)
  • src/rankers/grammar.ts (new) — GrammarAwareRanker
  • src/index.ts — exports GrammarAwareRanker
  • src/noroshi.ts — async pick, lastOutput fallback fix
  • examples/creative-coding-p5js/validator.ts — lex + parser emit errorOffset
  • examples/creative-coding-p5js/harness.mjs — mirrors the above for the browser path
  • examples/creative-coding-p5js/test-ranker.ts (new) — 6 cases
  • package.json — npm test runs the new suite alongside the others

When the LLM emits N candidates and none fully validate, "first-valid
wins" silently degenerates to "first candidate wins" — wasting the
self-consistency signal entirely. This change makes the rerank loop
aware of how far each candidate progressed through the grammar so that:

  - if any candidate validates, it wins;
  - otherwise, the candidate that parsed furthest before failing is
    chosen — its error feedback is the most specific, which the retry
    loop then folds back into the next prompt.

Changes:

- types.ts: ValidationResult now carries an optional errorOffset; the
  Ranker interface gets a `context?: { validations? }` parameter so
  rankers can score off the validator's progress without re-parsing.
- src/rankers/grammar.ts: GrammarAwareRanker — score is 0 for valid,
  else max(1, len - errorOffset). Exported as a public ranker.
- src/noroshi.ts:
  - pick() is now async and threads validations through to the ranker;
  - when out of retries the final return surfaces the picked output,
    not lastCandidates[0] (was dropping the ranker's choice).
- examples/creative-coding-p5js/validator.ts: lex + parser now report
  errorOffset on every failure path (uses a lastPos cursor for parser
  errors that don't have an obvious token-level position).
- examples/creative-coding-p5js/harness.mjs: mirrors the async pick,
  the lastOutput fix, and ships a GrammarAwareRanker for in-browser
  use.
- examples/creative-coding-p5js/test-ranker.ts: 6 cases — direct rank()
  ordering (valid < deep-fail < shallow-fail) and end-to-end through
  generate() (all-invalid → deepest wins; mixed → valid wins; no-ranker
  baseline → first-valid still wins).
- npm test runs the new suite alongside the existing three; total is
  now 38 assertions (9 + 8 + 15 + 6) per run.

Background: addresses the "small-LLM 1-shot miss" weakness identified
in the next-step research spike; closes the gap between Wang et al.'s
first-valid heuristic and what CRANE / GAD-style validity-aware
sampling buys at this scale, while staying logit-free.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@piroz
piroz merged commit c87c022 into main May 22, 2026
2 checks passed
@piroz
piroz deleted the feat/grammar-aware-ranker branch May 22, 2026 06:36
piroz added a commit that referenced this pull request May 22, 2026
Five-row ablation ladder (baseline / +grammar / +few-shot / +retry /
+rerank) over 20 creative-coding tasks distinct from the few-shot
bank, every output evaluated by the same CreativeCodingValidator.

Frozen run on Qwen2.5-1.5B-Instruct (Ollama, RTX 2060 Mobile):

  baseline      0/20  ( 0%)   1.0 attempts   492 ms
  grammar-only  1/20  ( 5%)   1.0 attempts   645 ms
  +few-shot    11/20  (55%)   1.0 attempts   443 ms
  +retry       15/20  (75%)   1.8 attempts   727 ms
  +rerank      17/20  (85%)   4.7 attempts  1539 ms

Files:
- examples/creative-coding-p5js/bench/tasks.ts: 20 novel tasks across
  the grammar surface (static / repeat / tick / mouse / time).
- examples/creative-coding-p5js/bench/run.ts: configurable runner —
  endpoint, model, API key via env vars; prints per-cell ✓/✗ trace
  and a summary table; writes results-<model>.json.
- examples/creative-coding-p5js/bench/results-qwen2.5_1.5b.json: the
  frozen run, checked in so the README numbers are reproducible.
- examples/creative-coding-p5js/bench/README.md: methodology, full
  table, and per-row commentary (where each component buys what).
- README.md (root): headline table — "1.5B on-device model goes from
  0% to 85% valid without fine-tuning."

Background: completes step 2 from the post-PR-#3 plan (small-LLM ×
novel-DSL public bench). The +rerank row is the GrammarAwareRanker
shipped in PR #3 doing its intended job.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant