Skip to content

Small-LLM × novel-DSL benchmark: Qwen2.5-1.5B goes 0% → 85% - #4

Merged
piroz merged 1 commit into
mainfrom
feat/bench-small-llm-dsl
May 22, 2026
Merged

piroz merged 1 commit into
mainfrom
feat/bench-small-llm-dsl

Conversation

@piroz

@piroz piroz commented May 22, 2026

Copy link
Copy Markdown
Contributor

Completes step 2 from the post-PR-#3 plan: a public ablation bench that shows what noroshi buys on a 1B-class on-device model.

Headline

20 novel creative-coding tasks (distinct from the few-shot bank), driven through 5 ablation cells with Qwen2.5-1.5B-Instruct via local Ollama on an RTX 2060 Mobile:

Ablation Valid Success Avg attempts Avg latency
baseline (task only) 0/20 0% 1.0 492 ms
+ BNF grammar in prompt 1/20 5% 1.0 645 ms
+ few-shot examples 11/20 55% 1.0 443 ms
+ retry-with-feedback (×3) 15/20 75% 1.8 727 ms
+ best-of-3 with GrammarAwareRanker 17/20 85% 4.7 1,539 ms

A 1.5B model on a 4-year-old laptop GPU clears 85% of novel DSL tasks without any fine-tuning.

What each row buys

  • 0 → 5%: grammar alone barely moves the needle on a model this size.
  • 5 → 55%: derivation-first few-shot is the big lever — biggest single jump.
  • 55 → 75%: retry-with-feedback recovers 4 of the 9 remaining misses by piping the validator's error back into the next prompt.
  • 75 → 85%: best-of-3 + GrammarAwareRanker (PR Add GrammarAwareRanker for best-of-N self-consistency #3) clears 2 more cases where retry's correction didn't generalise.

Latency tradeoff: +rerank triples wall time for +10 pt. Right when the human-noticing threshold matters less than success rate (one-shot prompts), wrong for keystroke-level interactivity. Documented in the bench README.

Files

  • examples/creative-coding-p5js/bench/tasks.ts — 20 tasks across static / repeat / tick / mouse / time.
  • examples/creative-coding-p5js/bench/run.ts — env-configurable (endpoint, model, API key), prints per-cell trace and a summary table, writes results-<model>.json.
  • examples/creative-coding-p5js/bench/results-qwen2.5_1.5b.json — frozen run, checked in so the README numbers are reproducible.
  • examples/creative-coding-p5js/bench/README.md — methodology, full table, per-row commentary, caveats.
  • README.md (root) — headline table mirrored, with a link to the full bench dir.

How to reproduce

ollama pull qwen2.5:1.5b
npx tsx examples/creative-coding-p5js/bench/run.ts

Other models / endpoints via env vars (NOROSHI_BENCH_MODEL, NOROSHI_BENCH_ENDPOINT, NOROSHI_BENCH_API_KEY).

Not in scope (follow-ups)

  • Second-model run (e.g. llama3.2:1b, gemma-2-2b-it) so we can factor model-vs-pipeline.
  • Best-of-N with N > 3 — diminishing returns at this size are likely but unmeasured.
  • CRANE-style derivation/output delimiter split (option A in the research spike). The +few-shot row already includes derivation as Wang et al. originally proposed; CRANE goes further and we have a separate issue for it.

Five-row ablation ladder (baseline / +grammar / +few-shot / +retry /
+rerank) over 20 creative-coding tasks distinct from the few-shot
bank, every output evaluated by the same CreativeCodingValidator.

Frozen run on Qwen2.5-1.5B-Instruct (Ollama, RTX 2060 Mobile):

  baseline      0/20  ( 0%)   1.0 attempts   492 ms
  grammar-only  1/20  ( 5%)   1.0 attempts   645 ms
  +few-shot    11/20  (55%)   1.0 attempts   443 ms
  +retry       15/20  (75%)   1.8 attempts   727 ms
  +rerank      17/20  (85%)   4.7 attempts  1539 ms

Files:
- examples/creative-coding-p5js/bench/tasks.ts: 20 novel tasks across
  the grammar surface (static / repeat / tick / mouse / time).
- examples/creative-coding-p5js/bench/run.ts: configurable runner —
  endpoint, model, API key via env vars; prints per-cell ✓/✗ trace
  and a summary table; writes results-<model>.json.
- examples/creative-coding-p5js/bench/results-qwen2.5_1.5b.json: the
  frozen run, checked in so the README numbers are reproducible.
- examples/creative-coding-p5js/bench/README.md: methodology, full
  table, and per-row commentary (where each component buys what).
- README.md (root): headline table — "1.5B on-device model goes from
  0% to 85% valid without fine-tuning."

Background: completes step 2 from the post-PR-#3 plan (small-LLM ×
novel-DSL public bench). The +rerank row is the GrammarAwareRanker
shipped in PR #3 doing its intended job.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@piroz
piroz merged commit fa00c39 into main May 22, 2026
2 checks passed
@piroz
piroz deleted the feat/bench-small-llm-dsl branch May 22, 2026 07:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant