Skip to content

Multi-model bench: noroshi pipeline vs model selection on small LLMs - #5

Merged
piroz merged 1 commit into
mainfrom
feat/bench-multi-model
May 22, 2026
Merged

piroz merged 1 commit into
mainfrom
feat/bench-multi-model

Conversation

@piroz

@piroz piroz commented May 22, 2026

Copy link
Copy Markdown
Contributor

Adds two more frozen bench runs alongside the existing qwen2.5:1.5b row from PR #4 and rewrites the bench tables around the cross-model comparison.

Result table

Same 5-ablation × 20-task bench, three on-device-class models via local Ollama on an RTX 2060 Mobile:

Model Size baseline +grammar +few-shot +retry +rerank +rerank ms
qwen2.5:1.5b 0.99 GB 0% 5% 55% 75% 85% 1,539
gemma2:2b 1.6 GB 0% 10% 30% 45% 75% 7,244
llama3.2:1b 1.3 GB 0% 0% 0% 0% 0% 17,839

What it teaches

Two takeaways, both worth surfacing in the README:

  1. The pipeline itself is roughly model-agnostic. qwen and gemma ride the same ablation curve shape — each additional component (grammar → few-shot → retry → rerank) buys success rate roughly in proportion. The pipeline isn't tuned to a specific model.

  2. Model selection still dominates. llama3.2:1b flatlines at 0% in every cell. Inspection of per-task error messages shows it consistently echoes the grammar's leading rule (start: block+ → output begins with the literal token start → lex fails at offset 0). retry-with-feedback doesn't shake it loose; best-of-3 doesn't either, because all N samples make the same mistake. noroshi can't amplify a model that doesn't have the DSL surface in its prior.

Practical guidance derived from this: for a 1-2B on-device creative-coding target, default to the Qwen-2.5 family.

Files

  • examples/creative-coding-p5js/bench/results-llama3.2_1b.json, results-gemma2_2b.json — frozen JSON outputs, checked in so the README numbers are reproducible.
  • examples/creative-coding-p5js/bench/README.md — now a three-model success table + three-model latency table + per-row reading on the best-performing model + a "what the multi-model row teaches us" section.
  • README.md (root) — Benchmark section's table swaps from single-model to three-model.

Not in scope (follow-ups)

  • Rescue attempt on llama3.2:1b (different grammar header naming, heavier derivation labels, larger Llama-family model) — only worth doing if a user actually plans to ship Llama-class.
  • Best-of-N sweep on the working models (N=5 / N=8) — diminishing returns likely, unmeasured.
  • CRANE-style derivation/output delimiter split (option A from the research spike).

Adds frozen runs for two more on-device-class models alongside the
existing qwen2.5:1.5b row, and rewrites the bench/root READMEs around
the comparison.

Frozen runs (Ollama on RTX 2060 Mobile, 2026-05-22):

  Model            base  +gram  +fs   +retry  +rerank   +rerank ms
  qwen2.5:1.5b      0%    5%   55%    75%      85%        1539
  gemma2:2b         0%   10%   30%    45%      75%        7244
  llama3.2:1b       0%    0%    0%     0%       0%       17839

What it shows:

- The pipeline itself is roughly model-agnostic — qwen and gemma ride
  the same shape of ablation curve, gemma just lower and ~5x slower.
- Model selection still dominates. llama3.2:1b flatlines at 0% across
  every cell because it consistently echoes the grammar's leading rule
  ("start: block+" → output begins with the literal token "start" →
  lex fails at offset 0). retry-with-feedback doesn't shake it loose;
  best-of-3 doesn't either, because all N samples make the same
  mistake. noroshi can't amplify a model that doesn't have the DSL
  surface in its prior.

Practical guidance derived from this: for a 1-2B on-device creative-
coding target, default to the Qwen-2.5 family.

Files:
- bench/results-llama3.2_1b.json, bench/results-gemma2_2b.json
- bench/README.md: triple-model table (success + latency), per-row
  reading, "what the multi-model row teaches us" section, caveats.
- README.md (root): the multi-model table replaces the single-model
  one in the Benchmark section.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@piroz
piroz merged commit 074ae7e into main May 22, 2026
2 checks passed
@piroz
piroz deleted the feat/bench-multi-model branch May 22, 2026 10:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant