Repository navigation
Multi-model bench: noroshi pipeline vs model selection on small LLMs - #5
Merged
Merged
Conversation
Adds frozen runs for two more on-device-class models alongside the
existing qwen2.5:1.5b row, and rewrites the bench/root READMEs around
the comparison.
Frozen runs (Ollama on RTX 2060 Mobile, 2026-05-22):
Model base +gram +fs +retry +rerank +rerank ms
qwen2.5:1.5b 0% 5% 55% 75% 85% 1539
gemma2:2b 0% 10% 30% 45% 75% 7244
llama3.2:1b 0% 0% 0% 0% 0% 17839
What it shows:
- The pipeline itself is roughly model-agnostic — qwen and gemma ride
the same shape of ablation curve, gemma just lower and ~5x slower.
- Model selection still dominates. llama3.2:1b flatlines at 0% across
every cell because it consistently echoes the grammar's leading rule
("start: block+" → output begins with the literal token "start" →
lex fails at offset 0). retry-with-feedback doesn't shake it loose;
best-of-3 doesn't either, because all N samples make the same
mistake. noroshi can't amplify a model that doesn't have the DSL
surface in its prior.
Practical guidance derived from this: for a 1-2B on-device creative-
coding target, default to the Qwen-2.5 family.
Files:
- bench/results-llama3.2_1b.json, bench/results-gemma2_2b.json
- bench/README.md: triple-model table (success + latency), per-row
reading, "what the multi-model row teaches us" section, caveats.
- README.md (root): the multi-model table replaces the single-model
one in the Benchmark section.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds two more frozen bench runs alongside the existing
qwen2.5:1.5brow from PR #4 and rewrites the bench tables around the cross-model comparison.Result table
Same 5-ablation × 20-task bench, three on-device-class models via local Ollama on an RTX 2060 Mobile:
qwen2.5:1.5bgemma2:2bllama3.2:1bWhat it teaches
Two takeaways, both worth surfacing in the README:
The pipeline itself is roughly model-agnostic.
qwenandgemmaride the same ablation curve shape — each additional component (grammar → few-shot → retry → rerank) buys success rate roughly in proportion. The pipeline isn't tuned to a specific model.Model selection still dominates.
llama3.2:1bflatlines at 0% in every cell. Inspection of per-task error messages shows it consistently echoes the grammar's leading rule (start: block+→ output begins with the literal tokenstart→ lex fails at offset 0). retry-with-feedback doesn't shake it loose; best-of-3 doesn't either, because all N samples make the same mistake. noroshi can't amplify a model that doesn't have the DSL surface in its prior.Practical guidance derived from this: for a 1-2B on-device creative-coding target, default to the Qwen-2.5 family.
Files
examples/creative-coding-p5js/bench/results-llama3.2_1b.json,results-gemma2_2b.json— frozen JSON outputs, checked in so the README numbers are reproducible.examples/creative-coding-p5js/bench/README.md— now a three-model success table + three-model latency table + per-row reading on the best-performing model + a "what the multi-model row teaches us" section.README.md(root) — Benchmark section's table swaps from single-model to three-model.Not in scope (follow-ups)
llama3.2:1b(different grammar header naming, heavier derivation labels, larger Llama-family model) — only worth doing if a user actually plans to ship Llama-class.