Zero-dependency prompt compression for LLM applications. One deterministic algorithm, three native SDKs — TypeScript/Bun, Rust and Go — with byte-identical output across all of them.
No ML model, no model download, no tokenizer dependency, no provider SDKs required. The local compressor, Extract, is a weights-free extractive algorithm: sentence segmentation → integer term-rarity scoring (boosted for numbers, entities, identifiers and constraints) → redundancy filtering → greedy selection to a keep-ratio, emitted in original order.
A heuristic classifier gates everything: code, stack traces, JSON blobs and tool-call blocks are never touched; prohibitions ("do not …") always survive; system messages and the most recent turns are never compressed.
| leanprompt | Neural compressors (e.g. LLMLingua-2) | |
|---|---|---|
| Model download | none | ~1 GB+ |
| Runtime dependencies | zero, in all 3 languages | GPU/ONNX runtime, tokenizer, model weights |
| Cold start | instant | model load (seconds, first call) |
| Output | deterministic, auditable | probabilistic, model-version-dependent |
| Cross-language parity | byte-identical by spec | not applicable — single implementation |
| Code/JSON/tool-call safety | classifier-gated, never touched | depends on wrapper logic |
| Language | Directory | Runtime deps | Quickstart |
|---|---|---|---|
| TypeScript (Node ≥ 18 / Bun) | ts/ |
zero | npm install leanprompt |
| Rust | rust/ |
zero ([dependencies] empty) |
cargo add leanprompt |
| Go | go/ |
zero (stdlib only) | go get github.com/itaides/leanprompt/go |
The TypeScript package is the reference implementation of
docs/parity-spec.md — a written, normative spec
(pinned character classes, integer-only scoring, explicit tiebreaks,
canonical JSON). The Rust and Go ports assert byte-equality against the
golden vectors in parity/, so identical inputs produce
identical compressed output in every language.
Full option reference (every field, every default, per-language) lives in
CONFIG.md. A few common patterns:
import OpenAI from "openai";
import { leanpromptFetch } from "leanprompt";
const client = new OpenAI({
fetch: leanpromptFetch({
mode: "on",
trigger: { thresholdTokens: 2000 },
routing: { prose: "extract" },
}),
});
const response = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [{ role: "user", content: LONG_DOCUMENT }],
});
console.log(response.usage.leanpromptTokensSaved);For SDKs (or SDK wrappers) that don't accept a custom fetch:
import Anthropic from "@anthropic-ai/sdk";
import { wrap } from "leanprompt";
const client = wrap(new Anthropic(), {
mode: "on",
routing: { prose: "extract" },
});
const response = await client.messages.create({
model: "claude-sonnet-4-6",
max_tokens: 1024,
messages: [{ role: "user", content: LONG_DOCUMENT }],
});import { OpenAI } from "leanprompt"; // leanprompt's own minimal client, not the `openai` package
const client = new OpenAI({
apiKey: process.env.OPENAI_API_KEY,
leanpromptConfig: { mode: "on", routing: { prose: "extract" } },
});
const response = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [{ role: "user", content: LONG_DOCUMENT }],
});Useful outside an HTTP request/response cycle (batch jobs, custom transports):
import { Middleware } from "leanprompt";
const mw = new Middleware({ mode: "on", routing: { prose: "extract" } });
const [compressed, stats] = mw.compressMessages(messages);
console.log(stats.inputTokens, stats.outputTokens, stats.method);use leanprompt::{json, Config, Middleware};
let messages = json::parse(r#"[{"role":"user","content":"..."}]"#)?;
let mw = Middleware::new(Config {
mode: "on".into(),
routing: vec![("prose".into(), "extract".into())],
extract_ratio_millis: 400,
..Config::default()
});
let (compressed, stats) = mw.compress_messages(messages.as_arr().unwrap());import leanprompt "github.com/itaides/leanprompt/go"
mw := leanprompt.NewMiddleware(leanprompt.Config{
Mode: "on",
Routing: map[string]string{"prose": "extract"},
ExtractRatioMillis: 400,
})
compressed, stats := mw.CompressMessages(messages) // []map[string]anyEvery SDK also exposes SelfLLM — delegating summarization to a cheap
model (Anthropic / OpenAI / Gemini) over raw HTTP instead of running the
local algorithm. See CONFIG.md
for its options and each package's README for the language-specific API.
LangChain.js: leanprompt/langchain (TypeScript only) wires
leanpromptFetch into ChatOpenAI/ChatAnthropic — see
ts/README.md. Adds no dependency to the core
leanprompt import.
Only PROSE is compressed. Everything the classifier flags as code, error or
structured data — and all tool/image blocks — passes through verbatim by
design. Total savings ≈ prose_token_share × (1 − keep_ratio) plus
dedup/purge wins:
- prose-heavy histories: roughly 40–50% at the default keep-ratio 0.5
- code/tool-heavy agent histories: substantially less — measure on your own traffic before quoting a number
Token counts (the trigger threshold and every leanprompt* telemetry field)
come from a zero-dependency heuristic estimator, not the provider's actual
BPE tokenizer — see ts/src/tokens.ts for the full spec.
It's tuned per-script: space-delimited text (English and other Latin-script
languages) uses a 4-chars/token divisor, and CJK/Hangul/Thai/Lao/Khmer/Myanmar
— scripts with no space-delimited word boundaries, where real tokenizers run
far denser — use a separate ~1.5-chars/token divisor. Both are still
estimates: expect the reported numbers to diverge from your provider's actual
token accounting, more so for scripts and content this estimator wasn't
tuned against.
| Language | Test | Lint / typecheck | Regenerate parity vectors |
|---|---|---|---|
| TypeScript | bun test |
bun x tsc --noEmit |
bun scripts/gen-parity.ts (from ts/) |
| Rust | cargo test |
cargo clippy --all-targets -- -D warnings |
— (asserts against parity/) |
| Go | go test ./... |
go vet ./... && gofmt -l . |
— (asserts against parity/) |
bun bench/run-quality.ts (Extract vs a naive baseline), bun bench/run-cross-language.ts (ts/rust/go agreement dashboard), bun bench/run-workload.ts --file <messages.json> (savings on your own
conversation export). See bench/README.md.
ts/ TypeScript SDK — reference implementation (bun test)
rust/ Rust crate (cargo test)
go/ Go module (go test ./...)
mcp-proxy/ MCP server wrapper compressing tool/task/sampling payloads (Claude Desktop and any MCP client)
parity/ golden vectors generated from ts/ (bun ts/scripts/gen-parity.ts)
bench/ quality, cross-language and real-workload measurement tooling
docs/ parity-spec.md — the normative cross-language spec
CONFIG.md every configuration option, all three languages
ROADMAP.md planned support for OpenAI Codex CLI
CHANGELOG.md notable changes per release
Claude Desktop support shipped as mcp-proxy/ — an
MCP server wrapper that compresses oversized tool results, task payloads, and
sampling requests for any MCP client, verified end-to-end in Claude Desktop.
See ROADMAP.md for what's still planned (OpenAI Codex CLI).
See CHANGELOG.md for notable changes per release.
See SECURITY.md for supported versions and how to report a vulnerability.
MIT. See LICENSE.