Skip to content

About

Zero-dependency prompt compression for LLM applications. One deterministic algorithm, three native SDKs (TypeScript/Bun, Rust, Go) with byte-identical output. No ML model, no downloads.

Topics

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

13 Commits

Folders and files

Repository files navigation

leanprompt

CI zero dependencies no ML model TypeScript Rust Go License: MIT

Zero-dependency prompt compression for LLM applications. One deterministic algorithm, three native SDKs — TypeScript/Bun, Rust and Go — with byte-identical output across all of them.

No ML model, no model download, no tokenizer dependency, no provider SDKs required. The local compressor, Extract, is a weights-free extractive algorithm: sentence segmentation → integer term-rarity scoring (boosted for numbers, entities, identifiers and constraints) → redundancy filtering → greedy selection to a keep-ratio, emitted in original order.

A heuristic classifier gates everything: code, stack traces, JSON blobs and tool-call blocks are never touched; prohibitions ("do not …") always survive; system messages and the most recent turns are never compressed.

How it's different

leanprompt Neural compressors (e.g. LLMLingua-2)
Model download none ~1 GB+
Runtime dependencies zero, in all 3 languages GPU/ONNX runtime, tokenizer, model weights
Cold start instant model load (seconds, first call)
Output deterministic, auditable probabilistic, model-version-dependent
Cross-language parity byte-identical by spec not applicable — single implementation
Code/JSON/tool-call safety classifier-gated, never touched depends on wrapper logic

The SDKs

Language Directory Runtime deps Quickstart
TypeScript (Node ≥ 18 / Bun) ts/ zero npm install leanprompt
Rust rust/ zero ([dependencies] empty) cargo add leanprompt
Go go/ zero (stdlib only) go get github.com/itaides/leanprompt/go

The TypeScript package is the reference implementation of docs/parity-spec.md — a written, normative spec (pinned character classes, integer-only scoring, explicit tiebreaks, canonical JSON). The Rust and Go ports assert byte-equality against the golden vectors in parity/, so identical inputs produce identical compressed output in every language.

Usage

Full option reference (every field, every default, per-language) lives in CONFIG.md. A few common patterns:

TypeScript — keep your official SDK, compress on the wire

import OpenAI from "openai";
import { leanpromptFetch } from "leanprompt";

const client = new OpenAI({
    fetch: leanpromptFetch({
        mode: "on",
        trigger: { thresholdTokens: 2000 },
        routing: { prose: "extract" },
    }),
});

const response = await client.chat.completions.create({
    model: "gpt-4o-mini",
    messages: [{ role: "user", content: LONG_DOCUMENT }],
});
console.log(response.usage.leanpromptTokensSaved);

TypeScript — wrap an existing client instance instead

For SDKs (or SDK wrappers) that don't accept a custom fetch:

import Anthropic from "@anthropic-ai/sdk";
import { wrap } from "leanprompt";

const client = wrap(new Anthropic(), {
    mode: "on",
    routing: { prose: "extract" },
});

const response = await client.messages.create({
    model: "claude-sonnet-4-6",
    max_tokens: 1024,
    messages: [{ role: "user", content: LONG_DOCUMENT }],
});

TypeScript — no official SDK at all (minimal built-in client)

import { OpenAI } from "leanprompt"; // leanprompt's own minimal client, not the `openai` package

const client = new OpenAI({
    apiKey: process.env.OPENAI_API_KEY,
    leanpromptConfig: { mode: "on", routing: { prose: "extract" } },
});
const response = await client.chat.completions.create({
    model: "gpt-4o-mini",
    messages: [{ role: "user", content: LONG_DOCUMENT }],
});

TypeScript — call the compression pipeline directly

Useful outside an HTTP request/response cycle (batch jobs, custom transports):

import { Middleware } from "leanprompt";

const mw = new Middleware({ mode: "on", routing: { prose: "extract" } });
const [compressed, stats] = mw.compressMessages(messages);
console.log(stats.inputTokens, stats.outputTokens, stats.method);

Rust

use leanprompt::{json, Config, Middleware};

let messages = json::parse(r#"[{"role":"user","content":"..."}]"#)?;
let mw = Middleware::new(Config {
    mode: "on".into(),
    routing: vec![("prose".into(), "extract".into())],
    extract_ratio_millis: 400,
    ..Config::default()
});
let (compressed, stats) = mw.compress_messages(messages.as_arr().unwrap());

Go

import leanprompt "github.com/itaides/leanprompt/go"

mw := leanprompt.NewMiddleware(leanprompt.Config{
    Mode:               "on",
    Routing:            map[string]string{"prose": "extract"},
    ExtractRatioMillis: 400,
})
compressed, stats := mw.CompressMessages(messages) // []map[string]any

Every SDK also exposes SelfLLM — delegating summarization to a cheap model (Anthropic / OpenAI / Gemini) over raw HTTP instead of running the local algorithm. See CONFIG.md for its options and each package's README for the language-specific API.

LangChain.js: leanprompt/langchain (TypeScript only) wires leanpromptFetch into ChatOpenAI/ChatAnthropic — see ts/README.md. Adds no dependency to the core leanprompt import.

What savings to expect — honest math

Only PROSE is compressed. Everything the classifier flags as code, error or structured data — and all tool/image blocks — passes through verbatim by design. Total savings ≈ prose_token_share × (1 − keep_ratio) plus dedup/purge wins:

  • prose-heavy histories: roughly 40–50% at the default keep-ratio 0.5
  • code/tool-heavy agent histories: substantially less — measure on your own traffic before quoting a number

Token counts (the trigger threshold and every leanprompt* telemetry field) come from a zero-dependency heuristic estimator, not the provider's actual BPE tokenizer — see ts/src/tokens.ts for the full spec. It's tuned per-script: space-delimited text (English and other Latin-script languages) uses a 4-chars/token divisor, and CJK/Hangul/Thai/Lao/Khmer/Myanmar — scripts with no space-delimited word boundaries, where real tokenizers run far denser — use a separate ~1.5-chars/token divisor. Both are still estimates: expect the reported numbers to diverge from your provider's actual token accounting, more so for scripts and content this estimator wasn't tuned against.

Commands / test reference

Language Test Lint / typecheck Regenerate parity vectors
TypeScript bun test bun x tsc --noEmit bun scripts/gen-parity.ts (from ts/)
Rust cargo test cargo clippy --all-targets -- -D warnings — (asserts against parity/)
Go go test ./... go vet ./... && gofmt -l . — (asserts against parity/)

Benchmarks

bun bench/run-quality.ts (Extract vs a naive baseline), bun bench/run-cross-language.ts (ts/rust/go agreement dashboard), bun bench/run-workload.ts --file <messages.json> (savings on your own conversation export). See bench/README.md.

Repository layout

ts/         TypeScript SDK — reference implementation (bun test)
rust/       Rust crate (cargo test)
go/         Go module (go test ./...)
mcp-proxy/  MCP server wrapper compressing tool/task/sampling payloads (Claude Desktop and any MCP client)
parity/     golden vectors generated from ts/ (bun ts/scripts/gen-parity.ts)
bench/      quality, cross-language and real-workload measurement tooling
docs/       parity-spec.md — the normative cross-language spec
CONFIG.md   every configuration option, all three languages
ROADMAP.md  planned support for OpenAI Codex CLI
CHANGELOG.md notable changes per release

Roadmap

Claude Desktop support shipped as mcp-proxy/ — an MCP server wrapper that compresses oversized tool results, task payloads, and sampling requests for any MCP client, verified end-to-end in Claude Desktop. See ROADMAP.md for what's still planned (OpenAI Codex CLI).

Changelog

See CHANGELOG.md for notable changes per release.

Security

See SECURITY.md for supported versions and how to report a vulnerability.

License

MIT. See LICENSE.

About

Zero-dependency prompt compression for LLM applications. One deterministic algorithm, three native SDKs (TypeScript/Bun, Rust, Go) with byte-identical output. No ML model, no downloads.

Topics

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages