A drop-in proxy that applies 28 optimisation techniques to every LLM call — sending fewer tokens for the same result, with a single line of code, no change to how your developers work, and quality preserved.
Already run an LLM gateway? Keep it — TokenLean sits in front of, or inside, the one you already have.
TokenLean speaks the OpenAI, Anthropic (Claude Code) and Gemini APIs natively. Swap your base URL — your code, prompts, SDK and even tool-calls stay exactly the same. TokenLean optimises every request before forwarding it to any of 10 first-class providers.
client = OpenAI( api_key="sk-openai-..." )
client = OpenAI( api_key=os.environ["TOKENLEAN_KEY"], base_url=os.environ["TOKENLEAN_ENDPOINT"] + "/v1" ) # that's it — fewer tokens, quality preserved
/v1/chat/completions
/v1/messages
:generateContent
TokenLean organises 28 techniques into 8 groups, layering a full pipeline of token-reduction strategies onto every call — each one measured, each one optional — and wraps every request in non-bypassable trust-&-safety and provider-failover layers. Because optimisation without measurement is guesswork.
LLMLingua-2 strips redundant tokens from prompts while preserving meaning.
20–50% savedExact, semantic and response caches return answers without ever hitting the model.
30–80% savedA confidence-checked cascade sends easy calls to cheaper models, hard ones to bigger ones.
40–70% savedReorders messages to hit provider prefix-caches — zero quality risk, large savings.
up to ~84%AST-aware compression of code, JSON and logs, plus collapsing of near-duplicate turns — structural savings that never touch your wording.
5–40% savedCompacts older history before a long conversation overflows the context window — prune, compress, then a cached summary. Recent turns are never touched.
20–60% on long threadsCompact encoding for structured payloads, plus an opt-in async lane claiming the provider's own 50% batch discount — for jobs that can wait.
25–60% savedClassifies how hard each request is and sets reasoning effort to match — a simple question never pays for extended thinking.
10–30% savedEvery response carries line-item token & dollar savings — fully auditable ROI.
Total visibilityInjection and PII scanned on the prompt — then again on whatever your RAG and memory inject. Non-bypassable, before any tokens are spent.
Non-bypassable · ×2Circuit breaker + retry + fallback provider keep requests serving when an upstream degrades — before any bytes reach the client.
Always-onRetrieval hit-rate, grounding coverage and schema-validation tracked apart from savings — so "cheaper" never quietly becomes "worse".
Measured signalsTokenLean isn't a black box. Every optimisation, every dollar and every guardrail event is visible to your team in a multi-tenant portal — from real-time savings to per-user chargeback and security you can audit.
🔒 Enterprise exclusive — this console ships with the managed Enterprise plan, not the free self-hosted version.
Billed requests, tokens saved, dollars saved and average saving % — live, per tenant. Savings are disclosed estimates; you're only ever billed on request count.
Month-to-date spend, a month-end projection, cost per request, savings ROI and cache hit-rate — with a daily tokens-saved trend. Your LLM unit economics, at a glance.
Break requests, tokens and cost down by user, model, provider or day. Hand finance a clean showback for every team and every bot — no spreadsheets required.
Toggle and dial each optimisation — compression aggressiveness, cache thresholds, routing confidence — and see changes apply within a minute. Savings vs. quality, on your terms.
Bring your own provider keys (stored encrypted), pick per-tier models for the routing cascade, and set a hard monthly spend ceiling that stops runaway or leaked-key costs cold.
A per-day trust-&-safety trend plus a live incident log — every PII redaction (G29) and blocked injection (G30), by type and action. Metadata only; the matched content is never stored.
Six months of invoice history, billed purely on request count — with the value you saved shown right alongside, and never added to the bill.
{
"baseline_tokens": 450,
"final_tokens_sent": 220,
"total_abs_saving": 230,
"total_pct_saving": 51.1,
"cache_hit": false,
"routed_model": "gpt-4o-mini"
}
Every request fires twice: once straight to the provider, once through TokenLean. We compare the provider's own billed usage — never our self-report — over recognised public datasets, on your key. Costs under a dollar.
| Workload | Public A/B | Reproduce with |
|---|---|---|
| cache · warm-repeat | ~90% | --workload cache |
| agentic · tool loop | ~20% | --workload agentic |
| prose · first-ask | ~8% | --workload standard |
| reasoning | ~0% | --workload standard |
# all workloads + the blend · OpenAI · <$1 ./examples/benchmark/run.sh --ab --workload full # 10-provider sweep ./examples/benchmark/run.sh --ab --providers all
ops profile of verbose DevOps payloads is production-shaped, not a recognised public benchmark, and is labelled as such in the repo. Agentic reads ~20% here against 46% internally because the largest lever structurally cannot fire in a live single-loop A/B — so we publish the lower number and show you why.
Self-host the open-source core at no cost, or let us run it for you — fully managed and billed on request count, never on tokens, so every token you save stays yours.
A short walkthrough of the savings and how the proxy works, plus the deck that lays it out.
We're opening TokenLean to a select group of early-access customers ahead of general availability. Come in now and you're not just a user — you're a founding partner shaping the product, on terms that reward you for it.
We're keeping the first cohort deliberately small so every partner gets real attention. In exchange for working with us to make TokenLean better, early-access customers lock in pricing — and access — that won't be offered again: