LLM Cost Control: Tokens, caching and model routing
Token spend compounds quietly until a finance review forces a panic. A production guide to unit economics, a worked cost model, token audits, caching, routing, batching and budget enforcement - before the bill, not after.
LLM cost scales with traffic and context length, per token, and compounds quietly until a finance review forces a scramble. Cost control is keeping spend predictable and proportionate to value - before the invoice, not after.
The two questions
Prototype: “Can we afford the model?” Production: “What’s cost per user, per feature and per month - and what happens to the bill at 10× traffic?”
A demo runs a handful of calls. Production runs millions, and a token count that looked trivial per request becomes a five-figure line item.
What breaks in production
- The surprise bill. Spend triples after a launch and no one saw it coming.
- No attribution. You can’t say which feature, team or customer drove the cost, so you can’t fix it.
- Context bloat. Every request stuffs the full knowledge base into the prompt, paying for tokens the model doesn’t need.
- One model for everything. A frontier model answers trivial queries a model a tenth the price would handle fine.
- No alerts. You learn about a runaway from finance, not a notification.
Know your unit economics
Cost per request is input tokens times the input price plus output tokens times the output price - but “input” is really three different things, priced differently:
- a cacheable prefix that repeats on every request (system prompt, tool definitions),
- dynamic input that changes per request (retrieved context, conversation history, the question),
- and output, which is usually several times more expensive per token than input.
Prompt caching changes the price of the first one only. On Anthropic’s list prices, for example, writing a prefix to the 5-minute cache costs 1.25× the base input price and reading it back costs 0.1× - so caching pays for itself after a single cache read. Cache-read rates also differ by model (5% of input on Claude Opus 5.5, 2.5% on Claude Fable 5.1), which is why the Cost Calculator uses each model’s own rate instead of a flat discount.
A worked example: one assistant, four changes
A support assistant handles 20,000 requests a day. Each request carries a 2,500-token prefix (system prompt and tool definitions), 3,700 tokens of dynamic input (five retrieved chunks, recent history, the question) and about 450 tokens of output. This models what each lever does to the monthly bill, using list prices checked on September 27, 2026:
# USD per 1M tokens: input, 5-minute cache write, cache read, output
# (Anthropic list prices, checked 2026-09-27 - re-check before you rely on them)
PRICES = {
"sonnet-5": {"in": 2.00, "write": 2.50, "read": 0.20, "out": 10.00},
"haiku-4.5": {"in": 1.00, "write": 1.25, "read": 0.10, "out": 5.00},
}
REQUESTS = 20_000 * 30 # a month of a support assistant at 20k requests/day
def request_cost(p, prefix, dynamic, output, hit_rate=0.0):
"""prefix: cacheable tokens (system prompt, tool definitions);
dynamic: tokens that change per request (retrieved context, history, question)."""
prefix_cost = prefix * (hit_rate * p["read"] + (1 - hit_rate) * (p["write"] if hit_rate else p["in"]))
return (prefix_cost + dynamic * p["in"] + output * p["out"]) / 1e6
def monthly(mix, **shape):
return REQUESTS * sum(share * request_cost(PRICES[m], **shape) for m, share in mix.items())
scenarios = [
("A baseline: one model, no caching",
{"sonnet-5": 1.0}, dict(prefix=2500, dynamic=3700, output=450)),
("B + cache the 2,500-token prefix (90% hits)",
{"sonnet-5": 1.0}, dict(prefix=2500, dynamic=3700, output=450, hit_rate=0.9)),
("C + route 60% of traffic to the small model",
{"sonnet-5": 0.4, "haiku-4.5": 0.6}, dict(prefix=2500, dynamic=3700, output=450, hit_rate=0.9)),
("D + top-k 5 -> 3 chunks, tighter output cap",
{"sonnet-5": 0.4, "haiku-4.5": 0.6}, dict(prefix=2500, dynamic=2500, output=350, hit_rate=0.9)),
]
base = None
for name, mix, shape in scenarios:
cost = monthly(mix, **shape)
base = base or cost
print(f"{name:<48} ${cost:>9,.0f}/month ({cost / base:.0%} of A)")
# A separate nightly job: summarise 20,000 tickets (2,000 tokens in, 150 out) on the small model
nightly = 20_000 * 30 * request_cost(PRICES["haiku-4.5"], prefix=0, dynamic=2000, output=150)
print(f"\nnightly summaries, synchronous ${nightly:>7,.0f}/month")
print(f"nightly summaries, batch API ${nightly * 0.5:>7,.0f}/month")
Its actual output:
A baseline: one model, no caching $ 10,140/month (100% of A)
B + cache the 2,500-token prefix (90% hits) $ 7,785/month (77% of A)
C + route 60% of traffic to the small model $ 5,450/month (54% of A)
D + top-k 5 -> 3 chunks, tighter output cap $ 4,022/month (40% of A)
nightly summaries, synchronous $ 1,650/month
nightly summaries, batch API $ 825/month
Three things to take from it:
- No single lever does it. Caching saves about a quarter, routing another quarter, trimming the rest. The combination cuts the bill by 60%.
- Each saving is only real if quality holds. Routing 60% of traffic to a smaller model and dropping from five retrieved chunks to three are both quality decisions. Run both through your eval set - and check retrieval recall for the second - before you bank the saving.
- The 90% hit rate is an assumption. It holds for a busy prefix that is requested well within the cache lifetime. A feature that sees a request every ten minutes will miss a 5-minute cache almost every time; measure the hit rate from your provider’s usage fields rather than guessing it.
Swap in your own shape and prices - or use the Cost Calculator, which does the same arithmetic per model.
Find where the tokens go
Before you pull any lever, measure. Log token counts per component, not just per request, and the fix usually becomes obvious:
| Component | Common bloat | Fix |
|---|---|---|
| System prompt + tools | Definitions for tools this request will never use | Send only the relevant tools; keep them in the cached prefix |
| Retrieved context | High top-k, whole documents instead of chunks | Lower k after checking recall; chunk and rerank |
| Conversation history | The full transcript resent on every turn | Keep the last few turns, summarise the rest |
| Tool results | Raw API payloads passed straight back to the model | Return only the fields the model needs |
| Output | No cap, verbose default style | Set max output tokens; ask for the format you need |
Tools cost more than their visible definitions. Anthropic, for instance, documents a tool-use system prompt that it adds automatically whenever tools are present - 354 tokens on Claude Sonnet 5 - on top of your tool names, descriptions and schemas. On a short request that overhead can be a large share of the input.
The levers, in more detail
- Caching. Cache what repeats and pay a fraction for cached reads. Caching matches on the prefix, so order the prompt from most stable to least: system prompt, tool definitions, long reference documents, then the per-request part. A single timestamp or user name at the top of the system prompt silently breaks every cache hit.
- Routing. Two patterns work in practice. Rules route on cheap signals - input length, detected intent, customer tier. A cascade sends everything to the small model first and escalates when its output fails a check (schema validation, a low grounding score, an explicit “I’m not sure”). Cascades cost a second call on escalation, so they pay off when most traffic stays small.
- Right-sizing. Cap max output tokens and context length per request - the cheapest token is the one you don’t send.
- Batching. Anything that doesn’t need an answer in seconds - nightly classification, eval runs, backfills, bulk summaries - can go through a batch API. OpenAI’s Batch API is 50% cheaper than its synchronous endpoints and completes within 24 hours; Anthropic lists batch requests at 50% off as well, and says its batch and caching discounts stack.
A minimal router:
# model ids live in one reviewed config, never at the call site
MODELS = {"frontier": "...", "small": "..."}
def choose_model(query, context_len):
if context_len > 8000 or looks_complex(query):
return MODELS["frontier"] # for the hard cases
return MODELS["small"] # cheaper, for the bulk of traffic
Any routing change affects quality, so gate it on your eval set - and keep measuring quality on both paths after launch, because the traffic mix drifts.
Attribute, alert, enforce
Tag every request with the feature, team and customer it serves, so spend can be grouped any way finance asks. Then use three kinds of rule, from soft to hard (a sketch - the helpers come from your metrics store):
def cost_guard(feature):
mtd, budget = month_to_date_spend(feature), monthly_budget(feature)
if mtd >= budget:
return "enforce" # gateway downgrades to the small model, or rejects
if mtd >= 0.8 * budget:
return "warn" # notify the owning team
if cost_per_request(feature, days=1) > 1.3 * median_cost_per_request(feature, days=7):
return "drift" # a prompt or retrieval change made every request dearer
return "ok"
The drift rule catches what budgets miss: a prompt edit that adds 2,000 tokens to every request won’t breach a monthly budget for weeks, but it shows up in cost per request the next day. Per-user caps belong here too - uncapped usage is a security problem as much as a cost one. Gateways such as LiteLLM and Portkey (see cost tracking) can enforce budgets per key before a request reaches the provider.
The token-usage signals from observability are what make all of this possible - you can only attribute what you measured.
Minimal vs mature
| Aspect | Minimal | Production-grade |
|---|---|---|
| Visibility | Total spend | Per request, per feature, per component |
| Caching | None | Stable prefix first, hit rate measured |
| Models | One for all | Routed or cascaded, quality tracked per path |
| Async work | Synchronous calls | Batch API for anything non-urgent |
| Limits | None | Capped output, context and per-user usage |
| Budget | Reviewed quarterly | Warn at 80%, enforce at 100%, alert on drift |
Where this lives in a real system
Routing, caching and per-provider cost belong in a control point - see the multi-provider gateway reference architecture. Model the trade-offs in the Cost Calculator, pick tooling from cost tracking, and clear the cost items in the Production Checklist.