← All articles
Cost control

LLM Cost Control: Tokens, caching and model routing

Token spend compounds quietly until a finance review forces a panic. A production guide to unit economics, a worked cost model, token audits, caching, routing, batching and budget enforcement - before the bill, not after.

May 29, 2026 · updated September 27, 2026 · 7 min read · cost · caching · routing

LLM cost scales with traffic and context length, per token, and compounds quietly until a finance review forces a scramble. Cost control is keeping spend predictable and proportionate to value - before the invoice, not after.

The two questions

Prototype: “Can we afford the model?” Production: “What’s cost per user, per feature and per month - and what happens to the bill at 10× traffic?”

A demo runs a handful of calls. Production runs millions, and a token count that looked trivial per request becomes a five-figure line item.

What breaks in production

  • The surprise bill. Spend triples after a launch and no one saw it coming.
  • No attribution. You can’t say which feature, team or customer drove the cost, so you can’t fix it.
  • Context bloat. Every request stuffs the full knowledge base into the prompt, paying for tokens the model doesn’t need.
  • One model for everything. A frontier model answers trivial queries a model a tenth the price would handle fine.
  • No alerts. You learn about a runaway from finance, not a notification.

Know your unit economics

Cost per request is input tokens times the input price plus output tokens times the output price - but “input” is really three different things, priced differently:

  • a cacheable prefix that repeats on every request (system prompt, tool definitions),
  • dynamic input that changes per request (retrieved context, conversation history, the question),
  • and output, which is usually several times more expensive per token than input.

Prompt caching changes the price of the first one only. On Anthropic’s list prices, for example, writing a prefix to the 5-minute cache costs 1.25× the base input price and reading it back costs 0.1× - so caching pays for itself after a single cache read. Cache-read rates also differ by model (5% of input on Claude Opus 5.5, 2.5% on Claude Fable 5.1), which is why the Cost Calculator uses each model’s own rate instead of a flat discount.

A worked example: one assistant, four changes

A support assistant handles 20,000 requests a day. Each request carries a 2,500-token prefix (system prompt and tool definitions), 3,700 tokens of dynamic input (five retrieved chunks, recent history, the question) and about 450 tokens of output. This models what each lever does to the monthly bill, using list prices checked on September 27, 2026:

# USD per 1M tokens: input, 5-minute cache write, cache read, output
# (Anthropic list prices, checked 2026-09-27 - re-check before you rely on them)
PRICES = {
    "sonnet-5":  {"in": 2.00, "write": 2.50, "read": 0.20, "out": 10.00},
    "haiku-4.5": {"in": 1.00, "write": 1.25, "read": 0.10, "out": 5.00},
}
REQUESTS = 20_000 * 30  # a month of a support assistant at 20k requests/day

def request_cost(p, prefix, dynamic, output, hit_rate=0.0):
    """prefix: cacheable tokens (system prompt, tool definitions);
    dynamic: tokens that change per request (retrieved context, history, question)."""
    prefix_cost = prefix * (hit_rate * p["read"] + (1 - hit_rate) * (p["write"] if hit_rate else p["in"]))
    return (prefix_cost + dynamic * p["in"] + output * p["out"]) / 1e6

def monthly(mix, **shape):
    return REQUESTS * sum(share * request_cost(PRICES[m], **shape) for m, share in mix.items())

scenarios = [
    ("A  baseline: one model, no caching",
     {"sonnet-5": 1.0}, dict(prefix=2500, dynamic=3700, output=450)),
    ("B  + cache the 2,500-token prefix (90% hits)",
     {"sonnet-5": 1.0}, dict(prefix=2500, dynamic=3700, output=450, hit_rate=0.9)),
    ("C  + route 60% of traffic to the small model",
     {"sonnet-5": 0.4, "haiku-4.5": 0.6}, dict(prefix=2500, dynamic=3700, output=450, hit_rate=0.9)),
    ("D  + top-k 5 -> 3 chunks, tighter output cap",
     {"sonnet-5": 0.4, "haiku-4.5": 0.6}, dict(prefix=2500, dynamic=2500, output=350, hit_rate=0.9)),
]
base = None
for name, mix, shape in scenarios:
    cost = monthly(mix, **shape)
    base = base or cost
    print(f"{name:<48} ${cost:>9,.0f}/month  ({cost / base:.0%} of A)")

# A separate nightly job: summarise 20,000 tickets (2,000 tokens in, 150 out) on the small model
nightly = 20_000 * 30 * request_cost(PRICES["haiku-4.5"], prefix=0, dynamic=2000, output=150)
print(f"\nnightly summaries, synchronous  ${nightly:>7,.0f}/month")
print(f"nightly summaries, batch API     ${nightly * 0.5:>7,.0f}/month")

Its actual output:

A  baseline: one model, no caching               $   10,140/month  (100% of A)
B  + cache the 2,500-token prefix (90% hits)     $    7,785/month  (77% of A)
C  + route 60% of traffic to the small model     $    5,450/month  (54% of A)
D  + top-k 5 -> 3 chunks, tighter output cap     $    4,022/month  (40% of A)

nightly summaries, synchronous  $  1,650/month
nightly summaries, batch API     $    825/month

Three things to take from it:

  • No single lever does it. Caching saves about a quarter, routing another quarter, trimming the rest. The combination cuts the bill by 60%.
  • Each saving is only real if quality holds. Routing 60% of traffic to a smaller model and dropping from five retrieved chunks to three are both quality decisions. Run both through your eval set - and check retrieval recall for the second - before you bank the saving.
  • The 90% hit rate is an assumption. It holds for a busy prefix that is requested well within the cache lifetime. A feature that sees a request every ten minutes will miss a 5-minute cache almost every time; measure the hit rate from your provider’s usage fields rather than guessing it.

Swap in your own shape and prices - or use the Cost Calculator, which does the same arithmetic per model.

Find where the tokens go

Before you pull any lever, measure. Log token counts per component, not just per request, and the fix usually becomes obvious:

ComponentCommon bloatFix
System prompt + toolsDefinitions for tools this request will never useSend only the relevant tools; keep them in the cached prefix
Retrieved contextHigh top-k, whole documents instead of chunksLower k after checking recall; chunk and rerank
Conversation historyThe full transcript resent on every turnKeep the last few turns, summarise the rest
Tool resultsRaw API payloads passed straight back to the modelReturn only the fields the model needs
OutputNo cap, verbose default styleSet max output tokens; ask for the format you need

Tools cost more than their visible definitions. Anthropic, for instance, documents a tool-use system prompt that it adds automatically whenever tools are present - 354 tokens on Claude Sonnet 5 - on top of your tool names, descriptions and schemas. On a short request that overhead can be a large share of the input.

The levers, in more detail

  • Caching. Cache what repeats and pay a fraction for cached reads. Caching matches on the prefix, so order the prompt from most stable to least: system prompt, tool definitions, long reference documents, then the per-request part. A single timestamp or user name at the top of the system prompt silently breaks every cache hit.
  • Routing. Two patterns work in practice. Rules route on cheap signals - input length, detected intent, customer tier. A cascade sends everything to the small model first and escalates when its output fails a check (schema validation, a low grounding score, an explicit “I’m not sure”). Cascades cost a second call on escalation, so they pay off when most traffic stays small.
  • Right-sizing. Cap max output tokens and context length per request - the cheapest token is the one you don’t send.
  • Batching. Anything that doesn’t need an answer in seconds - nightly classification, eval runs, backfills, bulk summaries - can go through a batch API. OpenAI’s Batch API is 50% cheaper than its synchronous endpoints and completes within 24 hours; Anthropic lists batch requests at 50% off as well, and says its batch and caching discounts stack.

A minimal router:

# model ids live in one reviewed config, never at the call site
MODELS = {"frontier": "...", "small": "..."}

def choose_model(query, context_len):
    if context_len > 8000 or looks_complex(query):
        return MODELS["frontier"]  # for the hard cases
    return MODELS["small"]         # cheaper, for the bulk of traffic

Any routing change affects quality, so gate it on your eval set - and keep measuring quality on both paths after launch, because the traffic mix drifts.

Attribute, alert, enforce

Tag every request with the feature, team and customer it serves, so spend can be grouped any way finance asks. Then use three kinds of rule, from soft to hard (a sketch - the helpers come from your metrics store):

def cost_guard(feature):
    mtd, budget = month_to_date_spend(feature), monthly_budget(feature)
    if mtd >= budget:
        return "enforce"   # gateway downgrades to the small model, or rejects
    if mtd >= 0.8 * budget:
        return "warn"      # notify the owning team
    if cost_per_request(feature, days=1) > 1.3 * median_cost_per_request(feature, days=7):
        return "drift"     # a prompt or retrieval change made every request dearer
    return "ok"

The drift rule catches what budgets miss: a prompt edit that adds 2,000 tokens to every request won’t breach a monthly budget for weeks, but it shows up in cost per request the next day. Per-user caps belong here too - uncapped usage is a security problem as much as a cost one. Gateways such as LiteLLM and Portkey (see cost tracking) can enforce budgets per key before a request reaches the provider.

The token-usage signals from observability are what make all of this possible - you can only attribute what you measured.

Minimal vs mature

AspectMinimalProduction-grade
VisibilityTotal spendPer request, per feature, per component
CachingNoneStable prefix first, hit rate measured
ModelsOne for allRouted or cascaded, quality tracked per path
Async workSynchronous callsBatch API for anything non-urgent
LimitsNoneCapped output, context and per-user usage
BudgetReviewed quarterlyWarn at 80%, enforce at 100%, alert on drift

Where this lives in a real system

Routing, caching and per-provider cost belong in a control point - see the multi-provider gateway reference architecture. Model the trade-offs in the Cost Calculator, pick tooling from cost tracking, and clear the cost items in the Production Checklist.

Get the Production Checklist → Explore the Stack →