Mini tool · cost control

LLM Cost Calculator

Model, traffic and cache hit rate in - cost per request, per month, per year out. See what prompt caching saves and whether a smaller model pays off.

Pricing last updated: September 2026. Always verify provider pricing before financial decisions.
Inputs
Preset prices last checked September 2026 · standard API rates / 1M tokens
0%
production · cost projection
Cost per request -
Daily≈ today -
Monthly× 30 days -
Annual× 12 months -
Per active user/ month -
Per 1,000 users/ month -
Saved by cachingper month -
-
Pick a model to compare savings.
all models · monthly at these settings
Model Vendor In / Out · /1M Cached · /1M Monthly

Preset prices are standard API rates per 1M tokens (input / output), checked September 2026 against provider pricing pages - they change often, so verify before you rely on a number. Each preset carries its vendor's own cached-input rate, and they differ: 10% of the input price for most Claude and OpenAI models, 2.5% on Claude Fable 5.1, 5% on Claude Opus 5.5, 25% on Grok 4.7. The cached rate is editable; for Custom it starts at 10% of your input price, which is an assumption - replace it. Cache-write premiums and Google's cache storage fee are not included, so heavy caching reads slightly optimistic. Monthly figures use 30 days and annual figures 12 of those months (360 days), so a year is not 365 days here. Edit any price field to model batch rates, regional premiums or your own negotiated pricing. - LLMOps.si