TokenPad
Cost

LLM API Cost Calculator

Model your monthly API bill before you commit to a model.

30,000 requests × (1,200 in + 400 out) = 48,000,000 tokens per month. Models without a published cached-input rate ignore the cache slider.

Monthly cost — GPT-5 nano is cheapest, 145× less than Claude Fable 5
ModelPer requestPer monthInput $/1MOutput $/1M
GPT-5 nanoOpenAI$0.000220$6.60$0.0500$0.4000
Ministral 3 8BMistral AI$0.000240$7.20$0.1500$0.1500
Gemini 2.5 Flash-LiteGoogle$0.000280$8.40$0.1000$0.4000
DeepSeek V4 FlashDeepSeek$0.000280$8.40$0.1400$0.2800
GPT-4o miniOpenAI$0.000420$12.60$0.1500$0.6000
Mistral Small 4Mistral AI$0.000420$12.60$0.1500$0.6000
GPT-5.6 LunaOpenAI$0.000720$21.60$0.2000$1.20
GPT-5.4 nanoOpenAI$0.000740$22.20$0.2000$1.25
DeepSeek V4 ProDeepSeek$0.000870$26.10$0.4350$0.8700
GPT-5 miniOpenAI$0.001100$33.00$0.2500$2.00
GPT-4.1 miniOpenAI$0.001120$33.60$0.4000$1.60
GPT-3.5 TurboOpenAI$0.001200$36.00$0.5000$1.50
Mistral Large 3Mistral AI$0.001200$36.00$0.5000$1.50
Devstral 2Mistral AI$0.001280$38.40$0.4000$2.00
Gemini 3.5 Flash-LiteGoogle$0.001360$40.80$0.3000$2.50
Gemini 2.5 FlashGoogle$0.001360$40.80$0.3000$2.50
Grok Build 0.1xAI$0.002000$60.00$1.00$2.00
Grok 4.3xAI$0.002500$75.00$1.25$2.50
GPT-5.4 miniOpenAI$0.002700$81.00$0.7500$4.50
o4-miniOpenAI$0.003080$92.40$1.10$4.40
Claude Haiku 4.5Anthropic$0.003200$96.00$1.00$5.00
Magistral MediumMistral AI$0.004400$132.00$2.00$5.00
Gemini 3.6 FlashGoogle$0.004800$144.00$1.50$7.50
Grok 4.5xAI$0.004800$144.00$2.00$6.00
Mistral Medium 3.5Mistral AI$0.004800$144.00$1.50$7.50
Gemini 3.5 FlashGoogle$0.005400$162.00$1.50$9.00
GPT-5.1OpenAI$0.005500$165.00$1.25$10.00
GPT-5OpenAI$0.005500$165.00$1.25$10.00
GPT-4.1OpenAI$0.005600$168.00$2.00$8.00
o3OpenAI$0.005600$168.00$2.00$8.00
Claude Sonnet 5Anthropic$0.006400$192.00$2.00$10.00
GPT-4oOpenAI$0.007000$210.00$2.50$10.00
GPT-5.6 TerraOpenAI$0.007200$216.00$2.00$12.00
GPT-5.4OpenAI$0.009000$270.00$2.50$15.00
Claude Sonnet 4.6Anthropic$0.009600$288.00$3.00$15.00
Claude Sonnet 4.5Anthropic$0.009600$288.00$3.00$15.00
Claude Opus 5Anthropic$0.0160$480.00$5.00$25.00
Claude Opus 4.8Anthropic$0.0160$480.00$5.00$25.00
Claude Opus 4.6Anthropic$0.0160$480.00$5.00$25.00
GPT-5.6 SolOpenAI$0.0180$540.00$5.00$30.00
GPT-5.5OpenAI$0.0180$540.00$5.00$30.00
Claude Fable 5Anthropic$0.0320$960.00$10.00$50.00

Token counts here are yours to supply. If you are guessing them, measure a real prompt first — the difference between a guessed 800 and an actual 1,400 is the difference between a budget that holds and one that does not.

What this tool tells you

Four numbers determine an LLM API bill: how many requests you make, how big the prompt is, how long the answer runs, and how much of the prompt the provider has already cached. Feed those in and this page prices every model we track side by side, using rates read from each provider’s own documentation on the date shown beside them.

The reason to compare rather than calculate a single model is the spread. Between the cheapest and dearest model in the same table, at identical volume, the monthly figure routinely differs by two orders of magnitude. That gap is a product decision, not a rounding error.

Getting the inputs right

Requests per month

Count API calls, not users. An agent that loops five times per user action makes five requests, and a retry on failure makes six. Teams underestimate this line more than any other.

Input tokens per request

This is the whole request, not the user’s message: system prompt, conversation history, tool definitions, retrieved documents. In a retrieval-augmented setup the retrieved chunks usually dominate everything else. Measure a real one in the token counter rather than guessing — guesses here are wrong in the expensive direction.

Output tokens per request

Note this is what the model actually generates, which is not your max_tokens ceiling. If you have production logs, use the median. If you do not, generate twenty representative answers and count them.

Cached share of input

If your requests share a stable prefix — a long system prompt, a fixed document, a tool schema — providers will charge roughly a tenth of the base input rate for that portion on subsequent calls. A chatbot with a 2,000 token system prompt and 200 tokens of user text is around 90% cacheable, and moving that slider is the single largest saving available on this page.

Reading the comparison

Sort by monthly cost and look at the shape of the table rather than only the top row. Three patterns show up repeatedly:

  • Output-heavy workloads reshuffle the ranking. Summarisation and classification are input-heavy, so cheap-input models win. Generation and code-writing are output-heavy, and the ranking inverts.
  • Caching flattens the field. At 90% cached input, models separate almost entirely on output price, and several expensive-looking options become competitive.
  • The cheapest model is rarely the cheapest solution. A weaker model that needs two attempts, a longer prompt, or a human correction costs more than the table shows. Price the workflow, not the call.

What this does not include

These are standard on-demand rates. Not modelled here: the roughly 50% discount most providers offer for asynchronous batch processing, negotiated enterprise rates, regional or data-residency premiums, per-search charges for server-side tools, and image or audio input priced on a separate scale. Reasoning models add a further wrinkle — their internal thinking tokens are billed as output even though you never see them, so a reasoning model’s real output count can be several times the visible answer. Treat the result as a well-founded upper bound on a straightforward text deployment.

On price freshness

Model prices change without notice, and a stale calculator is worse than none because it is confidently wrong. Every model entry here stores the URL it was read from and the date it was read; both are visible on the model price table. If a figure looks wrong, follow the source link and check — and if the provider has moved, that is a bug worth reporting.

Frequently asked questions

Where do these prices come from?
Each model entry stores the URL it was read from and the date it was read, both shown on the page. They come from the provider’s own pricing documentation, never from a third-party aggregator or from memory.
What does prompt caching change?
Most providers charge far less for input tokens they have already processed and held in cache — often around a tenth of the base input rate. If your requests share a long, stable prefix such as a system prompt or a document, marking the cached share moves that portion to the cheaper rate and can cut the input line of your bill by most of its value.
Why is output so much more expensive than input?
Input is processed in parallel in a single forward pass; output is generated one token at a time, each requiring a full pass over the model. Output typically prices at four to six times input, which is why a verbose model can cost more than a nominally pricier one that answers concisely.
Does this include batch or volume discounts?
No. The figures are standard on-demand rates. Several providers offer roughly 50% off for asynchronous batch processing, and negotiated enterprise rates exist above certain volumes. Treat the result as an upper bound on a straightforward deployment.

More cost tools