TokenPad

Anthropic

Claude Sonnet 5 pricing

$2.00 per million input tokens, $10.00 per million output. 1M token context window. Read from Anthropic’s own documentation on August 3, 2026.

Input
$2.00
per 1M tokens
Cached input
$0.2000
10% of base
Output
$10.00
5.0× input
Context window
1M
128K max output
Token counting
Estimate
o200k_base
Price verified
2026-08-03
2 days ago

Introductory pricing through 31 Aug 2026. From 1 Sep 2026: $3 input / $15 output per 1M tokens.

Source: Anthropic pricing documentation. Prices change without notice — verify before committing spend.

What Claude Sonnet 5 costs on real work

Four workload shapes at 100,000 requests a month. The point of showing four is that the ranking between models changes depending on which one describes you.

Claude Sonnet 5 cost by workload shape
WorkloadInOutPer requestPer month
ClassificationShort input, one-word answer. Input-dominated.50050$0.001500$150.00
Chat turnA system prompt plus a few turns of history.1,500300$0.006000$600.00
Document summaryA long document in, a paragraph out.20,000800$0.0480$4,800.00
Code generationOutput-heavy — where output pricing dominates.2,0001,500$0.0190$1,900.00

Put your own numbers in the cost calculator, or measure a real prompt first in the token counter. If your requests share a stable prefix, the cached rate applies to most of your input — check the structure in the cache checker.

Counting tokens for Claude Sonnet 5

Anthropic does not publish a tokenizer that runs in a browser, so any pre-flight count for Claude Sonnet 5 is an estimate rather than a measurement.

Anthropic does not publish a client-side tokenizer. Counted with o200k_base and scaled ~1.18x, the commonly reported gap between tiktoken and Claude's pre-4.7 tokenizer.

Treat it as accurate to within roughly ten to twenty percent. That is fine for budgeting and wrong for sizing a prompt right at a context window boundary — where precision matters, use Anthropic’s own token counting endpoint from your backend. The methodology page sets out every scaling factor used here.

Other Anthropic models

The tier question: is a cheaper model in the same family enough for your task?

Other Anthropic models compared with Claude Sonnet 5
ModelInputOutputContextChat turn
Claude Sonnet 5 — this page$2.00$10.001M$0.006000
Claude Fable 5$10.00$50.001M$0.0300
Claude Opus 5$5.00$25.001M$0.0150
Claude Opus 4.8$5.00$25.001M$0.0150
Claude Opus 4.6$5.00$25.001M$0.0150
Claude Sonnet 4.6$3.00$15.001M$0.009000
Claude Sonnet 4.5$3.00$15.00200K$0.009000

Alternatives from other providers

Models priced nearest to Claude Sonnet 5, not the cheapest on the market — those are the ones actually worth evaluating against it.

Side-by-side comparisons: Claude Sonnet 5 vs GPT-5.4 · Claude Sonnet 5 vs Gemini 3.5 Flash

Frequently asked questions

How much does Claude Sonnet 5 cost?
$2.00 per million input tokens and $10.00 per million output tokens, with cached input at $0.2000 per million. On a typical chat turn of 1,500 input and 300 output tokens that is $0.006000 per request, or $600.00 per month at 100,000 requests. Read from Anthropic's own documentation on August 3, 2026.
Can I count Claude Sonnet 5 tokens exactly?
No. Anthropic does not publish a tokenizer that runs in a browser, so any pre-flight count for Claude Sonnet 5 is an estimate. Anthropic does not publish a client-side tokenizer. Counted with o200k_base and scaled ~1.18x, the commonly reported gap between tiktoken and Claude's pre-4.7 tokenizer. Treat it as accurate to within roughly ten to twenty percent and never as the basis for sizing a prompt right at a context window boundary.
What is the context window of Claude Sonnet 5?
1,000,000 tokens, with a maximum of 128,000 output tokens in a single response. That budget covers everything in the request — system prompt, conversation history, tool definitions, documents — plus the response itself, not just your input.
Why is output more expensive than input on Claude Sonnet 5?
Output costs 5.0 times input here. Input is processed in a single parallel pass, while output is generated one token at a time with a full pass over the model for each. That is why a model that answers concisely can be cheaper in production than one with a lower headline rate.