TokenPad

8 comparisons

Model comparisons

Same-tier, cross-provider pairs — the decisions people are actually weighing. Each page prices both models across four workload shapes at a million requests a month, because a single sample request settles nothing when two providers price input and output on different ratios.

Frontier

The most capable models available, and the most expensive place to make a routing mistake at volume.

Workhorse

Where most production traffic belongs. Capable enough for open-ended work, priced to survive being used on all of it.

Budget

High-volume classification, extraction and routing. A price difference here compounds faster than anywhere else because the request count is highest.

Why these pairs and not every combination

This layer used to be a cross product of eight headline models, which produced twenty-two pages that all said roughly the same thing. Most of those pairs were not decisions anyone was making — nobody chooses between a frontier model and a budget one, because they are not solving the same problem.

So pairs are enumerated by hand now, and each has to earn its page: someone types it, both models sit in the same tier, they come from different providers, and something substantive differs between them. That last condition is checked at build time against the price data, so a pricing change that collapses a difference removes the page rather than leaving a comparison with nothing to compare.

What each page contains

A side-by-side table of input, cached input and output rates, context window and token-counting accuracy, each figure carrying the provider URL it came from and the date it was read. Then the part that is not on any pricing page: both models costed across four workload shapes at a million requests a month, so you can see whether the ranking holds when the work is output-heavy rather than input-heavy.

Where both providers publish a cached input rate, there is a second calculation with a realistic cache hit — because on repetitive traffic that rate matters more to the bill than the headline input price, and it occasionally reverses the answer entirely.

What they deliberately leave out

Which model is better. Published benchmark scores are run on public sets the models may have trained on, reported by parties with an interest in the result, and measured on data that is not yours. Repeating them here would add confident-sounding numbers with nothing behind them.

Price, context and tokenizer availability are facts with sources attached. Quality is a measurement you take yourself — thirty cases from your own traffic will tell you more than any leaderboard. The model migration checklist covers what else changes when you switch: tool calling formats, stop reasons, refusal behaviour, and the prompt re-tuning that usually follows.

If your pair is not here

The full pricing table carries all 42 models across 6 providers, with a filter for budget, context window and whether tokens can be counted exactly. For a specific workload, the cost calculator takes your own volume and prompt size rather than a representative shape. Prices last reviewed August 5, 2026.