TokenPad
Cost

Model Routing Savings Calculator

Send the easy 70% to a cheap model and see what it saves.

Settings
Routing SavingsExact
0Token change
Token change0no change
Input tokens0what you pasted
Output tokens0what you would send
Token cost of this result
Output tokens0
As input$0.00
× 100K requests$0.00

Everything on this page runs in your browser. Nothing you paste is transmitted, because there is no server here to transmit it to.

Result
 

Escalation is the cost people forget

An escalated request is paid twice: once on the cheap model that could not handle it, once on the flagship that could. That is the honest cost of routing and it is included here.

It is still usually a large net win, because the price gap between tiers is commonly twenty to one while escalation rates are in single digits.

The break-even rate

Above a certain escalation rate, routing stops paying and you should send everything to the flagship. The calculator reports that threshold for your prices.

If your measured rate is close to it, the fix is usually a better router — a confidence signal, or a cheap classifier deciding the tier — rather than abandoning the idea.

Frequently asked questions

How do I decide which requests are simple?
Start with a rule on an obvious signal — input length, endpoint, customer tier — before reaching for a classifier. Rules are free and are usually most of the win.
How do I detect that escalation is needed?
A confidence score if the task produces one, an explicit "I am not sure" instruction in the cheap model’s prompt, or a validation check on the output. All three are cheaper than the flagship request they avoid.

More cost tools