TokenPad
Cost

Annual LLM Budget Planner

Multiple workloads, compounding growth, twelve months.

One workload per line: name | requests/month | input tokens | output tokens

86 characters3 lines0 tokensor drop a file

Annual BudgetExact
0Token change
Token change0no change
Input tokens0what you pasted
Output tokens0what you would send
Token cost of this result
Output tokens0
As input$0.00
× 100K requests$0.00

Everything on this page runs in your browser. Nothing you paste is transmitted, because there is no server here to transmit it to.

Result
 

Growth compounds, and budgets rarely account for it

Eight percent monthly growth is a 2.5× increase over a year. A budget built on today’s traffic is wrong by that factor by December, and the conversation about it happens at the worst possible time.

The month twelve run rate is the number to take to a planning meeting, not the month one figure.

Optimise the dominant workload only

Spend is almost always concentrated: one workload is half the bill and three others share the rest. Optimising the small ones is satisfying and pointless.

The breakdown sorts by cost for exactly that reason. Work on the top line until it is no longer the top line.

Frequently asked questions

What growth rate should I assume?
Your measured rate over the last three months, not your target. If you have no history, model an optimistic and a pessimistic case and check that the optimistic one is survivable.
Why is a high-volume cheap workload sometimes the largest line?
Because volume multiplies. A classification endpoint at 400 tokens and 900,000 requests can outspend a chat endpoint at 1,800 tokens and 300,000. Rank by total, never by per-request cost.

More cost tools