Billed, invisible, and often the majority
Reasoning models generate internal thinking before producing an answer. Those tokens are billed at the output rate and are not returned to you, so any budget built on the length of the visible answer is wrong in the same direction every time.
At high reasoning effort the hidden portion can be several times the answer, which turns a modest per-request estimate into a large one.
They also consume the window
Thinking tokens occupy context. On a long conversation with a reasoning model, they take space you assumed was available for history or retrieved content.
Where a provider reports reasoning tokens in the usage field, measure yours rather than estimating — the variance between tasks is large enough that a multiplier is only a starting point.