AI API Bill Higher Than Your Token Estimate? Here's Why
An AI API bill runs above a token estimate because the estimate counts the text you wrote, while the provider bills every token it actually processed: a tool-use system prompt, tool definitions, the full conversation re-sent on each step of a tool loop, thinking tokens billed as output, cache writes at a premium, and, on Claude 4.7 and later, a tokenizer that produces about 30 percent more tokens for the same text. Cost trackers also miscount because Anthropic's input_tokens excludes cached tokens, OpenAI's input_tokens includes them, and Gemini reports thinking tokens outside candidatesTokenCount, so one formula across providers is wrong for at least two of them.
Your AI API bill comes in above your token estimate because the estimate counts the text you wrote, and the provider bills every token it processed. That second number includes a tool-use system prompt you never see, your tool definitions, the whole conversation re-sent on every step of a tool loop, and thinking tokens billed as output even when only a summary comes back. On Anthropic's Claude 4.7 and later models, the same text also tokenizes to about 30 percent more tokens than on earlier models, per Anthropic's token-counting docs as of October 2026.
There is a second, quieter problem. Even teams that log the real usage object often total it wrong, because the three big providers do not add their fields up the same way. Anthropic's input_tokens leaves cached tokens out. OpenAI's input_tokens already includes them. Gemini puts thinking in thoughtsTokenCount, outside candidatesTokenCount. A cost tracker that applies one formula to all three is wrong on at least two of them, and the error does not cancel out.
Where does the gap between estimate and bill come from?
Most "why is my bill so high" posts list causes in the abstract: context grows, retries happen, reasoning costs money. That is true and not much help when you are staring at a number. What helps is a ledger: each line is a specific source of billed tokens your estimate did not include, with the field that proves it.
| Gap source | What your estimate missed | Where it shows up |
|---|---|---|
| Tokenizer | You counted with a different model's tokenizer. Claude 4.7 and later produce about 30% more tokens for the same text than earlier Claude models | input_tokens higher than your count from day one |
| Tool-use system prompt | Anthropic adds a hidden system prompt whenever tools is present: 286 tokens on Claude Opus 5.5 and Sonnet 5.5, 675 on Opus 4.7 | Input, every request with tools |
| Tool definitions | Names, descriptions, and JSON schemas are input tokens. A dozen detailed tools can run past 1,000 tokens | Input, every request |
| Loop re-sends | Each step of an agent or tool loop is a new request carrying everything before it | Input grows step by step within one run |
| Thinking | Billed as output in full; you see a summary or nothing | reasoning_tokens, thinking_tokens, thoughtsTokenCount |
| Prior thinking | On Claude Opus 4.5 and 4.6-and-later models, earlier turns' thinking blocks stay in context and bill as input | Input on multi-turn conversations |
| Cache writes | Anthropic charges 1.25x base input for a 5-minute cache write, 2x for 1-hour | cache_creation_input_tokens |
| Retries | Your code re-runs a request after a parse error or timeout, and every attempt that completed is billed | Two or more requests in the logs for one job |
Two of these are the usual culprits in business automations. The loop re-send, because almost every useful AI step now calls at least one tool. And thinking, because reasoning models made it cheap to turn on and invisible to read.
How much bigger can one run get? A worked example
Here is a support-triage step we see often: read a customer email, look up the order, check the refund policy, write a reply. The numbers below are illustrative, built from the shape of flows we run, not a measurement of a specific account.
The estimate someone made in a spreadsheet: a system prompt plus the email, about 2,000 input tokens by a quick count, and a 300-token reply. One request.
What actually runs is three requests on a Claude model with three tools and thinking on:
- Request 1. System prompt and email, 2,600 tokens after the newer tokenizer's roughly 30% increase. Tool definitions, 1,200. Tool-use system prompt, 286. Input: 4,086. The model thinks for 900 tokens and calls the order lookup, 80 tokens. Output: 980.
- Request 2. Everything above, plus the assistant turn it just produced (980, thinking included, since it is the current turn), plus a 600-token order record. Input: 5,666. It thinks for 400 more and calls the policy tool. Output: 480.
- Request 3. All of that again, plus the 480 and a 600-token policy excerpt. Input: 6,746. Final reply: 350.
Total for one run: 16,498 input tokens and 1,810 output tokens. The spreadsheet said 2,000 and 300. That is about 8 times the input and 6 times the output, and nothing went wrong. No retry, no loop, no bug. This is the normal cost of a three-step tool run, and it is why we track cost per run rather than per request.
The fixes follow from the ledger. Cache the system prompt and tool definitions, which are identical on every request, so requests 2 and 3 read them at a tenth of the price. Return 150 tokens of the order record, not 600. Set the effort level or reasoning budget to the lowest that keeps quality, since thinking on request 1 is worth paying for and thinking on request 3 often is not.
Why do the usage fields add up differently by provider?
This is the part almost nobody writes down, and it is the reason two people looking at the same logs come up with different monthly totals. Each provider returns a usage object, and each one draws the boundaries differently.
| Provider | Total input | Full-price input | Output billed |
|---|---|---|---|
| Anthropic | input_tokens + cache_creation_input_tokens + cache_read_input_tokens | input_tokens only; the other two have their own prices | output_tokens (already includes output_tokens_details.thinking_tokens) |
| OpenAI | input_tokens (already includes input_tokens_details.cached_tokens) | input_tokens minus cached tokens | output_tokens (already includes output_tokens_details.reasoning_tokens) |
| Google Gemini | promptTokenCount (already includes cachedContentTokenCount) | promptTokenCount minus cached tokens | candidatesTokenCount + thoughtsTokenCount |
Read the table for the traps:
- Anthropic, undercount. If your tracker treats
input_tokensas total input, it misses every cached token. On a well-cached prompt,input_tokensmight read 50 while the request actually processed 100,050, the example Anthropic uses in its own caching docs. Worse, it misses cache writes, which are priced above base input. - OpenAI, double count. If you copy the Anthropic logic and add
cached_tokenstoinput_tokens, you count those tokens twice and price them at full rate. Your tracker then says caching made things more expensive. - Gemini, undercount on output. If you take
candidatesTokenCountas output, you skip the thinking tokens, which Google bills as output. On a reasoning-heavy step, thinking can be most of the output. - Thinking on OpenAI and Anthropic is a breakdown, not an extra.
reasoning_tokensandthinking_tokenssit insideoutput_tokens. Adding them on top double counts.
Our rule for any multi-provider automation: store the raw usage object as JSON on every call, next to the model ID, and compute cost in one function per provider. Never normalize at write time. When a provider adds a field, such as the cache_write_tokens count in OpenAI's current caching docs, you can recompute history instead of guessing at it.
How do you estimate cost before you ship?
Estimate from real requests, not from text length. The process that has held up for us:
- Count under the right tokenizer. Anthropic's
/v1/messages/count_tokensendpoint is free, accepts tools, images, and PDFs, and counts under the model you name. Anthropic's docs say outright not to reuse counts from an older model when moving to a 4.7-or-later one. Do not count Claude prompts with an OpenAI tokenizer library. - Run 20 real inputs end to end with thinking and tools on, exactly as production will. Record the full
usageobject for every request in every run. - Price per run, at the 90th percentile. The average hides the long emails and the runs that took five tool calls instead of three. Budget for the heavy tail.
- Put a cap on it.
max_tokens(ormax_output_tokens) bounds output for one request, thinking included, but not a whole tool loop. A per-run budget in your own code does that.
If a run then costs more than the estimate, the ledger above tells you which line moved.
What to do next
Pull the usage object from your last 50 AI calls and check one thing: does your cost formula match the provider row in the table above? If you run more than one provider, there is a good chance one of them is being counted wrong right now. Then split the bill by automation, which we cover in which AI automation is spending your budget, and add a per-run guard as described in setting a spending cap on an AI automation. If most of your input is a fixed system prompt and tool list, prompt caching is the first fix to try, and choosing when a reasoning model is worth it covers the thinking side.
We build AI steps with cost logging, caching, and per-run budgets wired in from the first deploy. If your AI bill has outgrown its estimate and you want a second set of eyes on the logs, see our AI systems work or get in touch.
Frequently Asked Questions
SOURCES & CITATIONS
- Token counting — Anthropichttps://platform.claude.com/docs/en/build-with-claude/token-counting
- Prompt caching — Anthropichttps://platform.claude.com/docs/en/build-with-claude/prompt-caching
- Reasoning models — OpenAIhttps://developers.openai.com/api/docs/guides/reasoning
- Gemini thinking — Googlehttps://ai.google.dev/gemini-api/docs/thinking
About Alexey Yushkin
Alexey is the founder of GENERAL INFORMATICS LLC. He designs and ships AI and automation systems for businesses and operators across the US.
Related reading
Want this kind of system in your business?
We build practical AI and automation systems for operators. Send us your current workflow and we will show you what to automate first.
