One workload at one healthcare company: 180,000 support conversations in a month. The tokens cost $2,520. The invoice was $49,840.
An airline fare is a real number, and it is not the number you pay. The fare gets you the seat; the fees get you the bag, the seat assignment, the change you had to make. Nobody budgets by the fare, because everyone has learned that the fare is the smallest line.
A metered AI bill works the same way. The token price is the fare. Everything the model does to produce an answer is the fees. The difference is that an airline prints its fees on the ticket.
The token price was never the bill. It is the one line a vendor quotes because it is the one line that sounds small.
Every line below is work the model did between the question and the answer — steps nobody watched and everybody paid for.
Base token fees 630M tokens × $4/M | $2,520 |
Context loading customer history, docs, policy re-sent each turn | +$1,440 |
Multi-step reasoning search, reason, check, respond | +$11,880 |
Agent loops each action is another metered call | +$8,000 |
Premium model routing complex questions sent to costlier models | +$15,000 |
Error retries regenerations and output quality filters | +$11,000 |
Total for the month of which $47,320 is not tokens | $49,840 |
Nothing carries between turns, so the history is re-sent every step.
Reasoning, retrieval and self-checking each bill separately.
Simple jobs get pointed at costlier models by default.
You cannot forecast a bill whose largest lines are decisions the model makes on its own.
On the meter: The End of “Per Token” Costs. On the lines beyond tokens: The Rest of the Invoice. On smaller models: Why Smaller Models Often Win. Worked example from the 12-page paper.
Full series — iterate.ai/partners/netapp/papers.
A joint educational series on private AI and the AIPod Mini. NetApp — the governed data-control layer. Iterate.ai — the private intelligence layer.