NetApp
/
What’s in Private AI Certified — Expert
/
Anatomy of a $47K Month LLM Bill
NetApp
Iterate.ai
NetApp Sellers & Partners
Joint Educational Series
Field Brief · Reading the Bill

Anatomy of a $47K Month LLM Bill

One workload at one healthcare company: 180,000 support conversations in a month. The tokens cost $2,520. The invoice was $49,840.

What you'll learn
  • The six lines on a metered AI bill, and which one the budget was built on
  • Why the four largest lines are all things the model did without being asked
  • Why the total does not transfer between companies — and the ratio does
The core argument

An airline fare is a real number, and it is not the number you pay. The fare gets you the seat; the fees get you the bag, the seat assignment, the change you had to make. Nobody budgets by the fare, because everyone has learned that the fare is the smallest line.

A metered AI bill works the same way. The token price is the fare. Everything the model does to produce an answer is the fees. The difference is that an airline prints its fees on the ticket.

$2,520
base token fees — the line the budget assumed
$47,320
everything else on the same invoice
19.8×
fees to fare, for one month of ordinary support work
$11,880Multi-step reasoning$8,000Agent loops$15,000Premium model routing$11,000Error retries$2,520 — base tokensthe line the budget was built onContext loading  $1,440Fare $2,520Fees $47,320One month. Total $49,840. Drawn to scale.
Your number will be bigger. This is one workload at one company. A large enterprise runs dozens at once, and Meta staff were reported to move 73.7 trillion tokens in a single month — roughly $1.1bn at list rates. The total is not what travels between companies. The ratio is.
What the line item hides

The token price was never the bill. It is the one line a vendor quotes because it is the one line that sounds small.

NetApp | Iterate.ai
Anatomy of a $47K Month LLM Bill · 01 / 02

Where the other $47,320 went

Every line below is work the model did between the question and the answer — steps nobody watched and everybody paid for.

Base token fees
630M tokens × $4/M
$2,520
Context loading
customer history, docs, policy re-sent each turn
+$1,440
Multi-step reasoning
search, reason, check, respond
+$11,880
Agent loops
each action is another metered call
+$8,000
Premium model routing
complex questions sent to costlier models
+$15,000
Error retries
regenerations and output quality filters
+$11,000
Total for the month
of which $47,320 is not tokens
$49,840
Why the big lines are big

The model forgets

Nothing carries between turns, so the history is re-sent every step.

The loop is invisible

Reasoning, retrieval and self-checking each bill separately.

Routing defaults up

Simple jobs get pointed at costlier models by default.

What owning changes. On an AIPod Mini the loops still happen — they are how the answer gets made — but on NetApp® hardware you own, they stop being billable events.
The cost you can't forecast

You cannot forecast a bill whose largest lines are decisions the model makes on its own.

Further reading

On the meter: The End of “Per Token” Costs. On the lines beyond tokens: The Rest of the Invoice. On smaller models: Why Smaller Models Often Win. Worked example from the 12-page paper.

Full series — iterate.ai/partners/netapp/papers.

About this series

A joint educational series on private AI and the AIPod Mini. NetApp — the governed data-control layer. Iterate.ai — the private intelligence layer.

AIPod Mini
NetApp
Iterate.ai
v1.1 · Aug 11, 2026
NetApp | Iterate.ai
Anatomy of a $47K Month LLM Bill · 02 / 02