The harder your AI works when using Public AI or shared models, the bigger the bill grows — and AI agents work a lot harder than one simple chat reply.
Public AI charges you like a taxi meter that never stops running. Every question, and every extra step the AI takes to think, adds a few more cents to the fare. The busier your AI gets, the faster that meter spins — so the more work it does, the bigger the bill. Own the car instead, and the meter disappears for good.
A 2023 chatbot task that cost $0.04 now runs ~$1.20 rebuilt as a 2026 agent — 30× (EY). Four cases:
The kicker: since June 15, agent tools bill you on a separate metered pool at API rates — outside your plan. The meter is now its own line item. (Anthropic)
Per-token prices have fallen up to 98% a year (÷200 since 2024) — yet total AI bills keep tripling. The price cut never reaches your invoice, because three forces multiply against it:
Illustrative — framework adapted from Agora Software; price data: Epoch AI. Usage multipliers vary by workload.
Cheaper tokens just unlock more use (the Jevons paradox): most of the spend is orchestration and checking, not the answer. In 21.8% of matchups the cheaper list price actually costs more — by up to 28×.
A chatbot answers once. An agent plans, fetches, checks, retries — every hidden step is another billed prompt.
Steps 2 to 4 repeat, often dozens of times. You see step 1 and step 5 and pay for everything in between — where nearly all the tokens go.
One request becomes dozens of model calls — 5–30× the tokens of a single chat, sometimes far more.
The model keeps no state between calls, so context is re-sent or re-fetched each step. Nothing is forgotten — you pay to reload it every turn.
Two habits pile on: simple jobs get sent to premium models that cost more than the task needs, and the AI’s answers (output tokens) bill several times higher than your prompt (input).
A session grows from ~2K to ~25K tokens per call (Mem0). You pay for the prompt, answer, and everything in between.
On hardware you own it stops mattering — the agent loops as often as the job needs, and the invoice does not move.
~$118K in year one (then ~$53K/yr) vs $564K on public AI — the meter is gone.
Illustrative math — Iterate.ai model for one mid-size deployment; year one includes one-time hardware. Assumptions in the full paper.
Full paper, 12 pages — iterate.ai/partners/netapp/papers. On the economics: “The Token Tax” and “The Death of the Token Tax” — Iterate.ai. Budget figures: FinOps Foundation, 2026 State of FinOps.
A joint educational series on private AI and the AIPod Mini — from NetApp and Iterate.ai.
