NetApp
/
What’s in Private AI Certified — Expert
/
The End of “Per Token” Costs (Deep Dive)
NetApp Iterate.ai
NetApp® Sellers & Partners
Joint Educational Series
Field Brief  ·  Paper Two of the Series

The End of “Per Token” Costs

How private AI changes the economics — and why the token tax punishes success while the AIPod Mini keeps costs flat.

What You’ll Learn
Why AI bills explode even as per-token prices fall — the “token tax” that ties cost directly to usage.
Where a $47,000 monthly bill actually comes from — context loading, multi-step reasoning, agent loops, and premium model routing.
Why agentic workflows multiply token usage 5–30× per interaction, and what that costs at enterprise scale.
How private AI replaces metered per-token billing with flat infrastructure cost — and the year-one and ongoing savings.
Why going private protects more than budget: your data, IP, and governance never leave infrastructure you own.
Version 1.2
Aug 11, 2026  ·  01 / 12
A True Story

AI Works. Usage Grows. Bills Explode.

These are real stories from companies we work with.

$43,000 in a single day
A finance department at a large fitness company deployed an agent. It burned the entire month’s token budget in a single day.
A full year’s budget gone by April
Uber’s developers adopted AI-assisted coding across engineering. Usage exploded. By the time finance caught up, the company had spent its entire 2026 AI coding allocation in four months.
$221M / month — $2.65B a year
Meta’s internal AI usage consumed 73.7 trillion tokens in 30 days — roughly 55 trillion words, or 12,000 English Wikipedias in a single month — tracked on an internal leaderboard called “Claudeonomics.” It is now imposing strict token budgets and building an “AI Gateway” to monitor spend in real time.

Uber® and Meta® are trademarks of their respective owners, referenced here for editorial and educational purposes only. This document is not affiliated with or endorsed by either company.

This is the token tax. And it is eating enterprise AI budgets alive.
The pattern is the same everywhere: AI works, usage grows, bills explode.

But the bigger problem is looming. As agents become prevalent — reasoning through multiple steps, calling tools, orchestrating workflows — token usage will multiply 5–30× per interaction. The crisis companies feel today with simple chatbots is about to get thirty times worse.

NetApp | Iterate.aiThe End of “Per Token” Costs  ·  02 / 12
The Basics

What Is a Token?

Think of a token as a word fragment. When you send text to an AI system, it breaks your message into tokens. When the AI responds, that response is measured in tokens too. You pay twice — once for what you send (input), once for what comes back (output). Developers call these input and output tokens.

“cat”
1 token
“understand”
~2 tokens
“customer support”
2 tokens

A typical support interaction uses about 3,500 tokens — maybe $0.01 at current rates. Cheap. But enterprise AI doesn’t run one interaction. It runs millions.

The Anatomy of a Token Bill

Where a $47,000 Month Comes From

A healthcare company’s support system handled 180,000 interactions, ~3,500 tokens each, at $4 per million tokens. That’s 630 million tokens — only $2,520 in raw token fees. So where did the other $44,480 come from?

Base token fees
630M tokens × $4/M
$2,520
Context loading
customer history, docs, policy — +2,000 tok/interaction
+$1,440
Multi-step reasoning
search, reason, check, respond — 3–5×
+$11,880
Agent loops
each action is another metered API call
+$8,000
Premium model routing
complex questions to costlier models
+$15,000
Error retries & quality checks
regenerations, output filters
+$11,000
Total per month — and climbing
$47,320

The CFO’s budget may have just assumed simple Q&A. Reality delivered an agentic system with memory, reasoning, and actions — all metered by the token.

NetApp | Iterate.aiThe End of “Per Token” Costs  ·  03 / 12
The Core Idea

Token Costs Scale With Success

Here is the cruel irony: the better your AI works, the more you pay. When that healthcare company’s AI started working well —

It got smarter by pulling more context — more tokens per interaction.
They added use cases — billing, scheduling, refills — more volume.
They integrated CRM and EHR systems — more agent actions.
And if they exposed it to customers, their clientele used it more — more interactions.

This is the opposite of normal technology economics. Usually, when you build something that works, marginal cost drops. With token-based AI, marginal cost stays constant — or grows. Per-token pricing ties your cost directly to your success. Success becomes a budget problem.

The Research Confirms It

Prices fell 98%. Budgets grew nearly 6×.

583%
Average enterprise AI budget growth — from $1.2M (2024) to $7M (2026) — even as per-token prices dropped 98%.
73%
of enterprises reported AI costs exceeded original projections in 2026 (FinOps Foundation).
71%
of companies experienced AI cost overruns in 2025 (Reuters).
−98%
Fall in LLM token prices. The problem isn’t the unit price — it’s the volume no one predicted.
“A token is cheap. A billion tokens a month is a line item the board will ask about.”
NetApp | Iterate.aiThe End of “Per Token” Costs  ·  04 / 12
The Core Idea

Agents Make Everything Worse

The token cost problem you see today is about to get 5–30× worse with AI agents. A simple chatbot uses ~3,500 tokens. An agentic workflow uses far more per task, because agents:

Reason through multiple steps — plan, execute, validate.
Maintain context across actions — memory, state, history.
Generate intermediate outputs — drafts, summaries, analysis.
Call external tools — databases, APIs, internal systems.
Self-correct and retry — error handling, quality checks.
Run multi-turn — each new turn re-sends the whole conversation, so cost compounds.

Industry analysis shows agentic AI systems require 5 to 30 times more tokens per task than standard conversational tools (Optimum Partners, 2026).

And agents increasingly run on their own schedule, not yours. At Iterate, three agents kick off automatically every morning — including weekends — with no one prompting them. That’s a productivity win: work gets done before the team logs on. But every one of those unattended runs burns tokens whether or not anyone reads the output. Autonomous agents don’t wait for a human to decide the cost is worth it — which is why their spend has to be monitored.
The Math Gets Brutal Fast
Chatbot
3,500
tokens / interaction
$0.014
per interaction
Agentic System
17.5K–105K
tokens / interaction
$0.07–$0.42
per interaction

That’s 5–30× higher cost for the same customer interaction. And agents are the future — every major vendor is pushing agentic workflows. Your customers are deploying them right now, often without understanding the token economics.

NetApp | Iterate.aiThe End of “Per Token” Costs  ·  05 / 12
The Real Cost at Enterprise Scale

What the Token Tax Costs Per Year

Three company sizes, all running agentic workflows — where everyone is heading — at $4 per million tokens.

Small Enterprise
5,000 employees · 50K interactions/mo · ~15K tokens each · 750M tokens/mo
$36K
/ year
Manageable — but already 4× higher than simple chatbots.
Mid-Sized Enterprise
15,000 employees · 500K interactions/mo · ~20K tokens each · 10B tokens/mo
$480K
/ year
Now a major budget line — and this is conservative.
Large Enterprise
50,000+ employees · 2M interactions/mo · ~25K tokens each · 50B tokens/mo
$2.4M
/ year
At this scale, the token tax is on the C-suite’s radar.
Real Example — AI Image Generation
Per image
$0.80 $0.04
20× cheaper on private AI
At ~$2,000 / day
~$700K/yr
on the shared model, for images alone
Per image
12s 2s
6× faster, optimized private AI
And $2.4M can be the low end. One large retailer — ~40,000 employees, 20M website visitors a month — was projected at $6M–$17M a year in token costs from a single use case: exposing conversational commerce to shoppers. Add internal chatbots and agents on top, and it climbs further. On private AI, Iterate could bring that same workload down to roughly $500K.
These figures also assume $4 per million tokens. Premium models (GPT-4 class) run $8–$15, and multimodal costs more — so many enterprises are already paying 2–3× the numbers above.
NetApp | Iterate.aiThe End of “Per Token” Costs  ·  06 / 12
The Shift

How Private AI Changes the Economics

With the AIPod Mini, the token meter disappears entirely. Instead of paying per token, you pay for infrastructure (one-time), software (annual), and energy (ongoing). No tokens. No metered usage. No surprise bills.

Public AI — Token-Based
180,000 interactions / month
$47,000 / month in token costs
Private AI — AIPod Mini
Hardware ~$65K (NetApp AFF)
Generate software ~$50K/yr · Energy ~$3K/yr
$564K
per year
$118K yr 1
then $53K/yr (software + energy)
$446K
saved in Year 1
$511K
saved every year after
100×
cheaper per token, open-source on-prem

And the best part: cost stays flat no matter how much you use it. Run 180,000 interactions or 1.8 million — same cost. When you own the infrastructure, the per-token charge goes to zero. Open models like Llama 3.1 8B cost $0.83 per million tokens on private infrastructure versus $90.00 for GPT-4 on public APIs.

The NetApp AIPod Mini appliance
The AIPod Mini — NetApp storage and Iterate’s Generate reasoning layer in one appliance you own and run on-prem.
NetApp | Iterate.aiThe End of “Per Token” Costs  ·  07 / 12
Real-World Proof

It Already Works, In Production

Two customers running private AI today.

Healthcare RCM
Agents process insurance claims, spot denial patterns, generate appeals. 50,000 claims/month.
Public AI~$180K/yr
Private AI~$55K/yr
69% saved
Document Intelligence
AI extracts data from contracts, invoices, compliance docs. 100,000 documents/month.
Public AI~$240K/yr
Private AI~$60K/yr
75% saved

Both told us the same thing: “We couldn’t afford to scale on public AI. Private AI made it economically viable.”

Your Talk Track

What This Means for Sellers

If they’re on public AI today, they’re paying per token and the bill is climbing; they’re afraid to expand use cases; they’re getting budget pushback from finance. Start the conversation here:

Ask Your Customer
“How much are you spending on AI today — and more importantly, how much are you not doing because of token costs?”
Most customers have a list of AI use cases they’ve shelved because the token economics don’t work. Private AI unlocks them.
NetApp | Iterate.aiThe End of “Per Token” Costs  ·  08 / 12
Beyond Cost

The Biggest Benefit Isn’t the Savings

Everything so far has been about money — and the money is real. But for many customers, the most important reason to move to private AI has nothing to do with the bill. It’s about what leaves the building when you use public AI, and what doesn’t when you don’t.

Every prompt to a public AI service carries your company’s intelligence with it — your data, your processes, your customer relationships, the accumulated know-how that makes you competitive. That information flows into a shared, third-party system: governed by someone else’s terms, trained on someone else’s roadmap, and ultimately controlled by people with their own commercial self-interests. Once it leaves, you can’t fully see where it goes or how it’s used.

With private AI, your company’s intelligence and IP never leave infrastructure you own — and never become raw material for a vendor you don’t control.
Public AI
Your data and IP travel into a shared third-party system.
Governed by the vendor’s terms and roadmap, not yours.
Limited visibility into how it’s stored, used, or retained.
Private AI — AIPod Mini
Intelligence and IP stay on infrastructure you own.
You set the governance, retention, and access rules.
What the AI learns stays yours — and compounds over time.
NetApp | Iterate.aiThe End of “Per Token” Costs  ·  09 / 12
The Value Proposition

The AIPod Mini Math, By Customer Type

Spending $100K+/yr on Tokens
AIPod Mini pays for itself in 6–12 months.
Then saving $200K–$500K+ annually.
Scale usage without budget anxiety.
Own the infrastructure and the models.
Just Starting With AI
Predictable costs from day one.
No surprise bills as usage grows.
Freedom to experiment without metering.
Built-in data sovereignty and compliance.
Production needs a different model

The Token Tax Punishes Success

Token-based AI made sense when AI was experimental. But AI is now production infrastructure — and production needs predictable, scalable economics. Private AI flips the model:

Metered per query.
Flat, unlimited usage.
Cost grows with success.
Pay for infrastructure, not queries.
Model owned by a third party.
The AI — and what it learns — is yours.

The companies that move to private AI now will hold a 3–5 year cost advantage over competitors still paying the token tax. The ones that don’t will be asking their CFOs for bigger AI budgets every quarter.

NetApp | Iterate.aiThe End of “Per Token” Costs  ·  10 / 12
NetApp’s AIPod Mini as the answer

This is what the AIPod Mini solves — a complete, validated solution that’s ready to deploy today:

Deploys in under 20 minutes. Start building your own agents within 30.
Keeps your data on NetApp storage you own. No egress, no cloud lock-in.
Runs AI models locally — no cloud, no token costs, no metered usage.
Integrates with ONTAP Snapshots, FlexClone, and SnapLock for backup, testing, and compliance.
Scales from a single department to enterprise-wide deployment.

Standard configurations: ~$45K–$70K annually for Iterate’s Generate software, with NetApp hardware from entry-level to enterprise scale.

More in This Series
What Makes AI Different from IT? The foundational paper on memory, reasoning, and learning.
Attack Surface & Memory Risks. Technical threat model and infrastructure concerns.
What is Private AI? For sovereignty and control.
About This Series

A joint educational series on private AI and the AIPod Mini. NetApp — the governed data-control layer. Iterate.ai — the private intelligence layer.

NetApp AIPod Mini
NetAppIterate.ai
v1.2 · Aug 11, 2026
NetApp | Iterate.aiThe End of “Per Token” Costs  ·  11 / 12
Appendix

A Quick Glossary for Sellers

Plain-language definitions for the technical terms in this paper — enough to hold the conversation with a customer’s technical team.

Token
A word fragment — the unit AI reads and writes in. Metered on both the input you send and the output that comes back.
Context
Everything handed to the model alongside the question — history, documents, policies. More context, more tokens.
Context Loading
Pulling that supporting material into each request. Convenient, but it silently adds thousands of tokens per interaction.
Reasoning
The model “thinking through” a problem in steps before answering. Each step generates tokens, multiplying cost.
Memory
What an agent carries between steps or sessions — state, history, prior results — re-sent as tokens each time.
Agent
An AI that doesn’t just answer — it plans, acts, and works toward a goal across steps, calling tools and checking itself.
Agent Loops
The repeated act–observe–act cycle an agent runs to finish a task. Each pass is another metered API call.
E.g. search records → read result → refine query → search again, until it lands the answer.
Multi-turn
An extended back-and-forth. Each turn re-sends the whole conversation, so cost compounds as it grows.
E.g. a 10-message troubleshooting chat re-sends all 10 messages on the final turn.
Tools
Capabilities an agent can invoke mid-task — search, calculators, functions. Each call routes back through the model.
E.g. a billing agent calls a calculator to total an invoice.
External Tools
Third-party systems beyond your own stack — databases, APIs, SaaS apps — an agent reaches into to get work done.
E.g. pulling a customer’s order history from Salesforce mid-conversation.
Self-correct
When an agent catches its own mistakes and retries. Every retry regenerates output — and costs more tokens.
Regenerations
Re-running a response to get a better result. Each regeneration is billed as a fresh output, doubling or tripling cost.
Output Filters
Quality, safety, or format checks run on responses before they ship — often another model call, another charge.
Model Routing
Sending harder questions to costlier, more capable models. Improves answers, but premium models cost far more per token.
Autonomous / scheduled
An agent that runs on its own schedule — even weekends — unprompted. Great for productivity; accrues cost unattended.
Token tax
Per-token pricing that ties cost to usage — the better AI works and the more it’s used, the higher the bill climbs.
Private AI
Running models on infrastructure you own — like the AIPod Mini — not a metered public API. Pay for hardware, not tokens.
Shared Models
Public AI models hosted by a vendor and used by many customers at once. Your prompts run through infrastructure — and terms — you don’t control.
NetApp | Iterate.aiThe End of “Per Token” Costs  ·  12 / 12