NetApp Iterate.ai
NetApp Sellers & Partners
Joint Educational Series
Field Brief · The Wider Economics

The Rest of the Invoice

Token costs are just the first line on the bill. Idle computers and stalled projects make up the rest of it.

The core argument

Picture a company renting a giant warehouse but only ever using one small corner of it. That is what is happening with AI computers today: on average, only 5 out of every 100 dollars of AI computer power actually gets used. The other 95 sits empty, still costing money every single day. Add the jump in cost when a test project finally goes live, and most of the real bill is empty warehouse space, not the tokens everyone watches.

Key facts

The average enterprise GPU fleet sits idle 95% of the time.
A $60K pilot can become a $250K production system.
AI can be costed the same way as everything else — per interaction, workflow, or outcome.
A credible AI ROI case is one finance will actually trust.

By the numbers

5%
Average enterprise GPU utilisation (Cast AI, 2026)
$401B
New AI infrastructure spending in 2026 (Gartner)
3–6×
Typical cost jump from pilot to production system

What ownership actually changes

Owning the hardware does not erase the invoice — production AI still has to be integrated, governed and staffed. It changes which lines move with usage and which ones stop moving.

Dimension Public cloud AI AIPod Mini
Utilisation risk You pay whether the GPUs are busy or idle. Yours to tune, not priced into someone else's rate card.
Scaling curve Roughly linear with usage, or worse. Front-loaded, then mostly fixed.
Governance Every prompt is a billed, logged, third-party event. ONTAP Snapshots, FlexClone and SnapLock apply as they already do to your data.
Rent less of the warehouse

Stop renting the whole warehouse to use one shelf. Own only what you need, and the empty space stops costing you money.

NetApp | Iterate.aiThe Rest of the Invoice · 01 / 02

What is actually on the bill

Everyone watches the token line because it is the one with a number next to it. It is rarely the biggest. Here is roughly where the money goes.

Tokens
The line everyone tracks. Real, but rarely the largest.
Idle compute
Hardware you rented and did not use. Bills the same either way.
Getting data ready
Finding, cleaning and connecting files. Mostly people, not machines.
Integration
Wiring AI into the systems your teams already work in.
Pilot to production
The 3–6× jump when a proof of concept has to hold up.
People
The team keeping all of it running once it is live.

Bars show relative weight, not precise figures. The shape is the point.

Three ways to cost it

Per interaction

What does one answer cost? Good for support and service work, where volume is easy to count.

Per workflow

What does one finished job cost, start to end? Good when an agent does several steps.

Per outcome

What did the recovered claim, or the avoided penalty, earn? The one a CFO recognises.

Why owning changes the shape

Owned hardware turns the two biggest lines into one fixed number. Idle time stops being a variable cost, and the pilot-to-production jump mostly disappears — the box that ran the pilot runs production.

Optimizing the wrong line

Arguing about token prices while 95% of the hardware sits idle is optimising the small line and ignoring the big one.

Further reading

Full paper, 9 pages — iterate.ai/partners/netapp/papers.

About this series

A joint educational series on private AI and the AIPod Mini. NetApp® — the governed data-control layer. Iterate.ai — the private intelligence layer.

NetApp AIPod Mini
NetAppIterate.ai
v1.1 · Aug 11, 2026
NetApp | Iterate.aiThe Rest of the Invoice · 02 / 02