Token costs are just the first line on the bill. Idle computers and stalled projects make up the rest of it.
Picture a company renting a giant warehouse but only ever using one small corner of it. That is what is happening with AI computers today: on average, only 5 out of every 100 dollars of AI computer power actually gets used. The other 95 sits empty, still costing money every single day. Add the jump in cost when a test project finally goes live, and most of the real bill is empty warehouse space, not the tokens everyone watches.
Owning the hardware does not erase the invoice — production AI still has to be integrated, governed and staffed. It changes which lines move with usage and which ones stop moving.
| Dimension | Public cloud AI | AIPod Mini |
|---|---|---|
| Utilisation risk | You pay whether the GPUs are busy or idle. | Yours to tune, not priced into someone else's rate card. |
| Scaling curve | Roughly linear with usage, or worse. | Front-loaded, then mostly fixed. |
| Governance | Every prompt is a billed, logged, third-party event. | ONTAP Snapshots, FlexClone and SnapLock apply as they already do to your data. |
Stop renting the whole warehouse to use one shelf. Own only what you need, and the empty space stops costing you money.
Everyone watches the token line because it is the one with a number next to it. It is rarely the biggest. Here is roughly where the money goes.
Bars show relative weight, not precise figures. The shape is the point.
What does one answer cost? Good for support and service work, where volume is easy to count.
What does one finished job cost, start to end? Good when an agent does several steps.
What did the recovered claim, or the avoided penalty, earn? The one a CFO recognises.
Owned hardware turns the two biggest lines into one fixed number. Idle time stops being a variable cost, and the pilot-to-production jump mostly disappears — the box that ran the pilot runs production.
Arguing about token prices while 95% of the hardware sits idle is optimising the small line and ignoring the big one.
Full paper, 9 pages — iterate.ai/partners/netapp/papers.
A joint educational series on private AI and the AIPod Mini. NetApp® — the governed data-control layer. Iterate.ai — the private intelligence layer.
