Token costs are the first line on the bill. Idle GPUs, stalled pilots, and unclear ROI are the rest of it — and why owning the infrastructure changes every line.
A joint educational white paper by NetApp® and Iterate.ai.
According to Cast AI’s 2026 State of Kubernetes Optimization Report — measured from production telemetry across 23,000 clusters, not a survey — the average enterprise GPU fleet runs at just 5% utilization. Ninety-five percent of the compute enterprises paid a premium to secure sits idle at any given moment.
That gap is expensive at scale: Gartner puts new AI infrastructure spending at $401 billion in 2026 alone. Enterprises spent two years racing to secure GPU capacity before competitors did. Now the capacity is sitting there, mostly unused, and the bill for it doesn’t care whether the silicon is working.
This is the part of the invoice sellers rarely walk a customer through. Token costs get the attention because they show up on a monthly bill with a dollar figure attached. Idle infrastructure is a cost too — it just doesn’t send an invoice. It just sits there, depreciating.
The End of “Per Token” Costs — elsewhere in this series — walks through how token-based pricing punishes agentic workflows: context loading, multi-step reasoning, and premium model routing that can multiply a bill 5–30× per interaction. That argument still holds, and it’s worth reading on its own. But token metering is only the first line item on a much longer invoice.
The rest of this paper walks through those remaining line items — and how owning the infrastructure changes what shows up on each one.
A proof of concept is built to impress a room, not to survive contact with production. Add reliability engineering, monitoring, scaling, and support on top of the pilot build, and a $60,000 proof of concept can easily become a $250,000 production system — a jump procurement rarely plans for because the pilot invoice never mentioned it.
That 3–6× jump shows up across independent cost analyses, not just one estimate: Gartner notes that raising accuracy requirements alone — from 90% to 99% — can multiply implementation effort three to five times. It’s a structural cost of production AI, not a budgeting failure by any one team.
It’s also part of why 95% of enterprise generative-AI pilots never reach measurable business impact (MIT, 2025) — many stall at the point where the invoice jumps and no one budgeted for it.
Fixed infrastructure changes the shape of that jump. When the hardware is already yours, scaling usage doesn’t re-price the deal the way scaling API calls does.
Most AI cost conversations stop at “dollars per token” or “dollars per GPU-hour” — a unit that means something to an engineer and nothing to a CFO. Every other function in the business is costed per unit of value delivered. AI deserves the same discipline, at four different levels of the same question.
None of these numbers exist by default. They have to be defined before the pilot starts — which is where most AI ROI cases go wrong.
Deloitte’s 2026 enterprise survey found that nearly three-quarters of companies report their most advanced AI initiatives met or exceeded ROI targets. The gap between those and the ones that stall isn’t luck — it’s visibility. Organizations that can trace AI cost to a specific business outcome manage it well. Organizations flying blind are spending without knowing what they’re getting back.
Owning the infrastructure doesn’t erase the invoice — production AI still has to be integrated, governed, and staffed. What ownership changes is which lines move with usage and which ones don’t. On a public cloud bill, almost every line scales with how much the business actually uses AI. On owned infrastructure, the biggest line is fixed the day it’s installed.
Every enterprise running AI at scale is already paying for idle GPUs, integration work, governance overhead, and the jump from pilot to production — whether or not those costs show up on a single line. The companies that win aren’t the ones spending the least. They’re the ones who can see every line on the invoice before it arrives, and who own enough of the infrastructure that scaling usage doesn’t mean re-pricing the deal.
A joint educational series on private AI and the AIPod Mini. NetApp — the governed data-control layer. Iterate.ai — the private intelligence layer.

Plain-language definitions for the cost and economics terms in this paper — enough to hold the conversation with a CFO or finance team.