NetApp
/
What’s in Private AI Certified — Expert
/
The Rest of the Invoice (Deep Dive)
NetApp Iterate.ai
NetApp Sellers & Partners
Joint Educational Series
Field Brief  ·  The Wider Economics

The Rest of the Invoice

Token costs are the first line on the bill. Idle GPUs, stalled pilots, and unclear ROI are the rest of it — and why owning the infrastructure changes every line.

A joint educational white paper by NetApp® and Iterate.ai.

What You’ll Learn
Why the average enterprise GPU fleet sits idle 95% of the time — and who pays for that idle capacity.
Why a $60K pilot can become a $250K production system, and how to budget for that jump before it happens.
A framework for costing AI the way you cost everything else — per interaction, per workflow, per employee, per outcome.
How to build an AI ROI case finance will actually trust.
NetApp | Iterate.aiThe Rest of the Invoice  ·  01 / 09
The Hidden Line Item

The Bill Nobody Sees Coming

According to Cast AI’s 2026 State of Kubernetes Optimization Report — measured from production telemetry across 23,000 clusters, not a survey — the average enterprise GPU fleet runs at just 5% utilization. Ninety-five percent of the compute enterprises paid a premium to secure sits idle at any given moment.

5%
Actually in use — running real workloads
95%
Idle — paid for, powered on, doing nothing

That gap is expensive at scale: Gartner puts new AI infrastructure spending at $401 billion in 2026 alone. Enterprises spent two years racing to secure GPU capacity before competitors did. Now the capacity is sitting there, mostly unused, and the bill for it doesn’t care whether the silicon is working.

This is the part of the invoice sellers rarely walk a customer through. Token costs get the attention because they show up on a monthly bill with a dollar figure attached. Idle infrastructure is a cost too — it just doesn’t send an invoice. It just sits there, depreciating.

Read: “GPU Utilization: Why 95% of Enterprise Capacity Sits Idle” — citing Cast AI’s 2026 State of Kubernetes Optimization Report (May 27, 2026).
NetApp | Iterate.aiThe Rest of the Invoice  ·  02 / 09
The Companion Question

Tokens Are the First Line, Not the Whole Bill

The End of “Per Token” Costs — elsewhere in this series — walks through how token-based pricing punishes agentic workflows: context loading, multi-step reasoning, and premium model routing that can multiply a bill 5–30× per interaction. That argument still holds, and it’s worth reading on its own. But token metering is only the first line item on a much longer invoice.

1
Idle infrastructure
The GPU utilization problem on the previous page — capacity billed whether or not it’s working.
2
Integration & professional services
Wiring AI into the systems, data sources, and workflows that already run the business. Rarely quoted up front.
3
Governance & compliance overhead
Audit trails, access controls, review cycles — the cost of being allowed to keep using it.
4
Talent & operating cost
The people who maintain, retrain, and babysit the system after the launch announcement.
5
The production surcharge
What changes, and what it costs, between a pilot that impressed a room and a system the whole company depends on. Covered next.

The rest of this paper walks through those remaining line items — and how owning the infrastructure changes what shows up on each one.

NetApp | Iterate.aiThe Rest of the Invoice  ·  03 / 09
The Production Surcharge

Why Pilots Look Cheap

A proof of concept is built to impress a room, not to survive contact with production. Add reliability engineering, monitoring, scaling, and support on top of the pilot build, and a $60,000 proof of concept can easily become a $250,000 production system — a jump procurement rarely plans for because the pilot invoice never mentioned it.

Pilot
$60K
Production
$250K+

That 3–6× jump shows up across independent cost analyses, not just one estimate: Gartner notes that raising accuracy requirements alone — from 90% to 99% — can multiply implementation effort three to five times. It’s a structural cost of production AI, not a budgeting failure by any one team.

It’s also part of why 95% of enterprise generative-AI pilots never reach measurable business impact (MIT, 2025) — many stall at the point where the invoice jumps and no one budgeted for it.

Fixed infrastructure changes the shape of that jump. When the hardware is already yours, scaling usage doesn’t re-price the deal the way scaling API calls does.

NetApp | Iterate.aiThe Rest of the Invoice  ·  04 / 09
The Core Idea

Costing AI the Way You Cost Everything Else

Most AI cost conversations stop at “dollars per token” or “dollars per GPU-hour” — a unit that means something to an engineer and nothing to a CFO. Every other function in the business is costed per unit of value delivered. AI deserves the same discipline, at four different levels of the same question.

1
Cost per interaction
What does one customer query or agent action actually cost, start to finish — not just the API call, but the infrastructure behind it?
2
Cost per workflow
What does automating one complete business process cost — claims review, contract intake, invoice matching — versus doing it manually?
3
Cost per employee
What does it cost to give one team member an AI-augmented workflow, and what is that worth in hours saved or output gained?
4
Cost per business outcome
What did it cost to produce one recovered claim, one closed deal, one avoided compliance failure — the number a CFO actually cares about?

None of these numbers exist by default. They have to be defined before the pilot starts — which is where most AI ROI cases go wrong.

NetApp | Iterate.aiThe Rest of the Invoice  ·  05 / 09
The Practical Step

Building a Credible AI ROI Case

Deloitte’s 2026 enterprise survey found that nearly three-quarters of companies report their most advanced AI initiatives met or exceeded ROI targets. The gap between those and the ones that stall isn’t luck — it’s visibility. Organizations that can trace AI cost to a specific business outcome manage it well. Organizations flying blind are spending without knowing what they’re getting back.

1
Define the business outcome first
Not “deploy AI” — “reduce claim-review time by 40%.” Everything else is costed against this number.
2
Price the production system, not the pilot
Budget for the 3–6× jump from page 4 before the pilot ships, not after it stalls.
3
Count every line on the invoice
Infrastructure, integration, governance, and talent — not just the model bill.
4
Track utilization, not just spend
A GPU fleet at 5% utilization is a cost problem even if the invoice looks reasonable.
5
Revisit the case every quarter
AI cost and AI value both move. A case built once at kickoff goes stale fast.
Read: “How Much Does AI Cost? The Complete Guide for 2026”, citing Deloitte’s 2026 State of AI in the Enterprise survey — CloudZero (Apr 27, 2026).
NetApp | Iterate.aiThe Rest of the Invoice  ·  06 / 09
The NetApp-Specific Advantage

What Ownership Changes

Owning the infrastructure doesn’t erase the invoice — production AI still has to be integrated, governed, and staffed. What ownership changes is which lines move with usage and which ones don’t. On a public cloud bill, almost every line scales with how much the business actually uses AI. On owned infrastructure, the biggest line is fixed the day it’s installed.

Dimension
Public Cloud AI
AIPod Mini
Utilization risk
You pay whether the vendor’s GPUs are busy or idle — utilization risk is priced into every API call.
You own the hardware, so utilization is a tuning problem you control, not a margin baked into someone else’s rate card.
Scaling cost curve
Cost scales roughly linearly (or worse) with usage — more agents, more tokens, more bill.
Cost is front-loaded and mostly fixed — scaling usage doesn’t re-price the deal the way scaling API calls does.
Production surcharge
Pilot-to-production jump still applies, plus the token-cost multiplier from agentic workflows.
The same jump applies to integration and governance work — but not to a re-priced infrastructure bill.
Governance overhead
Every prompt, response, and tool call is a billed, logged, third-party event.
ONTAP Snapshots, FlexClone, and SnapLock apply governance the same way they already do to your data.
NetApp | Iterate.aiThe Rest of the Invoice  ·  07 / 09
The costs already on your books

The Invoice Was Always Longer Than the Bill

Every enterprise running AI at scale is already paying for idle GPUs, integration work, governance overhead, and the jump from pilot to production — whether or not those costs show up on a single line. The companies that win aren’t the ones spending the least. They’re the ones who can see every line on the invoice before it arrives, and who own enough of the infrastructure that scaling usage doesn’t mean re-pricing the deal.

The AIPod Mini turns the biggest line item fixed.
Hardware you own, sized to your real workload, governed the same way as the rest of your data — so utilization becomes a tuning problem, not a bill that grows every time the business uses AI more.
More in This Series
The End of “Per Token” Costs. The companion piece to this paper — the token-billing mechanics behind the first line of the invoice.
From Empty Infrastructure to Working AI in 20 Minutes. Why data readiness, not model choice, is the real deployment bottleneck.
The Oohs, Awes, and Dangers of AI Memory. The five layers of AI memory, and why governing them privately is a competitive advantage.

About This Series

A joint educational series on private AI and the AIPod Mini. NetApp — the governed data-control layer. Iterate.ai — the private intelligence layer.

NetApp AIPod Mini
NetAppIterate.ai
v1.2 · Aug 11, 2026
NetApp | Iterate.aiThe Rest of the Invoice  ·  08 / 09
Appendix

A Quick Glossary of AI Economics

Plain-language definitions for the cost and economics terms in this paper — enough to hold the conversation with a CFO or finance team.

GPU Utilization
The share of provisioned GPU capacity actually running workloads at a given moment, as opposed to powered-on and idle.
E.g. a fleet billed for 100 GPU-hours a day that only runs real jobs 5 of those hours.
TCO (Total Cost of Ownership)
The full cost of a system across its lifetime — hardware or API fees, plus integration, governance, talent, and maintenance.
Unit Economics
Costing a system per unit of value delivered — per interaction, workflow, employee, or outcome — instead of per raw resource consumed.
E.g. knowing a contract-review agent costs $0.40 per contract, not just $4,000 a month.
Pilot vs. Production
A pilot proves a concept works for a small, forgiving audience. Production means it works reliably, securely, and at scale for everyone who depends on it.
Production Surcharge
The typical 3–6× cost increase between a working pilot and a production-grade system, driven by reliability, monitoring, and scale requirements the pilot never needed.
E.g. a $60K pilot that needs another $190K in reliability and monitoring work to go live.
Sunk Infrastructure Cost
Capacity already purchased that doesn’t recover its cost through use — the financial risk behind low GPU utilization.
E.g. GPUs already purchased and idling, which don’t get cheaper by staying unused.
ROI Case
A documented link between an AI investment and a specific, measured business outcome — built before the investment, not after it stalls.
Agentic Workflow
A multi-step AI process where an agent plans, calls tools, and iterates — each step is a separate, billable unit of work.
NetApp | Iterate.aiThe Rest of the Invoice  ·  09 / 09