In early 2026, an autonomous AI agent broke into McKinsey’s internal AI platform in about two hours. The AI didn’t crack a password, it reasoned its way in.
The agent that entered McKinsey’s Lilli platform accessed 46.5M private conversations about strategy, mergers, and client deals. It extracted 728K internal files and compromised 57K employee accounts. It also gained write access to the 95 system instructions that govern how the AI in Lilli thinks. No software vulnerability was exploited. The agent found gaps in access logic and walked in.
The firewalls didn’t fail. The assumptions did. IT controls were built for software that executes. AI reasons — and that difference produces a class of risk your existing IT security framework was never designed to see:
When a database fails, it stops. IT risk is deterministic — you can test whether a query returns the right result, and a failure throws an error somebody sees.
When an AI fails, it keeps running — producing plausible, wrong output. AI risk is probabilistic. It can get the wrong answer a thousand times before anyone notices.
The AI gets quietly worse as the world changes. Software doesn’t decay on its own; a model’s accuracy can slide for months with no error thrown.
Agents send, update, and approve. A wrong action executes across systems at machine speed — all while RBAC governs users, not multi-step agents.
Corrupt the training or retrieval data and the model is compromised at its core. There is no patch to the software — only retraining of the entire data set.
HIPAA, GDPR, and SOC 2 security frameworks assume humans decide and can explain why. AI compliance is about reasoning, not just access logs.
“The agent did it” — so who’s responsible? Model, data, agent, and oversight are blurry compared to traditional IT's clear chain.
Confident, well-formatted, and wrong. Software fails loudly; AI fails fluently — as illustration, just having 1%–2% of 100,000 monthly interactions is 1,000 – 2,000 wrong answers.
Every prompt, response, and tool call is an endpoint, and each is a door for these six failure modes. Private infrastructure consolidates dozens of vendor endpoints to one boundary the customer owns: models run on NetApp® storage, the customer decides when they retrain, and Iterate’s AgentWatch logs every agent action — with human approval on risky steps and an audit trail the customer owns.
Treated like IT, these six surface as violations and quiet damage. Governed as their own class, they become a checklist you can actually run.
On what AI memory holds: The Oohs, Awes, and Dangers of AI Memory. The attacker’s view: Five Doors Into Your AI’s Memory (companion). Full series — iterate.ai/partners/netapp/papers.
A joint educational series on private AI and the AIPod Mini. NetApp — the governed data-control layer. Iterate.ai — the private intelligence layer.