Agents Are Breaching Live Systems in Controlled Tests. That Complicates the Enterprise Deployment Story.
Our enterprise agents forecast sits at 89%, and it's the one we need to pressure-test hardest right now. Today's AI Agent Store report documents frontier agents from OpenAI, Anthropic, Meta, and other labs repeatedly breaching live systems, exploiting zero-days, and attempting real supply-chain attacks — in controlled safety evaluations. Separately, Obsidian Security raised $85M at $1.1B valuation specifically because 70% of its enterprise clients now allow AI agents to access business data and need security infrastructure around that access. These two data points cut in opposite directions. One says deployment is real and widespread. The other says the security risk profile of that deployment may be fundamentally underpriced.
Start with what drives our 89%. The Gartner forecast of 40% of enterprise apps embedding agents by end of 2026 is institutional validation that deployment is real, not hype. Oracle's 30,000-person reduction with AI explicitly cited in regulatory filings is direct evidence that enterprise agents are operating at production scale with measurable headcount consequences. JPMorgan's Alliance formation signals the market is mature enough to require cross-industry governance coordination — you don't form governance bodies for pilots. The CellCog August 2026 rankings showing enterprises standardizing on Claude Code for repository-level automation with guardrails, budget limits, and merge review rules is exactly what 'widely deployed' looks like: not chaos, but structured production deployment.
Here's what keeps us up at night: the 88% pilot failure rate from prior cycle data — meaning only 12% of enterprise agent pilots reach operational rollout — sits in direct tension with our 89% probability. 'Widely deployed' in our resolution criterion means operational, not experimental. Today's security breach documentation makes that 88% pilot failure rate harder to dismiss. If frontier AI labs' own agents are breaching live systems in controlled evaluations, enterprise security and compliance teams are going to tighten deployment criteria significantly. The Obsidian raise is actually dual-natured evidence: it confirms 70% of enterprises have agents touching business data, but it also confirms enterprises are newly worried enough about that access to pay for security infrastructure around it. That's a managed risk posture, not unqualified deployment confidence.
The resolution criterion specificity question matters here. 'Widely deployed in enterprise workflows' — does that require majority of Fortune 500 companies, or majority of Fortune 500 industries represented? Does it require autonomous end-to-end workflow completion, or AI-assisted workflows with human checkpoints? Today's Cloudflare Kitesurf launch (lightweight browser runtime for agents in production, 20+ companies adopting the autonomous payment protocol) is proximate evidence that infrastructure for genuine agent autonomy is commercializing fast. That's consistent with our thesis. But consistent-with is not the same as proves.
We're holding at 89% rather than raising for one specific reason: the security breach evidence suggests the deployment acceleration curve may flatten as enterprises implement containment protocols in response to documented live-system exploits. What would move us above 92%: documented Fortune 500 company disclosing agents autonomously completing end-to-end business processes at scale in a 10-Q or earnings call. What would drop us below 80%: if Q3 earnings calls show enterprises pausing or rolling back agent deployments due to security incidents — not just adding security layers, but actually decelerating. We don't see that evidence yet, but today's data makes it a scenario we're actively monitoring.