When Everything Works and It Still Fails
· Enterprise AI · 14 min read
When enterprise AI fails, it usually fails above the model — in trust, accountability and the messy business of actually running the thing.
The analysis
The thought experiment is deliberately unfair to the usual explanations. Assume the data is secure, privacy handled, the model accurate, bias tested and drift monitored, and take cost off the table. The agentic programme still fails. That is not hypothetical — it is roughly where the market is. The technical layer is the part enterprises already know how to buy: necessary, and nowhere near sufficient.
The evidence converges from five independent vantage points. Gartner expects more than 40% of agentic AI projects to be cancelled by end-2027. MIT's Project NANDA found around 95% of enterprise generative AI pilots delivered no measurable P&L impact, and was explicit that the divide is driven by approach rather than model quality or regulation. McKinsey describes the same gap as the gen AI paradox, and names the shift that makes agentic different: from systems that might say the wrong thing to systems that can do the wrong thing. Accenture found only about 13% of organisations reporting significant enterprise-level value — and that firms with responsible-AI governance across the lifecycle were roughly 2.7 times more likely to create it. Infosys puts the responsible-AI gold standard at about 2% of firms, with only around 14% having scaled agentic AI at all.
In regulated industries the gap has a shorter fuse. A wrong answer does not arrive as product friction; it arrives with a regulator attached. India's Reserve Bank made this concrete with its FREE-AI framework, placing explainability and accountability at the centre of AI in financial services. When an agent declines a loan, the customer has a right to reasons — a requirement a black box that is usually right cannot satisfy.
Trace each failure downward and stop only when the answer is no longer 'buy a better component'. What remains is a small set of structural fault lines. Accountability has been separated from the act: the agent takes a consequential action, and the human principal the law requires neither made nor could have foreseen it. Justification debt accumulates: the decision path was non-deterministic and state-dependent, so accuracy is available but a stable, contestable account of the decision is not. Tacit rules go unwritten: the agent follows the explicit instruction and violates the unwritten one nobody knew needed writing down. And tail risk governs trust — months of good behaviour end with a single egregious, often irreversible action, because autonomy raises the ceiling of harm per action and removes the human-paced friction that used to catch errors before they scaled.
The practical implication is a reordering of the work. The controls that matter are the ones that bind an answerable actor to each consequential action, produce a durable record of why an action was permitted, make the organisation's unwritten rules explicit enough to encode, and cap the blast radius of irreversible actions before scale rather than after an incident. None of that is bought from a model vendor, and none of it is a compliance overlay applied at the end. It is the operating design of the programme itself.
Full essay on Substack: When Everything Works and It Still Fails.
More on Enterprise AI
- Nobody Called It a Model
- What the Stanford 2026 AI Index Actually Says
- Our AI Coding Numbers Were Perfect. We Were Shipping Slower.
Discuss this with Prasanna: pw@prima-partners.com · book a 30-minute call.