The three numbers, and what they actually count
MIT's 95% comes from a 2025 study on generative AI pilots at large enterprises. It measured pilots that were expected to produce measurable P&L impact within six months and didn't. That's a narrow, aggressive bar. A pilot that improved internal support ticket resolution by 15% but never got a dollar figure attached to it would count as a failure under that definition, even if the team considered it a win.
Gartner's 40% is a different animal entirely: a forward-looking estimate that 40% of agentic AI projects will be cancelled by 2027, due to cost, unclear business value, or inadequate risk controls. It's a prediction about a specific category, agentic projects with multi-step autonomy, not a retrospective count of all AI pilots.
The 88% figure floating around vendor content usually traces back to surveys of proof-of-concept-to-production conversion rates across all AI projects, including basic automation and simple chatbots. It's the broadest net and the least useful number for a CTO deciding whether their specific pilot is in trouble.
None of these numbers tell you anything about your pilot. They tell you what population was studied. The useful move is to stop treating the stat as a verdict and start using it as a prompt to check five concrete gaps.