The number that should reorganize your thinking is not 74. It is 81 — the rollback rate among organizations with fully mature guardrails. Rollback is not a quality signal. It is a detection signal.
A survey of more than 2,500 AI decision-makers found that 74 percent of enterprises that deployed AI agents into customer communications subsequently rolled them back or shut them down. That is the number being quoted. It is the less interesting of the two.
Among organizations with fully mature guardrails, the rollback rate was 81 percent.
Read that twice. The companies that invested most heavily in AI governance pulled their agents back more often than the companies that did not. Every intuition about what a maturity model is supposed to do runs the other way.
The inversion, and what it actually means
There are two explanations and only one survives contact with the rest of the data.
The first is that governance makes deployments worse — that guardrails introduce friction, degrade the experience, and cause teams to give up. Nothing in the evidence supports this, and it would be a strange mechanism.
The second is that governance makes failures visible. An organization with mature guardrails has monitoring, evaluation, escalation paths and someone whose job includes noticing. An organization without them has an agent in production and a customer-complaint queue. Both have failures. Only one has a rollback event, because only one can see what it needs to roll back from.
The survey's own framing, from Sinch's chief product officer, lands on exactly this: "The most advanced organizations aren't failing less; they're seeing failures sooner."
Which means rollback rate is not measuring the quality of your AI. It is measuring the quality of your instrumentation. And the organizations reporting no rollbacks are not the winners of this cycle — they are the ones who have not looked.
The supporting figures point the same way. 84 percent of AI engineering teams spend at least half their time on safety infrastructure. And 75 percent now place trust, security and compliance in their top three areas of spend — ahead of AI development itself, at 63 percent. The industry has quietly reallocated the majority of its engineering effort from building the thing to containing it, and has not updated any of its public narratives to match.
The same pattern shows up in code
If this were confined to customer-facing agents it might be a domain quirk. It is not.
A survey of more than 200 enterprise technology leaders found 61 percent of code is now AI-generated or AI-assisted, and 81 percent report an increase in production issues linked to that code. The figure that makes it a story rather than a statistic: 92 percent were confident the code was production-ready before shipping.
An 81 percent failure-increase rate sitting next to a 92 percent confidence rate is not a competence problem. It is a feedback problem. The people shipping have no signal that would update their belief, so the belief does not update.
The downstream costs are quantified and rarely budgeted: 54 percent report significantly higher CI/CD infrastructure spend, 53 percent higher testing, security and deployment costs, and 70 percent say maintaining the test suite is now a bigger burden than writing the code. Meanwhile only 31 percent of AI spending connects to any measurable business result.
Visual 1 — What a rollback rate is actually telling you
Reported rollback rate | Naive reading | What it more likely means |
|---|---|---|
High, with mature governance | The technology is failing | Failures are being detected and acted on. This is the system working. |
High, with weak governance | The technology is failing | Failures reached customers before anyone noticed. Detection came from outside. |
Low, with mature governance | Success | Genuinely good — but verify the deployments are non-trivial before believing it. |
Low, with weak governance | Success | No signal at all. Indistinguishable from failing silently. The riskiest cell on this table. |
How to read it: The bottom row is where most organizations quietly sit, and it is the one that gets reported to boards as progress. You cannot distinguish "did not fail" from "did not notice" using any published rollback statistic, including the ones in this article.
Uniform governance is the error
Gartner's contribution to this is a prediction with a mechanism attached: 40 percent of enterprises will demote or decommission autonomous AI agents by 2027, because governance gaps get discovered after production incidents rather than before. Its recommendation is to stop applying one control regime across all agents and instead tier them by autonomy:
Four autonomy tiers
Observe — the agent watches and reports; no action taken.
Advise — the agent recommends; a human decides and acts.
Act with approval — the agent proposes a specific action; a human authorizes it.
Act autonomously — the agent executes without a checkpoint.
The practical value is that controls, monitoring and review cadence attach to the tier rather than to the system. Most enterprises today run a single policy across all four, which means it is simultaneously too heavy for the first tier and too light for the last.
The tiering also explains why rollbacks cluster where they do. An agent moved from "advise" to "act autonomously" as a quiet configuration change inherits the governance of the tier it left, not the one it entered. That transition is where most of the incidents live, and almost nobody treats it as a gated event.
A word about the numbers everyone is quoting
This story is unusually polluted, and readers deserve to know which figures survive checking.
Statistics to stop using
"95% of AI pilots fail." This is from MIT NANDA's The GenAI Divide, published August 2025. It is routinely reprinted under 2026 headlines with no date. There is no 2026 successor study. Cite it with its date or not at all.
"Gartner: 89% of AI agent pilots never scale." We could find no Gartner release containing this claim. It appears to have been invented in circulation.
"Over 40% of agentic AI projects will be canceled by end-2027." This one is real, but it dates to June 2025, not 2026 — and it is a different prediction from the May 2026 demote-or-decommission figure cited above. The two are frequently conflated into one inflated number.
The Sinch and CloudBees figures used here are vendor-commissioned surveys with disclosed sample sizes. Both companies sell into the problems they measured. We have cited them because the methodology is stated and the findings are internally consistent — not because a vendor survey is the same thing as independent research.
What to do differently
Replace rollback rate with time-to-detection. How long between an agent behaving wrongly and someone knowing? That number is comparable across teams, cannot be gamed by deploying less, and tells you the thing rollback rate only implies.
Tier your agents and gate the transitions. Moving from advise to autonomous should require the same ceremony as a production change, because it is one. Today it is usually a settings toggle owned by whoever built the thing.
Budget the safety infrastructure explicitly. If 84 percent of AI engineering teams spend half their time on it, then half your AI engineering cost is containment, whether or not your business case says so. A plan that funds only the build is under-funded by roughly the amount that determines whether it survives.
Make a rollback path an acceptance criterion. Not a runbook written afterward — a tested, exercised path, demonstrated before the agent handles real traffic. The organizations rolling back at 81 percent could do so because the path existed.
Ask what your quiet deployments are actually doing. If you have agents in production and no incidents, resist the conclusion that they are working. Sample the outputs. The absence of an alert is not evidence when nothing is configured to alert.
The prevailing story of this year is that enterprise AI is failing and the numbers prove it. The numbers prove something narrower and more useful: organizations that can see what their AI is doing keep finding reasons to pull it back, and organizations that cannot see keep reporting that everything is fine. Only one of those groups is learning anything, and it is the one currently generating the worst-looking statistics.
Sources and method. A BusinessInfomatics original. Rollback figures (74 percent overall; 81 percent among organizations with fully mature guardrails; 84 percent of AI engineering time on safety infrastructure; trust/security/compliance at 75 percent of top-three spend versus AI development at 63 percent) from a Sinch survey of 2,500+ AI decision-makers, reported by The Register, May 13, 2026 — vendor-commissioned research; Sinch sells adjacent products. Code figures (61 percent AI-generated or assisted; 81 percent reporting more production issues; 92 percent pre-ship confidence; 54/53 percent cost increases; 70 percent test-maintenance burden; 31 percent of AI spend tied to measurable results) from a CloudBees survey of 200+ enterprise technology leaders, reported by The Register, May 20, 2026 — also vendor-commissioned. Gartner's demote-or-decommission prediction and the four autonomy tiers per Gartner, May 26, 2026; the separate 40-percent-canceled prediction dates to Gartner, June 25, 2025. The "95% of AI pilots fail" figure is MIT NANDA, August 2025. We were unable to locate any Gartner publication containing the widely circulated "89% of agent pilots never scale" claim.



