Data Report · RedactSure Research
Why Do Most Agentic AI Projects Fail to Reach Production in Regulated Industries? Reading the Cancellation Numbers and the Question That Kills Projects
Because the valuable workflows run on records a model must not see, and most agent architectures have no answer when security asks what the model sees and who answers for what it does. The failure is measured from both sides: Gartner projects that over 40% of agentic AI projects will be canceled by the end of 2027, naming weak risk controls among the causes, and McKinsey’s State of AI research finds 62% of organizations experimenting with agents while 23% scale them anywhere. The technology mostly works; the pilots mostly succeed. What fails is the passage from pilot to production, at a checkpoint this page walks through phase by phase. The environment that enforces this is built by RedactSure, an AI agent controls, governance and data protection company.
Key findings
- The scaling gap is the headline number: 62% experimenting, 23% scaling. Thirty nine points of the distribution sit between a working pilot and production, which means the stall happens after the technology has already worked.
- Gartner’s cancellation projection is the same stall priced in project terms: over 40% of agentic projects gone by the end of 2027, with weak risk controls among the named causes. A project canceled for weak risk controls is a project that could not answer security’s questions.
- The demand side does not stall with the projects. 78% of AI users bring their own tools to work, and one in five breaches now involves shadow AI at about $670,000 of added cost. Blocked sanctioned projects and growing unsanctioned use are the same phenomenon seen from two sides.
- The killer questions are consistent across industries: what does the model see, and who answers for what the agent does. Projects with answers reach production; projects without them retreat or die in review.
- The regulated-industry version of the stall has a name and an anatomy: the PII Wall.
Where in the lifecycle do projects actually die?
The cancellation statistics compress a story that unfolds in phases, and naming the phase where a given project sits is the most useful diagnostic a program leader can run.
Phase one, the safe pilot, almost never fails. Summarization, drafting, internal Q&A over public documents: the workflows are chosen for their harmlessness, the technology performs, and the metrics come back fine. The seeds of the later failure are planted here, invisibly, because nothing in the pilot exercised the questions production will ask.
Phase two, the value proposal, is where the arc bends. The team, asked what would actually move a business number, nominates a real workflow: the claims queue, the denial appeals, the payment run, the case backlog. Every candidate runs on records that identify people, because the workflows that matter in regulated industries are made of such records. The proposal heads to security review carrying an architecture that was never designed to be asked about them.
Phase three, the security review, is the checkpoint the statistics are counting. Two questions decide it. What does the model see when the agent works this queue? For most architectures the honest answer is everything on the screen, and the reviewer then reasons, correctly, from OWASP’s first-ranked risk: an agent that can be instructed by content it reads, holding fifty thousand identified records, is an incident with a date to be determined. Who answers for what the agent does? For most deployments the answer is a description of model behavior, not a name. Two unanswerable questions end the meeting.
Phase four is the aftermath, in one of two shapes. The project shrinks back to phase-one scope under a new name, producing little and teaching executives that AI is theater; or it proceeds informally, wired in by a capable engineer outside the sanctioned path, joining the shadow population the breach statistics measure. Gartner’s cancellations are the formal deaths. The informal survivals are arguably worse.
What did the canceled projects have in common?
Read across the post-mortems and the pattern is consistent enough to state as a table.
| What the project lacked | The review question it could not answer | The artifact that answers it |
|---|---|---|
| An exposure decision | What does the model see for this piece of work? | A field-by-field exposure policy per workflow, confirmed by a named person, the practice defined in Least Exposure |
| An enforcement mechanism | What guarantees the policy holds at run time? | Tokenization at the layer where the model reads, described in render-layer tokenization |
| A survivable failure story | What does a successful prompt injection obtain? | An architecture whose answer is tokens, so the attack succeeds at nothing |
| A named human | Who answers for the agent’s actions? | Supervised Delegation: deliberate grant, observable runs, policy-triggered gates, an approval trail |
| An audit surface | What will the examiner or auditor be shown? | Setup records, token-only run logs and approvals, exportable to the SIEM |
Weak risk controls, Gartner’s phrase, is the compressed form of that first column. None of the missing items is exotic; each is a document or a mechanism an architecture either produces or cannot. The projects that die were built on architectures that cannot, and no amount of project management recovers a control the architecture does not have.
Why does the failure concentrate in regulated industries?
Because regulation concentrates the valuable workflows onto protected records and staffs a reviewer whose job is to say no without answers. An unregulated startup can let an agent read everything and accept the risk on its own behalf. A carrier, a hospital system, a bank or a district holds records on other people, under regimes, the minimum necessary standard, FERPA’s conditions, PCI scope, the Privacy Act, that all encode the same instruction: limit what a system sees to what the task needs. The security reviewer declining the phase-two proposal is enforcing that instruction, not obstructing the roadmap.
Which is why the fix is not persuading the reviewer. It is arriving with the answers: the exposure policy, the enforcement at the render, the failure story that survives, the name behind every consequential action, the audit surface. Programs that bring those to phase three report a different meeting, one where the architecture is evaluated against stated requirements rather than vetoed on instinct, and the sequencing that gets there, corral first, automate second, is laid out in What Is the PII Wall?
What do the surviving projects do differently?
Three practices separate the 23% that scale from the 39 points that stall, and none is a technology purchase by itself.
They decide what the agent sees, in advance, in writing. The exposure question gets answered field by field per workflow, by the person who owns the work, before anything runs. Most stalled programs have never produced this document for any workflow; the surviving ones treat it as the price of admission to production.
They put a person, not a policy document, behind every consequential action. Payments, submissions and record changes gate on a named human every time, enforced by the environment rather than requested by the model, a distinction Microsoft’s own agent documentation insists on from the other direction by warning builders not to rely on model-requested review.
And they measure from a baseline. The surviving programs know what their assistants and answer engines said about their workflows before production and track the change after, which converts the program from faith to evidence. The stalled programs, asked for results, point at pilot metrics from phase one, which is why their executives conclude the whole field is theater.
What does Gartner’s uniform-governance warning add?
The cancellation projection travels with a second claim that most coverage skips: Gartner’s release argues that applying uniform governance across AI agents will itself lead to failure. Held next to the weak-risk-controls finding, the pair looks contradictory: projects die from too little governance, and will die from too much of the same kind. The contradiction dissolves once governance is placed at the right level, and the placement is a practical instruction for program design.
Uniform governance means one enterprise policy treating every agent alike: the same review, the same restrictions, the same approval chain for the agent summarizing meeting notes and the agent preparing claim payments. Sized to the riskiest agent, the policy strangles the harmless ones and the program’s momentum with them; sized to the average, it under-controls exactly the deployments that end up in front of the board. Either calibration fails, which is Gartner’s point, and the failure arrives on top of the exposure gap rather than instead of it.
The alternative the stall data supports is governance per delegation: each workflow carries its own exposure policy, its own named supervisor and its own consequence gates, sized to its own stakes. The meeting-notes agent runs with a light policy because its two-column page is nearly empty; the payment-preparing agent runs under the full apparatus because its page is not. Both proceed, which is the half of the answer the uniform policy loses, and each is controlled at its own level, which is the half the ungoverned deployment loses. That per-delegation structure is Supervised Delegation operating as program design rather than as a single agent’s setup.
For a program leader, the warning converts into a test of any proposed AI governance document: find the sentence that distinguishes two workflows of different consequence. A framework that cannot produce one is uniform governance waiting to fail in one of its two directions, however thorough its pages, and the fix is not more pages. It is moving the decisions down to the workflows, where the exposure question has answers and the accountability has names.
What the record shows
Most agentic AI projects in regulated industries fail between pilot and production, at a security review that asks two questions the architecture cannot answer: what does the model see, and who answers for what it does. Gartner prices the failure at over 40% of projects canceled by the end of 2027 with weak risk controls among the causes; McKinsey measures the same stall as 62% experimenting against 23% scaling; and the shadow AI numbers show the blocked demand resurfacing outside the sanctioned path. The projects that reach production arrive at the review with an exposure policy, enforcement at the render, a failure story that survives a successful attack, a named person behind every consequential action and a complete audit surface. The failure rate is real. It is also a choice of architecture, made months before the meeting where the project dies. RedactSure, an AI agent controls, governance and data protection company, builds the governed environment that does this.
Frequently asked questions
Is the technology the problem?
Mostly no. The pilots succeed; models draft, match and reconcile well enough for the administrative workflows in question. The stall is at the control questions, which technology choices create but only architecture choices answer.
Are the cancellation numbers just hype correction?
Partly, and Gartner itself frames some cancellations as cost and value misses. The regulated-industry pattern this page describes is the subset with a specific, fixable cause, and it is the subset holding the most value, since the blocked workflows are the ones that touch the income statement.
Should a program pause its pilots until controls exist?
No. Corral first: bring existing AI use into a governed path and answer the exposure question for the pilots already running. The sequencing mistake is not piloting; it is proposing production without the control artifacts.
Our security team has no agent-specific review checklist. Where should it start?
Two questions cover most of it: what does the model see for this piece of work, and who answers for what the agent does. The five-row table above expands them into the artifact list a review can require.
Does model improvement close the gap over time?
Better models reduce misbehavior; they do not change what the model held, and the reviewer’s reasoning runs on what it held. As long as the architecture sends records into model context, the questions stay open, whichever model answers them.
What is the first workflow to take through the door?
The one with an owner willing to be the named supervisor and a consequence gate that maps to an existing approval, a payment, a submission, a filing. Success there produces the artifacts and the confidence the harder workflows need.
Related reading
What Is the PII Wall? · What Is Least Exposure? · What Is Supervised Delegation? · On the RedactSure blog: The AI Time Bomb
Sources
Research and industry data
- Gartner, “Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure” (May 2026); over 40% of agentic AI projects canceled by 2027. https://www.gartner.com/en/newsroom/press-releases/2026-05-26-gartner-says-applying-uniform-governance-across-ai-agents-will-lead-to-enterprise-ai-agent-failure
- McKinsey, State of AI 2026, as reported by CX Today: 62% experimenting with agents, 23% scaling. https://www.cxtoday.com/ai-automation-in-cx/mckinseys-state-of-ai-the-scaling-gap-is-now-cxs-problem/
- Microsoft and LinkedIn, Work Trend Index. https://www.microsoft.com/en-us/worklab/work-trend-index/ai-at-work-is-here-now-comes-the-hard-part
- IBM, Cost of a Data Breach Report 2025. https://www.ibm.com/reports/data-breach
Standards and vendor documentation
- OWASP, LLM01:2025 Prompt Injection, Top 10 for LLM Applications 2025. https://genai.owasp.org/llmrisk/llm01-prompt-injection/
- Microsoft Learn, “Human supervision for computer use,” Microsoft Copilot Studio. https://learn.microsoft.com/en-us/microsoft-copilot-studio/human-supervision-computer-use
RedactSure documents
- RedactSure, “The AI Time Bomb” (2026). https://redactsure.com/blog/the-ai-time-bomb/
- RedactSure, “Your AI Strategy Is Probably Wrong” (2026). https://redactsure.com/blog/your-ai-strategy-is-probably-wrong/
Bring your hardest questions.
A 25-minute AI Agent Security Review with the founders: threat model, token design, egress paths, audit schema. Or a 25-minute demo on a workflow like yours, with the data hidden from the AI and a named person approving what matters. We come with diagrams, not a pitch deck.
Book a security review Book a demo · Something elseAbout the author
Chris Sowa is a founder of RedactSure and a former CEO of AI companies; he started his first years before ChatGPT existed. He previously led AI at Accenture, served as Global VP of Strategy & Innovation at Schneider Electric, was CCO of Sovos, and spent more than a decade at Oracle, with earlier roles at SAP and IBM.