AI Agents Beyond the Experiment: Governance as the Lever for Scale
Published on 8/27/2026 · André Hellmann
AI agents in business have left the curiosity phase. 62 percent of organizations are at least experimenting with agents (Source: McKinsey, The State of AI, 2025). Very few reach production. The pattern is familiar — it is the pilot graveyard, this time with the power to act.
Discuss your next step in a free diagnosis call. Book a slot →
Contents
- The agent hype in numbers
- Why security and risk slow down scaling
- Governance in practice: decision rights, escalation, measurement
- Agents in production: the operations approach
- Conclusion: governance is the lever for scale
- Frequently asked questions about AI agents
- Sources
The agent hype in numbers
Curiosity runs high. 62 percent of surveyed organizations are at least experimenting with AI agents (Source: McKinsey, The State of AI, 2025). That is the number every deck quotes.
The second number appears far less often. Those who do scale mostly scale in one or two functions. In no single business function do more than 10 percent of respondents say their organization is scaling agents (Source: McKinsey, The State of AI, 2025).
A third number completes the picture. Average responsible AI maturity rose to 2.3 in 2026, up from 2.0 the year before. Only about 30 percent of organizations reach level three or higher in strategy, governance, and agentic controls (Source: McKinsey AI Trust Maturity Survey, 2026).
The order of these three values is the real finding. The distance between experiment and production is not a technology problem. It sits precisely where maturity in governance and controls is missing.
The free self-check shows where a company stands in a few minutes.
Why security and risk slow down scaling
Asked about the brake, nearly two thirds of respondents name security and risk concerns as the top barrier to fully scaling agentic AI — well ahead of regulatory uncertainty and technical limits (Source: McKinsey AI Trust Maturity Survey, 2026).
That is a meaningful shift. AI initiatives used to stall on missing capability. Agents stall on missing confidence that autonomous systems can be run safely.
The reason lies in what agents are. A chat assistant says the wrong thing. An agent does the wrong thing — it triggers an action, misuses a tool, or operates beyond the intended guardrails (Source: McKinsey AI Trust Maturity Survey, 2026).
The consequences are measurable. 51 percent of organizations using AI report at least one negative consequence, and nearly one third of all respondents report consequences from inaccurate output (Source: McKinsey, The State of AI, 2025).
One more gap runs through almost every risk category: companies recognize risks faster than they actively mitigate them (Source: McKinsey AI Trust Maturity Survey, 2026). The awareness is there. The controls are not.
An agent without decision rights is not a tool. It is a liability with an interface.
Governance in practice: decision rights, escalation, measurement
Governance sounds like committees and policy documents. What matters is smaller: four decisions per agent, not per company.
- Define the scope of action. Which actions may the agent trigger, and which may it only propose? Orders, pricing, and commitments to customers belong in the second category.
- Name the escalation path. In which cases does the agent stop and hand over — and to whom by name? Without a named owner, every exception becomes guesswork.
- Assign accountability. One responsible person per agent, as with any machine in operation. Organizations with clear accountability for responsible AI reach a maturity score of 2.6, those without 1.8 (Source: McKinsey AI Trust Maturity Survey, 2026).
- Measure the effect. Cases handled, abort rate, error rate. Without those three numbers, any discussion about agents stays a matter of opinion.
These four items take a few hours per agent. They decide whether an agent may go into production at all.
How autonomy can be dialed up in stages is covered in the article on governance, control, and autonomy. What an AI agent technically is, the glossary entry explains.
Agents in production: the operations approach
The lesson from the pilot graveyard applies to agents unchanged: what gets built for a test does not reach production. The difference is exposure — an agent acts, while a pilot only demonstrates.
That is why we build agents along the workflow they will run in. Three things come before the first line of configuration:
- The workflow. Which step costs time today, and who owns it? Without that answer, an agent automates disorder.
- The boundary. Where does the agent’s authority end, and what does the stop look like? That boundary is part of the configuration, not a rule in a handbook.
- The measurement. Which number proves after four weeks that the agent holds up? It is defined upfront, not searched for afterwards.
The effort pays off. Organizations investing substantially in responsible AI report material AI benefits far more often — including EBIT impact above 5 percent (Source: McKinsey AI Trust Maturity Survey, 2026). Governance is not a tax on innovation. It is its precondition.
Why we only ever build for production is covered in Production from Day One.
Which workflow suits a first agent is something we map out in the free diagnosis call.
Conclusion: governance is the lever for scale
The agent numbers repeat a familiar pattern. Plenty of experiments, little production — exactly what the Implementation Gap describes.
What is new is the cause. Capability is not the brake, confidence is. Nearly two thirds name security and risk as the main barrier (Source: McKinsey AI Trust Maturity Survey, 2026). That confidence does not come from better models. It comes from clear rules.
The practical consequence is unspectacular: scope of action, escalation path, named accountability, measured effect. Four decisions per agent. Without them, every agent stays an experiment — even two years in.
Frequently asked questions about AI agents
How many companies run AI agents in production?
62 percent are at least experimenting with agents. In no single business function do more than 10 percent of organizations scale them (Source: McKinsey, The State of AI, 2025). The gap between test and production is substantial.
What slows down the scaling of AI agents?
Nearly two thirds of respondents cite security and risk concerns — ahead of regulatory uncertainty and technical limits (Source: McKinsey AI Trust Maturity Survey, 2026). The brake sits in governance, not in technology.
How does agent governance differ from general AI governance?
An agent acts. So the question is no longer only about wrong statements but about wrong actions — triggered processes, misused tools, exceeded guardrails (Source: McKinsey AI Trust Maturity Survey, 2026). Governance therefore has to regulate rights to act, not just review content.
Who should be accountable for agents inside a company?
One named person per agent. Organizations with clear accountability for responsible AI reach a maturity score of 2.6, those without 1.8 (Source: McKinsey AI Trust Maturity Survey, 2026).
Where is the best place to start with a first agent?
With a recurring workflow that has a clear output and low external exposure — research, pre-qualification, data maintenance. Which one that is in a specific case is something we clarify in the free diagnosis call.
Sources
- McKinsey AI Trust Maturity Survey, 2026: State of AI trust in 2026 — Shifting to the agentic era, March 25, 2026
- McKinsey, The State of AI, 2025: The state of AI in 2025 — Agents, innovation, and transformation, 2025