Out in the wild
A week that will sound familiar.
None of these is a story about a broken agent. In two of the three the agent did something perfectly sensible, which is what makes them hard to talk about afterwards.
You would never give a new hire this much room.
Agents without rules means nobody has written down what an agent may decide on its own. What it must hand to a person, what it may never touch, and who watches that line.
Ask what your agent is allowed to decide on its own and you will get three different answers from four people. The line exists in somebody's head. This page is about getting it onto one page, in about twenty minutes.
How it usually surfacesIt refunded a customer nine hundred euros. It was right, actually. That is the part that bothered me.A CTO, 90 people, in June
Out in the wild
None of these is a story about a broken agent. In two of the three the agent did something perfectly sensible, which is what makes them hard to talk about afterwards.
In plain terms
The conversation usually gets framed as how much we trust the model, and that framing takes you nowhere. You do not trust a new account manager in the abstract either. You tell them they can discount up to ten per cent, that anything above that comes to you, and that they never promise a delivery date without checking.
Nobody finds that insulting. It is how delegation works when it is done well. An agent gets none of it, because it arrived through a tool and not through a hiring process, so the only real limit on it is what the software happens to allow.
It's like flying a plane at 10,000 feet, being told to climb to 12,000, replace both engines mid-flight and ensure zero turbulence. No one would choose to pilot that plane, but that's exactly what companies are doing today.Afonso Eça, Executive Board Member at Banco BPI1
He is describing the pressure, and the way out of it is duller than people expect. You do not need a governance programme. You need two sentences per agent that a new colleague could read and act on, and one person who says them out loud.
The evidence
That is the average number of agent incidents companies reported over the past year, where something unintended or harmful needed a human to correct it.1 Roughly one in six of those was serious enough to take more than four hours to contain.
Now the part that surprised me, and the reason this page is not a warning. In the same study, the companies that built limits into the system rather than watching manually were running sixteen times as many agents as everybody else.
Read that the way you would read it about people. The teams with the clearest boundaries are the ones you can hand the most to. It has never worked differently, and it turns out agents are no exception.
Why it happens
What is the most this thing may give away without asking. Fifty euros, five hundred, nine hundred. Somebody has to say a number out loud and then live with it. That is uncomfortable in a way that buying a licence is not, so it gets pushed to the next meeting, and the agent keeps running in the meantime with whatever limit the software came with.
The rule you never wrote down is still a rule. It is just whatever the software happened to allow.Paul Musters
Seven out of ten technology leaders say teams across the business are deploying faster than IT can track.1 An agent goes live in a marketing team on a Wednesday. Legal hears about it in October. By then the question is no longer what should this be allowed to do, it is what has it already been doing.
For CIOs and CTOs, the challenge now is scaling AI systems that operate continuously and autonomously, often within governance models and architectures designed for a far slower, more predictable environment.Matt Lyteson, CIO at IBM1
Most AI policies I read describe which tools are approved and what data may go in. Useful, and it answers none of this. An agent that acts needs a different kind of sentence: what it may decide, what it must hand over, what it may never touch. Only a third of companies have formally adopted a policy for agents at all.2 The rest have something about tools, which the agent has never read.
Three minutes, six questions, anonymous. It gives you your level, the price of staying on it, and what moves at the next one.
One thing to try
Take one agent that is running today. Get the person who owns it and the person who carries the risk in the same room, which is usually two people and not a committee. Fill in three columns.
Things you would be comfortable reading about after the fact. Put a number on anything with money in it.
Things where a person adds judgement rather than a click. Name who, not which team.
The short list. If it is long, you are describing a job the agent should not have.
Example lines. Yours will be different, and the middle column is where the real conversation happens.
Two rules about the exercise. Every line in the middle column needs a name, because “escalate to the team” means escalate to nobody. And if you cannot fill in the first column without a debate, you have found a decision the business never made, which the agent has been making for you in the meantime.
In the five levels
In the five AI Culture Levels we use with clients, this is the step where things stop living in people's heads. Level 3 is where the way of working gets written down, and an agent's boundaries are part of that.
At Level 2 the limits live in the head of whoever built the thing. That works until they are away, and it stops working entirely once a second person can change the agent. At Level 3 the way of working is written down, so a new colleague can read what the agent may do and act on it in their first week. That is the whole difference, and it is smaller than most transformation programmes make it sound.
Level 4 is where the limits stop being a document and start being part of the system, so the agent cannot exceed them even if somebody wants it to. That is what the sixteen times number is about. Companies at that level are not more cautious. They can hand over more precisely because the edges hold.
And most companies sit on more than one level at once. Your engineering team may be at 3 while the commercial side is at 1, and the agent with the loosest boundary is usually not in the team you are watching.
An Operating Profile in use. Personality type and AI level in one profile, with the agents that fit it.
A rule tells the agent where to stop. It does not tell you whether the person on the other side of that line can handle what arrives there. Escalating to somebody who will glance at it and click approve is not oversight, it is a delay with a signature on the end.
We measure it at the level of the person. How somebody thinks and works, and where they are with AI. Then the middle column of your decision line goes to a name that can actually carry it. In practice the right name is often not the most senior one in the room.
emaho measures one Operating Profile per person: personality type and AI level in a single profile. On that we build a personal set of AI agents that fit how that person works, inside the tools they already use. Fifteen minutes to complete, first profile free, built for companies between 20 and 500 people.
Fifteen minutes per person. No credit card, no strings.
What to do
Start with the agent that touches money or customers, not the one that is easiest. Draw the three columns for that one. Twenty minutes with two people beats a quarter with a working group, and you will learn more from the argument than from the document.
Then put a number on the middle column. This is the step everyone slides past. Somebody has to say the amount, the date range, the tone. If the room cannot agree, do not paper over it with a phrase like “use judgement”. Write down that it is unresolved and put a name and a date against it, because right now the agent is resolving it for you every day.
Make escalation land on a person, not a channel. A queue nobody owns is where escalations go to be ignored, and after a month people stop sending them there. One name, and a second name for when the first is away.
Then move what you can into the system itself. A limit that is written in a document depends on everyone remembering it. A limit that is coded into the agent holds on a Sunday. That is where the difference between the companies with a quarter fewer incidents comes from.1
And read the exceptions once a quarter. What got escalated, and what should have been and was not. That half hour tells you more about whether your line is in the right place than any dashboard will.
The middle column is where I can actually tell you something. Send it and I will tell you what I would move and what I would tighten. No deck, no call. One message back.
What waiting costs
The visible cost is the incidents themselves, and they are more mundane than the word suggests. Of the serious ones, over a third ended in data exposure or a security breach and a third caused failures that spread into other systems.1 Four hours to contain, on a day you had planned to do something else.
Underneath that sits a slower cost. Once a team has been surprised twice, they start checking everything the agent produces. Nobody asked them to. They do it because nobody told them where the edge was, and checking everything is the only safe answer to that. At that point you are paying for the agent and paying for the human review, and the gain you built the thing for has quietly gone.
And there is the one that shows up in a board meeting. Somebody asks what the agent is allowed to do and you need a week to answer. That week tells them how much of the business is now running on decisions nobody wrote down, and once a board sees that, the appetite for the next agent goes to nothing. More than 40% of agent projects are expected to be cancelled before the end of 2027, and inadequate risk controls sit in that same sentence.3
What people search for
Close by
A boundary only works if somebody is behind it and somebody is still looking. These three sit either side of this one.
A rule with no name behind it is a suggestion. Before you draw the line, somebody has to be responsible for what sits on either side of it.
Read this one 18You wrote the rule and put a person behind it. Six months later that person approves everything within a minute, because it always looked fine before.
Read this one 07Rules for the agents you know about do nothing for the ones running in personal accounts. Almost half of employees use AI nobody approved.
Read this oneTwenty-five things that break inside a company once people start using AI, each with the research behind it and the level where it starts to bite. This one starts to bite at Wild West, and 8 of the others start there too.
Getting going
The three columns take twenty minutes. What they cannot tell you is whether the person the middle column escalates to can judge what lands there. That is the part we measure, starting with you.
First profile free · no credit card · built for companies of 20 to 500 · you decide what your team gets to see
Not ready to put your team in anything yet? Start with the level of the company instead. The Culture Level scan is six questions, three minutes, and asks nothing of you.
Numbers are quoted as published. Where a figure is described as around or roughly, that is how the source states it. Nothing on this page is legal advice.