“I don't remember deciding to stop checking.”

Oversight erodes

Automation bias is when you stop checking what the machine hands you. Nobody announces it. The review stays in the calendar and turns into a look at the first paragraph.

It gets it right ten times, so the eleventh time you skim instead of read. Nobody decides that. It just happens, and one day you notice you stopped being the check a while ago.

8 min read 5 September 2026 Updated 7 September 2026
A workspace where people and their agents are working side by side
One incident, October 2025 They were hired to check somebody else's system. Nobody checked theirs. Deloitte Australia. The report quoted a court ruling that was never written.2

What happened

They were hired to check something. Nobody checked them.

Worth telling in full, this one, because it happened at a firm where checking other people's work is the actual product.

An Australian government department runs a computer system that hands out penalties to people on benefits. No person looks at it, the system decides. They wanted an outside opinion on whether it worked properly, so at the end of 2024 they hired Deloitte to go through it. About 440,000 Australian dollars, so roughly a quarter of a million euros. The report came in that July.2

Then Chris Rudge read it. He does research at the University of Sydney, the subject sits close to his own work, and he was reading it because he was curious. He found more than a dozen references to studies that do not exist. A few were credited to people he knows, who had never written them. There was also a quote from a court judgment. The judgment was real. The quote was not.2

Deloitte said they had used AI in places, published a corrected version in October and paid back about 97,000 of the fee. The department said the conclusions still held.2

The bit I keep coming back to is that nothing was broken. There were reviewers. There was a quality process. A partner's name was on the cover. All of it did exactly what it was built to do, and nobody clicked a footnote.

Oversight erodes means the checking stops while the process that describes it stays exactly where it was. Researchers at University College London put 1,401 people through this. Working next to a slightly skewed algorithm pulled their own judgement further off than a skewed colleague did, and they had no idea.1

In two minutes

  1. People on their own got it slightly wrong, one way, 53 times in 100. A model that learned from those same people got it wrong the same way 65 times in 100.1
  2. New people spent one session with that model. They started at 50 and ended at 61, on their own answers, given before they saw what the model said.1
  3. Run the same thing with a colleague who is off instead of a model and nothing happens. No drift, at any point.1
  4. Afterwards, people said the accurate algorithm was the one that had swayed them. It was the other one.1
  5. Deloitte sent out a report with invented sources in it and paid part of the fee back. Reviewers, quality process, partner sign-off, all in place.2

You will recognise at least one of these

  • There is a number in your deck and nobody can say where it came from any more
  • You approve things faster than you did a year ago and you are not sure that is good news
  • A client asked how something had been checked and it took you two days to give an honest answer
  • Somebody senior said “the AI said” in a meeting like that settled it, and nobody blinked

What you take away

  • A card you can fill in over coffee that tells you whether anybody is still checking
  • What happens to your team's judgement when they work next to a tool, with the numbers
  • Why a colleague who is wrong is less dangerous to you than a tool that is wrong
  • Where checking sits in the five AI Culture Levels, and what moves you there

Sound familiar

It never once felt like a decision.

Each of these is somebody doing their job properly, in a hurry. That is why they are hard to spot and easy to repeat.

Someone at work reviewing output alongside their agent
Thursday, four o'clock, it has to go out Marieke reads the proposal for typos and sends it. Ask her and she will say she reviewed it, and she is not lying. Three weeks later the client quotes a sentence back at her that nobody here recognises. All she can say is that she read it. True, and no help to anyone.
A colleague at a workstation with an agent beside him
The number that has been on slide four since May It came out of an analysis somebody ran in the spring and it has been in every deck since. Then a board member asks where it comes from. The trail runs back to a chat window that has been cleared. Now the whole slide is suspect, including the two numbers that were fine.
Two colleagues checking work together with their agents
“That's what came out of the system” Said calmly, as an answer, and everybody moves on. It used to be the opening of a conversation about whether the system was right. Once it becomes the end of one, you have lost a step you did not know you had, and the meeting looks exactly the same from outside.

What we mean by it

This gets filed as people being lazy. It is about judgement.

The easy version is that people got sloppy and need a reminder. That one comes with a cheap fix, everybody nods, and nothing changes. What actually happens is slower and a bit stranger. The tool moves where your people think normal sits, and they take that new normal into work the tool never touched.

Moshe Glickman and Tali Sharot at University College London tested it with 1,401 people. Simple task. You see twelve faces for half a second and say whether the group looks sadder or happier on average. Exactly half the sets were sadder, so 50 out of 100 is the right answer. On their own, people said sadder 53 times in 100. Barely a lean, and they corrected it themselves as they went.1

Interactions of other humans with this algorithm further increase the humans' initial bias levels, creating a feedback loop. Moshe Glickman and Tali Sharot, University College London, in Nature Human Behaviour1

Then they trained a model on those human answers. It came back saying sadder 65 times in 100. Then they sat new people down with that model. Those people started at 50 and finished at 61, on their own answers, given each round before they saw what the model said. They had picked up the bigger version and made it theirs.1

The loop

A small lean goes in. A big one comes back.

Same experiment, same task, and 50 out of 100 is the right answer every time. Read it as a circuit rather than three separate findings.1

Step 1
People, on their own A tiny lean towards sadder. They corrected it themselves as the session went on.
53%
Step 2
A model trained on those same people It took the lean and made it the rule. The messier the human answers it learns from, the bigger the distortion.
65%
Step 3
New people, after a session with that model Their own answers, given each round before they saw the model's. They began the session at 50.
61%
Control
New people, after a session with a person who was off Same setup, same disagreements, a colleague instead of a model. Nothing moved, at any point.
51%
How often each group said sadder, out of 100, when exactly half the sets were sadder. The number is that score. The bar shows how far it sits from 50, which is the right answer here. Steps 1 to 3 are the loop with an AI in it. The last row is the same test with a person.1
32.7%
of the time people changed their own answer when the AI disagreed. When a colleague disagreed: 11.3%. With nobody there at all: 4%
0
drift from working with a person who was off. Same setup, same disagreements, nothing learned at any point
Wrong
is what people said about which algorithm had swayed them. They named the accurate one. It was the other one

That last one is what I would put in front of a leadership team. Ask your people whether the tool is steering them and the answer runs roughly opposite to the truth. So asking them is not a check.

There is a good half to this that almost nobody quotes. The people who worked with an accurate algorithm got better at the task on their own. It runs both ways. What decides the direction is whether anyone is looking at the tool.

Why it happens

You stopped checking on a day you cannot name.

Checking costs time and being right is invisible

Nine times out of ten you check and it was fine, so it feels like wasted effort, because it was. Nobody has ever been thanked for confirming that a correct document was correct. So the checking quietly drops to whatever level has not gone wrong yet. Then it drops again.

Nobody schedules the moment they stop checking. It arrives as a run of afternoons where it turned out fine. Paul Musters

Work that looks finished switches off the careful reader in you

A colleague's draft has rough edges, so you read it for meaning. A machine's draft has none, so you read it for typos. That is the reflex you use on something that has already been past three people. Good reflex, wrong document.

Everybody thinks somebody further up already looked

This is where one person's slip becomes the company's problem. The analyst thinks the manager will check it. The manager thinks the analyst already did. Both are being sensible. Four people each doing a light pass is a document nobody read. That is the Deloitte story, and it is not really a story about Deloitte.

Oversight is a level question too

Six questions, three minutes, no name attached. Your level, what it costs you, and what one step up changes.

Do the Culture Level scan

Do this first

Five lines. Twenty minutes. Line four is the one that stings.

Take three things your team sent out last month with a lot of AI in them. A proposal, an analysis, an answer to a client. Fill this in about those three, on your own, before you tell anyone you are doing it.

Verification check Your team, last month
  1. The three things.Name them. Real documents that left the building, not categories.
  2. Who looked at each one.One name per item. If the answer is a team or a process, write that down, because that is the answer.
  3. What did they actually look at.Go and ask. It is usually narrower than either of you thinks, and usually about tone and typos.
  4. Which of the three would have got through with a made-up fact in it.A believable number, a source that does not exist, a date that is slightly off. One at a time, and be honest.
  5. What changes on Monday.One thing, small enough that it actually happens. One named checker on one kind of document beats a policy about all of them.

Your answers stay in this browser and are not sent anywhere.

There is a harder version I use with clients. Take a document that has already gone out, put a small error into a copy, and hand it back to the same reviewer with a decent reason to look again. What comes back tells you more than any policy will.

Level by level

Checking survives at Level 3, because that is where it is written down.

In the five AI Culture Levels we use with clients, Level 3 is where the way of working gets written down and the good method becomes the normal method. Checking is one of the first things that goes in there. That is why this is a Level 3 question and not a training one.

01Campfire60%
02Wild West25%
03Blueprint10%Checking gets written down here
04Engine4%
05Ecosystem1%
Share of companies per level. At Level 2 the checking depends on who happens to be careful. At Level 3 it is part of the method, which is where it starts surviving a busy week.3

At Level 2, Wild West, everyone works their own way. Some people check properly and some do not, and the difference is character rather than policy. It holds up until the first real deadline. Then the second one. Somewhere in there it goes, and there was never a meeting about it.

Level 3, Blueprint, is where somebody writes down what checked means for this kind of document. Which claims get looked up, by whom, and what they open to do it. About one in ten companies is there. Written out like that it sounds like paperwork. In practice it is two sentences per document type.

One thing before you place yourself. Most companies sit on different levels per department. Finance usually checks out of habit, marketing usually does not, and it goes first wherever the output is hardest to quickly prove wrong.

What a measurement shows that a policy does not

An Operating Profile in use. Personality type and AI level in one profile, with the agents that fit it.

A policy says output has to be reviewed. It cannot tell you whether the reviewer can tell a good answer from one that only looks finished, and that is the whole difference between a check and a signature. Somebody further along knows where these tools are confidently wrong and goes straight there. Somebody at the start reads it as a finished document, because that is what it looks like.

So the measuring happens per person. How somebody thinks and works, and their AI level. That shows you who your real checkers are, and it is often not the people the process appointed.

About emaho

emaho measures one Operating Profile per person: personality type and AI level in a single profile. On that we build a personal set of AI agents that fit how that person works, inside the tools they already use. Fifteen minutes to complete, first profile free, built for companies between 20 and 500 people.

Fifteen minutes per person. No credit card, no strings.

What actually helps

An hour is enough to start.

Decide which claims get checked, per kind of document. Not everything, because everything means nothing. Numbers that leave the company, anything with a source under it, anything a client could act on. Two sentences per document type, written where people already work and not in a policy folder.

Then put a name on it per document, the way you would with a signature. One person, and their job is the claims, not the writing. When the checking belongs to the process instead of a person, everybody does a light pass and it goes out unread.

Make it cheap enough to survive a busy Thursday. Four clicks to check a source and it happens. Fifteen and it stops. Keep the source next to the claim while the work is being made, so the reviewer never has to go hunting. Most of what I run into is friction, not attitude.

And test it twice a year. Put a believable error into a document on its way through and see where it stops. Say up front that you are going to do this at some point, so nobody feels tricked, then do it without warning. Until you know what your process actually catches, the rest of this list is a good intention.

Ran the card? Tell me about line four

Which of the three would have got through with a made-up fact in it, and what that turned out to be about. I will tell you what I usually see and what I would fix first. I reply myself. Nothing else happens unless you want it to.

Message me on WhatsApp

What it costs

On the day the machine is wrong, nobody is looking.

The visible cost is the one document that gets caught outside your own walls. Deloitte paid back 97,000 on a 440,000 contract, so the money was the small part. What it really cost them is every other report they have sitting with a client right now, because anyone who read that story went and looked at their own copy differently.2

Under that sits the drift, and nobody sends an invoice for that. Your people's sense of what a good answer looks like slides towards whatever the tool says most confidently, and they take it into work the tool never touched. In the research that happened inside a single session, to people who were certain it was not happening to them.1

Then the reversal, which is the expensive one. Get burned once and the reflex is to have everything checked by everyone, which is slower than not using the tools at all. Six months later somebody quietly stops because there is no time, and you are back at the start with a policy on the wall that says otherwise. The way out is picking what actually needs checking before an incident picks it for you.

Questions

Automation bias, asked and answered

What is automation bias?
Automation bias is when you stop checking what a machine gives you. Two halves to it: taking a wrong answer because the machine said it, and doing nothing because the machine did not flag anything. Moshe Glickman and Tali Sharot at University College London published a study in Nature Human Behaviour in December 2024 showing that people working next to a slightly skewed algorithm shifted their own judgement, while saying they had not been influenced.
Why do people stop checking what AI produces?
Because nine times out of ten you check and it was fine, so checking feels like wasted time. Machine output also arrives with no rough edges, so people read it for typos instead of meaning. And in a review with several layers, each layer assumes an earlier one looked properly. Four people doing a light pass is a document nobody read.
How do you find out whether the checking in your company still happens?
Put a plausible error into a document on its way through and see where it stops. Then read the organisation around it: emaho reads companies on five AI Culture Levels, and checking survives from Level 3, where the method is written down. About one in ten organisations is there. The scan takes three minutes.
Does working with AI actually change how people judge things?
Yes, and fast. In the Glickman and Sharot study of 1,401 people, those judging on their own said sadder 53 times in 100, where exactly half the sets were sadder. A model trained on those answers said sadder 65 times in 100. New people working with that model started at 50 and finished the session at 61, on their own answers, given before they saw the model's.
Is a biased colleague as risky as a biased AI?
The same study ran that comparison and found nothing moved when the partner was human. People changed their answer 32.7% of the time when an AI disagreed, against 11.3% when a colleague did, and the human version produced no drift at any point. A machine's answer carries more weight than a colleague's.
How do you know which people can spot a confident wrong answer?
Not from the org chart. emaho measures one Operating Profile per person, personality type and AI level in a single profile, which tells you who knows where these tools are confidently wrong and goes straight there, and who reads machine output as a finished document. Those are your real checkers, and often not the appointed ones.
What happened with the Deloitte report in Australia?
An Australian government department hired Deloitte to review the computer system that automatically hands out penalties in the welfare programme. Contract worth about 440,000 Australian dollars. The report cited studies that do not exist and quoted a court judgment that was never written. Chris Rudge, a researcher at the University of Sydney, found it, reading out of curiosity. Deloitte said it had used AI in places, published a corrected version in October 2025 and paid back just over 97,000.
Should one person be responsible for checking AI output?
Yes, per document rather than per process. When the checking belongs to a process, everybody does a light pass and assumes somebody else went deeper. One named person, whose job is the claims and not the writing, is what turns a review into a check instead of a signature.
Where can I read about the other AI challenges around control?
This is one of 25 AI challenges emaho documents. Automation bias sits next to agents nobody owns and advice that is right in general and wrong here.
How does this relate to AI maturity?
It is a Level 3 question in the five emaho AI Culture Levels. At Level 2, Wild West, the checking depends on who happens to be careful. At Level 3, Blueprint, the way of working is written down, so what gets checked and by whom is part of the method. About one in ten companies is at Level 3.
Does AI always make human judgement worse?
No, and the same study shows the opposite too. People who worked with an accurate algorithm got better at the task on their own. It runs in whichever direction the tool leans, which is exactly why somebody has to be checking which way that is.

Your next step

Your real checkers, and you can name them this week.

The card above tells you whether anybody checked. It cannot tell you who in your company can spot a confident wrong answer, and that is the person you want on the documents that matter. That is what we measure, starting with you.

  1. Ask who checked, and then ask howThe second half of that question is where most answers stop.
  2. Find your real checkersThe profiles tell you who reads critically and who forwards whatever looks finished.
  3. Give checking a place in the weekOversight that depends on somebody happening to feel uneasy is not oversight.

First profile free · no credit card · built for companies of 20 to 500 · you decide what your team gets to see

Not ready to put your team in anything yet? Start with the level of the company instead. The Culture Level scan is six questions, three minutes, and asks nothing of you.

Paul Musters

Paul Musters

Fifteen years of leadership and team development in Dutch scale-ups. That practice now sits in software: one Operating Profile per person, with agents that actually fit. He writes these pages from what he runs into with clients, not from a research summary.

LinkedIn · paul@emaho.world · WhatsApp

Sources and numbers used on this page
  1. Moshe Glickman and Tali Sharot, How human–AI feedback loops alter human perceptual, emotional and social judgements, Nature Human Behaviour, volume 9, pages 345 to 359, published online 18 December 2024. 1,401 participants across the experiments. Source of: humans judging alone classified 53.08% of face arrays as sadder where 50% were; a convolutional network trained on those human labels classified 65.33% as sadder; participants working with that network went from a 49.9% baseline to 56.3% overall, and from 50.72% in the first interaction block to 61.44% in the last; participants changed their own answer on 32.72% of trials where the AI disagreed, against 11.27% with a disagreeing human and 3.97% with no partner; human–human interaction produced no learned bias (51.45% against a 50.6% baseline, not significant); participants reported being more influenced by the accurate algorithm than by the biased one; interaction with an accurate algorithm improved participants' independent accuracy.
  2. Deloitte Australia and the Australian Department of Employment and Workplace Relations. Contract executed December 2024, valued at about A$440,000 on AusTender, for an independent assurance review of the IT system automating penalties in Australia's welfare programme. The report contained references to academic works that do not exist and a fabricated quote from a Federal Court judgment; the errors were identified by Chris Rudge of the University of Sydney and first reported by Paul Karp in the Australian Financial Review. Deloitte acknowledged limited use of generative AI, published a corrected version in October 2025 and refunded just over A$97,000, around US$63,000. Reported by CFO Dive, 21 October 2025, and the Associated Press.
  3. emaho AI Culture Levels. Share of organisations per level, calibrated against BCG 2025 and McKinsey 2025. Level 3, Blueprint, holds about one in ten organisations.

Numbers are quoted as published. The three moments in the recognition block are composites drawn from client situations rather than transcripts.