Future Skills and the Future of Work

What AI Agents Mean for Work: Why a Human in the Loop Is Not Enough

AI agents now do multi-step work with nobody watching. The decision facing leaders this term is not which agent to buy, but which work may run unattended, and whether "a human checks it" is a safeguard or a signature.

A blue paper pull cord handle hanging from a black cord and ceiling mount, with black motion arcs beside it, standing for the ability to stop work that is running on its own.

In brief

An AI agent does not just draft work, it carries it out across several steps with nobody watching, so the leadership decision is which work may run unattended. Research published in August 2026 by Mitchell, Ghosh and Passi found that overseeing agents degrades the overseer through approval fatigue, so "a human checks it" is not automatically a safeguard. Dan Fitzpatrick's Loop Test asks three questions before accepting it: would that person notice a wrong result, would they notice in time, and could the work be put back. Every task then sits in one of three places: hand it over, hold the send, or keep it human.

Somewhere in your organization this term, a piece of software will pick up a job, work through six or seven steps, and finish it, and nobody will watch it happen. Not because a leadership team decided that. Because the feature arrived switched on inside a platform you already use.

That is what an AI agent is, once you strip the vendor language off it: not a smarter chatbot, but software that carries out a sequence of actions toward a goal instead of handing you a draft and stopping. The change is not capability. It is attendance.

So the decision in front of you this term is not which agent to buy. It is which work may run without a person watching, and what has to be true before you say yes. Most leadership teams answer that in four words: a human checks it. Research published in August 2026 argues that those four words, on their own, are not a safeguard. They are a hope.

What an AI agent actually changes

An agent changes who is in the room while the work happens.

A chatbot produces something and waits. You read it, you fix it, you send it. An agent takes the goal and the permissions, and then reads, writes, decides and acts across several steps before anyone sees the result. Two examples, both reported by Government Technology on July 22, 2026: at Johnston Community College, agents sit inside the enrollment workflow and look at student records to move applications from first inquiry to completion, with a human verifying the final numbers. At Peninsula School District, agents handling reading coaching and literacy screening were being made available to all staff this fall.

This is arriving faster than most governance is. Microsoft's 2026 Work Trend Index, published on May 5, 2026, reported fifteen times year-on-year growth in active agents across Microsoft 365, and eighteen times in large enterprises. The same report found that only 26 percent of the AI users it surveyed said their leadership was clearly and consistently aligned on AI. Agents are multiplying inside organizations whose leaders have not yet agreed on a position.

Why "a human in the loop" is not automatically a safeguard

Because the job of watching an agent wears down the person doing the watching.

That is the argument in AI Agents Push Humans Out of the Loop, published on August 24, 2026 by Margaret Mitchell, Avijit Ghosh and Samir Passi of Hugging Face and Data & Society. Their case is uncomfortable and, I think, correct. Systems that require human oversight are currently built in ways that erode the capacities oversight depends on. Interfaces hand reviewers more output than anyone can meaningfully read. What follows is what the authors call approval fatigue: people start clicking approve as a reflex rather than a judgment. Worse, the systems learn from those approvals, so they are rewarded for "producing confident summaries, reducing friction, and surfacing oversimplified plans." The loop optimizes against the scrutiny it exists to provide. Oversight, in their phrase, degrades the overseer.

Sitting alongside that is a finding from Princeton about the machines themselves. In Towards a Science of AI Agent Reliability, published on February 18, 2026, Stephan Rabanser, Sayash Kapoor, Arvind Narayanan and colleagues separate an agent's accuracy from its reliability, and report that "reliability gains lag noticeably behind capability progress." Agents have become markedly more capable over eighteen months. They have not become correspondingly consistent, which is why the authors measure how far an agent's outcomes and its route to them vary across repeated runs of the same task.

Put those two findings beside each other and you get the practical problem. An agent that is wrong in the same way every time is annoying but supervisable, because a person learns where to look. An agent that is right most of the time and differently wrong the rest of it is what a tired reviewer is least equipped to catch. That is the combination arriving in your organization this term.

The Loop Test

The Loop Test is three questions I ask before accepting "a human checks it" as an answer: Would that person notice a wrong result without being told to look for it? Would they notice in time to stop it mattering? And if they missed it, could the work be put back? Oversight that survives all three is a safeguard. Oversight that fails any of them is a signature.

Would they notice? Not could they, if they read every line. Would they, at four o'clock on a Thursday, with the rest of the job still to do. If the answer depends on someone reading output at volume, you have built approval fatigue into your own controls. A reviewer who checks exceptions and samples will notice. A reviewer asked to read everything will notice nothing.

Would they notice in time? A wrong predicted grade caught in the same afternoon is an error. The same error caught after it has gone to a parent is an incident. Ask where the work lands and how long it sits there before it becomes real to somebody outside the building. That interval, not the accuracy rate, is your actual margin.

Could it be put back? A draft can be rewritten. An email cannot be unsent, a record written into a student information system is a job to unpick, and a message to a family cannot be recalled at all. Reversibility is the simplest safety feature you own, and the one most often given away to save a step.

This is my suggested way of thinking about agent oversight, not an established standard. It is deliberately narrow. It says nothing about whether a decision is yours to make, which is a separate question I have written about in which decisions AI should never make. The Loop Test is about work, and about whether anyone would notice if the work went wrong.

Where each piece of work should sit

Every task an agent could do belongs in one of three places, and naming the place is the decision.

Placement What it means When it fits
Hand it over The agent runs, unattended, and the output is used Wrong results are visible, quickly, and can be undone
Hold the send The agent does the work, a named person releases it Wrong results are recoverable only before they leave the building
Keep it human The agent does not touch it The judgment is the work, or the harm is not reversible

Most organizations run into trouble because they have only two settings, on and off, and then find that "on" quietly means "hand it over." Hold the send is the setting that is missing, and it is the one that does most of the work. It takes a few seconds an item and keeps a name attached to everything that leaves the building.

What I Tell Leadership Teams

The mistake I see most often is that leadership teams answer "who checks it?" with a role rather than a person.

Across the leadership teams I work with, in schools and in organizations outside education, the conversation usually runs the same way. Someone says the office team will review it. I ask which member of the office team, on which day, looking at what. The room goes quiet, because the honest answer is that everyone assumed someone else was reading it. That is not carelessness. It is what happens when oversight is assigned rather than designed.

The teams that get this right do one small, specific thing: they stop asking a person to read the output and start asking the agent to surface its own uncertainty. Flag what you were unsure about. List what you changed. Show me the three you would not bet on. A reviewer given twelve exceptions does a proper job. The same reviewer given four hundred approvals does not, and the research above explains why.

I write for Forbes about AI and work, and I spend most weeks in rooms with leaders whose software is further along than their policies are. That gap, between what an organization's tools can now do on their own and what its leaders have agreed they may do, is the widest I have seen it.

What agents do to jobs, and what to say to staff

They take tasks, not jobs, and your staff are less opposed to that than you may fear.

The most useful evidence here is the Stanford audit Future of Work with AI Agents, by Yijia Shao, Humishka Zope, Yucheng Jiang, Jiaxin Pei, David Nguyen, Erik Brynjolfsson and Diyi Yang, which surveyed 1,500 workers across 104 occupations and 844 tasks and was last revised on February 1, 2026. Workers wanted AI agents to take on 46.1 percent of the tasks they were asked about, most commonly to free up time for higher value work, cited by 69.4 percent. Where they resisted, the leading reason was not fear of replacement at 23 percent. It was lack of trust, at 45 percent. And on the researchers' Human Agency Scale, the preference workers expressed most often was equal partnership rather than full automation, dominant in 47 of the 104 occupations.

Read that as a leader and the staff conversation changes shape. The question your team is really asking is not "will this replace me." It is "will this be trusted with something I will be blamed for." Trust is a leadership answer, not a training answer, and it is built the same way AI confidence is built: permission, practice and protection.

There is a second consequence, and it is the one I would put in front of a governing board. When software does the doing, the scarce skill becomes noticing, which is not the skill most of your people were hired or trained for, and it belongs to the same family as the other skills that get scarcer as AI improves. Nobody gets better at catching a subtle error by sitting through a launch webinar. They get better by being given a small number of things to check and being told their catch rate matters.

What to do this term

Start by finding out what is already running, then give every piece of it a placement and a name.

  1. List the agents already running. Not the ones you chose. The ones that arrived switched on in tools you already use.
  2. Run the Loop Test on each one, out loud, in a leadership meeting. Write the three answers down.
  3. Give every agent a placement and a name: hand it over, hold the send, or keep it human, and one person accountable for it.
  4. Change what reviewers are asked to look at. Exceptions and flagged uncertainty, never the full output.
  5. Set a date to look again. Agents change under you when the vendor ships an update, so a decision made in September is a decision about September.

None of this requires you to know how the technology works. It requires you to know what your organization would notice, and how fast. Outsource the doing. Do not outsource the noticing.

If your leadership team is working out which work can now run without a person watching it, this is the kind of question I take into keynotes and leadership sessions with schools and organizations.

Sources and further reading

Dan Fitzpatrick is the founder of The AI Educator, a Forbes contributor, and a keynote speaker and advisor to leaders on AI strategy, readiness and organizational change.

Key takeaways

  • An AI agent differs from a chatbot in attendance, not intelligence: it carries out a sequence of actions toward a goal instead of handing a person a draft and stopping.
  • Microsoft's 2026 Work Trend Index (May 5, 2026) reported fifteen times year-on-year growth in active agents across Microsoft 365, while only 26 percent of AI users said their leadership was clearly and consistently aligned on AI.
  • Mitchell, Ghosh and Passi (arXiv, August 24, 2026) argue that human oversight of agents degrades the overseer, because interfaces produce more output than reviewers can read and approval becomes a reflex.
  • Princeton researchers reported in February 2026 that reliability gains lag noticeably behind capability progress, so agents are becoming more capable faster than they are becoming consistent.
  • The Loop Test is three questions before accepting that a human checks it: would they notice, would they notice in time, and could the work be put back.
  • Every task an agent could do belongs in one of three places: hand it over, hold the send, or keep it human. Most organizations are missing the middle setting.
  • Stanford's WORKBank audit found workers wanted agents to take on 46.1 percent of tasks, and where they resisted the leading reason was lack of trust at 45 percent, ahead of fear of replacement at 23 percent.

Frequently Asked Questions

What is an AI agent, and how is it different from a chatbot?

An AI agent carries out a sequence of actions toward a goal, while a chatbot produces something and waits for you. The difference is attendance rather than intelligence. An agent reads, writes, decides and acts across several steps before a person sees the result, which is why oversight has to be designed rather than assumed.

Is a human in the loop enough oversight for AI agents?

Not on its own. Research published in August 2026 by Margaret Mitchell, Avijit Ghosh and Samir Passi argues that overseeing agents degrades the overseer: interfaces produce more output than anyone can read, and approval becomes a reflex. Oversight counts only when the reviewer would notice, would notice in time, and could undo the work.

Which work can we let an AI agent do unattended?

Work where a wrong result would be visible quickly and could be undone. If a mistake only becomes visible after it reaches a family, a student record or a public channel, the agent should draft and a named person should release it. Where the judgment is the work, the agent should not touch it.

What is approval fatigue and why should leaders care?

Approval fatigue is what happens when reviewers face more agent output than they can meaningfully read and start clicking approve as a reflex. It matters because systems learn from those approvals and are rewarded for confident summaries and less friction, so the review loop steadily gets worse at catching errors.

Will AI agents replace jobs in schools and organizations?

Agents take tasks rather than whole jobs. Stanford's WORKBank audit of 1,500 workers across 104 occupations found they wanted agents to take on 46.1 percent of tasks, most often to free time for higher value work, and the most common preference was equal partnership rather than full automation.

Who should be accountable for an AI agent's work?

A named person, not a team or a role. Leadership teams commonly answer the question with a department, which means everyone assumes someone else is reading the output. Give each agent one accountable person, tell them what to look at, and make it exceptions and flagged uncertainty rather than the full output.

If your leadership team is working through these questions, this is the kind of work I support through AI strategy sessions and advisory work.

Learn more about working together
D
Dan Fitzpatrick

Delivered training to 150K+ educators | Founder of The AI Educator and AI Educator Tools | Forbes Contributor | International Keynote Speaker | 4 x #1 Bestselling Author