← Blog
Article

Human-in-the-Loop AI Agents: Where to Put the Gate

Human-in-the-loop AI agents done right: where to put approval gates, how to bind them to actions, plus evals and audit trails. With a real Fortnox example.
12 min readAI AgentsHuman in the LoopGovernanceMCP

By Assim ElHammouti, Founder and Lead Engineer, NordbeamPublished Updated

By Assim ElHammouti, founder and lead engineer, NordbeamPublished Updated

Short answer

What are human-in-the-loop AI agents?

A human-in-the-loop AI agent is an agent that cannot carry out a consequential action until a person has reviewed the exact operation and approved it. Nordbeam, an AI development studio in Gothenburg and Malmö, builds agents this way: reads run freely, hard-to-reverse writes wait for approval, and every decision is logged.

"Human in the loop" is on every AI agent slide, and in most products it means one of two things: a confirmation dialog that everybody clicks through, or a human who reads the output afterwards, when the action is already done. Neither is a control. Both are decoration.

This is the part of building agents we care about most, because it is what decides whether a finance team, an operations team or a compliance officer lets an agent anywhere near a production system. This article is about doing it properly: where the human goes, what they are shown, how you prove the agent is good enough, and what you keep for later.

The examples come from Nordsynk, a hosted Fortnox MCP and finance-operations platform that Nordbeam built and operates, and where these rules are not theory.

What are human-in-the-loop agents?

The phrase has two common meanings, and search results mix them up.

In machine learning, human-in-the-loop usually describes people labelling data and correcting a model so that it improves. That is a training-time process. This article is about the other meaning: a person who sits inside the execution path of a running agent and decides whether a specific action goes ahead.

Within that meaning there are two patterns worth separating:

  • In the loop: the agent proposes, a human approves, and only then does the system act. Nothing happens without a yes.
  • On the loop: the agent acts, and a human monitors and can intervene or roll back. The default is "yes unless stopped".

Neither is better in general. The rule we use is simple: the more expensive a mistake is to undo, the closer the human must be to the moment of execution. Reading data is on the loop at most. Sending an email to a customer, posting a ledger entry, changing a price, deleting a record: in the loop, with an approval before it happens.

In the loopOn the loop
Who acts firstThe agent proposes, a human decidesThe agent acts
Human roleApprove, edit or reject before executionMonitor, intervene or roll back
Use forHard-to-reverse writes: emails, ledger postings, price changes, deletionsReversible or low-cost actions
DefaultNo unless approvedYes unless stopped
Illustration · example data
Agent

Reads and drafts

Invoices, documents, suggestions

Proposal

Exact operation shown

Before and after, in plain language

System

Write executes

Only the approved request

Human approval

Post 3 payment reminders to the ledger?

Approve
  • ✓ evidence linked to every finding
  • ✓ decision and executed operation recorded
The human sits before the write, not after it.

Which actions need a human? Sort every tool by what it can break

Most agent safety problems are decided before a single prompt is written, at the point where you decide which tools the agent gets. OWASP's list of risks for LLM applications names the failure excessive agency, and gives it three root causes: excessive functionality, excessive permissions and excessive autonomy. The third is the one a human gate fixes. The first two you fix by giving the agent less to begin with.

A useful way to work is to classify every tool the agent can call:

ClassExampleGate
ReadFetch invoices, search documentsNone, but scoped to what the user may see
SuggestDraft a reply, propose a ledger entryNone; the output is text, not an effect
Write, reversibleTag a ticket, create a draftLight: confirm, easy undo
Write, hard to reversePost an entry, send an email, pay, deleteExplicit approval of the exact action

Two things follow. First, most of what an agent does is read, and a good design makes that fast and unobstructed, because friction there is wasted. Second, the approval gate belongs on a small, deliberately drawn set of write operations, not sprinkled across everything.

Nordsynk draws the line in the same place. Discovery, reads, suggested actions and writes are separate layers. Reads run through the authenticated Fortnox runtime. Suggestions are something the assistant can present. But any request that mutates Fortnox must carry explicit write intent and pass through an approval flow.

Tip

Avoid open-ended tools

A tool that runs arbitrary shell commands or arbitrary queries cannot be gated meaningfully, because the reviewer cannot tell what will happen. OWASP lists open-ended tools as something to avoid. Narrow tools, with typed arguments, are what make an approval readable.

What should an approval cover? Make it specific, not general

This is where most "human in the loop" designs quietly fail. There are three common shapes of a bad approval:

  1. The blanket approval. "Allow this agent to use your accounting system?" is a consent to a category, not to an action. After the first click, the agent has the keys.
  2. The summary approval. The user approves what the agent says it will do, and the system executes something slightly different, because the model's description and the actual request are separate things.
  3. The unreadable approval. The dialog shows an API payload, so the human clicks through without understanding it.

A good approval has four properties:

  • It is locked to the exact request. What is approved is the concrete operation that will run, and the approval is not valid for a different one. In Nordsynk, a mutating request must pass through an approval flow that is locked to the exact request being executed.
  • It is written in the user's language. Not "POST /vouchers" but what the accounting change is, with enough context to verify it. Nordsynk's approval experience is designed for non-technical users: explain the change, show enough context to check it, and only proceed when the human confirms that exact action.
  • It is enforced by the server, not the client. More on this below.
  • It is a real decision. A person must be able to refuse, and refusal must be as easy as agreement.

Why the server has to enforce it

The Model Context Protocol says, in its tools specification, that for trust and safety there should always be a human in the loop with the ability to deny tool invocations, and that applications should present confirmation prompts. Note the word: should. The protocol leaves the interaction model to the client application, and clients differ. The same specification also says clients must treat tool annotations as untrusted unless they come from trusted servers, which means a "this tool is read-only" hint is not a guarantee.

So if your safety depends on every client in the world showing a good confirmation dialog, you do not have safety. You have a hope. When we built Nordsynk, Claude, ChatGPT, Cursor and desktop clients turned out to have different constraints around authentication, tool discovery and interactive approval, and the product had to stay strict about tenant boundaries and write intent regardless. The rule we draw from that: the server should refuse any write that does not carry a valid approval for that exact request, whatever the client did or did not show.

The protocol itself is documented at modelcontextprotocol.io, and our MCP server development page describes how we build servers with this model.

Illustration · example data

Create 3 payment reminders for overdue customer invoices

Acme AB, Beta Konsult AB, Gamma HB · writes to the accounting system · expires if the request changes

RejectApprove
The approval names the exact action and what it changes, not a category.

What should the human be shown? The evidence

An approval is only as good as the information behind it. Agents summarise, and summaries can be wrong. The antidote is to make every recommendation carry a pointer back to the thing that supports it.

Nordsynk's read-only finance-control preview is built this way. A daily playbook checks payment changes, overdue invoice reminders, delivery readiness, cost movement anomalies and open accounting questions. It never writes to Fortnox. It turns review work into a queue of findings, decisions and feedback, and each finding carries evidence pointers back to the Fortnox or Nordsynk object that supports it. A reviewer does not have to trust the summary; they can open the source record.

The same product remembers human decisions. The design principle is stated in one sentence in the case study: read approved records, show the evidence, remember human decisions, and require an explicit approval before changing Fortnox.

You can apply this to any agent:

  • Every finding links to a source record, not a paraphrase.
  • Every proposed write shows before and after.
  • When the agent is unsure, it says so and routes to a person instead of guessing.

How do you know the agent and the gate are good enough? Evaluate both

A human in the loop does not replace testing. It is the backstop for what testing missed. Without evaluation you cannot tell whether the agent is getting better or worse, and you cannot tell whether the gate is catching errors or passing them.

Anthropic's engineering guidance on building agents is direct about this: run many example inputs, see what mistakes the model makes, and iterate; include stopping conditions such as a maximum number of iterations; and design tools so that mistakes are harder to make (for example, requiring absolute file paths after relative ones caused errors).

In practice, a minimal evaluation setup has:

  • A golden set. A few dozen to a few hundred real tasks with known-good outcomes, including the ugly edge cases. Start small and grow it every time something fails in production.
  • A regression run on every change. Prompt, model, tool description, tool schema: any of them can change behaviour. Re-run the set and compare.
  • Gate metrics. How often is a proposed action approved unchanged, edited, or rejected? A near-100% approval rate on a high-stakes action usually means either the agent is excellent or the reviewers have stopped looking. You need to know which, and the evaluation set is how you find out.
  • Stop conditions. Maximum steps, maximum spend, and an explicit "hand over to a human" outcome.

What should you keep afterwards? An audit trail

If something goes wrong in a system with an approval gate, the questions are always the same: who asked for this, what did the agent see, what did it propose, who approved it, and what was written?

The MCP tools specification recommends that clients log tool usage for audit purposes. We would go further and log it server-side as well, because the server is the only part you control. A useful audit record has:

  • the user, the company or tenant, and the workflow context;
  • the request, the tools called, and their results;
  • the proposal shown to the human and the human's decision;
  • the exact operation executed and its outcome.

An audit trail is also what makes the system improvable. The rejected proposals are your best source of new evaluation cases.

Illustration · example data
Audit log
  1. 10:42user asked: overdue invoices
  2. 10:42read 3 invoices (read-only tools)
  3. 10:43write proposed, exact request locked
  4. 10:44approved by finance
  5. 10:443 reminders created, outcome recorded

User, tenant, request, proposal, decision and executed operation.

Every step from request to write can be reconstructed afterwards.

What goes wrong with human-in-the-loop designs?

  • Approval fatigue. If every action needs a click, people stop reading. Keep the gate for the few writes that matter and make everything else frictionless.
  • The gate in the wrong place. A confirmation after the email is sent is a notification. Put it before.
  • Approve all. Batch approval is fine when the batch is reviewed as a batch, with exceptions highlighted. It is not fine as a way to avoid reading.
  • Automation bias. People over-trust system output. Article 14 of the EU AI Act, which requires human oversight for high-risk AI systems, explicitly asks that the people overseeing them stay aware of the tendency to over-rely on the system's output, that they can override it, and that there is a way to stop the system. Not every agent is a high-risk system under that regulation, and this is not legal advice, but the design list is a good one regardless of whether the law applies to you.
  • Untrusted content steering the agent. Text the agent reads (an email, an invoice, a web page) can contain instructions. Least-privilege tools plus an approval on writes limit the damage when a model is manipulated.

A short checklist

  1. Every tool is classified: read, suggest, reversible write, hard-to-reverse write.
  2. Hard-to-reverse writes require approval of the exact action, enforced on the server.
  3. The approval is written in the user's language and shows the before and after.
  4. Every finding links to evidence.
  5. A golden evaluation set runs on every change, and you track approve, edit and reject rates.
  6. The agent has stop conditions and a route to a human.
  7. You can reconstruct who asked, who approved and what was written.

Where this goes

The direction of travel for AI agents in business is toward more autonomy, and the right way to get there is to earn it action by action: a gate that starts tight, evaluation that shows where the agent is reliable, and an audit trail that proves it. When the data shows that a class of write is approved unchanged nearly every time, and the cost of an error is small, you can loosen that one gate. You can only do that responsibly if you built the measurement first.

If you are working out what this looks like for a specific workflow, our AI Workflow Sprint (from SEK 45,000) maps one workflow with its systems, risks and ROI. The Production Agent Pilot (from SEK 180,000) delivers one valuable workflow with integrations, approvals, evaluations, auditability, rollout and handover. And if you need senior product and engineering leadership around it, a Fractional AI Product Lead works one to three days a week.

What does human in the loop mean for AI agents?
It means a person sits inside the execution path of a running agent and decides whether a specific action goes ahead. The agent proposes, the human approves, and only then does the system act. This differs from the machine-learning sense of the term, where people label data to train a model.
What is the difference between human in the loop and human on the loop?
In the loop means an approval is required before the action. On the loop means the agent acts and a human monitors and can intervene afterwards. Use in the loop for actions that are expensive to undo, and on the loop for actions that are cheap to reverse.
Does human approval make AI agents too slow to be useful?
Not if the gate is in the right place. Most of what an agent does is reading and suggesting, which needs no approval. Only a small set of hard-to-reverse writes should be gated, and the time saved by the agent doing the searching and preparation is where the value comes from.
Can the AI agent approve its own actions?
No. The approval has to come from a person who can see the exact action and who can refuse. It should also be enforced by the server that executes the action, not left to the client application.
Does the Model Context Protocol require human approval?
The [MCP tools specification](https://modelcontextprotocol.io/specification/latest/server/tools) says there should always be a human in the loop with the ability to deny tool invocations, and that applications should present confirmation prompts. It words this as a recommendation and leaves the interaction model to the client, which is why server-side enforcement of approvals matters.
Is human oversight legally required for AI agents?
For high-risk AI systems, Article 14 of the EU AI Act requires human oversight measures such as the ability to understand the system's limits, override its output and stop it. Whether a particular agent is high-risk depends on its use case. This is not legal advice.
How do I know if my approval gate is working?
Measure it. Track how often proposed actions are approved unchanged, edited and rejected, and test with known-bad cases from an evaluation set. An approval rate that never moves is a sign reviewers may have stopped checking.

Designing an agent your team will trust

Tell us which workflow you want to automate and what would go wrong if it were wrong. We will tell you where the gate belongs, and what it takes to build it.

Start the Conversation →

Get in Touch

Which workflow should AI improve first?

Book a practical review of one high-value workflow. You'll speak directly with Assim, Nordbeam's founder and lead engineer.
Email directly
Email
hello@nordbeam.io
Locations
Gothenburg & Malmö, Sweden
Response time
Within 24 hours