← All posts

Human-in-the-loop for AI coding agents: where you actually belong

8 min readThe Yokka team

Most writing about human-in-the-loop AI is about enterprise workflows: an agent drafts a refund, a person clicks approve. For coding agents, the default version of that is the permission prompt. The agent wants to run a command, and you press yes.

That’s the wrong place for you. You end up approving npm test forty times a day while the decisions that actually matter, like what to build, what to do when the spec is ambiguous, and whether the change is right, happen without you.

This post is about putting yourself somewhere more useful. We think a coding agent needs a person at three checkpoints, plus one alarm for when things go quiet.

Why approving every step doesn’t work

Step-by-step approval feels safe. In practice it fails in three ways.

You stop reading. By the twentieth prompt you’re pressing yes without looking. The prompts are still there, but nobody is actually checking them.

It ties the agent to you. An agent that needs you every two minutes can only work while you’re sitting in front of it. You can’t run three agents in parallel, and you can’t leave one running over lunch.

It checks the wrong thing. “May I run the test suite?” is almost never the risky moment. The risk is in intent and outcome: did the agent understand what you wanted, and did it build that? No individual command answers either question.

Mechanical safety still matters, but it belongs in tooling, not in your attention. Sandboxing, permission modes, allowlists and scoped tokens make sure an agent can’t do certain things. Save your attention for judgment calls.

The three checkpoints

A coding task has three moments where a person adds more than they cost:

  1. Before: the brief. Deciding what to build and what “done” means.
  2. During: the blocking question. A decision only you can make, asked once, answered when you’re free.
  3. After: the review. Checking the outcome against the brief.

Everything between those is the agent’s job.

Checkpoint 1: write a brief the agent can finish

Most agent failures that look like bad code are really bad briefs. “Fix the date bug” makes the agent guess which bug, where, and how you’ll judge the fix.

A brief that works has three parts:

**What:** Week numbers in the report export are off by one for dates in late December.
**Why:** Finance runs year-end reports from this export. Week 53 shows as week 1 of the
wrong year, so totals land in the wrong year.
**Done when:**
- Dates from Dec 28 to Jan 3 get ISO 8601 week numbers.
- A test covers 2026-12-31 and 2027-01-01.
- The CSV header is unchanged.

What tells the agent where to look. Why lets it make sensible calls on the small decisions you didn’t spell out. Done when is the important part: every line is something the agent can check itself and you can check in review.

A good test of whether a card is ready: could someone who missed every conversation about it tell when it’s done? If not, it isn’t ready for an agent either.

On a Yokka board every card has this what, why and done-when brief. If you add checklist items, the agent treats them as acceptance criteria. It keeps them, works through them and ticks each off, and when it completes the card, the reply lists any still open.

Checkpoint 2: questions that don’t block you

Even with a good brief, agents hit real decisions: two reasonable readings of the spec, a migration that can’t be undone, a change that turns out bigger than planned. That’s when the agent should ask, and how it asks matters.

A question in the terminal blocks you both. The agent waits, and you only see the question if you happen to be looking at that terminal. With three sessions running in parallel, questions sit unanswered for an hour.

An asynchronous question works better. The agent asks once, parks the work and flags it for you. You answer when you get to it, and it picks up where it left off. That’s how request_input works on a Yokka board: the card moves to Needs you and is flagged, you reply on the card, and the agent reads your reply and carries on.

“When you get to it” works best when you don’t have to be at your desk. On Pro, an agent’s question sends a push to your phone, whichever agent asked: it goes to the person whose token the agent uses. If you start agents with Yokka’s runner, a permission prompt does too, and tapping the push opens the agent’s own session (the Claude app for Claude Code), where you answer. From the card you can pause, resume or stop the run.

Teach agents when to ask

Agents need rules for when to ask, or they either ask about everything or about nothing. Put yours in CLAUDE.md or AGENTS.md:

## When to ask me
Ask (one question, with the options and the one you'd pick) before you:
- change a database schema or delete data
- change a public API, a URL, or anything another team depends on
- go beyond the card's brief, or the change grows past ~300 lines
- touch auth, payments or permissions
- pick between two readings of the brief that lead to different code
Don't ask about naming, formatting, file layout or which library to use for small
things. Follow the codebase's conventions, pick, and mention it in your summary.

Notice the format: one question, the options, and the agent’s own recommendation. You can answer that in five seconds from your phone, without reading the whole diff.

Checkpoint 3: review the outcome, not the steps

When the agent says it’s done, check the result against the brief:

  • Does every done when line hold?
  • Did it stay in scope? Unrequested refactors are the most common way a small card becomes a risky PR.
  • Do the tests test the behavior, or only the implementation?

Two things make review manageable when several agents are shipping:

  • A summary with proof. Ask agents to finish with a one-line summary and the PR link. On Yokka, complete_card takes both, and the card keeps the whole history: claim, plan, progress and questions.
  • A review lane. In Yokka’s default Bugs swimlane, a completed card doesn’t go to Fixed. Flow rules send it to Verify, so a person checks every fix before it closes. You can add the same step to any swimlane.

Reviewing becomes the bottleneck quickly once agents are fast. A WIP limit on the working lane warns you when work is piling up faster than you can review it: the lane’s count turns amber, and agents see the limit on the board.

The alarm: agents that go quiet

The failure people rarely plan for is silence. The agent crashed, the laptop went to sleep, or the session hit a rate limit. The card still says In progress, and nobody is working on it.

You need two things. The agent should give work back explicitly when it can’t finish, rather than just stopping. On Yokka that’s release_card with a one-line reason, and the card goes back to be picked up again. And something should notice silence: if the agent holding a card goes quiet for longer than a threshold (30 minutes by default), the card is flagged as a problem.

Mechanical safety is a separate layer

None of the checkpoints above replace guardrails. They sit on top of them:

  • Permission modes and sandboxing in your agent limit what it can run on your machine.
  • Scoped tokens limit what it can do elsewhere. Give a planning agent a read-only token. Give a worker access to one project, with an expiry. By default only people can delete cards or change the board’s layout. See Agent access.
  • Branch protection means an agent’s PR still needs a person to approve the merge.
  • Who can start agents on your machine. If agents start from the board, the thing that starts them should be locked down too. Yokka’s runner is open source, connects outward only and can only be asked to start a run for a card, never sent a command or a path. Every run uses a permission mode you picked and never skips permissions, and only you can start runs on your machine: no teammate, not even an admin.

Guardrails decide what an agent can do. The checkpoints are where you make the judgment calls.

Putting it together on a board

Each checkpoint maps to a lane. Here’s Yokka’s default Bugs swimlane:

Lane Checkpoint Who acts
Triage Brief is written and ready You
Fixing Agent works, reports progress Agent
Needs you A blocking question You
Verify Review against done-when You
Fixed Done Nobody

The board moves cards between these lanes on its own as the agent claims, asks and completes, so you never have to drag a card to see where things stand.

How to tell if your loop is working

Count a few things for a couple of weeks:

  • Questions per card. Lots of questions mean the briefs are vague. Zero across many cards usually means the agent is guessing.
  • Time in Needs you. If cards wait hours for an answer, your response time is the bottleneck, not the agent.
  • Cards sent back from review. Frequent rejections usually point to done when lines that are too loose to check.
  • Released and stale cards. These often mean the cards are too big for one session.

Adjust the briefs and the rules, not the number of approvals.

FAQ

Do AI coding agents need human approval?

Yes, but not for every step. Approve the plan (the brief), answer blocking questions, and review the result. Let permission modes, sandboxing and scoped tokens handle safety at the command level.

How do I know when Claude Code needs my input?

On one machine, claude agents groups background sessions that are waiting on you under Needs input, and a Notification hook can alert you. Across machines and teammates, a board like Yokka flags the card and moves it to Needs you when the agent calls request_input. On Pro, you also get a push on your phone.

What makes a ticket ready for an AI agent?

A clear what, a why, and done-when criteria the agent can check itself. If someone who missed every discussion couldn’t tell when the ticket is done, it isn’t ready.

Should an AI agent merge its own pull requests?

For most teams, no. Let the agent open the PR and complete the card with a link, and keep a person on the merge, especially for bug fixes, which is why Yokka’s Bugs swimlane sends completed cards to Verify.

How do I stop agents asking too many questions?

Write down when to ask and when not to, and require every question to come with options and the agent’s recommendation. If one kind of question keeps coming up, answer it once in CLAUDE.md.