Human + AI operations

Put machines around your people before you take people out of the loop

The version of this that fails starts by removing judgement and discovers which parts were load-bearing afterwards, in production, with customers. The version that works removes the retyping, the searching and the re-explaining first — and leaves the judgement where it already was.

The gap between the pitch and the floor

Somebody senior has read that this category eliminates roles. Somebody who actually runs the operation knows that the roles contain a dozen unwritten judgements a week that nobody has ever documented, precisely because they are judgements. Both are partly right, and the argument between them has stalled your programme for two quarters.

Meanwhile the people doing the work spend a substantial part of their day on things that require no judgement at all — locating information that exists, moving it between systems that will not talk, re-explaining context to whoever picks the case up next, and writing summaries of what just happened.

That is the part worth automating, and it is almost never the part the pitch is about, because it does not produce a headcount slide.

Nobody has separated the judgement from the mechanics

Until you can say which decisions genuinely require a person, every automation proposal is negotiating in the dark and every objection is unfalsifiable.

The useful question is not "can AI do this job." It is "which parts of this job are decisions, which are mechanics, and which are decisions that have merely never been written down." Those three need different treatment and almost no organisation has separated them.

This is built around that separation. Mechanics move first: retrieval, summarisation, classification, drafting, updating a second system, initiating a follow-up. Decisions stay with a person by default and become delegated only where an explicit authority grant says so — with a ceiling, a budget and an approval threshold attached.

And the boundary is instrumented, so the conversation stops being philosophical. When a category of escalation happens every time, that is a decision that turned out to be mechanical. When an automated action is overridden consistently, that is a judgement the delegation should not have covered. Both are visible rather than argued.

What moves, and how you would know

Time spent on mechanics rather than judgement — measured by task-time split before and after, sampled the same way both times.

How often an automated action is overridden — measured by override rate by action class — the signal that a delegation was drawn wrong.

Escalations that were policy-required versus capability gaps — measured by escalation reason class, which is how a boundary earns its next change.

How long a new person takes to become useful — measured by ramp to team median by cohort, against cohorts before the change.

Whether tenure stops being the only source of institutional judgement — measured by variance in outcome quality between newest and longest-serving staff.

a productivity percentage. Productivity claims in this category are usually measured on the tasks that were automated rather than on the whole job, which is how a real number and a misleading one look identical.

It works inside the permissions your people already have

An assistance layer that needs broader access than the person it assists is a privilege-escalation surface wearing a productivity story. Here the platform operates inside the acting person’s own entitlements, so it cannot surface something they were not already permitted to see.

It joins the systems your people work in rather than becoming another window beside them, because a tool that reduces alt-tabbing by adding a tab has not reduced anything.

Identity comes from your identity provider, and the entitlement model is theirs rather than a second one maintained in parallel.

The ceiling is yours, and it is enforced rather than promised

Supervised autonomy means a specific thing here: an authority ceiling, an approval threshold, a budget, and explicit paths to pause, abstain, escalate, refuse or stop. A grant can be narrowed at any time without renegotiating anything.

Where a control is designed rather than independently attested, the page marks it. That distinction matters more on this page than most, because "human in the loop" is a phrase every vendor uses and almost none of them enforce structurally.

Security, and the review nobody puts on the list

Information security will ask about entitlements and data movement, and those answers are above. The review that gets missed is the one about employees: what is monitored, what a supervisor can see, and what happens to the record of somebody overriding the machine.

Those are policy settings rather than vendor defaults, and they should be agreed with whoever represents your people before a queue is touched.

Start with the mechanics, and let the boundary argue for itself

One team, and only the parts of their work that require no judgement — retrieval, summarisation, drafting, classification, updating a second system.

Automate nothing that involves a decision in the first phase, even where you suspect it could be. The point is to establish what the mechanics were actually costing, and that number is contaminated the moment a judgement is delegated alongside them.

Then let the instrumentation tell you where the boundary is wrong. A category escalating every single time was mechanical after all. An automated action overridden consistently was a judgement the grant should not have covered. Both are better than an argument.

Expand the ceiling on that evidence rather than on a roadmap date, and make it reversible — a delegation that cannot be withdrawn quietly becomes permanent whether or not it is working.

Questions buyers actually ask

Our executives want a headcount number and this page does not give them one.

Deliberately. A headcount number produced before the mechanics have been measured is a guess with a spreadsheet around it, and the floor will detect that it was decided in advance. Run the first phase, get the real time-split, and then have the conversation with a number that survives being questioned.

Every vendor says "human in the loop" and it means nothing.

Usually it means the interface has a confirm button. Ask a harder question: is the authority ceiling enforced at execution or implemented as a UI step, can a grant be narrowed without a contract change, and is an override recorded as evidence about the boundary or discarded. Those three separate the structural version from the reassuring one.

If we instrument overrides, our people will feel surveilled and stop overriding.

A real risk, and it is a design decision rather than an inevitability. Overrides can be recorded as evidence about the boundary without being attributed to an individual in supervisor-visible reporting, and what is visible to whom should be agreed with whoever represents your people before a queue is touched. If that agreement is not reachable, the instrumentation is not worth what it costs you.

The judgement in our work genuinely cannot be documented.

Some of it genuinely cannot, and this architecture assumes that rather than arguing with it — which is why judgement stays with a person by default and moves only on an explicit grant. What the instrumentation does reveal is the subset that turned out to be habit rather than judgement, and that subset is usually larger than the operator expects and smaller than the executive hopes.