Operational notes

Shadow mode first: how an AI system earns autonomy

The version of AI rollout that people fear looks like this: a switch flips on a Monday morning, and from that moment an agent runs part of your operation. Yesterday a person did the work; today software does, and everyone hopes.

Nobody sane signs off on that, and nobody should have to. Autonomy isn't something an AI system starts with. It's something it earns, in stages, with each stage producing the evidence for the next.

Stage one: shadow mode

The system goes live - but its output goes nowhere. It sees the real incoming work, the same leads or invoices or tickets your team handles, and it does the full job: categorizes, extracts, drafts, decides. Then its answer is set aside and compared with what your people actually did.

It touches nothing. It sends nothing. If it's wrong, the cost is zero.

What you get after a few weeks is something no demo can produce: a scorecard against reality. Agreement with your team on 91% of cases, disagreement on 9 - and each disagreement is worth reading, because it's one of three things: the system being wrong (fix it), the system catching something a person missed (it happens more than anyone expects), or two people on your own team handling the same case differently (which nobody knew until now). Shadow mode doesn't just test the system. It feeds the evaluation suite with real cases and real answers, and occasionally teaches you something about your own process.

Stage two: assisted

The system's output starts reaching people - as drafts. It proposes the categorization, pre-fills the record, writes the reply. A person reviews and releases, every time.

The work gets faster immediately: reviewing a prepared answer takes a fraction of producing one. But the real product of this stage is calibration. Every approval and every correction sharpens the picture of where the system is reliable and where it isn't - which case types sail through untouched and which get edited every time. Designed well, this stage is also where your team stops seeing the system as a threat and starts seeing it as the thing that does the boring half of their job.

Stage three: selective autonomy

Now the accumulated record does the deciding. The case types where the system has been consistently right for weeks - high confidence, low stakes - start flowing through without review. Everything else still goes to a person: the low-confidence calls, the novel input, and always, regardless of confidence, the irreversible ones.

The split isn't a guess. It's read directly off the shadow and assisted data: these categories, at this confidence, were right N weeks running. Autonomy is granted per case type, on evidence - not to the system as a whole, on faith.

Stage four: production, permanently supervised

Full operation looks like this: the routine majority runs automatically, the exceptions and the high-stakes calls route to people, monitoring watches the output continuously, and the audit trail records everything. If quality drifts - new input pattern, upstream change, model update - the affected slice drops back a stage. Autonomy stays earned, or it gets revoked.

Note what full operation doesn't look like: zero humans. The ramp doesn't end with people removed; it ends with people repositioned, from doing the routine work to supervising the edge of it.

Why the slow way is the fast way

The staged ramp feels cautious. In practice it's usually faster to real results than the flip-the-switch approach, because the switch approach fails: one early, visible mistake in week one and the team routes around the system, and the project dies with the trust. The ramp produces value from stage two onward, catches its mistakes while they're free, and never asks anyone for faith - at every step, the case for more autonomy is a page of numbers anyone can read.

A rollout plan with no shadow phase is a plan to discover the failure modes in production, with live customers. If someone proposes going straight to live - or you have a system that did, and you're now living with the results - talk to me. Building this ramp is a normal part of how I ship automation.


More notes from production

Get in touch

Got a workflow that needs to survive longer than a quarter?

Send me a short note about what you're trying to build and where it keeps breaking. I'll reply within a working day.