How do you decide when AI should act on its own and when it should ask?
Weigh the cost of being wrong against the cost of a human's attention. If an error is cheap and reversible, let it act and log it. If an error touches money, employment or a customer promise, require review. In between, combine a confidence threshold with hard rules: act when confident, escalate when not, and always show the evidence used.
Two axes decide almost every case
Reversibility and visibility. A wrong internal tag is reversible and invisible, so let the system act. A wrong text message to a customer is irreversible and highly visible, so a person releases it. Most decisions sort themselves once you plot them on those two axes rather than debating them abstractly.
A third factor matters at volume: how often you expect the case to occur. Requiring review on something that happens twice a day is fine. Requiring it on something that happens two thousand times a day guarantees the queue will be abandoned.
Calibrate from real errors, not intuition
Pick a threshold by sampling. Run the system on a few hundred real records, have a knowledgeable person label the correct answer, and look at where the model's confidence and correctness actually diverge. You will usually find a band where accuracy drops sharply, and that band is your escalation zone.
Then check staffing. If the threshold produces more escalations than your team can review in a day, the threshold is wrong regardless of what the accuracy numbers say. Capacity is part of the design, not an afterthought.
Confidence scores lie in a predictable way
Models tend to be well behaved on inputs that resemble their examples and confidently wrong on inputs that do not. That means confidence alone is a weak guard against exactly the cases you most want to catch: the unusual record, the garbled audio, the customer who does not fit any pattern.
So pair the score with deterministic rules that escalate regardless of confidence. Missing required data escalates. A value outside an expected range escalates. A record type the system has not seen before escalates. Rules catch the structural surprises; the score catches the ambiguous ones.
Watch the escalation rate as a health metric
Escalation rate is more informative than accuracy for a running system. A sudden rise usually means something upstream changed: a new location, a new service line, a vendor changing a data format. A sudden fall can mean the model became overconfident after a prompt or version change.
Track it, alarm on movement, and review a sample of both the escalated and the auto-approved items. Automatic decisions that nobody ever looks at drift, which is the same reason rubric evaluations are reviewed by a manager before they inform coaching. The point of operational AI is a better-prepared human decision, not an unattended one.
Topics: confidence · escalation · thresholds · governance
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.