Should an automation stop when something goes wrong, or keep going with a default?
Decide by comparing the cost of a wrong action to the cost of no action. Where a wrong action is expensive and irreversible, such as sending a customer message or writing an invoice, fail closed: stop and queue it. Where doing nothing is the expensive outcome, such as emergency routing, fail open toward a human. The one option you must never ship is failing silently, which is what most systems do by accident.
Three behaviors, only two of them chosen deliberately
Fail closed means the automation halts and does nothing further, ideally putting the work somewhere visible. Fail open means it proceeds using a safe default, which in a business context almost always means handing off to a person. Fail silent means it stops and tells nobody, which is what happens when nobody made a decision.
Every workflow you own is doing one of these three right now. The exercise worth running is going through your list and writing down which, because the third one is more common than anyone expects.
The question that settles it
Ask what a wrong action costs versus what no action costs, and be concrete about who bears each cost. An unnecessary transfer costs a minute of a CSR's time. A missed emergency call costs a customer and possibly more. Those are not comparable, so the emergency path fails open.
Now the other direction. An automated text sent in error is visible to a customer, cannot be recalled, and may violate messaging rules. A text not sent is a follow-up someone can make by hand. So outbound messaging fails closed.
Applied to the common home services workflows
- Emergency call routing. Fail open to a human, always, regardless of what else is broken.
- Outbound customer messaging. Fail closed. Queue it, do not guess.
- Writing to the financial record. Fail closed, with the item in a queue someone works daily.
- Appointment reminders. Fail closed on content errors, but escalate quickly, since a reminder that arrives after the appointment is worse than none.
- Lead capture and routing. Fail open into the general queue. A lead in the wrong bucket is recoverable; a lead in no bucket is lost.
Failing silently is a design defect, not bad luck
Silence happens because the success path was instrumented and the failure path was not. The automation logs when it works. It does not log when it never ran because a token expired or a webhook stopped being delivered.
The countermeasure is watching for absence, not just for errors. If an automation normally fires between twenty and sixty times a day, zero is an alarm condition and so is six hundred. Expected-volume monitoring catches the failures that produce no error at all, which are the ones that run undetected for weeks. More on that in operational AI, and on how it gets wired across systems in AI systems integration.
Topics: automation design · failure modes · risk · human in the loop
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.