Where does work go when an automation fails, and who is supposed to look at it?
It should go to a visible list with the failure reason, the original data, and a way to retry it after someone fixes the cause. That list needs a named owner and a daily review, or it becomes a landfill. The common anti-pattern is routing failures to an email alias nobody reads, which is the same as discarding them. The metric to watch is the age of the oldest unresolved item.
What the queue has to contain
A failure record that says "error" is useless. To be actionable it needs four things: what the automation was trying to do, the exact data it was trying to do it with, the specific error the downstream system returned, and a control that retries the same operation once the cause is addressed.
The retry control is the part usually missing. Without it, resolving a failure means someone recreating the work by hand in the target system, which they will do inconsistently and stop doing under load.
The organizational half is the hard half
The technical construct is a couple of days of work. Making it a functioning part of the operation is a management decision. Somebody owns the queue by name. It is reviewed on a defined cadence. There is a rule for what happens when items sit unresolved.
Where this consistently works, the review is attached to something that already happens daily, such as the dispatch or service manager's morning routine, rather than being a new meeting. Attaching it to an existing rhythm is why we usually deliver these through daily briefs rather than as a separate portal nobody opens.
Three numbers worth publishing
- Age of the oldest item. The single best health indicator. If it grows past a couple of days, the queue is not being worked.
- Failure rate by automation. Reveals which workflow is generating most of the load, which is usually one badly configured thing rather than general instability.
- Repeat causes. If the same error accounts for a large share of items, that is a bug to fix at the source, not a queue to work harder.
Not everything belongs in a queue
Some failures should interrupt someone immediately rather than wait for review. An emergency dispatch that failed to reach a human is not a queue item. A failed payment posting on a job closing today is not a queue item. The rule of thumb is whether the business consequence expires: if waiting until tomorrow makes the item worthless or harmful, it needs an alert with a person attached, not a list.
Everything else genuinely can wait, and it is worth being deliberate about which is which so the alerts stay meaningful. An alert channel that fires for routine failures gets muted within a month, and then the emergency one gets missed too. This is the same discipline discussed in operational AI, where the failure mode is almost always silence rather than noise.
Topics: error handling · operations · reliability · ownership · monitoring
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.