Why do AI automations end up creating duplicate records?
Because delivery is at-least-once. Webhooks retry, jobs re-run after a timeout, and a network hiccup can make a successful write look like a failure. If the automation has no idempotency key derived from the event rather than the moment it ran, every retry creates a new record. The fix is to make writes idempotent and to check before creating.
At-least-once is the normal contract
Nearly every webhook provider and queue guarantees that an event will be delivered at least once, not exactly once. Exactly-once delivery is expensive and rare, so the industry pushes the responsibility to the receiver. If your automation assumes each event arrives once, it is built on an assumption the platform never made.
The classic sequence: your endpoint writes a note, then takes too long to respond. The sender sees a timeout, marks it failed, and retries. You now have two notes and a log that says one delivery.
The same class of bug appears without any network problem. A person clicks a button twice, a scheduled job overlaps with its previous run, or a backfill is restarted after a partial failure.
What an idempotency key looks like
A good key is stable across retries and unique per logical action. The event id from the source system is ideal. Failing that, a hash of source system plus record id plus action type plus a period bucket works well.
A bad key is anything generated at run time: a timestamp, a random identifier, or a hash of the model output. That last one bites AI systems specifically, because a language model asked the same question twice may phrase the answer differently, so content hashing does not deduplicate reliably.
Three patterns that prevent duplicates
- Upsert on an external id. Store your key in a custom field on the destination record and update by that key instead of creating blindly.
- A processed-event ledger. Record every key you have handled with its result. Check first, act second. This also gives you replay: reprocessing becomes safe.
- A dedupe window. For actions with no natural key, suppress identical actions on the same record inside a defined time window. Cruder, but it prevents the worst outcomes.
Why AI pipelines get bitten harder than ordinary integrations
Model calls are slow relative to database writes, which pushes total processing time toward the sender's timeout and makes retries far more likely. They are also nondeterministic, so the duplicate is not identical and will not be caught by naive equality checks. And a single triggering event often fans out into several writes, so one retry can produce several duplicates at once.
The remedy is ordering: assign the key, write the ledger entry, then do the expensive work, then mark it complete. Combined with disciplined API integration and the write-back rules described in operational AI practice, this eliminates the class of bug entirely rather than reducing its frequency.
Topics: idempotency · duplicates · webhooks · data quality
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.