Why do integrations create duplicate records, and how do you stop it?
Duplicates come from retries after ambiguous failures and from at-least-once event delivery. The write succeeded but the response never got back, so the integration tried again. The fix is idempotency: derive a stable key from the source event, check for it before writing, and store it on the record. Matching later on name and phone is cleanup, not a design.
The ambiguous timeout
Almost every duplicate traces back to one situation. Your integration sends a write. The network drops the response, or the request times out at thirty seconds while the server keeps working. Your code has no way to distinguish "it failed" from "it worked and I did not hear back." It retries, because retrying is what resilient code does. Now there are two records.
You cannot eliminate this by writing more careful code. Distributed systems do not offer an exactly-once guarantee across a network. What you can do is make the second attempt harmless.
What makes a usable idempotency key
- Derived from the source, not generated by you. The source system's record id plus object type is stable across retries. A key you mint at call time is different on every attempt, which defeats the purpose.
- Includes a version when updates matter. Source id plus modified timestamp lets you apply a genuine update while rejecting a replayed one.
- Stored on the row, with a unique constraint. The database should refuse the duplicate. Application-level checks lose races when two workers process the same event at once.
- Sent to the vendor when supported. Some APIs accept an idempotency key header and will return the original result rather than creating a second record.
Fuzzy matching is a cleanup tool, not a strategy
Deduplicating after the fact on name, phone and address is genuinely useful for customer records, where the same person really does appear in several systems with slightly different spellings. It is the wrong tool for the retry problem, because it merges things that are legitimately distinct. Two service calls to the same house in the same week are not a duplicate. A landlord with twelve properties is not twelve duplicates.
Use exact source keys for integration safety and fuzzy matching only for identity resolution, and keep the two mechanisms separate so you can tell which one made a decision.
Why this shows up in the numbers
Duplicates rarely announce themselves. They inflate lead counts modestly, double-count revenue on the jobs that happened to be retried, and skew averages in ways that look like a good week. Any revenue reporting built on top will confidently repeat the error, and an AI summary written from that data will explain the growth in fluent, plausible sentences.
A cheap sanity check: count distinct source ids and compare it to your row count for the same window. If they differ, you have your answer before anyone argues about the dashboard.
Topics: duplicates · idempotency · retries · data quality
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.