Skip to main content

Why does our integration keep creating duplicate records, and how do you stop it?

AI Integration Published September 19, 2026
Short Answer

Usually because the sending system retried a request it never got confirmation for, and the receiving system had no way to recognize the retry as the same event. The fix is an idempotency key: a stable identifier derived from the source record, stored on the destination, and checked before every write. Retries then become updates instead of new rows.

Duplicates come from timeouts, not from carelessness

The classic sequence: your integration sends a create request, the destination creates the record, and then the network drops the response. Your side sees a timeout, assumes failure, and retries. Now there are two records and no error anywhere. This happens on a small fraction of requests, which is why it produces a slow trickle of duplicates rather than an obvious outage.

No amount of careful coding removes this. The only durable fix is making writes safe to repeat.

Where the key lives

An idempotency key is any value that is stable for a given source record and unique across records - typically the source system's own record ID, sometimes combined with an event type. Before writing, you check whether that key already exists on the destination. If it does, you update rather than insert.

You need somewhere to store it. In order of preference: a dedicated external ID field if the platform has one, a custom field you create for this purpose, or a mapping table in your own data store that pairs source IDs with destination IDs. The third option is more work but it is the only one available on platforms with rigid schemas, and it has the advantage of belonging to you.

Three different things all called duplicates

Before choosing a fix, work out which of these you have, because they need different treatment.

  • Retry duplicates. Identical records seconds apart from the same source. Solved with idempotency keys.
  • Cross-system duplicates. The same customer created independently by a web form and by a call. Solved with identity resolution, not with keys.
  • Human duplicates. A dispatcher creating a new customer because search did not find the existing one. Solved with better search and merge tooling, plus periodic detection.

Cleaning up what already exists

Deduplicating history is its own project and it is one-way, so it deserves care. Establish which record survives before you start - usually the oldest, because more history and more references point at it - and decide what happens to child records like jobs, invoices and calls attached to the loser.

Do it in reviewable batches with an export of every proposed merge, not as a single automated sweep. And fix the cause first. Merging ten thousand duplicates while the pipeline still creates new ones is work you will do twice. The same discipline applies whenever an integration layer writes into an operational system, which is why we bias toward resolving identity before writing anything at all.

Topics: duplicates · idempotency · data integrity · APIs

Have a version of this question about your own business?

The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.

Related Answers

People who read this also asked

Browse the Answer Hub →

AI is easy to access. Making it useful is hard.

Bluefrog makes AI useful by integrating it with the way your business actually works — your software, your calls, your customers, your marketing and your revenue.

Technology development since 1997 · AI integration platforms since 2001