How similar do two records have to be before we let the system merge them automatically?
Use two thresholds, not one. Above the high threshold the system merges automatically, below the low threshold it does nothing, and everything between goes to a human review queue. Set the high threshold by sampling candidate pairs at each score band and counting how many are genuinely the same customer. Normalized phone plus normalized address is usually safe; name similarity on its own almost never is.
One threshold forces a bad trade
With a single cutoff you are choosing between creating duplicates and destroying real customers, and you have to make that choice once for every pair in the database. Two thresholds turn it into three outcomes: confident merge, confident leave alone, and the ambiguous middle where a person spends thirty seconds and gets it right.
The middle band is the point. It is small if your scoring is decent, and it is where all the genuinely hard cases live. Sizing that band is also a useful early signal: if it contains a third of your candidate pairs, your scoring signals are too weak to automate anything yet and the first job is better normalization, not a better threshold.
Which signals carry weight
Match scoring works best on normalized components, not raw strings. Normalize before you compare or you will be measuring formatting, not identity.
- Phone, normalized to a single format. Strong, but shared across households more often than people expect.
- Address, parsed into components with the unit separated. Strong when combined with anything else.
- Email, exact match after lowercasing. Strong when present, frequently absent.
- Name similarity. Weak alone. Common surnames in a single service area produce constant false positives.
- Service history overlap. Same equipment serial, same job at the same address, same technician visit. Often the tiebreaker in the review band.
Calibrate on your own data
Vendor default thresholds are tuned on someone else's database. A company operating in one metro with a lot of tract housing and repeated street names has a very different false positive profile than a multi-state operator.
Take a few hundred candidate pairs, bucket them by score, and hand-check a sample from each bucket. Your high threshold is the lowest score band where you still find no false merges in the sample. That number will not match the vendor's suggestion, and yours is the one that counts.
The costs are not symmetric
A missed duplicate is an annoyance. A false merge blends two customers' service history, billing and contact preferences, and unmerging is difficult in most systems and impossible in some. Set the automatic threshold conservatively and let the review queue absorb the uncertainty. Operational systems should be biased toward the recoverable error, which is a principle that applies well beyond deduplication in operational AI, and it is worth building the same asymmetry into any platform that writes to your records.
Topics: deduplication · matching · thresholds · data hygiene
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.