How does an integration know that two records are the same customer?
By comparing a small set of stable identifiers - phone, email, service address - and applying rules about which combinations count as a match. None of it is certain. A sound setup produces three outcomes rather than two: confident match, confident non-match, and a review queue. Systems that force every pair into match or no-match will both merge different households and duplicate the same one.
The identifiers, ranked by how much weight they can carry
- Service address. In home services this is the strongest key available, because the work happens at a property and the property does not move. It fails when tenants change.
- Mobile number. Strong, but numbers get reassigned, and households share them. A landline on a rental property may belong to three different customers over five years.
- Email. Stable and precise when present, frequently absent or shared across a household, and often entered as a placeholder by whoever took the call.
- Name. The weakest. Spelling variation, nicknames, married names and common surnames make it useful only as a tiebreaker.
Home services has a structural advantage and a specific trap
Because service is delivered to a location, address-based matching works better here than in most industries. Normalizing addresses - unit designations, abbreviations, directional prefixes - does most of the work, and equipment history attached to a property is genuinely valuable across owners.
The trap is conflating the property with the person. When a house sells, the equipment history should follow the property and the billing, contact and communication history must not. Systems that key everything to the address will cheerfully text the previous owner about their water heater. Keeping property identity and customer identity as separate linked concepts is a design decision worth making early, and it shapes what customer intelligence can honestly report.
Match rate is a number you should be given
Any system doing this work knows what fraction of records it could resolve confidently. That number belongs in your reporting, not buried in a log. If you are told matching is handled and no one can state a rate or show you the unresolved queue, matching is being asserted rather than measured.
Watch the rate over time. A sudden drop usually means an upstream change - a form stopped collecting phone numbers, a field got renamed - long before anyone notices a downstream report looks wrong.
False merges cost more than duplicates
The two errors are not symmetric. A duplicate record is annoying and reversible: you merge it later. A false merge blends two real customers' histories, and unwinding it is often impossible because subsequent activity attached to the merged record cannot be reliably split apart.
That asymmetry should bias every threshold you set. When confidence is middling, queue it for a human rather than merging optimistically. In practice a short daily review queue is far cheaper than one bad merge in a billing system, which is why we build the queue before we build the automation in any integration project.
Topics: identity resolution · matching · customer data · data quality
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.