What should we monitor on a live integration so we find out before the client does?
Monitor outcomes, not uptime. Four alerts catch nearly everything: data freshness, meaning how long since the newest record arrived; volume anomaly, meaning today's row count against the normal range for that weekday; error rate by endpoint; and schema change detection on the fields you depend on. A green health check on a pipeline that is delivering zero rows is the most common way integrations fail quietly.
Successful responses are not successful syncs
The classic silent failure is an integration that runs on schedule, authenticates, calls the API, receives HTTP 200, and gets an empty result set because a filter parameter is now invalid or a permission was revoked from one object. Every technical indicator is green. No data has moved for eleven days.
So the primary signal must be about data, not about process. The question to alert on is "is the newest record I hold recent enough," and that question has no false positives worth arguing about.
The four alerts, with rules of thumb for thresholds
- Freshness. Alert when the newest record is older than about twice the normal lag plus one full sync cycle. Set it per entity, because calls arrive continuously and invoices arrive in bursts.
- Volume anomaly. Compare against the same weekday, not against yesterday. Service businesses have strong weekly shape, and a Monday-versus-Sunday comparison generates noise that trains people to ignore alerts.
- Error rate by endpoint. Aggregate error rates hide the endpoint that is failing completely while the others succeed. Break it out.
- Schema watch. Record the field names and types you depend on, and alert when a new required field appears, an expected field vanishes, or an enum grows a value you have no mapping for.
Alert the person who can act
An alert routed to a mailbox nobody reads is documentation of a failure, not detection of one. Route each alert to a named owner with a defined next action, and make the alert say what to check. "Job sync freshness exceeded threshold; last record 14 hours ago; check webhook subscription status" is actionable. "Pipeline error" is not.
Keep the volume low enough that alerts stay meaningful. Four well-tuned alerts that fire rarely beat twenty that fire daily.
The reconciliation report is the backstop
Monitoring catches breakage. It does not catch slow drift — a field that started arriving null, a mapping that stopped covering a renamed job type. For that you want a scheduled reconciliation that compares counts and key totals between source and destination and reports the difference in plain language.
Read it even when it is boring. That report is where you see the problem two weeks before it becomes a number somebody disputes in a meeting. We fold it into daily intelligence briefs so the health of the data pipeline is visible next to the numbers it produces.
Topics: monitoring · alerting · data freshness · integration maintenance
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.