Why do webhook events sometimes arrive out of order?
Because most vendors deliver each event independently from a pool of workers, with retries. An event that failed once and retried lands after events created later. Order is not guaranteed unless the vendor explicitly promises it. Handle it by treating every payload as a notification rather than a fact: compare the event's own timestamp or version against what you already stored, and re-fetch the current record before writing anything downstream.
Delivery is per-event, not per-stream
A webhook sender does not maintain one ordered queue per customer. It hands each event to whichever worker is free, and each worker retries on its own schedule. If the first attempt for job.created times out and retries ninety seconds later, job.updated for the same job has already landed. Nothing is broken. The system is behaving exactly as designed.
Some vendors do guarantee ordering per entity, and a few expose a monotonically increasing sequence number. Read the documentation rather than assuming, and assume no ordering when the documentation is silent. Most integration work that goes wrong here went wrong because someone assumed a guarantee that was never offered.
The failure looks like data corruption, not a timing bug
The symptom is a record that moves backwards. A job shows completed in the field service system and scheduled in your reporting. A customer's phone number reverts to an old value overnight. Somebody assumes the sync is dropping data, when in fact the sync applied a stale payload on top of a newer one.
The diagnostic is to compare the event's own creation timestamp against the time you received it. A consistent gap of seconds is normal delivery latency. A scattering of events whose creation times are minutes or hours behind their arrival times is retry traffic, and retry traffic is where reordering lives.
Three defenses that actually work
- Version guard. Store the event timestamp or version with the record and reject any incoming payload older than what you already have. This is a few lines of code and it eliminates most of the problem.
- Fetch on notify. Use the webhook only to learn that something changed, then call the API for the current state of that record. You trade a request for correctness, and the current state is never stale by definition.
- Idempotent upsert. Key every write on the source system's entity id so replays and duplicates collapse into one row instead of creating a second.
When ordering matters and when it does not
Ordering matters wherever a state machine is involved: appointment status, invoice lifecycle, membership activation. Applying those out of order produces reporting that contradicts the operational system, which is the fastest way to lose trust in a dashboard.
It matters much less when the event is only a cache invalidation signal. If every event triggers a re-fetch, order stops being a correctness question and becomes a cost question. That is often the right trade for field service integrations where record volume is modest and accuracy is worth more than request efficiency.
Topics: webhooks · event ordering · idempotency · sync design
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.