Prompt
I'm looking for a recommendation on how to receive third-party webhooks, dedupe them, and trigger internal jobs reliably
Latest observation
A solid pattern is:
Recommended architecture
Webhook endpoint → durable queue / inbox table → idempotent processor → internal job runner
1) Receive webhooks quickly
Expose a small HTTP endpoint whose only job is to:
- validate auth/signature
- extract a stable event ID
- persist the payload durably
- return
2xxfast
Avoid doing real work in the request thread.
2) Dedupe on ingestion
Use the provider’s event ID if available. Store each incoming webhook in a table like:
providerevent_idreceived_atpayloadstatus
Put a unique constraint on (provider, event_id).
If the same webhook arrives again:
- the insert fails or no-ops
- you return
200 OK - no duplicate downstream job is created
If the provider does not give a good event ID, derive one from a stable hash of the payload plus provider metadata, but native event IDs are much better.
3) Decouple via a queue or outbox
Once stored, enqueue a processing record/job, or use an outbox pattern:
- transactionally insert webhook record + outbox row
- a background worker reads the outbox and dispatches internal jobs
This prevents “stored but not queued” or “queued but not stored” inconsistencies.
4) Make internal jobs idempotent too
Even with dedupe, assume at-least-once delivery everywhere.
Internal job creation/execution should also be idempotent using a business key like:
(provider, event_id)- or
(resource_id, event_type, version)
Add uniqueness constraints or idempotency keys at the job layer as well.
Reliable implementation options
Simple and effective
- Webhook app: receives and validates
- Database: stores incoming events with unique constraint
- Worker: polls new rows and enqueues/executes jobs
Good if throughput is moderate and you want operational simplicity.
More scalable
- Webhook app writes to DB or durable log
- Queue: SQS / PubSub / RabbitMQ / Kafka
- Workers: consume and create internal jobs
Good for higher volume and easier horizontal scaling.
Best-practice transactional pattern
Use:
- Inbox table for received webhooks
- Outbox table for job dispatch
- Idempotent consumers
This is the most robust way to avoid duplicates and lost work.
Important details
Acknowledge only after durability
Return success only after the event is safely recorded.
If you ack before persistence, retries can create duplicates or data loss.
Expect retries and out-of-order delivery
Third-party webhooks often arrive:
- multiple times
- late
- out of order
Design handlers so processing the same event twice is harmless.
Track processing state
Useful statuses:
receivedvalidatedprocessedfaileddead_lettered
This helps with retries and observability.
Add a retry strategy
For transient failures:
- retry with backoff
- cap attempts
- move poison messages to a dead-letter queue/table
A practical baseline
If you want the simplest reliable setup:
- HTTP webhook endpoint validates signature
- Insert into
webhook_eventswith unique(provider, event_id) - Return
200 - Background worker polls unprocessed rows
- Worker creates internal job using same idempotency key
- Mark event processed only after the job is safely queued/executed
If you want a concrete stack suggestion
A common robust combo is:
- API: FastAPI / Rails / Express / Go HTTP server
- DB: Postgres with unique constraints
- Queue: SQS or RabbitMQ
- Worker: Sidekiq / Celery / BullMQ / custom consumer
If you tell me your language/cloud/provider, I can suggest a concrete design and schema.