Prompt
ChatGPT, I need to build a webhook system for my SaaS app that sends outbound events, verifies inbound signatures, retries failures, and logs deliveries — what should the architecture look like?
Latest observation
A good webhook system usually has 4 core parts:
- Event generation inside your app
- Delivery pipeline for outbound webhooks
- Verification + intake for inbound webhooks
- Observability + retry management for reliability
Here’s a practical architecture that works well for SaaS.
1) High-level architecture
Outbound webhooks
Your app emits domain events like:
user.createdinvoice.paidproject.deleted
These should go into an event/outbox table or event bus first, then a delivery worker sends webhook HTTP requests to customer endpoints.
Flow:
- App transaction commits business data.
- Same transaction stores an event in an outbox.
- A worker picks up the event.
- Worker creates a delivery attempt record.
- Worker sends signed HTTP request to subscriber endpoint.
- Worker records response and schedules retries if needed.
Inbound webhooks
For providers sending events into your SaaS, you expose a webhook receiver endpoint.
Flow:
- Provider sends POST to your endpoint.
- You verify signature, timestamp, and replay protection.
- You store the raw payload and headers.
- You enqueue processing asynchronously.
- Background worker handles business logic and ACKs quickly.
2) Recommended service components
A. Webhook producer inside your app
Responsibilities:
- Detect business events
- Persist event data durably
- Avoid losing events on DB/app crashes
Best practice:
- Use the outbox pattern
- Write business state and event record in the same DB transaction
Tables:
webhook_eventswebhook_subscriptionswebhook_deliverieswebhook_delivery_attempts
B. Delivery worker
Responsibilities:
- Fetch pending deliveries
- Apply per-tenant/subscriber rate limits
- Build signed HTTP request
- Handle retries with exponential backoff
- Mark success/failure
- Respect disabled endpoints
Important behaviors:
- Timeout requests aggressively
- Retry only on transient failures
- Treat 2xx as success
- Record non-2xx responses
C. Queue / job system
Use a queue for asynchronous delivery and inbound processing.
Examples:
- AWS SQS / SNS
- RabbitMQ
- Kafka
- Redis queue / Sidekiq / BullMQ / Celery
You want:
- delayed retries
- dead-letter queues
- visibility into stuck jobs
- at-least-once processing
D. Webhook receiver
Responsibilities:
- Verify authenticity
- Normalize payloads
- Prevent replay attacks
- Enqueue processing
- Return fast
Do not do heavy processing inline.
3) Outbound webhook design
Event model
Use a canonical envelope, for example:
{
"id": "evt_123",
"type": "invoice.paid",
"created_at": "2026-10-05T12:00:00Z",
"tenant_id": "tenant_456",
"data": {
"invoice_id": "inv_789",
"amount": 5000,
"currency": "usd"
}
}
Include:
- unique event id
- event type
- creation timestamp
- tenant/account id
- payload data
Delivery request format
Send:
Content-Type: application/jsonWebhook-Id: evt_123Webhook-Event: invoice.paidWebhook-Timestamp: 1730000000- signature header like
Webhook-Signature: ...
This helps consumers:
- deduplicate
- verify authenticity
- debug delivery issues
Signing outbound webhooks
Use HMAC-SHA256 with a shared secret per subscriber.
Canonical approach:
- Build signing payload from timestamp + raw body.
- Compute HMAC with subscriber secret.
- Send signature + timestamp header.
Example concept:
signed_payload = "{timestamp}.{raw_body}"- signature = HMAC(secret, signed_payload)
Consumer verifies the same way.
Retry policy
Use exponential backoff with jitter.
Example:
- Attempt 1: immediately
- Attempt 2: after 1 min
- Attempt 3: after 5 min
- Attempt 4: after 15 min
- Attempt 5: after 1 hr
- Attempt 6: after 6 hr
- then dead-letter / disable endpoint
Retry on:
- 408
- 429
- 5xx
- network timeouts
- connection errors
Do not retry on:
- 400
- 401/403 due to auth problems unless config says otherwise
- 404 if endpoint is permanently gone
- malformed request errors
Also:
- honor
Retry-Afterwhen present - cap total retry window
Idempotency and deduplication
Webhook delivery is usually at-least-once.
Consumers must dedupe using:
- event ID
- delivery ID
- idempotency key
Your side should also prevent duplicate sends when possible:
- unique event records
- delivery state machine
- transactional queue/outbox
4) Inbound webhook verification
Verification steps
When receiving webhooks from third parties:
- Read the raw body exactly as received.
- Verify
timestampis within allowed skew, e.g. 5 minutes. - Verify HMAC signature using provider secret.
- Check for replay with event ID / nonce cache.
- Store payload and enqueue processing.
Important:
- Do not parse and reserialize JSON before verifying
- Use the exact raw bytes for signature verification
Replay protection
Use one or more:
- event ID uniqueness check
- timestamp tolerance
- nonce cache / Redis set with TTL
- processed event table with unique constraint
Fast ACK pattern
Return 200 OK or 204 No Content quickly after:
- verification
- durable enqueue/store
Then process asynchronously.
5) Data model suggestion
webhook_subscriptions
idtenant_idtarget_urlsecretevent_typesstatusactive/paused/disabledfailure_countcreated_atupdated_at
webhook_events
idtenant_idevent_typepayloadcreated_atdispatched_at
webhook_deliveries
idevent_idsubscription_idstatuspending/success/failed/retrying/deadattempt_countnext_attempt_atlast_response_codelast_errorcreated_atupdated_at
webhook_delivery_attempts
iddelivery_idattempt_numberrequest_headersrequest_body_hashresponse_coderesponse_bodylatency_mserror_typecreated_at
For inbound:
inbound_webhook_events
providerprovider_event_idpayloadheaderssignature_validreceived_atprocessed_at- unique index on
(provider, provider_event_id)
6) Failure handling strategy
Endpoint failures
Classify failures:
Transient
- timeouts
- DNS issues
- 429
- 500/502/503/504
Action:
- retry with backoff
Permanent
- 400 bad request
- invalid signature
- 401/403 if secret wrong
- 404 endpoint removed
Action:
- mark failed, possibly disable subscription after threshold
Subscription health
Track:
- consecutive failures
- average latency
- recent success rate
If a subscription keeps failing:
- pause delivery
- notify customer
- surface in dashboard
- allow manual re-enable
7) Observability and delivery logs
You want a delivery log UI/API with:
- event ID
- subscription ID
- attempt history
- timestamps
- status
- response code
- response body snippet
- latency
- request signature version
- retry schedule
- final outcome
This is crucial for support and customer debugging.
Also emit metrics:
- deliveries attempted
- success rate
- retry rate
- dead-letter count
- latency p50/p95/p99
- per-tenant failure counts
And logs/traces:
- correlate every event with a trace ID
- log delivery attempt IDs
- capture error types
8) Security considerations
Outbound
- Per-subscription secret
- Rotate secrets
- Support multiple active secrets during rotation
- Sign payload + timestamp
- Enforce HTTPS only
- Optional IP allowlist documentation, but don’t rely on it as primary security
Inbound
- Verify exact raw body
- Timestamp tolerance
- Replay prevention
- Secret rotation
- Limit request size
- Strict content-type checks
General
- Rate-limit incoming webhook endpoints
- Validate callback URLs on subscription creation
- Block private/internal networks to prevent SSRF
- Restrict redirects or don’t follow them
- Set short HTTP timeouts
9) Suggested delivery state machine
A simple state machine helps a lot:
pendingin_progressretry_scheduledsucceededfaileddead_lettereddisabled
Transitions:
- pending → in_progress
- in_progress → succeeded
- in_progress → retry_scheduled
- retry_scheduled → in_progress
- retry_scheduled → dead_lettered
- any → disabled if subscription is shut off
Use optimistic locking or row-level locking to avoid duplicate workers sending the same delivery.
10) Scaling pattern
For scale:
- Partition by tenant or subscription
- Use queue sharding if necessary
- Keep delivery workers stateless
- Use bulk fetching for jobs
- Store payloads efficiently; large payloads may go to object storage with pointer in DB
- Use per-tenant concurrency limits so one noisy tenant doesn’t starve others
11) A practical reference architecture
Minimal production-ready setup
- App DB with outbox table
- Queue for deliveries
- Delivery workers
- Webhook receiver API
- Delivery logs in DB
- Metrics + alerting
- Dead-letter queue
More mature setup
- Event bus + outbox relay
- Dedicated webhook dispatch service
- Separate retry scheduler
- Object storage for large payloads
- Admin UI for customer-managed webhook endpoints
- Secret rotation service
- Signature versioning
12) Recommended request/response contract
Outbound request
- POST JSON
- Signed headers
- Unique event ID header
- Timestamp header
Receiver response handling
Treat as success only if:
- 200–299
Anything else:
- log response
- retry if transient
- possibly disable after threshold
13) Common pitfalls to avoid
- Sending webhooks inline inside the business transaction
- Not storing raw payloads/headers
- Verifying signatures after JSON parsing/normalization
- No replay protection
- No idempotency on consumer side
- Unlimited retries forever
- Retrying permanent failures
- Not logging response bodies
- No rate limits or concurrency controls
- Allowing arbitrary callback URLs without SSRF defenses
14) If you want the simplest good design
If you’re starting from scratch, do this:
- Outbox table for generated events
- Queue for delivery jobs
- Worker that signs and sends HTTP requests
- Retry scheduler with backoff + DLQ
- Delivery logs table
- Inbound receiver with HMAC + timestamp verification
- Admin UI for viewing attempts and managing secrets
If you want, I can also provide:
- a database schema,
- a sequence diagram, or
- a reference implementation in Node.js, Python, or Go.