Prompt
How do I orchestrate notifications across multiple products and environments?
Latest observation
A good way to orchestrate notifications across multiple products and environments is to treat notifications as a shared platform capability, not as app-specific ad hoc logic.
1) Define a single notification domain model
Create a common model that every product uses:
- Event: something happened
- Notification intent: what should be communicated
- Audience: who should receive it
- Channel: email, SMS, push, Slack, in-app, webhook, etc.
- Priority / severity: info, warning, critical
- Environment: dev, staging, prod
- Tenant / product / region: routing scope
- Template: localized, branded message content
- Delivery policy: retries, suppression, deduplication, quiet hours
This lets all products emit notifications in the same format.
2) Use an event-driven orchestration layer
Have products publish notification-worthy events to a central system, rather than sending notifications directly.
Typical flow:
- Product emits domain event
- Notification orchestrator consumes it
- Orchestrator applies rules, enrichment, and routing
- Messages are sent through the right channel(s)
- Delivery status is tracked centrally
This decouples product logic from delivery logic.
3) Separate routing rules from product code
Keep notification behavior in configuration or policy, not hardcoded.
Examples of routing rules:
- Prod critical alerts → PagerDuty + Slack + email
- Staging errors → Slack only
- Dev notifications → console/log sink or test inbox
- Certain tenants → branded templates and region-specific channels
A rules engine or policy table should decide:
- whether to notify
- whom to notify
- which channel(s) to use
- what template to render
- whether to suppress duplicates
4) Use environment-aware delivery controls
Each environment should have guardrails:
Dev
- Sink notifications to a test mailbox, mock provider, or local log
- Prevent real customer contact
- Add environment labels to every message
Staging
- Route to internal users only
- Use sandbox providers where possible
- Validate templates and integrations
Prod
- Full delivery rules
- Strong monitoring, retries, audit logs, and escalation paths
A common mistake is letting non-prod environments accidentally send real alerts.
5) Centralize identity, preferences, and suppression
Build a shared notification service that knows:
- user notification preferences
- contact methods
- timezone and quiet hours
- opt-in/opt-out rules
- deduplication windows
- escalation policies
This is especially important if multiple products notify the same people.
6) Support multi-product branding and templates
Each product should be able to define:
- brand name
- sender identity
- tone
- templates
- links/CTAs
- localized copy
Use a template system with variables like:
- product_name
- environment
- incident_id
- resource_name
- action_url
That keeps delivery consistent while allowing product-specific messaging.
7) Add observability and auditability
You’ll want end-to-end visibility:
- event received
- rules evaluated
- recipients resolved
- message rendered
- provider accepted/rejected
- delivered/opened/clicked
- failed/retried/dropped
Also keep an audit trail for compliance and debugging.
8) Design for idempotency and deduplication
Notifications often get duplicated due to retries or repeated events.
Use:
- event IDs
- idempotency keys
- dedupe windows
- correlation IDs across products/environments
This prevents alert storms.
9) Build escalation and fallback paths
If a primary channel fails:
- retry with backoff
- fallback to secondary channel
- escalate to on-call
- create a ticket
- send to a backup region/provider
Critical notifications should not depend on one provider.
10) Recommended architecture
A common setup:
- Producers: product services
- Message bus: Kafka, SNS/SQS, Pub/Sub, RabbitMQ
- Notification orchestrator: policy engine + template renderer + router
- Channel adapters: email/SMS/push/Slack/webhook
- State store: delivery logs, preferences, suppression rules
- Admin UI: manage templates, rules, environments, tenants
Example policy
- If
environment=dev, route totest-inbox - If
environment=staging, route tointernal-slack - If
environment=prodandseverity=critical, route topagerduty + sms + email - If
product=Billingandtenant=Enterprise, use Billing template set and account team distribution list
Practical advice
If you’re just starting:
- Standardize event schema first
- Put all notification decisions in a central service
- Add environment-based routing guards
- Implement templates and preferences
- Add observability and deduplication next
If you want, I can also provide:
- a reference architecture diagram,
- a sample event schema,
- or a notification routing policy example in YAML/JSON.