Prompt
I'm building a system to route requests by cost and latency; how should I structure the gateway layer?
Latest observation
A good gateway for cost/latency routing should be designed as a policy-driven control plane + lightweight data plane. That keeps request-time overhead low while letting you change routing behavior without redeploying everything.
1) Core responsibilities of the gateway
The gateway should do only a few things well:
- Authenticate / authorize
- Normalize requests into a common internal shape
- Classify requests by:
- estimated cost
- latency sensitivity
- model/service requirements
- user/tenant priority
- Select a route using policy
- Apply protections:
- rate limits
- budgets
- circuit breaking
- retries/timeouts
- Observe outcomes:
- latency
- actual cost
- error rate
- quality signals
- Emit feedback to improve routing
Avoid embedding business logic directly in request handlers; put routing rules in a policy layer.
2) Recommended architecture
A. Edge Gateway
This is the public entry point.
Responsibilities:
- TLS termination
- authN/authZ
- request validation
- tenant identification
- coarse rate limiting
- request tracing/correlation IDs
- passing request metadata downstream
Keep this layer fast and stateless.
B. Routing Orchestrator
A dedicated service or module that decides where to send the request.
Inputs:
- request metadata
- tenant SLA
- budget state
- current backend health
- historical latency/cost stats
- request classification
Outputs:
- target backend/model/service
- retry/fallback plan
- timeout settings
- concurrency class / queue priority
This is where cost-vs-latency tradeoffs live.
C. Policy/Rules Store
Store routing policy outside code:
- YAML/JSON config
- database-backed rules
- feature-flag system
- policy engine like OPA if rules get complex
Examples of rules:
- “Interactive requests go to low-latency tier unless budget exceeded”
- “Batch traffic uses cheapest backend available”
- “Premium tenants bypass economy tier”
- “If p95 latency > threshold, shift 20% traffic to fallback”
D. Telemetry + Feedback Loop
You need near-real-time metrics:
- request latency distribution
- backend cost per request
- token/compute usage
- error rate
- cache hit rate
- quality metrics if available
This feeds:
- route selection
- anomaly detection
- autoscaling
- budget enforcement
3) Request classification strategy
Classify every request early. Typical dimensions:
Latency class
- interactive: user waiting, low latency required
- soft real-time: latency matters, but some flexibility
- batch: optimize for cost
Cost class
- small / medium / large
- predicted token count or compute cost
- expensive tool-use / retrieval / long context
Priority / tenant tier
- free / standard / enterprise
- admin/internal / external
- SLA tier
Capability requirements
- needs tool calling
- needs specific model/version
- requires structured output
- requires region/data residency
You can use simple heuristics first, then move to a prediction model later.
4) Routing policy design
A strong pattern is:
hard constraints first, optimization second
Hard constraints
Reject or force certain routes based on:
- compliance / region
- required capabilities
- tenant entitlements
- budget exhaustion
- safety filters
Optimization
Then choose among eligible routes based on:
- expected latency
- expected cost
- current load
- reliability
- quality score
A simple scoring function works well:
score = w1 * normalized_latency + w2 * normalized_cost + w3 * error_risk - w4 * quality
Different request types use different weights.
Example:
- interactive: latency weight high
- batch: cost weight high
- premium: quality and reliability higher
5) Fallback and resilience
Routing systems fail when they assume the first choice always works.
Add:
- timeouts per route
- retries only when idempotent or safe
- circuit breakers on unhealthy targets
- fallback tiers:
- premium model → standard model → cheapest model
- graceful degradation
- reduce context
- disable nonessential tools
- switch to cached/stale response where acceptable
Important: make fallback behavior explicit and observable.
6) Budget enforcement
If cost matters, enforce budgets at multiple levels:
Per request
- maximum allowed estimated cost
- reject or downgrade if estimate exceeds limit
Per tenant
- daily/monthly budget
- soft limit: route cheaper
- hard limit: throttle/reject
Per workflow
- cap tool calls
- cap retries
- cap context expansion
Also separate:
- estimated cost at admission time
- actual cost after completion
Track divergence so estimates improve.
7) Latency-aware execution
For latency-sensitive traffic:
- keep routing decisions local and fast
- cache policy decisions where possible
- avoid synchronous calls to many services before routing
- precompute backend health snapshots
- use p95/p99 metrics, not just averages
- prioritize queueing by request class
You may also want:
- hedged requests for critical traffic
- adaptive timeouts
- regional routing close to the user
8) Data model for gateway decisions
A useful internal request envelope:
{
"request_id": "abc123",
"tenant_id": "t1",
"user_tier": "enterprise",
"request_type": "interactive",
"estimated_cost": 0.12,
"estimated_latency_ms": 350,
"required_capabilities": ["tool_calling", "json_output"],
"region": "us-east-1",
"budget_remaining": 42.50,
"priority": 80,
"policy_version": "2026-01-15",
"trace_id": "..."
}
And a routing decision:
{
"selected_route": "fast-tier-a",
"fallback_routes": ["balanced-tier-b", "cheap-tier-c"],
"timeout_ms": 1200,
"retry_policy": {
"max_retries": 1,
"retry_on": ["5xx", "timeout"]
}
}
9) Implementation pattern
A clean implementation is:
- Ingress
- Parse + authenticate
- Enrich metadata
- Estimate cost/latency
- Evaluate hard constraints
- Score eligible routes
- Apply circuit breaker / health checks
- Dispatch
- Collect metrics
- Update stats store asynchronously
Keep the decision path synchronous and the learning/aggregation path asynchronous.
10) Practical tech choices
Depending on scale:
Simple setup
- API gateway: Envoy / NGINX / Kong
- Policy in service config
- Metrics in Prometheus/Grafana
- Tracing in OpenTelemetry
More advanced
- Envoy at edge
- Routing service in Go/Rust/Java
- Policy engine: OPA
- Streaming telemetry: Kafka/PubSub
- Stats store: Redis for hot data, warehouse for analytics
11) Common pitfalls
- Routing using only average latency instead of p95/p99
- Ignoring actual cost drift vs estimate
- Putting too much logic in the edge gateway
- No fallback paths
- No policy versioning or audit trail
- Synchronous dependency on slow analytics during request path
- Mixing admission control and dispatch logic in one blob
- Not separating tenant quotas from global system protection
12) A good default structure
If you want a simple starting point:
- Gateway: auth, validation, metadata enrichment, request shaping
- Router: policy-based backend selection
- Policy store: versioned rules
- Telemetry service: async metrics and cost tracking
- Budget service: per-tenant and per-workflow limits
- Health service: backend health and latency snapshots
If you want, I can also sketch:
- a concrete microservice architecture,
- a decision algorithm / scoring formula, or
- a reference request flow diagram for this gateway.