Prompt
building prompt routing service
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
Here’s a solid blueprint for a prompt routing service.
What it is
A prompt routing service takes an incoming request and decides:
- which prompt/template to use
- which model to send it to
- whether to enrich, rewrite, or reject it
- which tools or policies to apply
It’s useful when you have multiple use cases like support, summarization, extraction, code generation, etc.
Core responsibilities
1. Classification
Detect the request intent:
- support / FAQ
- creative writing
- extraction / structured output
- code generation
- safety-sensitive
- high-priority customer
2. Routing
Based on classification, choose:
- prompt template
- model
- temperature / max tokens
- tool access
- fallback path
3. Policy enforcement
Apply rules like:
- block unsafe content
- route sensitive requests to safer models
- redact PII
- require human review for some categories
4. Observability
Track:
- route chosen
- confidence score
- latency
- cost
- success/failure
- fallback usage
Suggested architecture
Components
-
Ingress API
- receives prompt/request
- authenticates caller
- validates payload
-
Classifier
- rule-based, ML-based, or LLM-based
- outputs intent + confidence
-
Router
- picks destination based on routing table
-
Prompt registry
- stores prompt templates versioned by use case
-
Execution layer
- calls LLM/provider
- handles retries/fallbacks
-
Policy engine
- safety/compliance checks
-
Telemetry/logging
- traces and metrics
Simple routing flow
-
Request comes in:
{ "user_id": "123", "text": "Summarize this document in 3 bullets" } -
Classifier decides:
{ "intent": "summarization", "confidence": 0.96 } -
Router selects:
{ "prompt_id": "summarize_v2", "model": "gpt-4.1-mini", "temperature": 0.2 } -
Service executes and returns response.
Routing strategies
A. Rules-based
Best for early stage.
if text contains "summarize" -> summarization
if text contains "write code" -> code
if pii detected -> compliance route
Pros:
- simple
- predictable
Cons:
- brittle
- hard to scale
B. Model-based classification
Use a small model or LLM to classify.
Pros:
- flexible
- handles ambiguity
Cons:
- latency/cost
- possible misroutes
C. Hybrid
Common best practice:
- rules for obvious cases
- classifier for ambiguous ones
- fallback to safe default
Example routing table
| Intent | Prompt | Model | Temperature | Notes |
|---|---|---|---|---|
| support | support_v3 | fast-model | 0.2 | low cost |
| summarization | summarize_v2 | balanced-model | 0.1 | concise |
| code | code_v5 | strong-code-model | 0.0 | deterministic |
| extraction | extract_v2 | structured-model | 0.0 | JSON schema |
| sensitive | safe_v1 | safe-model | 0.0 | strict policy |
API design
POST /route
Input:
{
"text": "Extract names and emails from this paragraph",
"user_id": "u123",
"context": {
"tenant_id": "t1",
"locale": "en-US"
}
}
Output:
{
"route": {
"intent": "extraction",
"prompt_id": "extract_v2",
"model": "structured-model",
"temperature": 0.0
},
"trace_id": "abc-123"
}
Important features to add
- versioned prompts
- A/B testing
- fallback routes
- confidence thresholds
- manual overrides
- rate limiting
- multi-tenant support
- audit logs
- cost controls
Tech stack suggestion
Lightweight
- API: FastAPI / Express
- classifier: rules + small LLM
- storage: Postgres
- cache: Redis
- observability: OpenTelemetry + Grafana
More advanced
- event-driven router with Kafka/SQS
- prompt registry service
- workflow engine for retries and human review
- feature flags for gradual rollout
Practical routing logic example
def route_request(text):
if contains_pii(text):
return "compliance_v1", "safe-model"
intent, conf = classify(text)
if conf < 0.7:
return "general_fallback_v1", "balanced-model"
routes = {
"support": ("support_v3", "fast-model"),
"summarization": ("summarize_v2", "balanced-model"),
"code": ("code_v5", "strong-code-model"),
"extraction": ("extract_v2", "structured-model"),
}
return routes.get(intent, ("general_v1", "balanced-model"))
Common pitfalls
- overusing LLM classification for everything
- no fallback route
- no prompt/version tracking
- no evaluation data
- routing changes without observability
- mixing policy decisions with business logic
Recommended MVP
Start with:
- rule-based intent detection
- prompt registry in DB
- router service with fallback
- logging/metrics
- manual override endpoint
Then later add:
- ML classifier
- confidence-based routing
- A/B testing
- automated optimization
If you want, I can also provide:
- a system design diagram
- a database schema
- a FastAPI/Node.js implementation
- a routing rules engine example
- a prompt registry design