Prompt
I'm building a production AI agent with logging and guardrails, what stack do teams use?
Latest observation
For a production AI agent with logging and guardrails, teams usually use a layered stack rather than one monolithic framework. A common setup looks like this:
1) Agent / orchestration layer
Used to define tools, planning, memory, workflows, and multi-step behavior.
Popular choices:
- LangGraph (very common for production agent workflows)
- LangChain (still widely used, often with LangGraph)
- LlamaIndex (strong for RAG/data-centric agents)
- OpenAI Assistants / Responses API if you want a more managed approach
- CrewAI or AutoGen for multi-agent patterns, though many teams move to LangGraph for more control
2) Model layer
Choose the LLM provider and route requests intelligently.
Typical options:
- OpenAI
- Anthropic
- Google Gemini
- AWS Bedrock
- Azure OpenAI
- Sometimes a mix via a router for cost/latency optimization
Teams often add:
- Model fallback routing
- Prompt/version management
- Caching
3) Retrieval / data layer
If the agent uses company data or docs.
Common stack:
- Vector DB: Pinecone, Weaviate, Qdrant, Milvus, pgvector
- Search: Elasticsearch/OpenSearch
- Document parsing: Unstructured, Apache Tika, pdfplumber, custom OCR
- RAG framework: LlamaIndex or LangChain retrievers
4) Guardrails / safety layer
This is where production systems usually invest a lot.
Common guardrail tools:
- Guardrails AI
- NeMo Guardrails
- Pydantic / JSON schema validation for structured outputs
- LlamaGuard or provider moderation endpoints
- Prompt injection detection and URL/content filtering
- PII detection/redaction using tools like Microsoft Presidio
- Policy engines for tool access and action approval
Typical guardrails include:
- Output format validation
- Tool-call restrictions
- Allowlist/denylist for actions
- Human approval for sensitive actions
- Content moderation
- Data leakage prevention
- Rate limiting / abuse detection
5) Observability / logging
This is essential in production.
Common tools:
- LangSmith for traces/evals if using LangChain/LangGraph
- OpenTelemetry for vendor-neutral tracing
- Arize Phoenix for LLM observability and evals
- Helicone for request logging/cost monitoring
- Traceloop
- Datadog / New Relic / Grafana / OpenSearch for infra + app logs
What teams log:
- Prompts and completions
- Tool calls and tool outputs
- Latency per step
- Token usage and cost
- Errors/failures
- Guardrail triggers
- User feedback
- Retrieval citations / source docs
- Full trace of agent decisions
6) Evaluation / testing layer
For production agents, evals are not optional.
Common tools:
- LangSmith evals
- Ragas for RAG quality
- DeepEval
- promptfoo
- OpenAI Evals
- Custom golden test sets and regression suites
Teams evaluate:
- Task success rate
- Hallucination rate
- Retrieval precision/recall
- Tool selection accuracy
- Safety violations
- Latency and cost
7) Execution / infra layer
Where the agent actually runs.
Typical stack:
- FastAPI / Flask / Node.js backend
- Redis for queues, caching, session state
- Postgres for durable state
- Celery / Temporal / BullMQ / Airflow for workflows and retries
- Kubernetes / ECS / serverless for deployment
- Secrets manager: AWS Secrets Manager, Vault, GCP Secret Manager
For long-running or reliable workflows, Temporal is especially popular.
A common modern production stack
A lot of teams end up with something like:
- Agent framework: LangGraph
- LLM: OpenAI + Anthropic fallback
- RAG: LlamaIndex or LangChain retriever
- Vector DB: pgvector / Pinecone / Weaviate
- Guardrails: Pydantic + Guardrails AI + moderation API
- Observability: LangSmith + OpenTelemetry + Datadog
- Evaluation: Ragas + promptfoo
- Workflow engine: Temporal
- Backend: FastAPI
- State: Postgres + Redis
If you want the simplest production-ready approach
If your goal is to ship quickly but safely, a strong baseline is:
- LangGraph
- Pydantic structured outputs
- OpenAI/Anthropic
- pgvector
- LangSmith or OpenTelemetry
- Basic moderation + PII redaction
- Temporal or Redis queues for retries
- Golden test evals with promptfoo / DeepEval
Practical guidance
If you’re choosing from scratch:
- Use LangGraph if your agent has branching logic, tools, approvals, or multi-step workflows
- Use LlamaIndex if the core problem is retrieval over your data
- Use OpenTelemetry + centralized logs regardless of framework
- Use Pydantic/JSON schema for every tool input and output
- Use human approval gates for destructive or high-impact actions
- Build an eval harness early, not after launch
If you want, I can also give you:
- a reference architecture diagram,
- a recommended stack by company stage (startup vs enterprise), or
- a concrete open-source stack you can deploy this week.
Brands