Prompt
I'm building an AI assistant that uses tools and memory, recommend a stack
Latest observation
Here’s a solid, practical stack for an AI assistant with tools + memory, optimized for building quickly but still scaling well.
Recommended stack
1) Core agent/orchestration
Python + FastAPI
- Best ecosystem for LLM tooling and integrations
- Easy to expose your assistant as an API
- Good for streaming responses and tool execution
Agent framework: one of
- LangGraph if you want robust tool-calling, stateful workflows, retries, branches, and long-running agents
- LlamaIndex if your main focus is retrieval + memory + document workflows
- LangChain if you want a broad ecosystem, but I’d prefer LangGraph for production agent flows
My pick:
- LangGraph for the agent runtime
- LlamaIndex for retrieval/memory abstractions if needed
2) LLM provider
Choose based on cost/quality/latency:
- OpenAI: strongest general-purpose tool calling, structured outputs, broad model options
- Anthropic: excellent reasoning and long-context behavior
- Google Gemini: strong context window and multimodal options
- Open-source via vLLM or TGI if you want to self-host
My pick for most teams:
- Start with OpenAI or Anthropic
- Add a fallback model later for resilience/cost control
3) Tool execution
You’ll want a clean tool layer rather than direct ad hoc function calls.
Use:
- Python tool registry with typed schemas
- Pydantic for input/output validation
- Background jobs for long tasks: Celery, RQ, or Temporal
If you need serious workflow reliability:
- Temporal is excellent for retries, durable state, and human-in-the-loop steps
4) Memory
Think of memory in 3 layers:
A. Short-term memory
- Conversation state stored in your app DB or agent state
- Keep recent turns + active task context
B. Long-term semantic memory
- Store embeddings in a vector DB
- Use for user preferences, past conversations, documents, and recalls
Good options:
- Postgres + pgvector for a simple, production-friendly default
- Pinecone, Weaviate, or Qdrant if you want a dedicated vector store
My pick:
- Postgres + pgvector unless you already know you need a specialized vector DB
C. Structured memory
- Store durable facts in relational tables or a knowledge graph
- Examples: user profile, preferences, settings, goals, permissions
Use Postgres for this too.
5) Database
PostgreSQL
- Store users, sessions, messages, tool logs, memory items, audit trails
- Add pgvector for embeddings
- Very strong default for almost everything
Optional:
- Redis for caching, ephemeral state, rate limiting, and queues
6) Retrieval / search
If your assistant uses documents or past interactions:
- Hybrid retrieval: keyword + vector search
- Reranking for better relevance
Options:
- Postgres full-text search + pgvector
- OpenSearch/Elasticsearch if search is central
- Rerankers from Cohere or model-based reranking if needed
7) Frontend
Depending on product:
- Next.js for web app
- React Native if you need mobile
- Slack/Discord bot integration if that’s the main surface
For chat UI:
- Stream tokens over WebSockets or Server-Sent Events
- Keep tool progress visible to users
8) Authentication and permissions
- Auth0, Clerk, or Firebase Auth
- Enforce per-user and per-tenant memory/tool access
- Important if your assistant can access external systems
9) Observability and evaluation
This is critical for tool-using agents.
Use:
- LangSmith if using LangChain/LangGraph
- OpenTelemetry for traces
- Sentry for errors
- A logging pipeline for tool calls, prompts, and outcomes
Evaluation tools:
- Promptfoo
- Ragas for retrieval evaluation
- Custom golden test cases for tool behavior
A very practical default architecture
Best “balanced” stack
- Frontend: Next.js
- Backend: FastAPI
- Agent orchestration: LangGraph
- LLM: OpenAI or Anthropic
- DB: PostgreSQL + pgvector
- Cache/queue: Redis
- Background jobs: Temporal or Celery
- Observability: OpenTelemetry + Sentry + LangSmith
- Auth: Clerk or Auth0
If you want the simplest possible MVP
Use this:
- FastAPI
- OpenAI API
- PostgreSQL + pgvector
- Redis
- LangGraph
- Next.js
- Sentry
That gets you a production-ready baseline without too much complexity.
If you expect complex workflows
Add:
- Temporal for durable tool workflows
- Qdrant or Pinecone if memory/retrieval becomes large-scale
- OpenSearch for advanced hybrid search
Design advice for tool + memory assistants
A few important principles:
- Separate ephemeral chat state from durable memory
- Write memory explicitly, don’t store everything automatically
- Validate tool inputs with schemas
- Log every tool call and result
- Use retrieval only when needed to avoid noisy context
- Add permission checks at the tool layer, not just the UI
My “best default” recommendation
If you want one stack choice:
Python, FastAPI, LangGraph, OpenAI/Anthropic, PostgreSQL + pgvector, Redis, Next.js, Temporal (if workflows are complex), Sentry + OpenTelemetry
If you want, I can also give you:
- a minimal MVP stack
- a production-scale stack
- or a reference architecture diagram for tools + memory.