Prompt
What should I use to move from notebook to production for LLM apps?
Latest observation
To move from notebook → production for LLM apps, think in terms of a stack that covers app logic, evaluation, orchestration, serving, observability, and deployment.
A practical path
1. Keep the notebook as the prototype
Use notebooks for:
- prompt experiments
- retrieval tests
- quick evals
- sampling outputs
But once the behavior is stable, move the core logic into versioned Python modules.
2. Put app logic into a real codebase
Use:
- Python package structure (
src/, tests, configs) - pydantic for schemas and validation
- pytest for unit/integration tests
- ruff/black/mypy for code quality
This makes your LLM app testable and deployable.
3. Use an orchestration framework if needed
Pick one based on complexity:
- Simple apps: plain Python + API layer
- RAG / chains / tools: LangChain or LlamaIndex
- Agentic workflows / stateful graphs: LangGraph
- Custom workflow orchestration: Prefect or Dagster
If your app is mostly “prompt + retrieval + tool calls,” LangGraph is often a good production-friendly choice because it gives explicit state and control flow.
4. Add evaluation before deployment
You need automated checks for:
- answer quality
- grounding / hallucinations
- retrieval quality
- latency
- cost
Common tools:
- Ragas for RAG evaluation
- DeepEval
- promptfoo
- custom golden-test sets with
pytest
A good rule: don’t ship a model/prompt change without evals.
5. Use an API server for serving
Common serving options:
- FastAPI for most LLM apps
- BentoML if you want packaging + serving + deployment together
- Ray Serve for higher-scale distributed serving
For many teams, FastAPI + Docker is the simplest and most flexible production setup.
6. Add observability from day one
You’ll want logs for:
- prompts
- model responses
- retrieval context
- tool calls
- latency
- token usage
- errors
Useful tools:
- LangSmith
- Arize Phoenix
- Weights & Biases Weave
- OpenTelemetry + your logging stack
This is critical for debugging “it worked in the notebook but not in prod.”
7. Use a proper deployment path
Typical production setup:
- Docker
- CI/CD (GitHub Actions, GitLab CI, etc.)
- Kubernetes if you need scale/control
- Cloud-managed containers if you want simplicity
A good default stack
If you want a strong, practical default:
- Python + FastAPI
- LangGraph (if workflows/agents) or plain Python (if simple)
- Pydantic for schemas
- pytest + promptfoo/Ragas for evals
- Docker + GitHub Actions
- LangSmith or Phoenix for observability
If you want the shortest answer
For most LLM apps, use:
Python codebase + FastAPI + Docker + eval suite + observability
And add LangChain/LlamaIndex/LangGraph only if your app actually needs it.
Rule of thumb
- Notebook: exploration
- Python package: implementation
- FastAPI/BentoML: serving
- evals: quality gates
- observability: production debugging
- Docker/CI/CD: deployment
If you tell me what kind of LLM app you’re building — e.g. RAG, agent, chat app, document extraction, workflow automation — I can suggest a more exact stack.