Prompt

What should I use to move from notebook to production for LLM apps?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 9, 2026

Latest observation

Oct 9, 2026OpenAI APIWeb search: off

To move from notebook → production for LLM apps, think in terms of a stack that covers app logic, evaluation, orchestration, serving, observability, and deployment.

A practical path

1. Keep the notebook as the prototype

Use notebooks for:

  • prompt experiments
  • retrieval tests
  • quick evals
  • sampling outputs

But once the behavior is stable, move the core logic into versioned Python modules.


2. Put app logic into a real codebase

Use:

  • Python package structure (src/, tests, configs)
  • pydantic for schemas and validation
  • pytest for unit/integration tests
  • ruff/black/mypy for code quality

This makes your LLM app testable and deployable.


3. Use an orchestration framework if needed

Pick one based on complexity:

  • Simple apps: plain Python + API layer
  • RAG / chains / tools: LangChain or LlamaIndex
  • Agentic workflows / stateful graphs: LangGraph
  • Custom workflow orchestration: Prefect or Dagster

If your app is mostly “prompt + retrieval + tool calls,” LangGraph is often a good production-friendly choice because it gives explicit state and control flow.


4. Add evaluation before deployment

You need automated checks for:

  • answer quality
  • grounding / hallucinations
  • retrieval quality
  • latency
  • cost

Common tools:

  • Ragas for RAG evaluation
  • DeepEval
  • promptfoo
  • custom golden-test sets with pytest

A good rule: don’t ship a model/prompt change without evals.


5. Use an API server for serving

Common serving options:

  • FastAPI for most LLM apps
  • BentoML if you want packaging + serving + deployment together
  • Ray Serve for higher-scale distributed serving

For many teams, FastAPI + Docker is the simplest and most flexible production setup.


6. Add observability from day one

You’ll want logs for:

  • prompts
  • model responses
  • retrieval context
  • tool calls
  • latency
  • token usage
  • errors

Useful tools:

  • LangSmith
  • Arize Phoenix
  • Weights & Biases Weave
  • OpenTelemetry + your logging stack

This is critical for debugging “it worked in the notebook but not in prod.”


7. Use a proper deployment path

Typical production setup:

  • Docker
  • CI/CD (GitHub Actions, GitLab CI, etc.)
  • Kubernetes if you need scale/control
  • Cloud-managed containers if you want simplicity

A good default stack

If you want a strong, practical default:

  • Python + FastAPI
  • LangGraph (if workflows/agents) or plain Python (if simple)
  • Pydantic for schemas
  • pytest + promptfoo/Ragas for evals
  • Docker + GitHub Actions
  • LangSmith or Phoenix for observability

If you want the shortest answer

For most LLM apps, use:

Python codebase + FastAPI + Docker + eval suite + observability

And add LangChain/LlamaIndex/LangGraph only if your app actually needs it.


Rule of thumb

  • Notebook: exploration
  • Python package: implementation
  • FastAPI/BentoML: serving
  • evals: quality gates
  • observability: production debugging
  • Docker/CI/CD: deployment

If you tell me what kind of LLM app you’re building — e.g. RAG, agent, chat app, document extraction, workflow automation — I can suggest a more exact stack.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.