Prompt
I need a path from demo notebook to production service without rebuilding everything
Latest observation
Here’s a pragmatic path from a demo notebook to a production service without throwing everything away.
Guiding principle
Treat the notebook as the prototype of the model logic, not the final application.
Your goal is to preserve:
- data prep logic
- feature engineering
- model training/inference steps
- evaluation logic
…and move only the parts that need production hardening:
- packaging
- config management
- testing
- API/service layer
- observability
- deployment
A low-friction migration path
1) Freeze the notebook into a reproducible baseline
Before refactoring, make the notebook reproducible.
Do this:
- Parameterize paths and constants
- Restart kernel and run all cells top to bottom
- Pin dependencies in
requirements.txtorpyproject.toml - Save sample input/output artifacts
- Capture model version, dataset version, and metrics
Helpful output:
train.pyequivalent or notebook export- a saved model file (
.pkl,.joblib,.onnx, etc.) - a small test dataset and expected predictions
2) Extract notebook code into importable modules
Move code from cells into a real Python package gradually.
Suggested structure:
project/
notebooks/
demo.ipynb
src/
my_service/
__init__.py
data.py
features.py
model.py
predict.py
api.py
tests/
models/
configs/
Start by extracting:
- preprocessing functions into
features.py - training/loading logic into
model.py - prediction logic into
predict.py
Then make the notebook call those functions instead of containing everything inline.
This gives you a soft transition:
- notebook remains usable
- code becomes reusable and testable
3) Create one clean inference entry point
Production needs a stable interface.
Build a single function like:
def predict(input_data: dict) -> dict:
...
or for batches:
def predict(df: pd.DataFrame) -> pd.DataFrame:
...
Keep all notebook-specific exploration out of this path.
This function should:
- validate inputs
- apply preprocessing
- load the model
- return predictions in a consistent schema
4) Separate training from inference
A common mistake is to ship notebook training code directly into the service.
Instead:
- training code runs offline
- service only loads a trained artifact
So you end up with:
train.py→ produces model artifactpredict.py→ consumes model artifactapi.py→ exposes prediction endpoint
That makes production stable and cheaper.
5) Add tests around the extracted logic
Before deploying anything, add tests for the code you pulled out.
Focus on:
- preprocessing correctness
- schema validation
- model loading
- prediction output shape/type
- edge cases and bad inputs
A few high-value tests can prevent a lot of notebook-to-service bugs.
6) Wrap the inference function in a service layer
Once prediction logic is clean, expose it.
Typical options:
- FastAPI for HTTP inference
- gRPC if you need lower latency / strong contracts
- Batch job / scheduled job if real-time isn’t necessary
Minimal FastAPI pattern:
from fastapi import FastAPI
from pydantic import BaseModel
app = FastAPI()
class Input(BaseModel):
feature1: float
feature2: int
@app.post("/predict")
def predict(item: Input):
result = run_prediction(item.model_dump())
return result
This lets you keep the notebook-derived logic intact while adding a production interface.
7) Package the app for deployment
Use one of these packaging targets:
- Docker container
- serverless function
- VM/service unit
- managed ML platform endpoint
Docker is usually the easiest bridge from notebook to production.
Example benefits:
- reproducible environment
- same dependencies everywhere
- easy CI/CD integration
8) Add lightweight production concerns
Don’t overbuild, but do include:
- logging
- metrics
- request validation
- error handling
- model/version metadata
You want visibility into:
- input failures
- latency
- prediction counts
- drift signals if relevant
What not to do
To avoid rebuilding everything, don’t:
- rewrite the notebook from scratch in a new framework immediately
- bundle training and serving together in one giant script
- optimize prematurely for perfect architecture
- move code before identifying the stable prediction flow
Best incremental strategy
If you want the fastest path:
- Clean up notebook
- Extract preprocessing and prediction functions
- Save/load model artifact
- Add tests
- Wrap in FastAPI
- Dockerize
- Deploy
This keeps the notebook as the source of truth while productionizing the core logic.
A practical rule of thumb
If a cell:
- is deterministic,
- reused across runs,
- and needed for inference,
then it belongs in src/.
If a cell is:
- exploratory,
- plotting,
- debugging,
- or one-off analysis,
leave it in the notebook.
If you want the simplest possible production path
Use this stack:
- Notebook for exploration
- Python package for shared logic
- FastAPI for serving
- Docker for deployment
- pytest for validation
That’s usually the shortest route to production without a rewrite.
If you want, I can also give you:
- a step-by-step migration checklist, or
- a sample folder structure + code skeleton for notebook-to-service conversion.