Prompt
How to deploy model to API endpoint on GPU
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To deploy a model behind an API endpoint on a GPU, the general flow is:
- Package the model
- Serve it with an API framework
- Enable GPU inference
- Containerize
- Deploy to a GPU-capable host
- Expose the endpoint securely
1) Pick a serving stack
Common options:
- FastAPI + PyTorch/TensorFlow: simplest for custom APIs
- TorchServe: good for PyTorch model serving
- TensorFlow Serving: best for TensorFlow models
- NVIDIA Triton Inference Server: strong choice for high performance and multi-framework GPU serving
- vLLM / TGI / Text Generation Inference: best for LLMs
If you want a simple custom API, FastAPI is often easiest.
2) Example: FastAPI on GPU with PyTorch
Install dependencies
pip install fastapi uvicorn torch torchvision
Create app.py
from fastapi import FastAPI
from pydantic import BaseModel
import torch
app = FastAPI()
# Load model once at startup
model = torch.load("model.pt", map_location="cuda")
model.eval()
class InputData(BaseModel):
text: str
@app.post("/predict")
def predict(data: InputData):
with torch.no_grad():
# Replace this with your preprocessing and inference
input_tensor = torch.tensor([1.0], device="cuda")
output = model(input_tensor)
return {"prediction": output.tolist()}
Run it
uvicorn app:app --host 0.0.0.0 --port 8000
3) Dockerize it with GPU support
Dockerfile
FROM pytorch/pytorch:2.4.0-cuda12.1-cudnn9-runtime
WORKDIR /app
COPY . /app
RUN pip install fastapi uvicorn
EXPOSE 8000
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]
Build
docker build -t my-gpu-api .
Run with GPU
Make sure you have:
- NVIDIA driver installed
- NVIDIA Container Toolkit installed
Then:
docker run --gpus all -p 8000:8000 my-gpu-api
4) Deploy to a GPU-capable platform
You need a machine or service with a GPU, such as:
- AWS EC2 GPU instances (g4dn, g5, p4d, p5)
- GCP GPU VMs
- Azure GPU VMs
- RunPod
- Lambda Labs
- Kubernetes with GPU nodes
For Kubernetes, install:
- NVIDIA device plugin
- GPU node pool
- Your deployment with GPU resource requests
Example pod spec:
resources:
limits:
nvidia.com/gpu: 1
5) Expose the endpoint
Usually you place a reverse proxy/load balancer in front:
- Nginx
- Cloud Load Balancer
- API Gateway
Add:
- HTTPS
- Authentication
- Rate limiting
- Logging/monitoring
6) Production tips
- Load the model once at startup, not per request
- Use batching if requests are frequent
- Use half precision (
fp16) or quantization if supported - Keep the model on GPU memory
- Add health check endpoint:
@app.get("/health")
def health():
return {"status": "ok"}
- Monitor GPU memory and latency
7) If you’re serving an LLM
Consider using:
- vLLM for high-throughput generation
- TGI for Hugging Face models
- Triton for optimized serving
Example with vLLM:
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-3-8B-Instruct \
--host 0.0.0.0 \
--port 8000
This gives you an OpenAI-compatible API.
If you want, I can give you:
- a full FastAPI + Docker + GPU deployment example, or
- a cloud-specific guide for AWS/GCP/Azure, or
- a LLM-specific setup using vLLM or Triton.