Prompt
How do I serve a model in a private VPC?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
Serving a model in a private VPC usually means:
- the model endpoint is not publicly reachable
- only workloads inside your VPC, or connected networks, can call it
- you control egress/ingress with subnets, security groups, firewall rules, and private DNS
A good setup depends on where the model runs:
- self-hosted on VMs/Kubernetes in your VPC
- managed model serving inside a cloud provider’s private networking option
- hybrid with a private gateway/load balancer in front
Common architecture
-
Place the model server in private subnets
- No public IPs
- Attach only internal load balancers or service mesh ingress
-
Expose it privately
- Internal load balancer
- Private API gateway / private endpoint
- VPC peering / Transit Gateway / PrivateLink / PSC, depending on cloud
-
Restrict access
- Security groups / network ACLs / firewall rules
- Only allow your app subnets, bastion, or peered VPCs
-
Control outbound traffic
- Use NAT only if needed for downloads/updates
- Prefer VPC endpoints/private service access for dependencies like object storage, container registry, logs
-
Set up identity/auth
- mTLS, JWT, IAM auth, or service-to-service auth
- Don’t rely on network isolation alone
-
Add observability
- Private logging/metrics/tracing
- Health checks from inside the VPC
If you’re self-hosting
Typical pattern:
- Inference container (vLLM, TGI, Triton, TorchServe, FastAPI, etc.)
- Run in:
- Kubernetes with a private internal ingress
- EC2/VMs behind an internal load balancer
- Store model weights in:
- private object storage bucket
- or bake into image if small enough
Example flow
- App in private subnet calls
internal-model.company.local - DNS resolves to private load balancer
- Load balancer routes to model pods/instances
- Model instances have no public IPs
If you’re using a managed cloud service
Most clouds offer a private connectivity option:
- AWS: VPC endpoints / PrivateLink, internal load balancers, SageMaker private networking options
- Azure: Private Link, private endpoints, internal load balancers, Azure ML managed online endpoints with private networking
- GCP: Private Service Connect, private Google access, internal load balancing, Vertex AI private service connectivity options
The exact steps vary, but the pattern is the same: create the endpoint so it is reachable only from your VPC or connected networks.
Security checklist
- No public IPs on model hosts
- Internal-only load balancer or private endpoint
- Security groups allow only trusted CIDRs or SG references
- Authenticated requests required
- Encryption in transit (TLS)
- Encryption at rest for weights and artifacts
- Private access to storage, registries, and secrets
- Audit logs enabled
- Rate limiting and request size limits
- Separate dev/stage/prod VPCs or accounts/projects
Operational considerations
- Model downloads: if the instance has no internet access, pre-stage weights in private storage
- Cold start: large models may take time to load; use readiness probes
- GPU scheduling: ensure GPU-capable private subnets/instances
- Scaling: autoscale on queue length, GPU utilization, latency, or QPS
- Failure domains: spread across AZs if available
A simple reference setup
If you want the simplest secure setup:
- private Kubernetes cluster
- model server in pods
- internal load balancer
- private DNS name
- service-to-service auth
- private object storage for model artifacts
- VPC endpoints for storage, logs, and registry
If you tell me which cloud you’re on and whether you’re self-hosting or using a managed service, I can give you exact steps and a concrete architecture.