Prompt

What is the best way to serve a custom LLM without running Kubernetes?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

If you want to serve a custom LLM without Kubernetes, the best approach is usually:

1) Use a simple containerized serving stack on a single VM or a small set of VMs

For most teams, the sweet spot is:

  • Docker for packaging
  • A purpose-built model server like:
    • vLLM — great for high-throughput GPU inference
    • Hugging Face TGI — solid production option
    • TensorRT-LLM — best if you’re heavily optimizing for NVIDIA GPUs
    • llama.cpp — good for CPU or smaller GPU deployments / quantized models
  • A lightweight process manager / reverse proxy:
    • systemd, Docker Compose, or Nginx/Traefik
  • Optional autoscaling handled outside Kubernetes:
    • cloud load balancer + VM autoscaling group, or just manual horizontal scaling

This is the most common “production without K8s” pattern.


2) Choose the serving engine based on your model and hardware

Best general choice: vLLM

Use this if:

  • you have an open-weight transformer model
  • you want high throughput and lower latency
  • you’re serving multiple concurrent requests
  • you have NVIDIA GPUs

Why it’s often the default:

  • continuous batching
  • efficient KV cache management
  • OpenAI-compatible server mode
  • easy Docker deployment

Best for Hugging Face ecosystem: TGI

Use this if:

  • you want a polished, widely used production server
  • you prefer Hugging Face-native tooling
  • you want simpler deployment with strong defaults

Best for maximum NVIDIA optimization: TensorRT-LLM

Use this if:

  • you’re willing to invest in optimization work
  • latency and throughput matter a lot
  • you’re deploying on NVIDIA GPUs only

Best for edge/CPU/quantized deployment: llama.cpp

Use this if:

  • the model is small enough or quantized
  • you want CPU inference or lightweight GPU use
  • you want very simple operational overhead

3) Run it behind a minimal production stack

A practical setup on a single server:

  • Nginx or Traefik as reverse proxy
  • vLLM/TGI in Docker
  • systemd or Docker restart policies for resilience
  • Prometheus + Grafana for monitoring
  • OpenTelemetry / structured logs if you need tracing

This gives you:

  • TLS termination
  • request routing
  • health checks
  • restarts on failure
  • observability

4) If you need scaling, use VMs and a load balancer instead of Kubernetes

For multiple replicas:

  • run identical model-serving containers on separate GPU VMs
  • put them behind:
    • AWS ALB/NLB, GCP Load Balancer, Azure Load Balancer, or
    • Nginx/HAProxy
  • scale using:
    • cloud auto scaling groups
    • custom scripts
    • Terraform + your CI/CD pipeline

This is much simpler than Kubernetes if your workload is straightforward.


5) Recommended “best overall” option

If you want a single answer:

Use Docker + vLLM on a GPU VM, fronted by Nginx or a cloud load balancer

That setup is usually the best balance of:

  • simplicity
  • performance
  • ease of deployment
  • production readiness

6) What not to do

Avoid:

  • running the model directly on the host without isolation
  • building a custom inference server unless you have a strong reason
  • over-engineering with orchestration tools if you only need one or two GPUs
  • using Kubernetes just for “best practice” if your scale doesn’t justify it

7) A good decision rule

  • Single model, one or a few GPUs, moderate traffic → Docker + vLLM/TGI
  • Heavy optimization needs → TensorRT-LLM
  • Small model / CPU / quantized edge use → llama.cpp
  • Need easy scaling across many nodes → VMs + load balancer, still no Kubernetes required

If you want, I can also give you:

  1. a reference architecture,
  2. a Docker Compose example, or
  3. a comparison of vLLM vs TGI vs llama.cpp for your model.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.