Prompt
How do I set up model serving for a RAG app connected to our internal docs?
Latest observation
Oct 1, 2026
OpenAI APIWeb search: off
Here’s a practical way to set up model serving for a RAG app that answers questions over your internal docs.
1) Use this basic RAG serving architecture
User → API → Retrieval → Prompt assembly → LLM inference → Response
Typical components:
- Document store: internal docs in SharePoint, Confluence, Google Drive, Git, PDFs, etc.
- Ingestion pipeline: cleans, chunks, tags, and embeds docs
- Vector store: stores embeddings for semantic search
- Retrieval service: fetches top-k relevant chunks for each question
- LLM serving layer: hosts the model that writes the answer
- App/API layer: chat UI or backend endpoint
- Auth/ACL layer: ensures users only retrieve docs they’re allowed to see
2) Choose your model serving approach
Option A: Hosted API model
Use a managed model endpoint from a provider if:
- you want fastest setup
- you don’t need strict on-prem hosting
- your data handling policy allows it
Pros:
- low ops burden
- easy scaling
- simple deployment
Cons:
- data residency/compliance concerns
- cost can grow with usage
- less control over latency/caching
Option B: Self-hosted model
Host an open-weight model in your environment if:
- docs are sensitive
- you need network isolation
- you want tighter cost/control
Common serving tools:
- vLLM
- TGI (Text Generation Inference)
- Hugging Face Transformers + FastAPI
- SGLang
- NVIDIA Triton in some setups
Pros:
- stronger control over data/security
- can run inside VPC/on-prem
- predictable behavior
Cons:
- GPU ops complexity
- scaling and reliability are on you
- model upgrades/testing needed
3) Recommended reference architecture
Ingestion path
- Pull documents from internal sources
- Extract text and metadata
- Chunk into passages
- Generate embeddings
- Store in vector DB plus metadata store
Query path
- User asks a question
- Authenticate user
- Apply ACL-aware retrieval
- Retrieve top chunks
- Build prompt with:
- system instructions
- user question
- retrieved context
- citation metadata
- Send prompt to served LLM
- Return answer with citations
4) Pick your serving stack
If you want the simplest self-hosted serving
- vLLM for model serving
- FastAPI for your app layer
- Postgres + pgvector, Pinecone, Weaviate, Milvus, or OpenSearch for retrieval
- Redis for caching
- Kubernetes if you need scaling; otherwise start with a single VM or ECS task
If you want enterprise control
- Put the LLM server in a private subnet
- Use service-to-service auth
- Log prompts/outputs carefully
- Add access controls at retrieval time, not only at answer time
5) Important design decisions
A. Don’t only secure the model endpoint
The bigger risk is usually retrieval leakage.
You need ACL-aware retrieval:
- tag every chunk with source doc permissions
- filter retrieval results by user identity/groups
- do not send unauthorized chunks to the model
B. Keep prompts short and relevant
For internal docs, retrieval quality matters a lot.
- use chunk sizes around 300–800 tokens as a starting point
- retrieve top 3–8 chunks
- re-rank if needed
- include only the most relevant snippets
C. Add citations
Always return:
- source document title
- section/page
- link if possible
This builds trust and makes the system auditable.
D. Separate embedding and generation models
Usually:
- one embedding model for search
- one chat/generation model for answers
This lets you optimize each independently.
6) Minimal deployment blueprint
For a small internal pilot
- LLM server: vLLM on a GPU VM
- Vector DB: Postgres + pgvector
- Backend: FastAPI
- Auth: SSO via Okta/Azure AD/Google Workspace
- Docs: Confluence/SharePoint ingestion jobs
- Deployment: Docker Compose or a single Kubernetes namespace
For production
- Kubernetes
- autoscaled LLM inference pods
- separate ingestion workers
- vector DB with backups and HA
- observability stack (Prometheus/Grafana/OpenTelemetry)
- secrets manager
- per-request audit logs
7) Example serving flow
Inference request flow
1. User asks: "What is our travel reimbursement policy?"
2. Backend verifies user identity
3. Backend queries vector DB using user ACL filters
4. Top relevant policy chunks are retrieved
5. Backend formats prompt:
- system: answer only from provided context
- context: policy snippets
- question: user query
6. LLM server generates response
7. Backend returns answer + citations
8) Good operational safeguards
- Rate limiting
- Prompt and output logging with redaction
- Cache frequent questions
- Timeouts and fallback responses
- Hallucination guardrails
- Evaluation set with real internal questions
- Human feedback loop for incorrect answers
9) What to do first
If you’re starting from scratch, do this in order:
- Pick one internal source (e.g. Confluence)
- Build ingestion + chunking
- Store embeddings in a vector DB
- Serve a model with vLLM or a hosted API
- Add ACL filtering
- Add citations
- Test with 20–50 real questions
- Measure answer quality and retrieval accuracy
- Scale to more doc sources
10) If you want a concrete starting stack
A strong default stack for many teams:
- Backend/API: FastAPI
- LLM serving: vLLM
- Embedding model: e5-large, bge-large, or similar
- Vector DB: pgvector or Weaviate
- Auth: Azure AD / Okta
- Deployment: Kubernetes or a GPU VM for pilot
- Observability: OpenTelemetry + Grafana
- Secrets: Vault or cloud secrets manager
If you want, I can also give you:
- a reference architecture diagram,
- a Kubernetes deployment example, or
- a step-by-step implementation plan for a specific stack like Azure, AWS, or on-prem.