Prompt

How do I set up model serving for a RAG app connected to our internal docs?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 1, 2026

Latest observation

Oct 1, 2026OpenAI APIWeb search: off

Here’s a practical way to set up model serving for a RAG app that answers questions over your internal docs.

1) Use this basic RAG serving architecture

User → API → Retrieval → Prompt assembly → LLM inference → Response

Typical components:

  • Document store: internal docs in SharePoint, Confluence, Google Drive, Git, PDFs, etc.
  • Ingestion pipeline: cleans, chunks, tags, and embeds docs
  • Vector store: stores embeddings for semantic search
  • Retrieval service: fetches top-k relevant chunks for each question
  • LLM serving layer: hosts the model that writes the answer
  • App/API layer: chat UI or backend endpoint
  • Auth/ACL layer: ensures users only retrieve docs they’re allowed to see

2) Choose your model serving approach

Option A: Hosted API model

Use a managed model endpoint from a provider if:

  • you want fastest setup
  • you don’t need strict on-prem hosting
  • your data handling policy allows it

Pros:

  • low ops burden
  • easy scaling
  • simple deployment

Cons:

  • data residency/compliance concerns
  • cost can grow with usage
  • less control over latency/caching

Option B: Self-hosted model

Host an open-weight model in your environment if:

  • docs are sensitive
  • you need network isolation
  • you want tighter cost/control

Common serving tools:

  • vLLM
  • TGI (Text Generation Inference)
  • Hugging Face Transformers + FastAPI
  • SGLang
  • NVIDIA Triton in some setups

Pros:

  • stronger control over data/security
  • can run inside VPC/on-prem
  • predictable behavior

Cons:

  • GPU ops complexity
  • scaling and reliability are on you
  • model upgrades/testing needed

3) Recommended reference architecture

Ingestion path

  1. Pull documents from internal sources
  2. Extract text and metadata
  3. Chunk into passages
  4. Generate embeddings
  5. Store in vector DB plus metadata store

Query path

  1. User asks a question
  2. Authenticate user
  3. Apply ACL-aware retrieval
  4. Retrieve top chunks
  5. Build prompt with:
    • system instructions
    • user question
    • retrieved context
    • citation metadata
  6. Send prompt to served LLM
  7. Return answer with citations

4) Pick your serving stack

If you want the simplest self-hosted serving

  • vLLM for model serving
  • FastAPI for your app layer
  • Postgres + pgvector, Pinecone, Weaviate, Milvus, or OpenSearch for retrieval
  • Redis for caching
  • Kubernetes if you need scaling; otherwise start with a single VM or ECS task

If you want enterprise control

  • Put the LLM server in a private subnet
  • Use service-to-service auth
  • Log prompts/outputs carefully
  • Add access controls at retrieval time, not only at answer time

5) Important design decisions

A. Don’t only secure the model endpoint

The bigger risk is usually retrieval leakage.

You need ACL-aware retrieval:

  • tag every chunk with source doc permissions
  • filter retrieval results by user identity/groups
  • do not send unauthorized chunks to the model

B. Keep prompts short and relevant

For internal docs, retrieval quality matters a lot.

  • use chunk sizes around 300–800 tokens as a starting point
  • retrieve top 3–8 chunks
  • re-rank if needed
  • include only the most relevant snippets

C. Add citations

Always return:

  • source document title
  • section/page
  • link if possible

This builds trust and makes the system auditable.

D. Separate embedding and generation models

Usually:

  • one embedding model for search
  • one chat/generation model for answers

This lets you optimize each independently.

6) Minimal deployment blueprint

For a small internal pilot

  • LLM server: vLLM on a GPU VM
  • Vector DB: Postgres + pgvector
  • Backend: FastAPI
  • Auth: SSO via Okta/Azure AD/Google Workspace
  • Docs: Confluence/SharePoint ingestion jobs
  • Deployment: Docker Compose or a single Kubernetes namespace

For production

  • Kubernetes
  • autoscaled LLM inference pods
  • separate ingestion workers
  • vector DB with backups and HA
  • observability stack (Prometheus/Grafana/OpenTelemetry)
  • secrets manager
  • per-request audit logs

7) Example serving flow

Inference request flow

1. User asks: "What is our travel reimbursement policy?"
2. Backend verifies user identity
3. Backend queries vector DB using user ACL filters
4. Top relevant policy chunks are retrieved
5. Backend formats prompt:
   - system: answer only from provided context
   - context: policy snippets
   - question: user query
6. LLM server generates response
7. Backend returns answer + citations

8) Good operational safeguards

  • Rate limiting
  • Prompt and output logging with redaction
  • Cache frequent questions
  • Timeouts and fallback responses
  • Hallucination guardrails
  • Evaluation set with real internal questions
  • Human feedback loop for incorrect answers

9) What to do first

If you’re starting from scratch, do this in order:

  1. Pick one internal source (e.g. Confluence)
  2. Build ingestion + chunking
  3. Store embeddings in a vector DB
  4. Serve a model with vLLM or a hosted API
  5. Add ACL filtering
  6. Add citations
  7. Test with 20–50 real questions
  8. Measure answer quality and retrieval accuracy
  9. Scale to more doc sources

10) If you want a concrete starting stack

A strong default stack for many teams:

  • Backend/API: FastAPI
  • LLM serving: vLLM
  • Embedding model: e5-large, bge-large, or similar
  • Vector DB: pgvector or Weaviate
  • Auth: Azure AD / Okta
  • Deployment: Kubernetes or a GPU VM for pilot
  • Observability: OpenTelemetry + Grafana
  • Secrets: Vault or cloud secrets manager

If you want, I can also give you:

  1. a reference architecture diagram,
  2. a Kubernetes deployment example, or
  3. a step-by-step implementation plan for a specific stack like Azure, AWS, or on-prem.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.