Prompt

How do I host embeddings and a chat model together?

Artificial Intelligence / AI Infrastructure1 observationLast seen Oct 10, 2026

Latest observation

Oct 10, 2026OpenAI APIWeb search: off

To host embeddings and a chat model together, the usual pattern is:

  1. Run both models behind one API service
  2. Route requests by endpoint:
    • /embeddings → embedding model
    • /chat or /v1/chat/completions → chat model
  3. Optionally add a shared vector store so the chat model can use retrieved context

Common deployment options

1) Single server, two model workers

Best if you want simplicity.

  • One API layer (FastAPI, Flask, etc.)
  • One embedding model process
  • One chat model process
  • API routes forward to the right worker

Good for:

  • Small/medium traffic
  • Different hardware requirements for each model
  • Easy maintenance

2) One inference server that supports multiple models

If your serving stack supports it, you can load both models into one system.

Examples:

  • vLLM: great for chat/completions, not ideal for embeddings unless you use a separate embedding model service
  • TGI (Text Generation Inference): primarily generation
  • Ollama: can host multiple models locally, including embedding and chat models
  • LM Studio: local dev, not production-focused

3) Separate microservices

Best for production and scaling.

  • embeddings-service
  • chat-service
  • retrieval-service / vector DB
  • API gateway in front

This is usually the cleanest architecture.

Recommended architecture

For a RAG-style app:

Client
  ↓
API Gateway
  ├── /embeddings → Embedding service
  ├── /chat       → Chat model service
  └── /query      → Retrieve docs from vector DB → chat model

Flow

  • When indexing documents:
    • send text to embedding service
    • store vectors in vector DB
  • When chatting:
    • embed the user query
    • search vector DB
    • send retrieved context + user prompt to chat model

Example with FastAPI

from fastapi import FastAPI
from pydantic import BaseModel

app = FastAPI()

class EmbedRequest(BaseModel):
    text: str

class ChatRequest(BaseModel):
    messages: list

@app.post("/embeddings")
def embeddings(req: EmbedRequest):
    # call embedding model
    vector = [0.1, 0.2, 0.3]
    return {"embedding": vector}

@app.post("/chat")
def chat(req: ChatRequest):
    # call chat model
    response = "Hello! How can I help?"
    return {"response": response}

Then internally:

  • /embeddings uses an embedding pipeline
  • /chat uses your LLM inference backend

Practical notes

Hardware

  • Embedding models are usually much smaller
  • Chat models need more GPU memory
  • You may want:
    • embedding model on CPU or smaller GPU
    • chat model on GPU
    • or separate GPUs

Latency

  • Keep embedding requests fast
  • Batch embeddings if indexing lots of data
  • Cache frequent embeddings if appropriate

API compatibility

If you want OpenAI-style endpoints:

  • /v1/embeddings
  • /v1/chat/completions

You can implement both in one service, which makes integration easier.

If you want the easiest self-hosted setup

Use:

  • Ollama for both chat + embeddings locally
  • or FastAPI + two model backends
  • plus Qdrant / Chroma / pgvector for vector search

If you want production scaling

Use:

  • Separate embedding and chat services
  • Kubernetes or Docker Compose
  • Vector DB
  • Load balancer / API gateway

If you want, I can give you:

  1. a Docker Compose example,
  2. a FastAPI reference implementation, or
  3. a production architecture diagram.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.