Prompt
I'm building an internal knowledge base search and want the simplest embedding stack
Latest observation
If you want the simplest practical embedding stack for an internal knowledge base search, I’d recommend:
1) Pick a single embedding model
Use a general-purpose text embedding model that’s easy to call and works well out of the box.
Good simple choices:
- OpenAI
text-embedding-3-small— very easy, strong quality, low cost - If you want self-hosted/open-source:
sentence-transformers/all-MiniLM-L6-v2— lightweight and simple
For “simplest overall,” I’d choose:
- Hosted/API:
text-embedding-3-small - Local:
all-MiniLM-L6-v2
2) Store vectors in one basic vector index
For a small or medium internal KB, keep it straightforward:
- FAISS if you want local and minimal dependencies
- Postgres + pgvector if you already use Postgres and want fewer systems
- SQLite + vector extension only if you want ultra-minimal experimentation
Simplest operationally for many teams:
- Postgres + pgvector
Because it keeps documents + metadata + vectors in one place.
3) Chunk documents before embedding
Don’t embed whole docs blindly. Split them into small chunks:
- Around 300–800 tokens per chunk
- Some overlap, like 50–100 tokens
This usually improves retrieval a lot.
4) Retrieval flow
Very simple pipeline:
- User asks a question
- Embed the query
- Find nearest chunks by cosine similarity
- Return top 3–10 chunks
- Optionally pass them to an LLM for answer synthesis
5) Metadata matters
Store:
- title
- source URL/path
- section heading
- created/updated date
- access control labels
That makes results much more useful than vectors alone.
6) A minimal stack I’d actually use
If I were optimizing for simplicity:
- Embedding model:
text-embedding-3-small - Chunking: simple token-based splitter
- Vector store:
pgvector - Search: cosine similarity
- Answering: optional LLM on retrieved chunks
7) If you want the absolute simplest MVP
Even simpler:
- Use one embedding model
- Store in Postgres
- Embed only the chunks from markdown/PDF/text files
- Retrieve top 5 chunks
- Show snippets with links
No reranker, no hybrid search, no complicated orchestration at first.
My recommendation
For an internal knowledge base, start with:
Postgres + pgvector + text-embedding-3-small + basic chunking
That’s probably the best balance of:
- simplicity
- quality
- maintainability
- low ops burden
If you want, I can sketch a minimal architecture diagram or give you a tiny Python implementation for indexing and search.