Prompt
I’m building a semantic search feature over internal docs; what’s the simplest architecture?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
The simplest solid architecture for semantic search over internal docs is:
1) Ingest docs
Pull from your sources:
- Google Drive / SharePoint / Confluence / Notion / file shares
- PDFs, DOCX, HTML, Markdown, plain text
2) Extract text
Convert each document into clean text:
- OCR if needed for scans
- Keep metadata like:
- title
- source URL/path
- author
- date
- doc type
- permissions/access group
3) Chunk the text
Split docs into small overlapping chunks:
- Typical size: 200–500 tokens
- Overlap: 20–100 tokens
- Store chunk boundaries so you can cite the source later
4) Create embeddings
Run each chunk through an embedding model and store the vector.
Good default:
- Use a general-purpose embedding model
- Normalize embeddings if your vector DB expects it
5) Store in a vector index
Use a vector database or vector-capable search store:
- Pinecone, Weaviate, Milvus, Qdrant, pgvector, Elasticsearch/OpenSearch kNN
- Store:
- embedding
- chunk text
- doc metadata
- ACL/permissions info
6) Query flow
When a user searches:
- Embed the query
- Vector search top-k similar chunks
- Filter by permissions
- Optionally rerank results with a cross-encoder or LLM reranker
- Return:
- top snippets
- source document links
- highlighted passages
7) Add a hybrid search layer
For better quality, combine:
- semantic search
- keyword/BM25 search
This helps with:
- exact names
- acronyms
- IDs
- rare terms
A common simple approach:
- retrieve from both vector and keyword indexes
- merge results
- rerank
8) Keep permissions enforced
Important for internal docs:
- filter at retrieval time by user/group ACL
- do not rely on the LLM to hide unauthorized content
Minimal stack
If you want the easiest production-friendly setup:
- Ingestion/extraction: Python + Apache Tika / unstructured
- Embeddings: OpenAI or open-source embedding model
- Storage: Postgres + pgvector or Qdrant
- Search: hybrid vector + BM25
- Reranking: optional, but useful
- UI/API: simple backend endpoint returning passages + links
Simplest viable architecture
If you want the fewest moving parts:
Docs → text extraction → chunking → embeddings → pgvector/Qdrant → query embedding → top-k retrieval → return snippets
That’s enough to ship an MVP.
Common pitfalls
- Chunks too large or too small
- Ignoring ACLs
- Not storing source metadata
- Using only vector search and missing exact-term matches
- No reranking, leading to noisy top results
- Not updating embeddings when docs change
If you want, I can also give you:
- a reference architecture diagram,
- a Python implementation outline, or
- a minimal stack recommendation based on your cloud/provider.