Prompt
I'm building a private AI assistant over our company files. What should I use for indexing and search?
Latest observation
For a private AI assistant over company files, the best default is usually:
Use a vector database for semantic retrieval + a keyword search engine for exact matches + reranking on top.
That gives you the best balance of accuracy, speed, and explainability.
Recommended setup
1) Document storage / source of truth
Keep the original files in:
- S3 / GCS / Azure Blob, or
- your existing file store
Then extract text/metadata into an indexing pipeline.
2) Indexing layer
Split documents into chunks and index both:
- Dense embeddings for semantic search
- BM25 / inverted index for keyword search
This is often called hybrid search.
3) Search layer
Use:
- Vector search for “find things like this”
- Keyword search for exact terms, IDs, names, policy numbers, code symbols
- Reranker to improve final top results
Good technology choices
If you want something simple and production-friendly
Postgres + pgvector + full-text search
- Great if your corpus is not huge
- Easy to operate if you already use Postgres
- Good for internal tools and early production
Best when:
- under a few million chunks
- you want minimal infra
- your team prefers SQL and simplicity
If you want a dedicated search engine
OpenSearch / Elasticsearch
- Excellent for hybrid search
- Strong keyword search, filters, analytics
- Can add vector search too
- Good for enterprise document search
Best when:
- lots of metadata filtering
- compliance/audit needs
- large-scale enterprise search
- need mature operational tooling
If you want a vector-first system
Qdrant, Weaviate, Pinecone
- Strong vector retrieval
- Easy APIs
- Good performance
Best when:
- your main need is semantic retrieval
- you want faster implementation than self-building search infra
Notes:
- Qdrant: strong open-source choice, simple, fast
- Weaviate: feature-rich, good hybrid capabilities
- Pinecone: managed, low-ops, good for scale
My practical recommendation
For most private company assistants:
Hybrid search with:
- Postgres or OpenSearch for metadata + keyword search
- Qdrant or pgvector for embeddings
- Reranker like Cohere Rerank, Jina Reranker, or a local cross-encoder
If you want the simplest stack:
- Postgres + pgvector + tsvector full-text search
If you want the best search experience at enterprise scale:
- OpenSearch + vector search + reranker
If you want easiest vector infra:
- Qdrant + separate keyword search
What matters more than the database
Search quality usually depends more on:
Chunking
- Split by structure, not just fixed token length
- Preserve headings, sections, tables, and metadata
- Keep chunks small enough for retrieval but large enough for context
Metadata
Store:
- file name
- author
- department
- created/updated date
- ACL / permissions
- document type
- source path
- tags
Access control
This is critical for company files:
- enforce document-level and chunk-level permissions at retrieval time
- don’t rely on the model to “not reveal” restricted content
Reranking
A reranker often improves results more than changing vector DBs.
Incremental updates
Make sure the index can handle:
- file edits
- deletions
- permission changes
- versioning
If you’re choosing one right now
Small to medium internal assistant
Use:
- Postgres + pgvector + full-text search
Medium to large enterprise assistant
Use:
- OpenSearch for keyword + hybrid search
- add a vector index
- add a reranker
If you want a dedicated vector DB
Use:
- Qdrant
- plus a keyword engine if exact match matters
Avoid these mistakes
- Using only embeddings with no keyword search
- Ignoring permissions in retrieval
- Chunking everything into fixed-size blobs
- Not storing source metadata and offsets
- Skipping reranking
- Re-indexing everything from scratch instead of incremental updates
Bottom line
If I had to pick a default today:
For most teams:
- Postgres + pgvector + full-text search + reranker
For enterprise-grade search:
- OpenSearch + vector search + reranker
If you want, I can suggest a concrete architecture based on:
- your file types,
- approximate corpus size,
- whether you need on-prem/private cloud,
- and your stack preference (Python, Node, Java, etc.).