Prompt
semantic search api for documents
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
Here’s a practical design for a semantic search API for documents.
Core idea
Instead of matching keywords only, semantic search:
- Breaks documents into chunks
- Converts chunks into embeddings
- Stores embeddings in a vector database
- On query, embeds the query and finds the most similar chunks/documents
Typical API endpoints
1. Ingest documents
POST /documents
Uploads a document and indexes it for search.
Request
{
"title": "Employee Handbook",
"content": "Full document text here...",
"metadata": {
"department": "HR",
"author": "Jane"
}
}
Response
{
"document_id": "doc_123",
"status": "indexed"
}
2. Search documents
POST /search
Searches semantically across indexed documents.
Request
{
"query": "What is the vacation policy?",
"top_k": 5,
"filters": {
"department": "HR"
}
}
Response
{
"query": "What is the vacation policy?",
"results": [
{
"document_id": "doc_123",
"chunk_id": "chunk_9",
"title": "Employee Handbook",
"score": 0.92,
"snippet": "Employees accrue 20 days of paid vacation..."
}
]
}
3. Get a document
GET /documents/{document_id}
Returns document metadata and optionally full text.
4. Delete a document
DELETE /documents/{document_id}
Removes document and its vectors from the index.
Suggested architecture
Components
- API server: FastAPI / Flask / Node.js
- Embedding model: OpenAI embeddings, sentence-transformers, etc.
- Vector database: Pinecone, Weaviate, Qdrant, Milvus, FAISS
- Document storage: PostgreSQL, S3, Elasticsearch, or plain DB
- Chunking service: splits long docs into searchable segments
Indexing flow
- Receive document
- Extract text
- Chunk text
- Generate embeddings for each chunk
- Store chunks + embeddings + metadata
Search flow
- Receive query
- Generate query embedding
- Vector similarity search
- Apply metadata filters
- Return top matches
Example FastAPI shape
from fastapi import FastAPI
from pydantic import BaseModel
from typing import Optional, Dict
app = FastAPI()
class DocumentCreate(BaseModel):
title: str
content: str
metadata: Optional[Dict] = {}
class SearchRequest(BaseModel):
query: str
top_k: int = 5
filters: Optional[Dict] = {}
@app.post("/documents")
def add_document(doc: DocumentCreate):
# chunk -> embed -> store
return {"document_id": "doc_123", "status": "indexed"}
@app.post("/search")
def search(req: SearchRequest):
# embed query -> vector search
return {
"query": req.query,
"results": []
}
Useful features to include
- Filters: by date, department, file type, tags
- Hybrid search: combine keyword + vector search
- Reranking: re-score top results with a cross-encoder/LLM
- Pagination
- Access control: per-user/per-role document permissions
- Chunk previews: return matching snippet
- Highlighting: show matched semantic passage
If you want a production-ready schema
I can provide:
- a full OpenAPI spec
- a FastAPI implementation
- a Node.js/Express implementation
- a design using Postgres + pgvector
- a design using Pinecone/Qdrant
If you want, I can generate a complete API spec next.