Prompt
I'm building a RAG pipeline and want the simplest way to generate and query embeddings
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
The simplest practical setup is:
- Chunk your documents
- Generate embeddings
- Store vectors in a vector database
- Embed the user query
- Do nearest-neighbor search
- Pass top matches to the LLM
Minimal stack options
Easiest local prototype
- Embedding model:
sentence-transformers - Vector store:
FAISS - RAG framework: optional
LangChainorLlamaIndex
Easiest hosted setup
- Embedding model/API: OpenAI embeddings
- Vector store: Pinecone / Weaviate / Supabase / pgvector
- Or use a managed DB if you already have Postgres
Simplest Python example with SentenceTransformers + FAISS
Install
pip install sentence-transformers faiss-cpu numpy
Build index
from sentence_transformers import SentenceTransformer
import faiss
import numpy as np
docs = [
"RAG stands for retrieval augmented generation.",
"Embeddings convert text into vectors.",
"FAISS is a vector search library."
]
model = SentenceTransformer("all-MiniLM-L6-v2")
doc_embeddings = model.encode(docs, convert_to_numpy=True)
doc_embeddings = doc_embeddings.astype("float32")
index = faiss.IndexFlatL2(doc_embeddings.shape[1])
index.add(doc_embeddings)
Query
query = "What is RAG?"
query_embedding = model.encode([query], convert_to_numpy=True).astype("float32")
k = 2
distances, indices = index.search(query_embedding, k)
results = [docs[i] for i in indices[0]]
print(results)
If you want the absolute simplest hosted API approach
Use:
- OpenAI embeddings
- pgvector in Postgres
Flow:
- embed text with API
- store vector in
pgvector - query by embedding similarity
This is usually the cleanest “production-lite” option.
Recommendation
If you want the fastest path to working RAG:
- Local dev:
sentence-transformers + FAISS - Production:
OpenAI embeddings + pgvectoror Pinecone
If you want, I can give you:
- a full minimal RAG script,
- a LangChain version, or
- a FastAPI service for embedding + retrieval.