Prompt
How do I create embeddings for product descriptions and FAQ pages?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
To create embeddings for product descriptions and FAQ pages, the general process is:
- Collect the text
- Clean and chunk it
- Convert each chunk into an embedding vector
- Store vectors in a database/index
- Use them for search, recommendations, or Q&A
1) Decide what text to embed
For product pages, you might include:
- Product title
- Short description
- Full description
- Key features
- Specs
- Category
For FAQ pages, you usually want:
- Each question + answer pair as one unit
- Or each FAQ item separately
Example:
Question: What is the warranty? Answer: 2 years...Question: Does it support fast charging? Answer: Yes...
2) Clean and prepare the text
Before embedding:
- Remove HTML
- Normalize whitespace
- Fix encoding issues
- Keep important structured info together
Example product text:
Title: Wireless Noise-Cancelling Headphones
Description: Over-ear headphones with active noise cancellation...
Features: 30-hour battery, Bluetooth 5.3, USB-C charging
3) Chunk long text
Embedding models have input limits and work better with manageable chunks.
For:
- Product pages: chunk by section if long
- FAQ pages: usually embed each Q&A separately
Good chunk size:
- ~200–500 words per chunk, depending on the model
4) Generate embeddings
Use an embedding model such as:
- OpenAI embeddings
- SentenceTransformers
- Cohere embeddings
- Hugging Face models
Example with Python and SentenceTransformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
texts = [
"Wireless Noise-Cancelling Headphones. 30-hour battery, Bluetooth 5.3...",
"Q: What is the warranty? A: Two years from date of purchase."
]
embeddings = model.encode(texts)
print(embeddings.shape) # e.g. (2, 384)
5) Store embeddings with metadata
Store:
- embedding vector
- original text
- product ID / FAQ ID
- page URL
- category
- language
Example metadata:
{
"id": "faq_123",
"text": "Q: What is the warranty? A: Two years...",
"source": "faq",
"url": "/support/warranty",
"category": "warranty"
}
6) Use them for retrieval
When a user searches:
- Embed the query
- Compare it with stored embeddings
- Return the most similar items
Common similarity methods:
- Cosine similarity
- Dot product
- Euclidean distance
Best practices
- Embed meaningful text, not raw HTML or navigation menus
- Keep FAQ Q&A together
- For product data, combine title + key specs + description
- Rebuild embeddings when content changes
- Use metadata filtering to improve search results
Simple recommendation
For ecommerce:
- Product embeddings: one embedding per product, plus extra embeddings for major sections if the page is long
- FAQ embeddings: one embedding per question-answer pair
If you want, I can also show:
- a Python example using OpenAI embeddings, or
- how to store and query them in a vector database like FAISS, Pinecone, or Chroma.