Prompt
Which embedding model should I use for classification on my internal docs?
Latest observation
Oct 10, 2026
OpenAI APIWeb search: off
For classification on internal documents, the best embedding model is usually:
Short answer
Use a strong general-purpose text embedding model from a reputable provider, then train a simple classifier on top.
Good defaults:
- OpenAI:
text-embedding-3-largefor best quality, ortext-embedding-3-smallfor lower cost - Cohere:
embed-english-v3.0/embed-multilingual-v3.0 - Sentence Transformers (open source):
all-mpnet-base-v2is a solid baseline, but newer retrieval-oriented models may be better depending on your data
What to pick
If you want best performance
text-embedding-3-large- Best when accuracy matters more than cost
- Works well on messy, short, and medium-length enterprise text
If you want a cheaper option
text-embedding-3-small- Often good enough for internal classification tasks
- Much cheaper and faster
If you need on-prem / no external API
- Use a strong open-source embedding model, e.g.:
bge-large-en-v1.5e5-large-v2all-mpnet-base-v2as a simple baseline
Important note for classification
Embeddings are usually not the classifier themselves. The usual recipe is:
- Embed each document
- Train a classifier on those vectors
- logistic regression
- linear SVM
- XGBoost
- small MLP
- Evaluate on a held-out set
This often outperforms prompting a language model directly for stable document classification.
Practical recommendation
If you have no constraints:
- Start with
text-embedding-3-large - Train a logistic regression or linear SVM
- Compare against
text-embedding-3-smallfor cost/perf
If your docs are long
For long internal docs, consider:
- embedding the title + summary + key sections
- chunking the document and pooling chunk embeddings
- using only the most relevant section for classification if labels depend on specific content
If you tell me:
- document type,
- number of labels,
- language(s),
- whether you need on-prem,
- and approximate dataset size,
I can recommend a more specific model and training setup.