Prompt
How do I test prompts before shipping an LLM feature?
Latest observation
A good way to think about prompt testing is: treat prompts like production code. You want a repeatable process that catches failures before users do.
1) Define what “good” means
Before testing, write down success criteria for the feature:
- Task success: Did it do the right thing?
- Accuracy / factuality: Is it correct?
- Style / tone: Does it sound right?
- Safety / policy compliance: Does it avoid disallowed content?
- Format validity: Is the output parseable / schema-compliant?
- Latency / cost: Is it fast and cheap enough?
If you don’t define these, prompt testing gets subjective.
2) Build a representative test set
Create a small but realistic dataset of inputs your feature will face:
- Common “happy path” examples
- Edge cases
- Ambiguous requests
- Adversarial / jailbreak attempts
- Empty, malformed, or noisy inputs
- Domain-specific tricky cases
Aim for 30–100 cases to start, then expand based on failures.
3) Use a fixed evaluation rubric
For each test case, score outputs against a rubric such as:
- Pass / fail
- Or a 1–5 scale on:
- correctness
- completeness
- adherence to instructions
- formatting
- safety
Keep the rubric explicit so different reviewers are consistent.
4) Automate what you can
Automated checks are great for objective constraints:
- JSON parses successfully
- Required fields are present
- Output length is within bounds
- No forbidden phrases / PII
- Regex or schema validation passes
- Classification label is in the allowed set
For subjective quality, use human review or a judge model as a supplement.
5) Compare prompt versions head-to-head
Don’t test prompts in isolation. Run:
- Baseline prompt vs new prompt
- Same inputs
- Same model settings
- Same decoding params
Then compare:
- win rate
- regression count
- average score
- cost and latency
This makes improvements and regressions obvious.
6) Test with model variability
LLMs are stochastic, so one run isn’t enough.
- Run each test multiple times, especially if temperature > 0
- Check consistency across runs
- Test at the exact temperature/top-p you plan to ship
If determinism matters, set temperature low and verify outputs remain stable.
7) Include adversarial and “weird” inputs
Many prompt failures show up only under stress:
- Prompt injection
- User asks to ignore instructions
- Conflicting instructions
- Extremely long inputs
- Unicode weirdness
- Inputs with misleading formatting
- Requests outside scope
These are essential if your feature touches user-provided text or tools.
8) Evaluate tool and structured-output behavior separately
If the LLM calls tools or returns structured data, test:
- Tool selection accuracy
- Correct argument formation
- Retry behavior on tool failure
- Whether it hallucinates tool results
- Whether the final answer matches the tool output
For structured output, enforce:
- schema validation
- strict JSON mode if available
- fallback behavior when invalid
9) Track regressions over time
Set up a prompt evaluation suite that runs on every change:
- prompt edits
- model version changes
- decoding parameter changes
- tool/API changes
Store results so you can see if a new change improves one metric while hurting another.
10) Do a small canary before full rollout
Even good offline tests miss real-world behavior.
- Ship to a small percentage of traffic
- Log failures and user feedback
- Monitor refusal rates, error rates, latency, cost
- Compare against the previous version
A practical workflow
A simple process that works well:
- Write prompt v1
- Build a test set of 50–100 cases
- Define a rubric
- Run baseline eval
- Edit prompt
- Re-run eval and compare
- Add failing cases to the test set
- Repeat until stable
- Canary deploy
- Monitor and iterate
Helpful tools
You can do this with:
- Spreadsheet + manual review for early stage
- Python scripts + pytest for automation
- OpenAI Evals / custom eval harnesses
- LangSmith, promptfoo, Weights & Biases, or similar eval tools
Rule of thumb
If the feature matters, don’t ship a prompt until it has:
- a test set
- a rubric
- automated checks
- regression comparison
- at least some adversarial coverage
If you want, I can also give you:
- a prompt testing checklist,
- a Python eval harness template, or
- a rubric for scoring LLM outputs.