Prompt
How do I tell if my chatbot got worse after a prompt change?
Latest observation
A good way is to compare the chatbot before vs. after the prompt change on the same set of test conversations and measure whether key outcomes got worse.
Practical ways to tell
1. Create a small benchmark set
Use 20–100 representative prompts your users actually ask:
- easy questions
- tricky edge cases
- multi-turn conversations
- unsafe/ambiguous requests
- formatting-sensitive tasks
Run the old prompt and the new prompt on exactly the same set.
2. Compare with clear metrics
Pick metrics that matter for your bot:
- Task success rate: Did it answer correctly or complete the task?
- Helpfulness / relevance: Is the answer on-topic and useful?
- Hallucination rate: Did it invent facts?
- Policy/safety compliance: Did it refuse appropriately?
- Format adherence: Did it follow the required structure?
- Conversation quality: Did it maintain context over turns?
If the new prompt is worse, you’ll usually see:
- more incorrect answers
- more refusals when it shouldn’t refuse
- more verbose or less structured outputs
- more context loss in longer chats
3. Use pairwise human review
Have a person compare old vs. new outputs side by side and answer:
- Which is better?
- Which is more correct?
- Which is more helpful?
- Which is safer?
This is often more reliable than trying to judge each answer independently.
4. Track user-facing signals
If this is a live bot, watch for changes in:
- user thumbs-down rate
- conversation abandonment
- escalation to human support
- re-ask rate (“That didn’t answer my question”)
- time to resolution
If these worsen after the prompt update, that’s a strong signal.
5. Check for regressions in specific categories
Prompts often improve one thing while harming another. Test separately for:
- factual QA
- creative writing
- tool use / function calling
- JSON or schema output
- long-context memory
- instruction following
- safety behavior
A simple rule of thumb
Your chatbot likely got worse if the new prompt:
- fails more often on your benchmark set,
- produces less consistent outputs,
- gets worse ratings in side-by-side review,
- or shows a negative shift in live metrics.
Best practice
Before changing the prompt, save:
- the old prompt
- a test set
- a scoring rubric
Then after the change, run both prompts on the same test set and compare results. That’s the clearest way to detect regressions.
If you want, I can give you a simple evaluation template or a spreadsheet rubric you can use to test prompt changes.