Prompt
I'm building a brand monitoring workflow for Claude and Gemini answers, how should I set it up?
Latest observation
Here’s a practical way to set up a brand monitoring workflow for Claude and Gemini answers.
1) Define what you want to monitor
Start by being explicit about the types of issues you care about. Common categories:
-
Brand mention / presence
Does the model mention your brand at all when it should? -
Correctness
Is the information about your brand accurate? (pricing, features, availability, policies) -
Ranking / recommendation
If the model recommends options, is your brand included appropriately? -
Tone / sentiment
Is the answer neutral, positive, or negative toward your brand? -
Competitive displacement
Are competitors being recommended instead of you, and why? -
Policy / compliance
Is the model making claims you wouldn’t want repeated publicly?
2) Build a question set
Create a stable prompt set that covers likely user intents. Use a mix of:
- Head terms: “best [category] software”
- Problem-based: “How do I solve [pain point]?”
- Comparison prompts: “[Brand] vs [competitor]”
- Purchase intent: “What’s the best tool for [use case]?”
- Support questions: “How do I cancel / refund / upgrade?”
- Reputation questions: “Is [brand] reliable / safe / legit?”
A good starting set is 50–200 prompts.
Keep prompts controlled
For monitoring, keep the prompts:
- stable over time
- written in natural language
- consistent enough to compare runs week over week
3) Decide your evaluation dimensions
For each model response, score a few things:
Suggested labels
- Mentioned brand? yes/no
- Brand accuracy: correct / partially correct / incorrect
- Recommendation quality: favorable / neutral / unfavorable
- Competitor included? yes/no
- Hallucinated claims? yes/no
- Citation quality if applicable: good / weak / missing
Useful metrics
- Mention rate
- Positive mention rate
- Misstatement rate
- Share of voice vs competitors
- Rank position in recommendation lists
- Response consistency across repeated runs
4) Capture responses from both models in a controlled way
For Claude and Gemini, use the same prompt set and the same collection rules.
Control variables
- model version
- temperature
- system prompt
- tool use on/off
- date/time of run
- region/language, if relevant
If your goal is monitoring over time, reduce randomness:
- temperature near 0 or a low fixed value
- same prompt wording every run
5) Store data in a structured format
Log every prompt/response pair in a table or database.
Recommended fields
- run_id
- timestamp
- model_name
- model_version
- prompt_id
- prompt_text
- raw_response
- extracted_brand_mentions
- evaluator_scores
- flags
- reviewer_notes
If you can, store the full raw output and a normalized evaluation record.
6) Use a two-stage evaluation approach
Stage A: Automated triage
Use rules or an LLM judge to quickly label responses:
- brand mention detection
- topic classification
- sentiment
- hallucination detection
- answer length / format checks
Stage B: Human review
Have a reviewer inspect:
- low-confidence cases
- high-impact prompts
- negative or risky answers
- anything with possible factual errors
This hybrid approach is much more reliable than fully automated scoring alone.
7) Create a scoring rubric
Make the rubric explicit so the workflow is repeatable.
Example for “brand accuracy”:
- 2 = correct
- 1 = partly correct / outdated
- 0 = incorrect or misleading
Example for “recommendation quality”:
- 2 = strongly favorable and appropriate
- 1 = mentions brand but not prioritized
- 0 = omitted when expected
- -1 = unfairly negative or competitor-biased
8) Compare Claude and Gemini side by side
You’ll want direct comparisons on the same prompts.
Compare:
- whether the brand appears
- whether facts are accurate
- whether recommendations are consistent
- whether one model is more likely to hallucinate
- whether one model is more cautious or more promotional
A simple dashboard can show:
- model A vs model B
- prompt category
- trend over time
- top failure types
9) Add alerts
Set alerts for important changes.
Examples:
- brand mention rate drops by >20%
- misstatement rate exceeds threshold
- competitor share of voice spikes
- new harmful claim appears
- model begins recommending a competitor for a core use case
10) Refresh prompts regularly
Models and user behavior change, so update the test set periodically:
- add new product launches
- add new competitors
- add new customer pain points
- remove obsolete prompts
A quarterly review is a good default.
A simple workflow you can implement
Daily or weekly job
- Load prompt set
- Query Claude and Gemini
- Save raw outputs
- Run automated scoring
- Send low-confidence/high-risk cases to human review
- Update dashboard
- Trigger alerts on threshold breaches
Recommended stack
If you want a lightweight implementation:
- Prompt storage: CSV / Airtable / database
- Execution: Python script or workflow runner
- Model calls: Claude API + Gemini API
- Evaluation: rules + LLM judge + human review
- Dashboard: Looker, Metabase, Streamlit, or even Sheets initially
- Alerting: Slack, email, PagerDuty
Best practices
- Use the same prompt set across both models
- Keep temperature low for comparability
- Don’t rely only on sentiment; track factual correctness
- Separate brand monitoring from SEO monitoring
- Include competitor prompts
- Keep a human audit trail for major decisions
- Re-test after major model version changes
Example output schema
{
"run_id": "2026-10-05-cl1",
"model_name": "claude",
"model_version": "claude-3.x",
"prompt_id": "cmp_014",
"prompt_text": "What is the best tool for team note-taking?",
"raw_response": "...",
"brand_mentioned": true,
"brand_accuracy": "correct",
"sentiment": "neutral",
"competitors_mentioned": ["Notion", "Evernote"],
"risk_flags": ["none"],
"human_review_needed": false
}
If you want, I can also help you with one of these next:
- a prompt set template
- a scoring rubric
- a database schema
- a Python pipeline example for Claude + Gemini