Prompt
I'm building a program to compare our AI visibility against competitors, any advice?
Latest observation
Yes — if you want to compare your AI visibility against competitors, treat it like a measurement problem first and a marketing problem second.
1) Define “AI visibility” precisely
Before building anything, decide what you’re measuring. Common dimensions:
- Presence: Are you mentioned at all in AI answers?
- Position: If mentioned, are you first, second, or buried?
- Sentiment/role: Are you recommended, compared neutrally, or criticized?
- Share of voice: How often do you appear vs competitors across a prompt set?
- Source influence: Which pages/domains are driving the mention?
- Citation rate: Do AI systems cite your site or third-party pages?
A lot of teams mix these together and get noisy results.
2) Build a prompt set, not just keyword queries
AI visibility depends heavily on prompt phrasing. Create a prompt library across intent types, such as:
- Discovery: “Best tools for X”
- Comparison: “X vs Y”
- Category: “Top platforms for Z”
- Problem/solution: “How do I solve [problem]?”
- Purchase intent: “What should I buy for…”
- Enterprise: “Best vendor for large teams”
- Regional: “Best options in [country]”
For each prompt, include variations in tone and specificity. This helps reduce overfitting to one answer style.
3) Track across multiple models and surfaces
Visibility differs by system. If possible, compare across:
- ChatGPT
- Gemini
- Claude
- Perplexity
- Bing/Copilot
- AI Overviews / search-generated answers
Each system has different retrieval behavior, citations, and phrasing. Don’t assume one model represents “AI visibility” overall.
4) Use a consistent evaluation rubric
For every prompt-response pair, score:
- Mentioned? yes/no
- Rank/order
- Recommendation strength: strong / moderate / weak / absent
- Accuracy: is the answer correct?
- Competitor count: how many rivals appear?
- Citation presence: cited / uncited
- Brand framing: positive / neutral / negative
You can automate some of this, but keep a human review layer for quality checks.
5) Separate retrieval from generation
If your program can detect sources, try to determine:
- Did the model mention you because it “knows” you from training data?
- Or because it retrieved current web sources?
- Which source pages were used?
This distinction matters because your strategy changes:
- Training-data presence is harder to influence directly.
- Retrieval presence can be improved through SEO, structured data, PR, and citations on authoritative pages.
6) Use normalized competitor comparison
Raw mention counts are misleading. Normalize by:
- prompt volume,
- prompt category,
- model,
- region,
- time window.
Then compute something like:
- Visibility score per competitor
- Weighted share of voice
- Category dominance
- Citation share
This lets you compare fairly instead of just counting mentions.
7) Watch for prompt drift and model variability
AI answers vary from run to run. To reduce noise:
- Run each prompt multiple times
- Use temperature as low as the platform allows
- Store timestamps and model versions
- Re-test periodically
A single response is not enough for benchmarking.
8) Tie visibility to business outcomes
Visibility is only useful if it predicts something:
- referral traffic,
- branded search lift,
- demo requests,
- conversion rate,
- share of pipeline in categories where AI is influential.
If you can, correlate AI visibility with downstream outcomes over time.
9) Build a source attribution layer
When your brand appears, capture:
- cited URLs,
- domain types,
- whether it’s your own content or third-party content,
- mention context.
Then identify patterns:
- Are review sites driving visibility?
- Are listicles outperforming your homepage?
- Are competitor-owned assets suppressing you?
10) Practical architecture for your program
A simple pipeline:
-
Prompt store
Maintain categorized prompts and competitor lists. -
Query runner
Send prompts to selected AI systems on a schedule. -
Response parser
Extract mentions, citations, rankings, sentiment. -
Scoring engine
Normalize and compute visibility metrics. -
Dashboard
Show trends by competitor, category, and model. -
Alerting
Notify when visibility drops or a competitor gains share.
11) Common pitfalls
- Comparing different prompt types as if they’re equivalent
- Using only one model
- Ignoring citations and source quality
- Treating one response as representative
- Over-indexing on “mentions” without business context
- Not accounting for regional or language differences
12) What I’d recommend as a minimum viable setup
If you want to start lean:
- 50–200 prompts
- 5–10 competitors
- 2–4 AI surfaces
- 3–5 repeated runs per prompt
- Human review for a subset
- A weighted visibility score + citation tracking
That’s enough to get a useful benchmark without overengineering.
If you want, I can help you design:
- a scoring model,
- a database schema,
- or a sample Python workflow for running and analyzing the comparisons.