Prompt
How do I build a prompt tracking system for ChatGPT and Perplexity?
Latest observation
I'm assuming you mean tracking how ChatGPT and Perplexity answer a fixed set of prompts over time (brand mentions, citations, competitors), which is often call…Designing a prompt tracking system for AI search platforms.Designing a prompt tracking system for AI search platforms.I'm assuming you mean tracking how ChatGPT and Perplexity answer a fixed set of prompts over time (brand mentions, citations, competitors), which is often called AI visibility or GEO tracking. If you meant logging your own prompt history, tell me and I'll adjust. Core architecture
- Build your prompt set. Start with 30-100 prompts that reflect how real users ask about your space, grouped by intent (e.g. "best X for Y", comparisons, "alternatives to Z", how-to questions). Tag each prompt with topic, funnel stage, and any persona or location. Keep the set stable so trends are comparable, and add new prompts rather than editing old ones.
- Collect responses. You have three options:
Perplexity: The Sonar API is the easiest source. It returns the answer along with its citations and search results, and it behaves similarly to the consumer product. ChatGPT via API: Use the OpenAI Responses API with the web search tool enabled, otherwise you're measuring the model's training memory rather than what users see. Note that API answers differ from the consumer app (no memory or personalization, different system prompt, different model routing). ChatGPT via the actual UI: Headless browser automation (Playwright) gets you closer to what users see, but it's brittle, can conflict with terms of service, and is hard to scale. Many teams use a commercial tracker for this part instead of building it.
- Handle non-determinism. Answers vary run to run. Run each prompt 3-5 times per platform per cycle and report rates ("mentioned in 60% of runs") instead of single results. Log the model name, timestamp, and any parameters with every run, since model updates will shift your baselines.
- Store everything raw. A simple relational schema works:
prompts (id, text, topic, intent, tags) runs (id, prompt_id, platform, model, run_at, raw_response, raw_json) citations (run_id, url, domain, position) mentions (run_id, entity, position_in_answer, sentiment, is_recommended)
Postgres or even SQLite is fine to start. Keep raw responses so you can re-parse later when your extraction logic improves. 5. Extract signals. Use string matching with an alias list for brand and competitor names, then optionally a cheap LLM pass to classify sentiment, whether you were recommended vs. merely listed, and rank order. Parse citation URLs into domains to see which sources drive answers. 6. Schedule and monitor. Run daily or weekly via cron, GitHub Actions, or Airflow. Add retries with backoff, rate-limit handling, and alerts when a run fails or returns empty. 7. Report. Key metrics are mention rate, share of voice vs. competitors, average position, citation share by domain, and week-over-week change. Metabase, Streamlit, or Looker Studio on top of your database is enough. Pitfalls to plan for
API results and consumer-app results can diverge, so treat API data as a proxy, not ground truth. Location and personalization affect results, so keep test conditions consistent. Model changes cause step-changes in your data; annotate them in your charts. Don't over-read small samples. A single run flipping is noise.
Model names and API features change often, so check the current OpenAI and Perplexity docs before you build. I can write a starter Python script that runs your prompts against both APIs, stores results in SQLite, and extracts mentions and citations, if you'd like.