Prompt

how to track if ChatGPT is using our docs

Technology / Seo aeo tools1 observationLast seen Sep 24, 2026

Latest observation

Sep 24, 2026OpenAI APIWeb search: off

If you want to know whether ChatGPT is using your docs, the key question is: which kind of “using” do you mean?

  1. Your docs are being used as a source in your app/assistant
    Then you can track this directly.

  2. You mean the public ChatGPT model may have learned from your docs
    Then you generally cannot reliably track that per query.

If you control the app

The best approach is to instrument retrieval:

  • Log every document retrieval
    • query
    • retrieved chunks/doc IDs
    • similarity scores
    • timestamp
    • user/session ID
  • Tag every chunk with metadata
    • doc name
    • version
    • section
    • source URL
  • Return citations
    • show which docs were used to answer
  • Measure hit rate
    • how often a doc gets retrieved
    • how often it’s cited
  • Use canary docs
    • add unique phrases to internal docs
    • see whether they appear in outputs

If you use RAG

Track the pipeline:

  • user question
  • search query sent to vector DB
  • top-k results
  • which chunks were actually passed into the model
  • final answer
  • citations

This lets you see:

  • whether the docs were retrieved
  • whether the model used them
  • whether the answer matches the source

If you mean OpenAI/ChatGPT training use

There is no simple per-request “trace” showing that the base ChatGPT model learned from your specific docs.

What you can do instead:

  • check your data-sharing/settings with the provider
  • use enterprise/business controls if available
  • keep sensitive docs out of public ingestion paths
  • use retrieval instead of fine-tuning when you need source control
  • add watermark/canary text to detect leakage

Practical setup

A simple logging schema:

{
  "session_id": "...",
  "user_query": "...",
  "retrieved_docs": [
    {"doc_id": "abc", "chunk_id": "abc-12", "score": 0.82}
  ],
  "prompt_sent_to_model": "...",
  "answer": "...",
  "citations": ["abc-12"]
}

Tools/metrics you can use

  • LangChain / LlamaIndex tracing
  • OpenTelemetry
  • Prompt logs
  • Vector DB logs
  • Evaluation sets with known answers from your docs

Best signal

If you want a strong answer to “is ChatGPT using our docs?”:

  • build a RAG assistant
  • log retrieval and citations
  • include unique canary phrases
  • test with controlled prompts

If you want, I can give you:

  1. a simple logging architecture,
  2. a Python example for tracing doc usage, or
  3. a method to detect whether answers came from your docs.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.