Prompt
How do I build a way to track branded and non-branded prompts?
Latest observation
To track branded vs. non-branded prompts, you want a system that can:
- Identify the brand terms in each prompt
- Classify the prompt as branded or non-branded
- Store the result so you can report on it later
- Measure trends over time
Here’s a practical way to build it.
1) Define what “branded” means
First, create a clear rule set.
A prompt is branded if it contains any of:
- Your company name
- Product names
- Trademarks
- Campaign names
- Common misspellings or abbreviations of those terms
A prompt is non-branded if it does not contain any of those terms.
Example:
- “Best shoes for running” → non-branded
- “Best Nike shoes for running” → branded
- “Compare Acme Cloud vs AWS” → branded if Acme Cloud is your brand/product
2) Maintain a brand dictionary
Create a list of brand terms and variants.
Example structure:
{
"brand_terms": [
"Acme",
"Acme Cloud",
"AcmeAI",
"Acmé",
"ACME"
]
}
You can store this in:
- A database table
- A config file
- A feature flag / CMS
- A lookup service
Helpful additions:
- Lowercased versions
- Aliases
- Misspellings
- Internationalized versions
3) Classify prompts with a simple rules engine
Start with keyword matching.
Logic
- Normalize the prompt: lowercase, remove punctuation if needed
- Check whether any brand term appears
- Mark as branded if yes
- Otherwise mark as non-branded
Example pseudocode
def classify_prompt(prompt, brand_terms):
normalized = prompt.lower()
for term in brand_terms:
if term.lower() in normalized:
return "branded"
return "non-branded"
This is the easiest and fastest MVP.
4) Store tracking data
For each prompt, store:
- Prompt text
- Timestamp
- User/session ID
- Classification: branded/non-branded
- Matched brand term
- Source channel
- Optional: language, geography, campaign
Example schema:
CREATE TABLE prompt_tracking (
id SERIAL PRIMARY KEY,
prompt ტექst TEXT NOT NULL,
created_at TIMESTAMP NOT NULL DEFAULT NOW(),
user_id TEXT,
session_id TEXT,
channel TEXT,
classification TEXT NOT NULL,
matched_brand_term TEXT
);
5) Add reporting
Once you have data, build metrics like:
- % branded prompts
- % non-branded prompts
- branded prompt volume by day/week/month
- most common brand terms
- branded vs non-branded by channel or campaign
- conversion or engagement by prompt type
Example questions:
- Are branded prompts increasing?
- Which product names are most searched?
- Do branded prompts convert better?
6) Improve accuracy
Keyword matching is a good start, but it can produce false positives/negatives.
Problems
- “apple” can mean fruit or brand
- “acme” might appear in unrelated contexts
- Misspellings and plural forms may be missed
Ways to improve
- Use exact phrase matching with boundaries
- Add context rules
- Use NER / entity recognition
- Train a classifier if you have enough labeled data
- Use a hybrid approach: rules first, model second
Example hybrid strategy:
- Exact brand match → branded
- Ambiguous term → model or manual review
- No match → non-branded
7) Add confidence and review for edge cases
For ambiguous prompts, store:
- confidence score
- reason for classification
- human review status
Example:
- “apple watch repair” → branded with high confidence
- “apple pie recipe” → non-branded or ambiguous depending on your rules
8) Build an event pipeline if volume is high
If prompts come in at scale:
- Send prompt events to Kafka / Kinesis / PubSub
- Classify in a stream processor or worker
- Write results to analytics DB / warehouse
This lets you track in near real time.
9) Consider compliance and privacy
If prompts may contain sensitive data:
- redact PII
- hash user IDs
- set retention policies
- log only what you need
10) Suggested MVP architecture
Frontend/app
→ sends prompt event
API
→ normalizes and classifies
DB
→ stores classified prompt
Dashboard
→ aggregates branded vs non-branded metrics
Example end-to-end flow
Prompt:
“Best Acme Cloud pricing for startups”
- Receive prompt
- Normalize text
- Match “Acme Cloud”
- Classify as branded
- Store:
- prompt
- branded
- matched term = “Acme Cloud”
- timestamp
- Show in dashboard/report
Recommended starting point
If you want the fastest path:
- Build a brand-term dictionary
- Use keyword/phrase matching
- Store classifications in a table
- Add dashboards
- Later, improve with NLP or ML
If you want, I can also give you:
- a database schema
- a Python/JavaScript implementation
- a SQL dashboard query
- or a more advanced ML approach for branded prompt detection
Brands