Prompt
How do I use click and purchase data to rank recommendations?
Latest observation
A common way to rank recommendations with click and purchase data is to treat them as different strengths of feedback and train a model to predict a weighted relevance score.
1) Decide what “good” means
Usually:
- Purchase = strong positive signal
- Click = weaker positive signal
- No interaction = unknown or negative, depending on exposure
A simple ranking target might be:
[ \text{score} = 1 \cdot \text{click} + 5 \cdot \text{purchase} ]
or any weighting that matches your business value.
2) Build training examples
For each user-item impression, create features such as:
- user features
- item features
- context features
- historical interactions
And label the outcome:
- clicked = 1/0
- purchased = 1/0
If both happen, the item gets a stronger label.
3) Use a weighted loss
You can train a ranking model with:
- weighted classification: purchase examples get higher weight
- pairwise ranking: items purchased/clicked rank above non-clicked items
- learning-to-rank methods like LambdaMART, XGBoost ranker, LightGBM ranker
Example weights:
- purchase: 5
- click only: 1
- no click/no purchase: 0 or small negative weight
4) Predict a recommendation score
At serving time, produce a score like:
[ \text{final score} = P(\text{click}) + \alpha \cdot P(\text{purchase}) ]
or train one model to directly estimate expected value:
[ \text{expected value} = P(\text{click}) \times v_{click} + P(\text{purchase}) \times v_{purchase} ]
where (v) is business value.
5) Rank by the score
Sort candidate items by the final score descending.
Practical approaches
A. Simple weighted labels
Best if you want something easy:
- label purchase = 5
- click = 1
- no interaction = 0
Train a regression or ranking model.
B. Two-stage model
- Candidate generation using clicks
- Re-rank using purchases more heavily
This works well because purchases are rarer and more valuable.
C. Multi-objective ranking
Model click probability and purchase probability separately, then combine: [ \text{rank score} = w_1 P(click) + w_2 P(purchase) ]
Important caveat: exposure bias
A missing click does not always mean the item was bad — the user may never have seen it. Ideally, train only on impressions or use logged exposure data.
Good starting recipe
If you want a strong baseline:
- Collect impression-level data
- Label click and purchase separately
- Train a ranker with:
- purchase weight = 5
- click weight = 1
- Evaluate with offline ranking metrics:
- NDCG
- MAP
- Recall@K
- Tune the weights based on online A/B tests
If you want, I can also show:
- a Python example with LightGBM/XGBoost ranking
- a scoring formula for click + purchase data
- or a full pipeline for recommendation ranking.