Prompt
Help me pick a dataset for churn prediction with recent labels, low missingness, and commercial rights
Latest observation
Oct 3, 2026
OpenAI APIWeb search: off
Here’s a practical way to choose a churn dataset that fits your constraints: recent labels, low missingness, and commercial rights.
Best-fit public options to consider
1) Telco Customer Churn
- Why it fits: Very common churn benchmark, relatively clean, easy to use.
- Data quality: Low to moderate missingness; usually manageable.
- Recent labels: Not truly “recent” in a live-business sense, but labels are straightforward and usable.
- Commercial rights: Depends on the source/license. You need to verify the dataset license carefully before commercial use.
- Good for: Fast prototyping and baseline churn modeling.
2) KKBox Churn Prediction Dataset
- Why it fits: Real-world subscription churn from a music streaming service.
- Data quality: Generally decent; some tables may require preprocessing, but missingness is not typically the main issue.
- Recent labels: Better than many older churn datasets because it reflects subscription behavior over time.
- Commercial rights: Often available through competition/research platforms, but commercial usage may be restricted. Check the exact terms.
- Good for: More realistic churn modeling with event/history data.
3) Online retail / subscription SaaS datasets from Kaggle or GitHub
- Why it fits: Sometimes have recent timestamps and lower missingness.
- Data quality: Varies a lot.
- Commercial rights: Often not safe for commercial use unless explicitly licensed for it.
- Good for: Exploration only, unless licensing is very clear.
What to prioritize
If your requirements are strict, use this filter:
-
License allows commercial use
- Look for CC0, CC-BY 4.0, or a clearly commercial-friendly license.
- Avoid datasets with “research only,” “non-commercial,” or unclear provenance.
-
Recent labels
- Prefer datasets with labels from the last 1–3 years if you need current behavior.
- Check whether churn is defined by a recent observation window and follow-up period.
-
Low missingness
- Aim for <5–10% missingness in key features.
- Avoid datasets where churn labels are present but core predictors are heavily incomplete.
My recommendation
If you need something usable quickly:
- Start with Telco Customer Churn for modeling workflow validation.
- But only use it commercially if the source license explicitly permits it.
If you need stronger realism:
- Consider KKBox, but verify whether the competition/data terms permit commercial use.
If commercial use is non-negotiable:
- You may be better off using a commercially licensed provider or building your own churn dataset from CRM/product logs, because many public churn datasets have licensing ambiguity.
Quick checklist before choosing
- Is the license explicit and commercial-friendly?
- Are the labels recent enough for your business case?
- Is missingness low in the main predictive features?
- Are there any privacy/PII restrictions?
- Does the dataset match your churn definition?
If you want, I can give you a shortlist of specific datasets with license notes or help you find one by industry (telecom, SaaS, retail, fintech, etc.).