Prompt
I’m building an attribution model for multi-device customer journeys, what usually breaks?
Latest observation
A lot breaks in multi-device attribution—usually not the math first, but the identity, data, and timing assumptions underneath it. The most common failure points:
1) Identity resolution is weaker than people think
- Same person appears as multiple users/devices.
- Household/shared devices get merged incorrectly.
- Logged-in and logged-out behavior don’t stitch cleanly.
- Cross-device matching is incomplete, delayed, or biased toward certain users.
Symptom: attribution overstates channels that happen late in journeys or on logged-in devices.
2) Missing or biased event capture
- Ad blockers, browser privacy changes, app tracking limits, and cookie loss cause gaps.
- One device/channel may be tracked better than another.
- Offline conversions or in-store touchpoints are often missing.
Symptom: “last touch” or certain channels look artificially strong because they’re the ones with better observability.
3) Time-window mismatch
- Touchpoints and conversions aren’t aligned consistently.
- Lookback windows differ by platform or device.
- Delayed conversions can get assigned incorrectly or dropped.
Symptom: attribution shifts when you change the window, often dramatically.
4) Cross-device path ordering is noisy
- You may know a touch happened, but not its exact sequence across devices.
- Event clocks can differ.
- Sessions get fragmented across device switches.
Symptom: path-based models mis-rank early vs. late touches.
5) Selection bias in logged-in users
- Cross-device data often only exists for users who log in.
- Logged-in users are usually more engaged, higher value, and have different behavior.
Symptom: model learned on stitched identities doesn’t generalize to the broader population.
6) Channel duplication and overlap
- The same impression/click can be reported by multiple systems.
- Retargeting, organic search, direct, email, and app push can all overlap in unclear ways.
- Walled gardens may report conversion credit differently than your internal data.
Symptom: double counting or inflated credit for multiple channels.
7) Attribution model assumptions don’t match reality
- Markov/Shapley/MTA models often assume observable touchpoints explain conversion.
- They usually struggle with:
- incrementality vs. correlation
- latent intent
- seasonality and promotions
- offline/upper-funnel effects
Symptom: model is “accurate” in backtests but fails business sanity checks or experiments.
8) Low-volume paths and sparse data
- Many device-path combinations are rare.
- Model estimates become unstable, especially with many channels/features.
Symptom: credit swings wildly with small data changes.
9) Causal inference gets confused with prediction
- A channel may correlate with conversion without causing it.
- Retargeting is especially prone to this problem.
Symptom: the model over-credits channels that target users already likely to convert.
10) Evaluation is hard
- You rarely have a true ground truth for credit assignment.
- Offline metrics can look good while business lift doesn’t improve.
Symptom: disagreement between attribution output and holdout/geo-test results.
Practical ways to reduce breakage
- Build a clear identity hierarchy: deterministic first, probabilistic second, and keep confidence scores.
- Track data coverage by device, browser, app, and channel.
- Standardize event timestamps and lookback windows.
- Separate attribution for logged-in vs. anonymous users.
- Validate against experiments, geo tests, or incrementality studies.
- Add guardrails for shared devices and duplicate events.
- Report uncertainty, not just point estimates.
If you want, I can also give you:
- a failure-mode checklist for implementation,
- a data schema for multi-device attribution, or
- a recommended modeling approach depending on whether you have login IDs, cookies, or probabilistic identity.