Prompt
How do I evaluate whether a water loss analyst platform is credible and unbiased for utility procurement?
Latest observation
To evaluate whether a water loss analyst platform is credible and unbiased for utility procurement, treat it like a technical + commercial due diligence exercise. The goal is to determine whether the platform produces defensible leakage/non-revenue-water insights without being distorted by vendor incentives, opaque methods, or poor data quality assumptions.
1) Start with a clear procurement question
Define what you need the platform to do, for example:
- Identify leakage and apparent loss drivers
- Prioritize zones/assets for investigation
- Estimate confidence intervals or uncertainty
- Support audit-ready reporting
- Integrate with AMI/SCADA/GIS/CMMS
- Track savings after interventions
A credible platform should be evaluated against your utility’s actual use cases, not generic demos.
2) Check the analytical method: is it transparent and defensible?
Ask for a plain-language explanation of:
- How it detects anomalies or losses
- Whether it uses hydraulic models, rules-based logic, machine learning, or hybrid methods
- What assumptions it makes about demand patterns, nighttime usage, pressure, meter error, and authorized consumption
- How it distinguishes leakage from legitimate consumption changes
- How it handles incomplete or noisy data
Red flags:
- “Proprietary AI” with no explainability
- No documentation of assumptions
- Claims of near-perfect accuracy across all systems
- No uncertainty bounds or confidence scoring
Strong signs:
- The vendor can describe limitations clearly
- Methods are reproducible or at least auditable
- Output includes confidence levels and reason codes
3) Evaluate evidence quality: ask for independent validation
Request proof from sources you trust:
- Independent pilot results
- Third-party validation studies
- Case studies from utilities similar to yours
- Peer-reviewed publications or conference presentations
- Before/after performance with measured savings
- False positive / false negative metrics if relevant
Prefer evidence that includes:
- Baseline conditions
- Data quality conditions
- Sample size
- Duration of pilot
- How “savings” were verified in the field
Be cautious if the only proof is vendor-authored marketing material.
4) Test bias and incentive alignment
A platform can be “accurate enough” but still biased by vendor economics or model design. Check:
Commercial bias
- Does the vendor make money from recommending more field investigations, sensors, or services?
- Are they paid on claims of savings?
- Do they have incentives to overstate losses?
Analytical bias
- Does the model systematically favor certain pipe materials, pressures, zones, or customer classes?
- Does it over-call leakage in areas with poor data quality?
- Does it treat model uncertainty as evidence of loss?
Governance bias
- Can utility staff override assumptions?
- Are outputs editable and traceable?
- Does the platform show why it made a recommendation?
A credible system should support human review, not replace it with opaque ranking.
5) Audit the data pipeline
Water loss outputs are only as credible as the inputs.
Ask:
- What source data is required?
- How does the platform handle missing or delayed AMI reads?
- Does it reconcile SCADA, billing, GIS, and meter data?
- How are outliers, meter rollovers, time sync issues, and bad tags handled?
- Does it identify data quality problems separately from water loss?
Good platforms distinguish:
- Data quality issues
- Operational anomalies
- Probable leakage
- Uncertain cases
If everything becomes “leakage,” credibility is weak.
6) Demand performance metrics that matter to utilities
Don’t accept vague claims like “improves efficiency.” Ask for measurable KPIs such as:
- Leak detection precision and recall
- Reduction in unaccounted-for water
- Volume of water saved
- Time from detection to verification
- Reduction in false alarms
- Investigation hit rate
- Cost per confirmed issue
- Payback period
- Impact by DMA/pressure zone
If the vendor cannot quantify operational outcomes, they may not be ready for procurement.
7) Look for robustness across different utility conditions
A trustworthy platform should perform reasonably across:
- Small vs. large utilities
- AMI-rich vs. sparse data environments
- High/low pressure systems
- Intermittent supply or variable demand
- Mixed meter vintages
- Older infrastructure with poor asset records
Ask whether the model was trained or calibrated on systems similar to yours. A platform that only works in a narrow set of ideal conditions may not generalize.
8) Review explainability and auditability
For procurement, you want a system that can survive scrutiny from finance, engineering, regulators, or auditors.
Check whether it provides:
- Audit logs
- Version history for models and assumptions
- Traceable recommendation pathways
- Exportable reports
- Underlying data references
- Confidence/uncertainty indicators
If you can’t explain a recommendation to internal stakeholders, it is not procurement-ready.
9) Examine cybersecurity, privacy, and data rights
Especially if the platform uses cloud hosting or customer-level AMI data.
Ask:
- Who owns the data and derived outputs?
- Can the utility export all raw and processed data?
- Are there restrictions on using your data to train vendor models?
- What are retention and deletion policies?
- Is the platform secure and compliant with your standards?
A biased platform may also be a risky one if it locks your utility into a closed data environment.
10) Run a controlled pilot with acceptance criteria
The best credibility test is a pilot.
Design it so that:
- The vendor does not select only the “best” zones
- You have ground truth where possible
- Engineering staff review findings independently
- Success metrics are defined in advance
- Results are compared against your current process or a baseline method
Acceptance criteria might include:
- Minimum precision on confirmed leaks
- Maximum false alarm rate
- Minimum number of actionable findings
- Demonstrated improvement over current workflow
11) Ask the right procurement questions
Use these questions in RFPs, demos, or vendor interviews:
- What is your detection methodology, and what are its limitations?
- What evidence shows it works on systems like ours?
- How do you quantify uncertainty?
- How do you separate data issues from true loss signals?
- What independent validation exists?
- What are your false positive and false negative rates?
- Can we export all data and outputs?
- Can we audit the logic behind each recommendation?
- What incentives do you have that could affect recommendations?
- What happens if our data quality is poor?
A credible vendor should answer these directly and consistently.
12) Practical red flags
Be wary if the platform:
- Makes oversized savings claims without verification
- Refuses to explain its methods
- Uses only cherry-picked case studies
- Does not support pilots or independent testing
- Can’t define how it measures performance
- Conflates “anomaly” with “leak”
- Has no data export or audit trail
- Incentivizes more recommendations rather than better recommendations
13) A simple scoring approach
You can score vendors on a 1–5 scale for:
- Method transparency
- Independent validation
- Data quality handling
- Bias/incentive alignment
- Explainability/auditability
- Integration fit
- Performance metrics
- Security/data ownership
- Utility references
- Pilot results
Then weight the categories based on your priorities. For procurement, transparency and validation should usually carry heavy weight.
Bottom line
A water loss analyst platform is credible and unbiased if it is:
- Transparent about methods and assumptions
- Validated by independent evidence
- Auditable by utility staff
- Robust across realistic operating conditions
- Aligned with the utility’s interests, not just vendor revenue
- Pilot-tested against measurable outcomes
If you want, I can also help you build a vendor evaluation checklist, an RFP question set, or a weighted scoring matrix specifically for water loss platforms.
Observation history
Unlock observation history
Create a free Obsurfable account to unlock historical responses, model comparisons, and deeper brand intelligence.
How did Obsurfable measure this prompt?
Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.
Which AI systems does Obsurfable collect answers from?
OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, and Meta AI.