Product Metrics and Experimentation
Version: 0.2
1. Measurement position
Measurement supports judgement; it does not replace it. A metric is a representation of an outcome under defined conditions, not the outcome itself.
HIF requires a chain:
human or organisational goal
→ observable signal
→ operational metric
→ decision rule
→ intervention
→ evaluation of intended and unintended effects
A product MUST NOT optimise a proxy while concealing harm to task success, accessibility, trust, safety or human agency.
2. Goals, signals and metrics
Define a goal in human and product terms before choosing data. For each goal record:
- intended population and context;
- desired and unacceptable outcomes;
- leading and lagging signals;
- metric definition, denominator and observation window;
- data provenance and known exclusions;
- baseline, uncertainty and expected practical effect;
- owner and decision enabled by the metric;
- guardrails and review date.
Metrics without an associated decision SHOULD be removed or reclassified as exploratory.
3. HEART dimensions
The HEART framework provides five reusable perspectives:
- Happiness — satisfaction, perceived ease, confidence and sentiment;
- Engagement — meaningful depth or frequency of use;
- Adoption — uptake by an intended population;
- Retention — continued value over a relevant interval;
- Task success — effectiveness, efficiency and error behaviour.
Not every product needs every dimension. Engagement MUST NOT be treated as intrinsically beneficial; increased use may indicate friction, compulsion or failure to complete.
4. Task metrics
For a defined task corpus, useful measures include:
- successful completion and completion with assistance;
- time on task, interpreted with task type and expertise;
- error rate, severity and recovery;
- first-attempt success;
- path deviation and repeated work;
- comprehension and consequence prediction;
- confidence calibrated against actual success;
- accessibility barriers by input and assistive technology.
Average values MUST be accompanied by distribution, sample size and relevant segments. A faster result is not better when it reduces understanding or increases consequential error.
5. Experience and trust
Self-report measures SHOULD use validated instruments when the construct and context fit. A single satisfaction score MUST NOT stand in for usability, accessibility or trust.
Trust measures SHOULD distinguish:
- perceived competence;
- predictability and transparency;
- control and correctability;
- privacy and security expectations;
- justified reliance.
High trust is not always desirable. The target is calibrated trust: reliance proportionate to actual capability and uncertainty.
6. Operational quality
Interface operations SHOULD monitor:
- availability of primary journeys;
- latency and responsiveness at relevant percentiles;
- client errors, failed submissions and abandoned recovery;
- data loss and duplicate action;
- accessibility regressions;
- browser, device, locale and input disparities;
- support contacts and complaint themes;
- rollback and remediation time.
Technical health metrics are diagnostic signals, not proof of successful human outcomes.
7. Segmentation and equity
Aggregate improvement MAY conceal harm. Analyse segments justified by the product and collected ethically, including device capability, connection quality, locale, new versus experienced users and accessibility modes.
Small samples MUST be reported with uncertainty and MUST NOT be used to expose or stereotype individuals. Absence of demographic data MUST NOT be presented as evidence of equal outcomes.
8. Experiment design
Before a controlled experiment record:
- hypothesis and mechanism;
- primary outcome;
- guardrail metrics;
- unit of assignment and contamination risk;
- target population and exclusions;
- minimum practically important effect;
- sample-size and duration rationale;
- stopping and analysis plan;
- novelty, learning and carry-over risks;
- accessibility, privacy, security and safety review.
Do not expose people to variants that violate established safety, accessibility, privacy or informed-choice requirements merely to measure the damage.
9. Interpretation
Statistical significance does not establish practical value, causality outside the design, long-term benefit or transfer to another population. Report:
- absolute and relative effect;
- confidence interval or another uncertainty account;
- missing data and exclusions;
- implementation fidelity;
- guardrail outcomes;
- plausible alternative explanations;
- duration and generalisation limits.
Repeated peeking, metric switching and selective segmentation MUST be disclosed.
10. Qualitative and quantitative triangulation
Telemetry explains what occurred in the instrumented system. It rarely explains why, what was attempted but impossible, or who never reached the product.
Combine operational data with representative task studies, interviews, accessibility evaluation, support evidence and expert model review. Conflicting evidence is a research result requiring explanation, not a reason to discard the inconvenient method.
11. Audit baseline and remediation value
A client audit SHOULD establish a reproducible baseline before remediation. For each accepted finding identify:
- affected task and population;
- baseline evidence;
- expected mechanism of improvement;
- acceptance criterion;
- measure and observation window;
- regression guardrail;
- residual risk.
Commercial value MUST NOT be fabricated from generic conversion multipliers. Financial estimates SHOULD present assumptions and a range, and MUST remain separate from verified usability or conformance findings.
12. Data ethics
Measurement MUST follow data minimisation, purpose limitation and appropriate retention. People MUST NOT be deceived into consequential research participation. Session replay, sensitive-field capture and cross-context tracking require explicit risk review and lawful authority.
Teams MUST document who can access raw evidence and how participant identity is protected in reports, recordings and issue trackers.