Design Process and Research
Status: normative method
Version: 0.1
This document defines the HIF lifecycle for discovering needs, designing interactive systems, evaluating claims and learning from operation. It applies to new products, redesigns and independent client audits.
1. Governing principles
Human-centred design MUST:
- begin with an explicit understanding of people, tasks and context;
- involve intended and affected people throughout the lifecycle;
- use evaluation to drive and revise design;
- iterate as understanding, implementation and context change;
- address the whole experience and work system;
- include multidisciplinary perspectives and skills.
The process is not a linear hand-off from research to design to engineering. Each cycle reduces a named uncertainty or risk and creates traceable evidence.
Commercial pressure, delivery speed and technical feasibility are constraints, not substitutes for human-centred evidence.
2. Lifecycle
frame
→ understand context
→ specify needs and outcomes
→ model the system and risks
→ generate alternatives
→ prototype at suitable fidelity
→ evaluate claims
→ implement and verify
→ release progressively
→ observe operation
→ revise or retire
The team MAY enter at any stage, but MUST identify missing upstream evidence.
2.1. Frame
Define:
- decision to be made and who has authority to make it;
- product, service and organisational boundaries;
- intended, affected and potentially excluded groups;
- important tasks, harms and business outcomes;
- applicable HIF profile, standards and law;
- existing evidence, assumptions and conflicts of interest;
- time, access, data and recruitment constraints.
The output is a research and design brief, not a solution description.
2.2. Understand context
Study actual or credibly simulated activity. Describe people, goals, tools, artefacts, dependencies, workarounds, interruptions, environment, authority and consequences.
Do not infer actual practice solely from:
- stakeholder descriptions;
- analytics events;
- formal process maps;
- feature requests;
- competitor screens;
- what participants say they usually do.
These are evidence streams, but each omits parts of situated behaviour.
2.3. Specify needs and outcome criteria
A user need is an outcome or capability required in context, not a feature.
When [context],
[person/role] needs to [goal or resolve a condition]
so that [meaningful outcome],
within [risk, accessibility and performance constraints].
For each need specify:
- present evidence and confidence;
- success and failure conditions;
- populations and contexts covered;
- observable outcome measures;
- applicable HIF requirements;
- unresolved questions.
2.4. Model system and risk
Create only the models that change decisions:
- actors, authority and division of labour;
- objects, identity and provenance;
- commands, preconditions and consequences;
- state transitions and persistence;
- information flows and transformations;
- journey or service dependencies;
- hazards, abuse cases and failure recovery;
- accessibility barriers and assistive-technology paths.
Models are versioned hypotheses. Contradictions between a model and observed work are findings.
2.5. Generate alternatives
At least two materially distinct alternatives SHOULD be considered for consequential design decisions. Alternatives MUST be compared against the same needs, risks and outcome criteria.
The team SHOULD include people who implement, operate, support and use the system. Participation does not transfer design accountability to participants.
2.6. Prototype at question-matched fidelity
Fidelity is selected by the claim:
| Question | Minimum useful artefact |
|---|---|
| Is the concept meaningful? | narrative, storyboard, service scenario |
| Is the information structure intelligible? | labelled content model or linked wireframe |
| Can the interaction be completed? | interactive prototype with realistic data and states |
| Does focus/semantics work? | coded prototype in the target accessibility stack |
| Is performance acceptable? | representative implementation and infrastructure |
| Is failure recoverable? | executable failure states and persistence behaviour |
A prototype MUST NOT be used to claim properties it cannot instantiate.
2.7. Implement, release, operate and retire
Design intent MUST survive implementation through semantic contracts, state models, content, test cases and acceptance criteria. Release SHOULD be progressive when risk or uncertainty is material.
Operational monitoring MUST include:
- task outcomes and failure signals;
- accessibility and compatibility regressions;
- support themes and workarounds;
- incidents, near misses and recovery;
- distributional effects across relevant groups;
- automation overrides and corrections;
- evidence review triggers.
Retirement is part of design. Export, migration, notice, continuity and deletion MUST match the persistence and rights model.
3. Research planning
Every study MUST have a protocol created before data collection:
RESEARCH-ID and owner
decision and research question
primary and secondary claims
method and rationale
population, sampling and recruitment
context, tasks and materials
measures and decision thresholds
procedure and moderator guidance
analysis plan
ethics, consent, privacy and compensation
accessibility and accommodations
pilot plan
stopping rule
limitations expected
artifact and retention plan
Exploratory findings MAY emerge, but MUST be labelled exploratory. Changing a primary outcome or analysis after observing data MUST be disclosed.
4. Selecting methods
Select methods from the decision, not habit.
| Need to know | Appropriate methods |
|---|---|
| What people do and why in context | field observation, contextual inquiry, diary, artefact analysis |
| How people understand concepts or language | interview, concept elicitation, comprehension task |
| How information should be grouped | open/closed card study plus task validation |
| Where a task fails | moderated or unmoderated task-based usability evaluation |
| Whether an expert path is efficient | KLM/GOMS where assumptions hold, instrumented task study |
| Whether one alternative causes a difference | controlled experiment or defensible quasi-experiment |
| Whether accessibility requirements hold | conformance, manual, AT and disabled-user evaluation |
| How behaviour changes over time | longitudinal study, field deployment, cohort or telemetry analysis |
| Why operational data changed | logs plus qualitative and causal investigation |
| Whether rare harm is controlled | hazard analysis, simulation, adversarial and recovery testing |
Method choice MUST identify what the method cannot establish.
5. Generative and contextual methods
5.1. Observation and contextual inquiry
Observe work where it occurs when environment, collaboration or tacit practice matters. Record actions, artefacts, interruptions and constraints separately from interpretation. Ask about concrete recent events while the relevant artefact or action is available.
The researcher MUST avoid turning observation into performance appraisal. Where observation changes risk or behaviour, record the effect.
5.2. Interviews
Interviews are evidence about accounts, meanings, expectations and remembered experience. They are not direct measurements of behaviour.
Use neutral, open questions before probes. Avoid:
- leading or compound questions;
- hypothetical preference standing in for task evidence;
- asking participants to design the solution;
- treating fluent explanation as proof of successful action;
- losing disconfirming cases during synthesis.
5.3. Diaries and longitudinal sampling
Use diaries when events are intermittent, private, mobile or spread over time. Minimise reporting burden, define event triggers and distinguish missing entries from absence of events.
5.4. Participatory design
Affected people can contribute situated expertise, evaluate priorities and create alternatives. Participation MUST be accessible, compensated fairly and clear about decision authority. It MUST NOT be used to legitimise a decision already made.
5.5. Artefact and workflow analysis
Study forms, messages, spreadsheets, physical notes, histories and unofficial tools. They often carry coordination and memory functions that a replacement system must preserve. Treat workarounds as evidence about a system relation, not automatically as non-compliance.
6. Evaluative methods
6.1. Expert review
An expert review applies the HIF Constitution, Product Profile, domain standards and task corpus. Each finding MUST include an observed condition, affected task, consequence, applicable criterion, evidence and reproducible steps.
Heuristic review finds plausible problems; it does not measure prevalence, task success or user understanding.
6.2. Cognitive walkthrough
Use a cognitive walkthrough for learnability of a specified first-use or infrequent-use path. At each action ask whether the intended person is likely to:
- form the right sub-goal;
- notice the correct action;
- associate the action with the goal;
- interpret the feedback as progress.
Record the assumed knowledge. The method does not replace observation with representative people.
6.3. Task-based usability evaluation
Tasks MUST express goals and realistic starting states without revealing the interaction sequence. Include relevant permissions, data, interruptions and failure states.
Collect:
- task outcome and quality;
- critical and non-critical errors;
- assistance and recovery;
- time where it is meaningful;
- path and state transitions;
- comprehension of consequences and persistence;
- participant-reported experience after behaviour is observed.
The moderator MUST remain neutral, use a defined assistance policy and avoid teaching one participant what later participants are expected to discover. Iteration between sessions is permitted when the study is explicitly formative; versions and exposure MUST be recorded.
6.4. Controlled comparison
For causal comparison:
- define the primary outcome and smallest meaningful effect;
- assign conditions credibly and control order/learning effects;
- keep non-target differences stable;
- estimate required sample or precision before collection;
- specify exclusions, missing data and stopping rules;
- report effect sizes and uncertainty, not only p-values;
- distinguish confirmatory from exploratory analysis.
Statistical significance is not practical importance, and a p-value is not the probability that a hypothesis is true.
6.5. Field experiments and telemetry
Telemetry observes instrumented events, not intentions or complete experience. Event semantics, exposure, identity, missingness and version MUST be defined.
A/B testing is appropriate only where:
- randomisation and interference assumptions are credible;
- the metric represents a legitimate outcome;
- foreseeable harm is bounded;
- novelty, learning and long-term effects are considered;
- accessibility and minority effects are not hidden by an aggregate;
- stopping and rollback rules exist.
Dark-pattern optimisation and experiments that manipulate material choice
without appropriate consent or governance violate HIF-TRU-001 and
HIF-TRU-005.
6.6. Accessibility evaluation
Combine:
- normative conformance evaluation;
- automated checks for machine-detectable conditions;
- manual keyboard, zoom, reflow, contrast, motion and input checks;
- assistive-technology testing with defined configurations;
- task evaluation with disabled people relevant to the context.
No layer replaces another. A sample audit cannot prove that every page conforms, and a few disabled participants cannot represent all disabilities. Report exact scope, sample, technologies, user agents and unresolved coverage.
6.7. Questionnaires
Use a questionnaire only for a defined construct and population. Prefer a validated instrument when its construct fits the decision. Preserve wording, scale, scoring and administration unless validating a modification.
Self-report workload, satisfaction, trust and preference are different constructs. None alone demonstrates task success or safety. NASA-TLX, for example, is a subjective workload instrument, not a general usability score.
7. Sampling and recruitment
There is no universal “five users” rule.
Sampling MUST follow the claim:
- purposive variation for discovering mechanisms and barriers;
- criterion sampling for a defined capability or context;
- stratified or quota sampling for known important groups;
- probability sampling when population estimates are required;
- sufficient experimental sample for the intended precision or power;
- sufficient qualitative information power for the aim, sample specificity, theory, dialogue quality and analysis strategy.
Recruitment MUST record inclusion, exclusion, channel, incentive and non-participation. Convenient colleagues and professional test participants MUST NOT be presented as representative customers without evidence.
Accessibility studies SHOULD cover relevant combinations of disability, assistive technology experience, device and task rather than diagnostic labels alone.
8. Analysis
8.1. Data integrity
Preserve:
- protocol and study version;
- raw observations or authorised recordings;
- event definitions and query/code versions;
- transformation and exclusion log;
- analysis decisions;
- anonymised evidence excerpts linked to findings.
8.2. Qualitative analysis
State the analytic approach and the researcher’s role. Keep a chain from raw material to code, theme, claim and decision. Search for negative cases and alternative explanations. More quotations do not make a theme prevalent.
Double coding MAY reveal ambiguity, but agreement is not a universal validity test. Reflexivity, context and transparent interpretation remain necessary.
8.3. Quantitative analysis
Report denominators, missing data, distributions, effect estimates and uncertainty. Avoid averaging away meaningful task or group differences. A metric definition MUST remain stable across compared versions, or the break MUST be declared.
Multiplicity, repeated looks, optional stopping and post-hoc subgroup searches increase false discovery. Pre-specify confirmatory analyses where the decision requires them.
8.4. Severity
Severity is a judgement about consequence, reach, frequency, recoverability and evidence confidence. It is not visual prominence.
severity = consequence × exposure × likelihood × recovery difficulty
confidence = evidence quality and coverage
Keep severity and confidence separate. A high-severity, low-confidence hazard requires investigation or precaution, not silent downgrading.
9. Research ethics and governance
Research with people MUST respect persons, minimise harm, distribute burdens and benefits fairly, and comply with applicable law and ethics review.
Before collection:
- determine whether independent ethics review is required;
- obtain informed, voluntary and accessible consent;
- explain purpose, procedure, recording, data use, retention and withdrawal;
- collect only data necessary for the question;
- separate research consent from product terms and employment authority;
- protect confidentiality and plan disclosure control;
- provide safe withdrawal without loss of entitled compensation;
- plan for distress, sensitive disclosure and safeguarding;
- disclose sponsor, incentives and material conflicts.
Consent is a process, not a checkbox. Deception requires exceptional justification, prior ethical approval where applicable, risk control and debriefing.
Publicly observable data is not automatically ethically unrestricted. Researchers MUST consider reasonable expectations, identifiability, group harm and platform context.
10. Independent client audit
An audit sold together with remediation creates a conflict of interest: the evaluator may benefit from finding more or larger defects. HIF requires:
- scope, criteria, evidence threshold and pricing basis agreed before findings;
- reproducible findings linked to standards or observed task harm;
- severity independent of remediation revenue;
- explicit uncertainty and excluded scope;
- client access to evidence and the right to seek independent review;
- no conformance claim based on sampled evidence beyond its justified scope;
- re-test criteria defined before remediation;
- separate reporting of defects, opportunities and untested hypotheses.
The audit report MUST make the current product understandable without forcing the client to buy implementation. Remediation proposals MUST state which finding, HIF requirement and expected outcome each change addresses.
11. Synthesis and decision
A finding is not a design instruction. Synthesis connects evidence to action:
observation
→ interpretation
→ bounded claim
→ affected task and people
→ mechanism
→ HIF requirement
→ severity and confidence
→ alternatives
→ decision
→ verification
Conflicting evidence MUST remain visible. The decision record states why one interpretation or trade-off was selected and what evidence would change it.
12. Required outputs and gates
Discovery gate
- decision and scope defined;
- context and important variation described;
- evidence gaps and ethics reviewed;
- initial needs and risks traceable.
Design gate
- alternatives considered;
- object, command, state and authority models coherent;
- critical paths and failures prototyped;
- applicable HIF requirements mapped.
Validation gate
- claims matched to suitable methods;
- accessibility layers completed;
- evidence, limitations and residual risk recorded;
- high-risk findings resolved or formally excepted.
Release gate
- implementation preserves validated semantics;
- automated and manual verification passes;
- monitoring, rollback and support are ready;
- conformance statement is no broader than the evidence.
Learning gate
- operational outcomes and incidents reviewed;
- distributional and accessibility effects checked;
- evidence records updated;
- new work, exceptions or retirement decisions owned.
13. Reporting
A research or audit report MUST include:
- executive decision summary;
- scope, exclusions and dates;
- product/version and environment;
- people, tasks and sampling;
- method and procedure;
- findings with evidence, severity and confidence;
- measures with denominators and uncertainty;
- limitations and alternative explanations;
- HIF/standard traceability;
- recommendations and verification criteria;
- ethics, privacy and conflicts;
- artefact locations and retention.
Use the Common Industry Format principles so another qualified reader can judge validity and, where appropriate, reproduce the evaluation.
Sources
- ISO 9241-210:2019 — Human-centred design for interactive systems
- ISO 9241-11:2018 — Usability: Definitions and concepts
- NIST — Industry Usability Reporting and Common Industry Formats
- NIST — Common Industry Format for Usability Test Reports
- Holtzblatt and Beyer, Contextual Design: Design for Life
- Nielsen and Landauer (1993), A mathematical model of the finding of usability problems
- Malterud, Siersma and Guassora (2016), Sample size in qualitative interview studies: Guided by information power
- NASA Task Load Index
- American Statistical Association statement on p-values
- APA Journal Article Reporting Standards
- The Belmont Report — US Office for Human Research Protections
- ACM Publications Policy on Research Involving Human Participants and Subjects
- ACM Artifact Review and Badging
- W3C — Web Content Accessibility Guidelines 2.2
- W3C — Website Accessibility Conformance Evaluation Methodology 1.0
- W3C — Involving Users in Evaluating Web Accessibility