Foundations and Evidence
Status: normative method
Version: 0.1
This document defines the intellectual foundations on which HIF requirements and design decisions are to be justified. It is not a catalogue of fashionable “UX laws”. It is a method for moving from knowledge about people and systems to bounded, testable claims about an interface.
1. Epistemic position
HIF treats a design decision as a hypothesis about a human–technology system. The decision is not validated by taste, precedent, a screenshot, a stakeholder vote or the reputation of its author.
A defensible claim identifies:
- the people, activity and context to which it applies;
- the intended outcome and unacceptable harm;
- the mechanism by which the design is expected to affect that outcome;
- the evidence already available;
- the assumptions that make the evidence transferable;
- the method by which the claim can be falsified;
- the remaining uncertainty and the owner of that uncertainty.
Evidence changes confidence; it does not manufacture certainty. A positive finding in one population, task, device or organisational setting MUST NOT be silently generalised to another.
2. The unit of analysis
2.1. The interactive system
The unit of analysis is not a screen and not an isolated “user”. It is a system:
people and roles
↕
goals, tasks and practices
↕
representations, controls and automation
↕
data, rules, infrastructure and other actors
↕
physical, social, organisational and temporal context
ISO 9241-11 treats usability as an outcome of use, relative to specified users, goals and contexts. ISO 9241-210 makes human-centred work a lifecycle activity. ISO 6385 places human, social and technical requirements within one work-system design problem. HIF therefore rejects decontextualised claims such as “this component is usable” or “this pattern is intuitive”.
Every evaluation MUST state its context of use. At minimum:
- relevant user groups and meaningful variation within them;
- goals, tasks and consequences;
- equipment, input methods and assistive technologies;
- physical and social environment;
- knowledge, training and frequency of use;
- time pressure, interruption and fatigue conditions;
- organisational rules, incentives and authority;
- technical dependencies and failure modes.
2.2. Outcome dimensions
Usability is multi-dimensional. HIF distinguishes:
- effectiveness: whether intended goals are achieved accurately and completely;
- efficiency: resources expended in relation to the achieved result;
- satisfaction: the person’s responses arising from use;
- accessibility: whether people with relevant disabilities can perceive, understand, navigate, operate and complete the activity;
- safety and resilience: whether error, failure and recovery remain within acceptable harm;
- agency and trustworthiness: whether choices, authority, provenance and consequences are represented honestly;
- learnability and development: whether competence can be acquired, retained and extended;
- collective performance: whether coordination across people, artefacts and time succeeds.
No single metric is a substitute for this outcome set. In particular, speed, conversion, engagement, preference and task completion each admit harmful or misleading interpretations when considered alone.
3. Human capacities are constraints and resources
HIF does not model people as defective processors. Human capacities are finite, variable, adaptive and supported by cultural and material resources.
3.1. Perception is active and task-dependent
Perception is not a faithful recording of a display. It is selective, goal-dependent and influenced by expectation, context, contrast, motion, crowding and prior learning.
Design implications:
- information needed for a decision SHOULD be perceptually available where the decision is made;
- state differences MUST survive relevant visual, auditory and tactile limitations;
- salience MUST correspond to priority rather than decoration;
- a change that matters MUST remain discoverable after the transient signal has passed;
- tests MUST include realistic density, motion, lighting, zoom, display and distraction conditions.
Change blindness and inattentional blindness demonstrate that visible content can remain unnoticed when attention is occupied elsewhere. They do not justify a universal animation, colour or notification rule. They require the team to test whether the relevant signal is detected in the actual task.
3.2. Attention is allocative, not an engagement resource
Attention is allocated across competing goals and stimuli. Interruptions impose orientation and resumption costs, and urgent-looking signals can displace higher-value work.
HIF therefore treats attention as belonging to the person:
- interruption requires a proportionate, stated reason;
- urgency MUST reflect actual urgency;
- background events MUST be recoverably discoverable;
- a critical signal MUST be distinguishable from routine activity;
- notification volume MUST be evaluated as a system, not notification by notification.
3.3. Working memory is limited and strategy-dependent
The capacity and organisation of working memory vary with material, rehearsal, expertise, grouping and concurrent activity. Miller’s “seven plus or minus two” paper is historically important but MUST NOT be converted into a universal maximum number of menu items, steps or options. Later work likewise does not license a fixed “four-item rule” for interface composition.
Design SHOULD:
- preserve task-relevant state outside memory;
- support recognition and comparison without forcing serial recall;
- keep entered values and partial work visible or retrievable;
- externalise dependencies, scope and consequences;
- avoid requiring a person to remember an arbitrary code, state or instruction across an interruption;
- allow experts to form larger meaningful units rather than flattening all work into novice-sized steps.
Cognitive load theory distinguishes complexity inherent in the material from load introduced by its presentation and from resources devoted to learning. HIF uses that distinction diagnostically, not as a licence to hide necessary complexity.
3.4. Choice time depends on information, not option count alone
Hick and Hyman investigated choice reaction time under controlled conditions. Their results concern the information carried by alternatives and their probabilities, not a universal instruction to minimise every set of choices.
Application requires checking:
- whether choices are familiar or newly learned;
- whether probabilities are stable;
- whether options can be grouped or searched;
- whether comparison quality matters more than immediate response;
- whether removing an option moves work into memory or navigation;
- whether the task is a speeded, single-response choice at all.
Measure the real task. Do not infer task difficulty from the number of visible items.
3.5. Pointing performance has boundary conditions
Fitts’s original experiment modelled rapid, aimed movement under specific conditions. Its family of models can predict aspects of pointing time when distance, target width and movement conditions are defined.
It supports such hypotheses as “a larger effective target may reduce pointing time”. It does not by itself prove:
- that a target is noticed or understood;
- that touch, gaze, switch input and mouse input are equivalent;
- that adjacent destructive targets are safe;
- that a minimum target size is sufficient for every person or context;
- that movement time dominates the complete task.
Target design MUST also consider spacing, occlusion, posture, tremor, mobility, device precision, error cost and non-pointing alternatives.
3.6. Skilled action differs from novice action
Practice changes perception, memory chunks, strategy and movement. Predictive models such as GOMS and the Keystroke-Level Model are useful for stable, well-learned, error-light routine tasks. They are weak evidence for exploration, collaboration, creative work, safety-critical judgement, recovery or unfamiliar systems.
HIF requires separate analysis of:
- first-use comprehension;
- occasional-use recognition;
- expert throughput;
- transfer from related systems;
- recovery after interruption or error;
- relearning after absence.
Optimising the predicted expert path MUST NOT destroy orientation, explanation or recovery for other use conditions.
4. Cognition beyond the individual
4.1. External cognition
People reason with representations: lists, maps, diagrams, histories, labels, spatial arrangements and physical objects. An interface can transform a hard memory or inference problem into a perceptual comparison, but a representation can also conceal assumptions or make important relations difficult to inspect.
Evaluate a representation by the operations it enables:
- locating and identifying;
- comparing and sorting;
- tracing cause, provenance and history;
- coordinating with another person;
- detecting conflict and anomaly;
- projecting possible outcomes;
- preserving state across time.
The information architecture MUST follow these operations rather than a component taxonomy alone.
4.2. Distributed cognition
Distributed cognition treats the cognitive system as extending across people, artefacts, representations and time. Hutchins’s analysis of ship navigation shows that system-level cognitive properties cannot be inferred from one operator in isolation.
For HIF this means:
- hand-offs, shared displays, records and conventions are interface architecture;
- provenance and temporal order can be as important as current state;
- local optimisation can damage team-level performance;
- automation changes the distribution of knowledge and authority;
- evaluation SHOULD follow information transformations across the entire activity, including recovery from a missing or incorrect contribution.
4.3. Situated action
Situated action rejects the idea that a prior plan exhaustively determines conduct. Plans organise and account for action, but people also interpret circumstances and revise action through feedback.
An interface MUST support:
- inspection before commitment;
- incremental, observable action;
- interruption and safe exit;
- correction and change of goal;
- recovery when the actual situation differs from the planned path;
- explanation of system action in the current context.
A “happy-path” process model is not evidence that work will follow that path.
4.4. Activity theory
Activity theory analyses purposeful activity through subjects, objects, mediating tools, community, rules and division of labour. Contradictions among these elements can explain recurring workarounds that a screen-level review misclassifies as user error.
Use this lens when:
- the product changes responsibilities or professional practice;
- several organisations or roles pursue partially conflicting goals;
- formal procedure and actual work diverge;
- learning and adaptation occur over a long period;
- a local interaction problem persists after repeated interface changes.
Do not treat an activity-system diagram as proof. It is an analytic model whose claims require field evidence.
5. Affordance, signification and direct manipulation
Gibson’s ecological account concerns action possibilities in a relation between an organism and an environment. Norman’s design account distinguishes what an artefact affords from the perceptible cues that communicate possible action. HIF uses:
- affordance for a real action possibility;
- signifier for perceivable information about where or how to act;
- constraint for a relation that limits possible action;
- mapping for the relation between control and effect;
- feedback for perceivable information about action acceptance, progress and outcome.
Visual resemblance alone does not create an affordance, and an affordance can exist without being discoverable.
Shneiderman’s direct manipulation account joins continuous representation, physical or labelled action, and rapid, incremental, reversible, visible effects. These properties are valuable when the object and action can be represented faithfully. Indirect commands are preferable when they improve precision, repeatability, accessibility, scale, auditability or automation.
6. Models, regularities and heuristics
6.1. Classification
Every invoked design rule MUST be classified:
- standard requirement: a requirement from an identified standard and version;
- empirical result: an observation from a stated study;
- predictive model: a formal relation with parameters and fit;
- explanatory theory: an account of mechanisms and relations;
- heuristic: an expert prompt for finding possible problems;
- convention: a learned practice in a community or platform;
- product policy: a deliberate local rule;
- hypothesis: an untested expectation.
Calling all seven categories “laws” obscures their authority and boundary conditions.
6.2. Required model record
Before a model or heuristic can justify a consequential decision, record:
claim
classification
original source
population and sample
task and apparatus
independent and dependent variables
model form or reasoning mechanism
boundary conditions
replication or contrary evidence
similarity to the product context
decision sensitivity
planned verification
6.3. Prohibited shortcuts
HIF rejects the following arguments:
- “research says” without an identifiable claim and source;
- “users prefer” without sampling and procedure;
- “best practice” without context and failure conditions;
- “the average user” when meaningful variation is unexamined;
- a fixed cognitive capacity translated directly into a component limit;
- a correlation interpreted as a causal mechanism;
- statistical significance interpreted as practical importance;
- absence of observed problems interpreted as proof of absence;
- a heuristic inspection presented as user validation;
- one participant with a disability treated as representative of a disability class;
- a laboratory result treated as field performance without a transfer argument.
7. Evidence quality
7.1. Fit before rank
The strongest method is the method that can answer the decision question. Evidence quality has at least six independent dimensions:
- relevance: similarity of people, task, system and context;
- validity: whether the method supports the claimed inference;
- precision: uncertainty around the estimate or interpretation;
- credibility: protection against bias, confounding and selective reporting;
- coverage: inclusion of important populations, states and failure modes;
- independence: whether evidence comes from genuinely distinct data, analysts, methods or operational periods.
A large irrelevant experiment can be less useful than a small, well-targeted field study. A vivid quotation can reveal a mechanism without estimating its frequency.
7.2. HIF evidence ladder
The levels in EVALUATION.md express maturity, not automatic truth:
E0 Declarationrecords intent only;E1 Design reviewtests internal coherence and known criteria;E2 Static/automated verificationprovides repeatable evidence for properties a machine can observe;E3 Interactive expert reviewexercises complete behaviour and specialist workflows;E4 Representative user evaluationobserves intended people performing meaningful tasks;E5 Operational evidenceobserves the system over time in use.
Higher levels do not subsume lower ones. Production telemetry cannot determine semantic accessibility; a user study cannot exhaustively prove a code invariant; an automated checker cannot determine whether a workflow supports a real goal.
7.3. Claim–method matching
| Claim | Suitable evidence | Common overclaim |
|---|---|---|
| A semantic invariant always holds | type, contract, state-machine and property tests | a few successful user sessions |
| People notice and understand a state | task-based observation plus comprehension probe | pixel presence or stakeholder review |
| One design is faster for a defined task | controlled comparison with uncertainty and effect size | preference vote |
| A workflow fits real practice | contextual inquiry, field observation and operational follow-up | scripted laboratory task alone |
| Accessibility requirement is met | standards conformance, manual and assistive-technology testing, disabled-user evaluation | automated scan alone |
| A change causes an operational outcome | experiment or credible quasi-experimental design | before/after correlation alone |
| A rare critical failure is acceptably controlled | hazard analysis, simulation, incident evidence and defence testing | absence in a small usability study |
7.4. Negative and null evidence
“No problem observed” is meaningful only with detection power:
- what opportunities existed for the problem to occur;
- which people, tasks and states were covered;
- how the problem would have been detected;
- what sample or operational exposure was observed;
- what magnitude remains compatible with the data.
A null result MUST NOT be rewritten as proof of equivalence. An equivalence claim requires an equivalence margin and a design capable of testing it.
8. Triangulation
Triangulation is the deliberate comparison of evidence with different failure modes. It is not the accumulation of many similar opinions.
HIF recognises:
- method triangulation: for example observation, interview, task metrics and logs;
- data triangulation: different roles, contexts, devices and time periods;
- investigator triangulation: more than one analyst with recorded reconciliation;
- theory triangulation: alternative explanations tested against the same evidence;
- source triangulation: research, standards, incidents, support and product data.
A triangulation record states:
claim
evidence streams
independence of streams
agreement
conflict
plausible explanations
decision and confidence
next discriminating test
Disagreement is a result. It SHOULD trigger inspection of sampling, context, measurement and competing mechanisms rather than averaging the evidence away.
9. From evidence to HIF
9.1. Traceability chain
source or observed need
→ bounded claim
→ mechanism and context
→ risk or desired outcome
→ HIF principle/requirement
→ Product Profile interpretation
→ product requirement
→ design and implementation
→ verification method
→ evidence record
→ residual uncertainty or exception
Every MUST in a Product Profile MUST have a verification method. Every
high-risk design claim MUST identify the evidence and assumptions on which it
depends.
9.2. Foundation-to-requirement map
| Foundation | Principal HIF consequences |
|---|---|
| Contextual usability and work-system ergonomics | HIF-CTX-001, HIF-CTX-003, profile context and task corpus |
| Selective attention and change detection | HIF-STA-001, HIF-FBK-004, HIF-VIS-001 |
| Working memory and external cognition | HIF-STA-001, HIF-CNT-002, HIF-ERR-004 |
| Motor variability and speed–accuracy trade-off | HIF-INP-001, HIF-INP-002, HIF-TRU-004 |
| Distributed cognition | HIF-OBJ-004, HIF-STA-003, HIF-STA-005, HIF-AUT-004 |
| Situated action | HIF-FBK-001–005, HIF-ERR-003, HIF-NAV-005 |
| Activity theory and sociotechnical analysis | HIF-CTX-001, HIF-AUT-002, HIF-TRU-001–005 |
| Affordance, signification and direct manipulation | HIF-CMD-002–004, HIF-FBK-001–003, HIF-VIS-003 |
| Human variability | HIF-A11Y-001–004, HIF-CNT-003 |
| Evidence discipline | HIF conformance, traceability, release gates and exceptions |
This map is explanatory, not exclusive. A requirement can have several independent foundations.
10. Evidence record
Each material decision SHOULD have an evidence record:
EVIDENCE-ID
decision and owner
date and review date
applicable HIF requirements
people/tasks/context
claim and decision threshold
known mechanism
sources and prior evidence
method and protocol
sample/recruitment or operational exposure
measures and analysis
result with uncertainty
limitations and excluded populations
conflicts with other evidence
decision taken
residual risk
data/artifact location and retention
The record MUST distinguish observations, participant statements, analyst interpretations and product decisions.
11. Review rules
An evidence claim MUST be reviewed when:
- the target population, task or context materially changes;
- a new input method, device class or assistive technology becomes relevant;
- automation changes who knows, decides or acts;
- operational evidence contradicts the design model;
- a cited standard is revised or withdrawn;
- the cost of error changes;
- the evidence has exceeded its declared review date.
HIF requirements remain stable only while their rationale remains credible. Traceability exists so that the framework can be corrected rather than ritualised.
Sources
- ISO 9241-11:2018 — Usability: Definitions and concepts
- ISO 9241-210:2019 — Human-centred design for interactive systems
- ISO 6385:2016 — Ergonomics principles in the design of work systems
- ISO 10075-2:2024 — Ergonomic principles related to mental workload: Design principles
- Fitts (1954), The information capacity of the human motor system in controlling the amplitude of movement
- Hick (1952), On the rate of gain of information
- Hyman (1953), Stimulus information as a determinant of reaction time
- Miller (1956), The magical number seven, plus or minus two
- Cowan (2001), The magical number 4 in short-term memory
- Sweller (1988), Cognitive load during problem solving
- Treisman and Gelade (1980), A feature-integration theory of attention
- Rensink, O’Regan and Clark (1997), To see or not to see: The need for attention to perceive changes in scenes
- Simons and Chabris (1999), Gorillas in our midst: Sustained inattentional blindness for dynamic events
- Card, Moran and Newell (1980), The Keystroke-Level Model for user performance time with interactive systems
- Shneiderman (1983), Direct manipulation: A step beyond programming languages
- Gibson, The Ecological Approach to Visual Perception
- Norman, The Design of Everyday Things
- Hutchins, Cognition in the Wild
- Suchman, Human–Machine Reconfigurations: Plans and Situated Actions
- Engeström, Learning by Expanding
- Scaife and Rogers (1996), External cognition: How do graphical representations work?
- W3C WAI — Involving Users in Evaluating Web Accessibility
- ACM Code of Ethics and Professional Conduct