Design Review and Critique
Status: normative operational method
Version: 0.1
This document defines how a team examines design work, distinguishes evidence from preference, makes accountable decisions and converts findings into verifiable action.
1. Purpose
A design review is not a taste tribunal, presentation ritual or stakeholder vote. Its purposes are to:
- expose assumptions and failure modes early;
- test a design against user needs, system semantics and constraints;
- compare alternatives and their consequences;
- determine what evidence is sufficient for the next commitment;
- record decisions, dissent, exceptions and residual risk;
- improve the design and the system that produced it.
Critique evaluates the work, not the worth or talent of its author.
2. Four statement classes
Every review comment SHOULD identify its class:
- Observation — what is directly present or what occurred.
- Inference — a proposed interpretation or mechanism.
- Judgement — comparison with an explicit objective, requirement or criterion.
- Recommendation — a proposed change or next test.
Example:
Observation:
The confirmation control and cancellation link use the same weight and colour.
Inference:
People may not distinguish commitment from safe exit under time pressure.
Judgement:
This conflicts with the requirement to represent consequence before commitment.
Recommendation:
Prototype two hierarchy treatments and test action identification and error rate
with the defined critical-task sample.
“It feels wrong” may start an inquiry, but it is not sufficient as a release finding.
3. Inputs
The review owner provides an input packet appropriate to the decision:
- problem statement, people, tasks and context;
- intended outcomes, guardrails and unacceptable harms;
- applicable HIF requirements and Product Profile;
- content model and realistic content;
- object, command, state and permission models;
- journey, states and edge cases;
- alternatives considered and decisions already fixed;
- prototype and its fidelity limits;
- relevant research, analytics, incidents and previous findings;
- accessibility, localisation, platform, performance, security and legal constraints;
- specific questions the review must answer.
Missing input is itself a finding when it prevents a defensible judgement. The panel MUST NOT invent user needs or policy to fill the gap.
4. Roles and independence
Typical roles:
- owner: defines the decision and acts on the outcome;
- facilitator: protects scope, evidence discipline and participation;
- presenter: explains intent, constraints and known uncertainty;
- reviewers: examine the work through relevant disciplines;
- recorder: captures claims, decisions and actions;
- decision authority: accepts, rejects or escalates the result.
The presenter SHOULD NOT have to facilitate and record simultaneously. Consequential reviews SHOULD include reviewers not responsible for the design. Accessibility, content, research and engineering are not optional perspectives to be represented by a visual designer’s guess.
Conflicts of interest, including an auditor who may sell the remediation, MUST be disclosed.
5. Review levels
| Level | Purpose | Suitable artefact | Output |
|---|---|---|---|
| studio critique | explore direction and improve craft | sketches or alternatives | hypotheses and next iteration |
| design review | test coherence and requirements | realistic flow or prototype | findings and decisions |
| specialist review | inspect a defined discipline or hazard | accessible implementation or specification | specialist evidence |
| implementation review | compare built product with contract | integrated build | defects and deviations |
| release review | decide readiness and residual risk | release candidate plus evidence | gate decision |
| operational review | learn from actual use | telemetry, incidents, research and support | corrective changes |
A low-fidelity artefact MUST NOT be used to close questions that require rendering, interaction, assistive technology, performance or field use.
6. Multi-pass protocol
Review in passes so that visual polish does not conceal structural defects.
Pass 1: intent and context
- Who is acting, for what goal and under what conditions?
- What is the decision being reviewed?
- Which constraints are fixed and which remain negotiable?
- What would success, failure and harm look like?
Pass 2: meaning and authority
- Are objects, actions, states and ownership represented truthfully?
- Are scope, permissions, provenance and consequences inspectable?
- Does the interface promise an operation the system cannot guarantee?
- Are automated or AI-mediated actions distinguishable and controllable?
Pass 3: content and information architecture
- Does terminology match the domain and audience?
- Can people locate, compare and understand the information needed?
- Are labels, instructions, errors and recovery specific?
- Are search, navigation and history coherent across the journey?
Pass 4: interaction and state
- Are commands discoverable, predictable and available at the right point?
- Are loading, empty, partial, offline, error, success and stale states covered?
- Are interruption, back, undo, retry and safe exit defined?
- Do keyboard, pointer, touch and applicable alternative inputs preserve function?
Pass 5: perception and expression
- Does hierarchy correspond to importance and current task?
- Do grouping, alignment, spacing, typography, colour and imagery communicate the intended relations?
- Does expression support the desired character without false urgency, authority or affordance?
- Are aesthetic comments tied to an objective or explicitly labelled as exploratory preference?
Pass 6: accessibility and adaptation
- Do semantics and reading order agree with visual structure?
- Does the work survive zoom, reflow, text spacing, forced colours, dark and high-contrast modes, reduced motion and user overrides?
- Are all supported languages, writing directions and content extremes covered?
- Has relevant assistive technology and disabled-user evidence been planned or obtained?
Pass 7: implementation and operations
- Are components, tokens, states and responsive rules specified?
- Are assets licensed, attributable, performant and traceable?
- Are security, privacy, failure, observability and support states designed?
- Can the result be tested, maintained, migrated and rolled back?
Pass 8: evidence and decision
- Which observations are facts and which are hypotheses?
- Which thresholds have been met?
- What remains unknown, and how sensitive is the decision to it?
- What action, owner, date and verification method follow?
7. Finding anatomy
Every actionable finding has:
FINDING-ID
scope and affected state
observation and reproduction
expected criterion
user/system consequence
evidence and confidence
severity and rationale
recommended outcome, not compulsory cosmetic solution
owner and due condition
verification method
related requirement, component or token
Screenshots MAY support a finding but MUST NOT replace reproduction, interaction state or accessible evidence.
8. Severity is not preference
Severity combines:
- consequence to people or system;
- probability or exposure;
- detectability before harm;
- recoverability;
- breadth of affected people, tasks and states;
- legal, safety, security or trust obligation.
Visual novelty, stakeholder seniority, implementation effort and expected remediation revenue MUST NOT increase severity. Effort and commercial value are separate prioritisation dimensions.
Suggested outcome classes:
- blocker — unacceptable safety, rights, security, accessibility or core task failure;
- major — material failure with substantial consequence or no reasonable recovery;
- moderate — repeated friction, error or comprehension loss with recovery;
- minor — limited inconsistency or craft defect with low user consequence;
- hypothesis — plausible issue needing discriminating evidence;
- preference — non-binding aesthetic choice within the permitted design space.
9. Critique language
Useful form:
For [person/task/context],
the current [observable property]
may cause [consequence]
because [mechanism/evidence].
This conflicts with [criterion].
We should [desired outcome or next test],
verified by [method and threshold].
Reviewers SHOULD:
- ask before assuming intent or constraint;
- refer to artefacts and outcomes, not personal ability;
- explain the boundary of their expertise;
- distinguish a convention from a requirement;
- make uncertainty visible;
- offer alternatives proportional to the finding;
- preserve a recorded minority view when material disagreement remains.
Reviewers MUST NOT:
- invoke unnamed “users”, “research” or “best practice”;
- use personal preference as an accessibility claim;
- prescribe a component before agreeing on the problem;
- perform design work by committee through disconnected comments;
- reopen a fixed decision without new evidence or changed constraints;
- equate silence, hierarchy or consensus with validity.
10. Comparative review
A material direction SHOULD be compared with at least one credible alternative. Use the same:
- realistic content and states;
- task and participant criteria;
- device and environment;
- measures and thresholds;
- implementation assumptions.
A comparison table MAY record:
| Criterion | Weight or priority | Option A | Option B | Evidence | Uncertainty |
|---|
Weights MUST NOT be manipulated after results are known. A score helps expose a trade-off; it does not automate the decision or erase a hard constraint.
11. Red-team questions
Before approval, ask:
- How could this representation be misunderstood?
- What happens with the longest, shortest, missing, stale or hostile content?
- What happens with slow, partial, offline or contradictory system state?
- Who is excluded by the assumed device, ability, language or expertise?
- Can a persuasive hierarchy pressure a person into an unintended choice?
- Can colour, imagery, motion or sound create false status or urgency?
- What does an expert lose? What does a first-time user have to remember?
- What happens after interruption, undo, expiry or transfer to another person?
- Which claim would change the decision if it were false?
- How will production reveal that the decision is failing?
12. Decision outcomes
The decision authority records one outcome:
- approve — thresholds met within stated scope;
- approve with actions — non-blocking actions have owners and verification;
- iterate — a design change is required before another review;
- test — uncertainty is too consequential; obtain specified evidence;
- escalate — authority, policy or risk ownership lies elsewhere;
- reject — the direction cannot meet a constraint or outcome;
- exception — a named requirement is not met, with owner, evidence, mitigation, expiry and review date.
“Looks good” is not an approval record.
13. Review record
REVIEW-ID
date, scope and decision question
owner, facilitator, reviewers and authority
declared conflicts
inputs and prototype fidelity
applicable requirements and thresholds
observations
inferences
judgements
findings and severities
alternatives and trade-offs
dissent and unresolved uncertainty
decision and rationale
actions, owners and due conditions
verification and review trigger
linked evidence and artefact versions
The record SHOULD be stored with the product decision history and linked to requirements, implementation and tests.
14. Design QA hand-off
Before implementation:
- component and token contracts are named;
- supported states, inputs, breakpoints and preferences are specified;
- content rules and localisation cases are available;
- accessibility semantics and expected keyboard behaviour are explicit;
- asset sources, licences, transformations and fallbacks are recorded;
- measurable acceptance criteria are attached to work items.
Before release:
- implementation is compared with the approved artefact;
- deviations are accepted or corrected explicitly;
- real content and state combinations are exercised;
- visual regression is supplemented by semantic and interactive tests;
- required manual, assistive-technology and participant evidence exists;
- known residual risk has an owner and review condition.
15. Review quality indicators
Monitor the review system itself:
- proportion of findings linked to an explicit criterion;
- time from finding to verified resolution;
- recurrence of the same failure across products;
- findings discovered only after release;
- accessibility and localisation defects by lifecycle stage;
- rate and age of exceptions;
- participation and unresolved dissent;
- false-positive and duplicate findings;
- decisions later reversed because assumptions were not recorded.
Do not reward reviewers for comment volume. The desired outcome is better, safer decisions with less repeated waste.
Sources
- ISO 9241-210:2019 — Human-centred design for interactive systems
- ISO 9241-11:2018 — Usability: Definitions and concepts
- Nielsen and Molich (1990), Heuristic evaluation of user interfaces
- Lewis et al. (1990), Testing a walkthrough methodology for theory-based design of walk-up-and-use interfaces
- Rittel and Webber (1973), Dilemmas in a General Theory of Planning
- NASA Human Systems Integration Handbook
- GOV.UK Service Standard
- GOV.UK — What happens at a service assessment
- W3C WAI — Evaluating Web Accessibility Overview
- W3C WAI — Involving Users in Evaluating Web Accessibility
- W3C Web Content Accessibility Guidelines 2.2
- ACM Code of Ethics and Professional Conduct