HIF / Documentation / Client practice

HIF Web Audit Playbook

Status: operational method, version 0.1

Purpose: independent, evidence-led evaluation of client websites and web applications

This playbook turns the HIF Constitution and HIF Evaluation into a repeatable commercial service. It is designed for discovery audits, pre-release reviews, redesign baselines and post-remediation verification.

An audit finds and explains risk. It does not manufacture certainty, certify matters outside its scope or guarantee that no defect remains.

1. Governing principles

The audit MUST:

  1. evaluate meaningful tasks, not merely inspect screenshots;
  2. define scope, sample, environment and limitations before making claims;
  3. distinguish observation, interpretation, hypothesis and verified defect;
  4. preserve a reproducible chain from requirement to evidence to finding;
  5. combine methods because no single tool or reviewer covers the whole system;
  6. minimise collection of personal, confidential and security-sensitive data;
  7. report exclusion and loss of agency as product risk, not cosmetic polish;
  8. leave the client free to implement recommendations with any supplier;
  9. state uncertainty and conflicts of interest;
  10. make no legal, security or accessibility guarantee.

HIF requirements remain the product-level authority. External standards and platform guidance supply specialised criteria. Where criteria conflict, record the conflict and apply the hierarchy in HIF.

2. Engagement and authorisation

2.1. Required engagement record

Work MUST NOT begin until the following are recorded:

  • commissioning organisation and authorised contact;
  • audit purpose and decisions the report will support;
  • included domains, subdomains, applications, environments and journeys;
  • excluded systems, third parties and data;
  • permitted accounts, roles and test data;
  • permitted methods and prohibited actions;
  • dates, test window, rate limits and incident contact;
  • confidentiality, evidence retention and deletion arrangements;
  • requested standards, HIF profile and target accessibility level;
  • deliverables, review rounds, retest terms and acceptance route;
  • known releases or content changes during the audit;
  • explicit authority for any action that could alter data or service state.

The default posture is non-destructive observation using synthetic data and dedicated test accounts. Load generation, vulnerability exploitation, social engineering, production data modification and attempts to bypass controls are outside this playbook unless separately scoped, authorised and performed by appropriately qualified specialists.

2.2. Scope change

A newly discovered host or workflow is not automatically in scope. Record it as an inventory observation, assess whether it changes risk and obtain written approval before active testing.

2.3. Stop conditions

Pause testing and contact the nominated person if:

  • testing may harm people, data or service availability;
  • unexpected personal or confidential data is exposed;
  • an apparent security weakness could be worsened by continued interaction;
  • the production environment differs materially from the authorised target;
  • the site changes enough to invalidate collected evidence;
  • access, consent or authority becomes uncertain.

3. Audit questions

Every engagement translates its purpose into answerable questions:

  • Can each priority audience complete each priority task?
  • Is the current object, state, consequence and next action understandable?
  • Can people prevent, recognise and recover from errors?
  • Are primary tasks available across required input, viewport and browser conditions?
  • Does the experience preserve accessibility, privacy, security, trust and agency?
  • Is performance acceptable in real use and under the agreed test conditions?
  • Does content help people find, understand and act on the right information?
  • Which changes reduce the greatest validated risk?

Questions determine methods and evidence. A method is not performed merely because a tool supports it.

3.1. HIF traceability baseline

At minimum, the coverage matrix maps:

Audit concernConstitution baseline
purpose, boundaries and riskHIF-CTX-001..003
objects, identity and provenanceHIF-OBJ-001..004
command meaning, authority and repetitionHIF-CMD-001..005
state, persistence and concurrencyHIF-STA-001..005
acknowledgement, progress and completionHIF-FBK-001..005
prevention, error and recoveryHIF-ERR-001..005
orientation, focus and escapeHIF-NAV-001..005
input independence and equivalenceHIF-INP-001..004
accessibility and adaptationHIF-A11Y-001..004
content and localisationHIF-CNT-001..004
privacy, security, choice and trustHIF-TRU-001..005
automation and AIHIF-AUT-001..006
performance and resilienceHIF-PERF-001..004
visual hierarchy and semanticsHIF-VIS-001..005

The range notation is an index, not evidence that every requirement applies. Applicability and exceptions are decided and recorded individually.

4. Inventory and task model

4.1. Inventory

Create a dated inventory of:

  • hosts, applications, locales and authentication boundaries;
  • templates and page types;
  • priority journeys and service transactions;
  • roles, permissions and account states;
  • forms, search, navigation and transactional components;
  • documents, media and embedded third-party content;
  • responsive layouts, themes and personalisation;
  • error, empty, loading, offline, timeout and completion states;
  • supported browsers, devices and assistive technologies;
  • analytics or support signals supplied by the client;
  • releases, experiments and feature flags affecting the sample.

Do not treat distinct URLs as the only unit. A client-rendered state, modal, validation branch or role-specific view may be a separate audit unit.

4.2. Task corpus

For each priority task record:

TASK-ID
audience and context
goal
starting state
role and authority
preconditions and test data
success state
critical errors
important branches and interruptions
applicable HIF requirements
required evidence level

Include orientation, find, compare, create, change, submit, pay, preserve, delete, recover, refuse, leave, authenticate and obtain help where applicable.

5. Sampling

5.1. When to sample

Test the complete scope when it is small enough. Otherwise use a documented sample. Following the structure of WCAG-EM, the audit SHOULD combine:

  • a structured sample covering key functionality, templates, content types, technologies, states and user groups;
  • a random sample capable of exposing unanticipated repetition or variation.

The sample MUST include:

  • the highest-consequence and highest-volume tasks;
  • entry, completion, error and recovery states;
  • authentication and permission boundaries;
  • each materially different template and interactive pattern;
  • each supported locale and responsive regime;
  • content or components supplied by third parties where they affect the task;
  • known complaints, incidents and high-change areas.

5.2. Sample register

For every unit record inclusion reason, discovered variants, tested state, result and evidence reference. A finding observed in a sample MUST NOT be silently generalised to the whole site. State whether its reach is measured, inferred or unknown.

5.3. Discovery expansion

Expand the sample when a failure suggests a systemic pattern, when a template contains untested variants or when new technology changes the applicable checks. Record the reason so the final scope remains auditable.

6. Method stack

An audit plan selects the smallest combination that can answer the questions.

6.1. Document and model review

Review the product proposition, audience evidence, information architecture, content model, design system, supported environments, analytics definitions and known constraints. Map objects, commands, states, permissions and persistence against HIF.

6.2. Functional test design

Derive cases using the techniques in HIF Test Design:

  • equivalence partitioning;
  • boundary value analysis;
  • decision tables;
  • state transition testing;
  • combinatorial or pairwise coverage;
  • scenario and acceptance-criteria coverage;
  • error guessing;
  • exploratory testing with time-boxed charters.

Positive paths alone are insufficient. Include invalid, absent, repeated, interrupted, stale, unauthorised and recovered states.

6.3. Heuristic evaluation

Use at least the HIF Constitution, the applicable Product Profile and the ten usability heuristics published by Jakob Nielsen as review lenses. Heuristics identify plausible problems; they do not prove task failure, prevalence or user impact.

For consequential work, use more than one evaluator where feasible, reconcile duplicates and preserve disagreements. Separate:

  • criterion breach supported by direct evidence;
  • usability concern supported by expert judgement;
  • hypothesis requiring user research or operational data.

6.4. Cognitive walkthrough

For first-use and learnability questions, examine each action:

  1. Will the intended person try to achieve the right sub-goal?
  2. Will they notice the available action?
  3. Will they connect the action with the desired outcome?
  4. After acting, will feedback show useful progress?

Record assumed knowledge. A walkthrough predicts difficulty; it is not a substitute for observation of representative users.

6.5. Usability evaluation with participants

Define research questions, participant characteristics, recruitment, accessibility arrangements, tasks, consent, moderation, recording, analysis and stopping rules before sessions.

Tasks SHOULD be believable, goal-led and neutral: they must not reveal the control or wording under test. Record task success, critical errors, recovery, assistance, time where meaningful and qualitative evidence. Do not turn a small formative study into a population statistic or a competitive ranking.

6.6. Accessibility

Use WCAG 2.2 as the normative web-content criterion requested by the engagement and WCAG-EM for evaluation structure. Combine:

  • automated checks;
  • semantic and code inspection;
  • keyboard-only operation;
  • focus order, visibility and restoration;
  • zoom, text spacing, reflow and orientation;
  • colour, non-colour cues and forced-colour modes;
  • reduced motion and animation controls;
  • names, roles, values, status and error announcements;
  • screen-reader task completion in the agreed combinations;
  • pointer, touch and alternative-input checks;
  • representative disabled-user research where the claim and risk require it.

Native HTML is preferred where it provides the required semantics and behaviour. WAI-ARIA requirements apply when ARIA is used. The ARIA Authoring Practices Guide is useful implementation guidance, but its examples are not a normative standard or a production-ready design system.

Automated tools cannot establish WCAG conformance. A sample-based HIF audit MUST NOT claim that an entire site conforms unless the complete conformance scope and methodology support that statement.

Complete the Accessibility Assurance and Conformance Record when the engagement includes accessibility conformance, a public statement, procurement evidence or a legal-risk hand-off.

6.6.1. UK public-sector site protocol

For a UK public-sector site, the audit MUST additionally:

  1. identify the responsible public body and the exact service owner;
  2. capture the current accessibility statement, its scope, date, known limitations, contact route and enforcement route;
  3. record whether the audited material is a website, app, document, intranet, third-party service or potentially excluded content;
  4. use the current technical baseline stated by official UK guidance;
  5. distinguish a sampled technical finding from whole-site conformance;
  6. distinguish WCAG failure from a legal conclusion under the Public Sector Bodies Accessibility Regulations, Equality Act or other law;
  7. record any claimed exemption or disproportionate-burden assessment without accepting or rejecting its legal validity;
  8. report the barrier through an accessible, evidence-led route and preserve the response and retest history.

Official UK guidance currently identifies WCAG 2.2 Level AA and an accessibility statement as the baseline, while also describing scope, exceptions, monitoring and enforcement. A technical auditor MAY report observed facts and standards results. The auditor MUST NOT represent that an automated result proves a statutory breach, a regulator's decision, negligence, damages or entitlement to payment.

Commercial remediation MAY be offered separately and transparently. Threats, inflated liability claims, manufactured urgency or payment demands based on an unadjudicated allegation are prohibited.

6.7. Performance and resilience

Measure both lived performance and controlled diagnostics:

  • field data where valid and sufficiently representative;
  • laboratory tests with recorded hardware, network, cache and run conditions;
  • repeated runs and distributions, not a preferred single run;
  • critical-task responsiveness, not landing-page load alone;
  • behaviour under slow, interrupted and restored connections;
  • stability of focus, layout, input and acknowledged results.

Current Core Web Vitals are LCP, INP and CLS. Google classifies “good” at the 75th percentile as LCP no more than 2.5 seconds, INP no more than 200 milliseconds and CLS no more than 0.1. Record the data source, date, page group, device class and sample coverage. Lab measurements are diagnostic estimates; they are not field Core Web Vitals.

Budgets SHOULD also cover transferred bytes, request count, main-thread work, third parties and failure behaviour when these affect the task.

6.8. Responsive and cross-browser behaviour

Derive the matrix from actual audiences, required platforms, browser support policy and risk. Cover:

  • narrow through wide viewports, including intermediate breakpoints;
  • portrait and landscape where relevant;
  • zoom and text enlargement independently of device presets;
  • mouse, keyboard, touch and required alternative input;
  • current supported rendering engines and documented older versions;
  • high and low pixel density where imagery or targets are affected;
  • no-hover, coarse-pointer, reduced-motion, contrast and colour preferences;
  • virtual keyboard, safe areas, browser chrome and dynamic viewport changes;
  • printing or installed-app mode when part of the task.

The objective is equivalent access and consequence, not pixel identity. Record unsupported combinations explicitly.

6.9. Content and information architecture

Evaluate:

  • whether navigation and labels match audience language and task intent;
  • hierarchy, headings, landmarks, reading order and link purpose;
  • scent from entry point to goal and recovery from a wrong route;
  • search terms, zero results, filters and result relevance;
  • ownership, accuracy, freshness and duplication;
  • prerequisites, cost, time, eligibility and consequence before commitment;
  • plain language without loss of necessary precision;
  • consistent names for the same object, state and command;
  • localisation of meaning, layout and interaction rather than strings alone;
  • metadata, titles and sharing text where they affect recognition.

Card sorting, tree testing, search-log analysis and content testing MAY be commissioned when expert inspection cannot answer findability questions.

6.10. Forms and transactions

For every requested datum, establish why it is required, who needs it, how it is validated, retained, protected and kept current. Then test:

  • eligibility and prerequisites before avoidable effort;
  • logical grouping and question order;
  • persistent labels, instructions and examples;
  • required/optional status and format constraints;
  • autofill, paste, password managers and international input;
  • client- and server-side validation;
  • errors connected to fields and a useful summary;
  • preservation after error, back navigation, timeout and interruption;
  • review before consequential submission;
  • duplicate submission and idempotence;
  • truthful progress, receipt, reference and next step;
  • correction, cancellation and recovery routes.

6.11. Privacy, security and trust at the UI boundary

This playbook reviews user-visible controls and browser-observable behaviour, not the security of the whole system. Examine:

  • whether data use and disclosure are explained before collection or transfer;
  • whether non-essential acceptance and refusal are comparably available;
  • whether permission requests are contextual and least-privilege;
  • whether authentication, recovery and session expiry preserve agency;
  • whether password managers, paste and secure autofill are unnecessarily obstructed;
  • whether sensitive values appear in URLs, page source, browser storage, analytics payloads, errors or screenshots;
  • whether destructive and privileged actions disclose scope and consequence;
  • whether errors reveal sensitive account or system information;
  • whether cross-origin embeds and third-party scripts are visible in scope;
  • whether logout and account/device state are understandable;
  • whether the interface uses concealment, false urgency or asymmetric friction.

Use the OWASP Web Security Testing Guide to classify observations and refer suspected vulnerabilities for separately authorised security testing. Do not exploit a suspected weakness to increase the apparent value of the audit.

7. Evidence and reproducibility

7.1. Minimum evidence record

Each finding MUST contain or reference:

finding ID and title
UTC date and time
scope unit, route and state
build/release if known
browser, OS, device/viewport, input and AT
role, permissions and anonymised account state
preconditions and synthetic test data
exact reproduction steps
expected result and its authority
observed result
task/user consequence
evidence attachments
applicable HIF and external criteria
occurrence count and sample denominator
limitations and confidence

Capture the smallest evidence that proves the point: annotated screenshot, short recording, DOM/accessibility-tree excerpt, console/network record, performance trace or session note. Preserve original evidence separately from annotations.

7.2. Evidence handling

  • redact tokens, credentials, personal data and unnecessary query values;
  • use stable evidence identifiers rather than embedding secrets in reports;
  • preserve source resolution and timestamps;
  • record transformations, compression and redaction;
  • store access-controlled evidence only for the agreed retention period;
  • delete or return it according to the engagement record;
  • never fabricate an unavailable state for a more persuasive screenshot.

Evidence may be hashed when chain-of-custody matters, but a hash does not prove that the interpretation is correct.

8. Finding model

8.1. Finding states

  • Verified: reproducible and supported by sufficient evidence.
  • Intermittent: observed but not reliably reproduced; conditions recorded.
  • Systemic: recurrence demonstrated across a defined pattern or population.
  • Hypothesis: expert or participant signal requiring further evidence.
  • Not reproduced: reported signal not observed under recorded conditions.
  • Resolved: acceptance criteria passed in retest.
  • Accepted risk: client-owned decision with rationale and review date.

8.2. Severity

Use the HIF scale:

  • S0 Blocker: task impossible, loss of control/data or immediate critical risk;
  • S1 Critical: high likelihood of material harm or exclusion;
  • S2 Major: task possible only with serious difficulty or unreliable recovery;
  • S3 Moderate: material comprehension, efficiency or consistency problem;
  • S4 Minor: local friction or inconsistency with limited consequence;
  • Observation: useful signal not established as a defect.

Severity describes consequence, not implementation effort, stakeholder status or whether the auditor is selling a fix.

8.3. Reach, frequency and confidence

Record each independently:

Reach

  • R1 Isolated: one known state or rare audience;
  • R2 Limited: one template, segment or secondary task;
  • R3 Broad: several templates, roles or a priority journey;
  • R4 Systemic: site-wide pattern or primary capability.

Frequency

  • F1 Rare: exceptional conditions;
  • F2 Occasional: some realistic attempts;
  • F3 Frequent: many realistic attempts;
  • F4 Inevitable: every applicable attempt.

Confidence

  • C1 Low: plausible hypothesis or incomplete reproduction;
  • C2 Medium: direct evidence with material uncertainty;
  • C3 High: repeatable evidence and clear criterion/consequence.

Do not multiply labels into a pseudo-scientific truth. A sortable planning score MAY be used as:

risk score = severity weight × reach × frequency × confidence factor

severity weight: S0=5, S1=4, S2=3, S3=2, S4=1
reach/frequency: 1..4
confidence factor: C1=0.5, C2=0.75, C3=1

S0 is always escalated immediately. The score is a queueing aid, not a measurement of harm, legal exposure or conformance.

8.4. Remediation priority

After risk classification, add:

  • strategic value or deadline;
  • dependency and affected owner;
  • estimated effort range and assumptions;
  • confidence in the proposed remedy;
  • opportunity to fix a shared root cause;
  • validation and regression cost.

Recommend an order, not a promise. Prefer root-cause changes that remove several verified barriers. Never downgrade accessibility or safety because the affected group appears commercially small.

9. Recommendation and acceptance

Every proposed change SHOULD include:

  • problem and affected task;
  • design principle or requirement;
  • recommended outcome, not only a prescribed widget;
  • viable options and trade-offs;
  • affected content, components, analytics and operations;
  • assumptions and dependencies;
  • acceptance criteria;
  • test cases and regression surface;
  • evidence needed to close the finding.

An estimate is a range tied to assumptions, access and definition of done. Discovery, design, implementation, content migration, QA, accessibility verification and release work MUST NOT be silently collapsed into one number.

10. Retest and regression

Retest uses the original environment and steps where still valid, plus the acceptance criteria. Record:

  • build and date;
  • exact cases rerun;
  • result and new evidence;
  • remaining limitation;
  • introduced or exposed regressions;
  • finding state.

A changed screenshot is not proof of resolution. For shared components, retest representative instances and affected tasks. For high-risk changes, include negative paths, input alternatives and recovery.

11. Client deliverables

Unless the engagement states otherwise, deliver:

  1. executive decision brief;
  2. scope, authorisation, environment and limitations;
  3. inventory and sampling register;
  4. task and coverage matrix;
  5. prioritised finding register;
  6. detailed findings with reproducible evidence;
  7. remediation roadmap with assumptions;
  8. accessibility and performance appendices;
  9. raw evidence index and retention date;
  10. retest and regression plan;
  11. conformance and statement limits.

Use HIF Audit Report Template for the report and HIF Test Design for the case register.

12. Ethical commercial model

Apply Legal, IP and Platform-Reference Governance to the engagement and use the Rights Clearance and Release Record for client material, screenshots, code, assets and remediation deliverables. Written authority MUST cover the tested environments and evidence methods. Unless the auditor is retained as qualified legal counsel, a legal-risk observation MUST be reported as bounded exposure requiring specialist review, not as a conclusion that the client has or has not broken the law.

The service MAY offer design and implementation after the audit, subject to:

  • disclose that the auditor may benefit from remediation work;
  • price and scope the audit so its conclusions stand independently;
  • give the client the complete usable evidence and acceptance criteria;
  • permit the client or another supplier to implement;
  • do not exaggerate severity, certainty or regulatory exposure;
  • separate verified findings from improvement opportunities;
  • do not make remediation purchase a condition of releasing evidence;
  • disclose material product, affiliate or tool relationships;
  • require a client decision for scope, trade-offs and accepted risk;
  • report when a proposed fix fails retest, including when produced by the auditor;
  • never promise certification, immunity, conversion uplift or defect-free operation without a separately supportable basis.

The durable commercial value is trustworthy diagnosis and verifiable improvement, not dependency on the auditor.

13. Statement limits

Reports MUST use bounded language:

  • “No failure was observed in the tested sample and environment”, not “there are no failures”.
  • “The sampled pages met the tested success criteria”, not “the site is WCAG compliant”.
  • “Laboratory result under the recorded conditions”, not “real users load the site in this time”.
  • “Expert evaluation indicates”, not “users cannot”.
  • “Suspected UI-boundary security issue; specialist verification required”, not “the system is insecure”.
  • “Estimate based on listed assumptions”, not “fixed cost” when discovery remains.

Only a complete, appropriately expert evaluation of a defined conformance scope can support an accessibility conformance statement. HIF audit levels are not legal certification.

14. Quality gate

Before issue, an independent reviewer SHOULD confirm:

  • authority and final scope are recorded;
  • every headline claim traces to evidence;
  • sample limits and environment are visible;
  • duplicates are reconciled;
  • severity reflects consequence;
  • confidence reflects the evidence;
  • accessibility statements use correct scope;
  • lab and field performance are separated;
  • sensitive evidence is redacted;
  • recommendations have testable acceptance criteria;
  • conflicts and commercial incentives are disclosed;
  • language is understandable to both decision-makers and implementers.

15. Official English-language sources