Skip to main content
All articles
Accessibility testing6 min readPublished Updated

Manual vs Automated Accessibility Testing: A Practical Plan

Combine automated checks with manual review using a coverage matrix, complete journeys, and reproducible findings that show what was actually tested.

By OnChange

A useful accessibility test plan combines repeatable checks with direct evaluation of how people complete a task. Automated testing can find certain implementation problems quickly. Manual testing lets an evaluator inspect meaning, interaction, and the states that a scan did not reach. The choice is where each method contributes evidence to the same review.

Start with the question you need to answer: can someone find an appointment, choose a time, correct an error, and receive confirmation? A list of scanned URLs does not explain whether that process works. This guide shows how to plan both methods, record their coverage, and turn results into a useful handoff.

What automated accessibility testing can establish

An automated check evaluates conditions defined by its rules against the content it can inspect. Depending on the tool and its configuration, it may identify missing accessible names, invalid relationships, or some contrast problems. Check the tool's documentation before assuming that a particular rule runs or that every page state is included.

W3C WAI's guidance on evaluation tools explains that tools assist evaluation, cannot automatically check every aspect of accessibility, and can produce misleading results. A report with no detected issues therefore describes the checks performed on the captured content. It does not establish WCAG conformance.

Keep the rule identifier, tool version, affected element, and reproduction state with each result. If two tools disagree, compare what they actually tested. A scan of the initial page and a scan after opening a dialog are different observations, even if both reports show the same URL.

What manual testing adds

Manual evaluation starts with a task and follows the interaction. An evaluator can ask whether a field label makes sense, whether focus moves predictably, whether instructions explain the required input, and whether an error can be understood and corrected.

The methods used should match the question. A keyboard review examines operation without a pointer. A screen reader review examines the information exposed through an agreed browser and assistive technology combination. A visual review can examine focus indicators, layout, and presentation. Record these methods separately so a reader knows which evidence supports which conclusion.

For detailed procedures, use the keyboard accessibility testing checklist and the form testing guide. Their observations still need to be assessed against the applicable success criteria; completing a checklist is not a whole-site conformance decision.

Plan the remaining methods with the screen reader testing guide, zoom, reflow, and text spacing checks, color contrast workflow, and image alternative review. Give each method its own tested states and evidence. A page that works with a screen reader still needs the applicable visual and enlargement checks.

Build a coverage matrix before running checks

Use a compact matrix to assign an observation to a method and a state. The following is an illustrative plan for a fictional appointment service, rather than a prescribed WCAG evaluation sample.

Task or stateAutomated contributionManual contributionEvidence to retain
Search for an appointmentInspect the loaded search interfaceEnter a query, change filters, inspect the result announcementQuery, filter values, result state, observations
Open the booking dialogRun checks after the dialog is openInspect entry focus, operation, dismissal, and return focusTrigger, dialog state, keyboard sequence
Submit incomplete detailsInspect the rendered error stateFind each error, understand the correction, and continueInvalid test data, message, field relationship
Receive confirmationInspect the final rendered stateDetermine whether completion is perceivable in the tested setupFinal state and assistive technology observation

Add a coverage status such as tested, blocked, or not reached. Assign an owner to every blocked state. For example, a missing test account is an access gap that needs resolution; it is not a successful authentication test.

Write the product version, account role, language, browser, and relevant viewport next to the matrix. These details make it possible to repeat the review after a release. The audit scope template explains how to agree those boundaries before evaluation starts.

Follow complete processes, including failures

A booking form can be accessible in its initial state and fail when validation inserts new content. Plan both the intended route and meaningful branches: an unavailable time, incomplete data, a cancelled dialog, and a confirmation reached after correction.

WCAG 2.2's conformance requirements for complete processes require all pages in a process to conform at the specified level for conformance within that process. This is why an initial-page scan cannot settle a conclusion about the entire booking journey.

Keep a short state log during testing. Record what opened a panel, which input produced an error, and which controls were available afterward. If a branch depends on a third-party service or an account permission you cannot obtain, describe the limitation and the next action needed to inspect it.

Interpret reports without turning counts into a score

Separate a tool observation from an evaluated finding. A tool may flag one repeated component on several pages. The remediation team needs to know the shared cause and affected occurrences, while the coverage record needs to show which occurrences were actually inspected.

Use a finding record with these fields:

  1. The task, URL, state, and tested environment.
  2. Steps another reviewer can repeat.
  3. The observed result and its effect on completing the task.
  4. The relevant requirement, with the evaluator's reasoning.
  5. The affected component, owner, and proposed correction.
  6. The evidence and conditions needed for a retest.

A declining issue count can be useful for tracking a particular review, but do not relabel it as a percentage of WCAG conformance. A change in rules, pages, or reachable states can alter the count without improving the same user journey. Explain what was held constant before comparing reports.

Include disabled people in evaluation

Standards-based evaluation and evaluation with disabled people contribute different evidence. The latter can reveal where an interface is confusing or difficult in actual use, including problems a technical review did not identify.

W3C WAI's guidance on involving users recommends combining user involvement with standards evaluation and cautions against generalizing from a small number of participants. A successful task completed by one screen reader user does not establish that the service works for all disabled people.

Choose realistic tasks, explain the study's boundaries, and document the participant and technology characteristics relevant to the findings. When reporting a barrier, distinguish an observed user difficulty from a formally evaluated WCAG failure. Both can inform a fix without making the same claim.

Preserve a review that can survive the next release

Deliver the coverage matrix with the findings, rather than sending only a scan export. Include unresolved access gaps and the versions of the tools used. The accessibility evidence handoff guide provides a structure for connecting that material to ownership and retesting.

Saved HTML, styles, and screenshots can help reviewers investigate what changed between releases. They do not independently prove runtime focus behavior, screen reader output, or a completed authenticated journey. OnChange can support that evidence workflow, while the evaluator remains responsible for the interaction tests and their conclusions.

When a release changes a shared form or navigation component, return to the tasks that relied on it. Use the accessibility regression testing guide to plan that review. Keep untested behavior visible in the record so the next reviewer can finish the work.

Sources and how to use this guide

The standards and guidance linked above were checked on 6 October 2026. WCAG 2.2 defines the conformance requirements. WAI's evaluation resources explain testing approaches. The appointment example, coverage matrix, and handoff fields are practical planning suggestions from OnChange, not additional WCAG requirements.

Keep useful evidence of what changed

Start with a page you care about and review its changes with OnChange.

Get started free

Keep reading