How to review an accessibility pull request
Open-source accessibility fixes need more than a visual approval. Maintainers need a review record that connects the reported barrier, affected users, test method, environment, and evidence.
What the first review found
Seven real accessibility pull requests were scored against a 12-point rubric. None reached the mature band, and the most heavily reviewed pull request tied for the lowest score. Review volume alone did not demonstrate review quality.
The seven scored pull requests, each written up as a case study:
- Bootstrap — pull request #41607 (opens in a new tab)
- Bootstrap — pull request #42500 (opens in a new tab)
- Bootstrap — pull request #42524 (opens in a new tab)
- Bootstrap — pull request #42539 (opens in a new tab)
- MUI Material UI — pull request #48572 (opens in a new tab)
- Storybook — pull request #35321 (opens in a new tab)
- VS Code — pull request #324192 (opens in a new tab)
The seven pull requests explained in plain language (opens in a new tab) · The review rubric (opens in a new tab)
The question is not whether contributors cared. It is whether a maintainer could determine what barrier was fixed, reproduce the test, understand the affected assistive technology, and know what remained untested.
For maintainers
- Ask for the affected user and task.
- Require reproducible test conditions.
- Separate automated checks from user evidence.
- Record limitations and regressions.
Reusable artifacts
- ACCESSIBILITY.md template
- Accessibility issue template
- Pull-request template
- Review rubric: six criteria, 12 points
The review rubric
The rubric scores how the review process handled an accessibility pull request, not whether the fix itself was technically correct. It uses only the public record: the pull request description, every comment, and linked issues. Intent that isn’t written down is not inferred.
Six criteria, each scored 0 (absent), 1 (partial or implicit) or 2 (explicit and complete):
- User impact stated. Does the record name the affected users and the task they couldn’t complete, rather than “fixes an a11y issue”?
- WCAG mapping. Is the change tied to a specific WCAG success criterion anywhere in the thread?
- Verification evidence. Is there a record of the fix being checked with real assistive technology or a real device: which one, which version, and what was observed? Automated unit tests don’t count; they verify the mechanism, not the outcome. Full marks need a second person, before merge.
- Reviewer confidence signal. Did a reviewer engage with the accessibility claim itself (verify it, probe it, cite a specification), rather than approve without comment or self-merge?
- Direct language. Does the discussion name who is affected (for example, “NVDA users” or “keyboard-only users”) rather than “accessibility” in general?
- Outcome clarity. Could a future contributor tell, from the record alone, why the pull request landed, stalled or was rejected?
Score bands: 10–12, a review an accessibility-mature project would recognise as normal; 6–9, some signals present but reliant on individual expertise; 0–5, no structured accessibility review, with the outcome depending on who happened to touch the pull request. None of the seven pull requests studied reached 10.
Read the full rubric, with scoring guidance (opens in a new tab) · Help validate it if you use assistive technology
Use the review infrastructure
Read the published analysis
Can your project actually review an accessibility contribution? (opens in a new tab) explains the scoring method and the difference between extensive discussion and review evidence.
The overlay question security is rarely asked (opens in a new tab) examines accessibility overlays as a supply-chain and trust problem.