
What does the thirtieth study tell you that the first does not?
I ran accessibility research for years before I ran enough of it to find out that the people I was learning from disagreed with each other.
Ask one person for directions and you get a route. Ask thirty and you find out the town has two stations, and that most people assumed you meant the other one.
For a long stretch of my career, accessibility research was the first kind of asking.
What does one session actually buy you?
Less than we pretend, and it is worth being exact about why.
Something real does happen the first time a team includes a disabled participant. The room stops arguing from imagination. I would not talk anyone out of that, and for most teams it is the single biggest change available.
But one session cannot tell you whether you watched a pattern or a person. And we handle that uncertainty in a way that should embarrass us, because we handle it differently depending on who the participant is.
Nobody watches one sighted participant miss a button and writes "sighted users miss buttons" in the summary. Everybody watches one blind participant struggle with a table and writes "blind users need." One person becomes a population as soon as the population is small and unfamiliar to us. That is not a finding. It is a stereotype with a recording attached.
The opposite failure is quieter and looks more rigorous. A real finding gets set aside because it came from one person, in a study of one, with nothing to weigh it against. Either way the session produced a feeling rather than a direction.
What did thirty change?
At the Rhonda Weiss Center, under a multiyear Department of Education grant, I led design for a system that publishes IDEA data: what states report about how children with disabilities are served. We ran more than thirty usability studies with disabled participants, including blind and neurodivergent users, around a stakeholder group of about eighty people.
What I did not expect, and what I now think is the entire point, is that they disagreed with each other.
Not about whether the thing worked. About what a good answer looked like. Participants using screen readers did not all want the same treatment of the same table. People navigating by keyboard did not want what people using magnification wanted. Participants who process dense text quickly and participants who do not asked for opposite defaults, and each of them was right about themselves.
Past a certain number of sessions, what do blind users need stops being a question with an answer. It resolves into a range, and the range is the finding.
Why is a range harder to hear than a defect list?
Because a defect list can be finished.
If thirty people all hit a control with no name, you name it and the work is done. If thirty people want three different reading orders, there is nothing to fix. There is a decision about which default ships, what becomes adjustable, and what gets offered as a second route to the same meaning.
That is more expensive, and it is also just design. We already know how to do it everywhere else. Nobody ships one font size and considers the argument settled. We have simply been permitted, in accessibility alone, to treat a category of people as having one set of needs, mostly because we had met one of them.
Plurality is the word I would put on it. The reason to include disabled people in research is not to acquire a spokesperson. It is to stop designing for a composite who does not exist.
What did it cost?
Posts like this usually skip this part, so let me not.
Thirty sessions is mostly logistics. Recruitment that reaches past the same few people who always say yes. Scheduling around the assistive technology the participant actually uses rather than the one we happen to own. Paying people, because unpaid consultation quietly selects for those with spare time. Sessions that overrun because the setup did.
A multiyear federal grant paid for that. A product team on a quarterly cycle has a harder case to make, and I am not going to pretend otherwise.
What I will say is that the expensive part was almost never the sessions themselves. It was everything arranged around them, and most of that is a budgeting and scheduling problem rather than a research one. Teams rarely stop at one study because they learned enough. They stop because nobody owned the logistics of the second one.
What would I do with a tenth of that budget?
- Three sessions, not one. Three is not statistically anything. It is enough to surface one disagreement, and one disagreement changes how a team talks for the rest of the project.
- Recruit for difference inside the group. Two screen reader users with different tools and different tenure will teach you more than two who match.
- Pay everyone, at the rate you pay any other participant. This is the line most likely to be cut, and it decides who can afford to show up.
- Stop writing findings as "blind users." Write "two of the three participants using screen readers." It is longer, it is true, and it stops a person being promoted to a population.
- Book the second study before the first one runs. The highest-value scheduling decision available to you, and in January it costs nothing.
- Keep the disagreements in the report. A summary that resolves every conflict into a single recommendation has thrown away the most expensive thing you bought.
The part I keep carrying
Our defaults encode an assumption about who is in the room. The correction most teams reach for is to put one person in it, once, and treat what that person says as the answer.
That is better than nothing, and it is still a version of the same mistake: one voice standing in for everybody, selected by us.
Look at the last research plan you wrote and count the sessions that included disabled participants. If the number is one, the useful question is not what you learned from it. It is what the second one would have told you.