AI personas are not user testing. Why lived experience can’t be simulated

I want to start by saying I am not against Artificial Intelligence (AI) in principle. There are huge opportunities in how it could benefit our lives including, as we’ve written before, in supporting accessibility work. But its rapid roll out, the lack of meaningful policy around it, and its concentration in the hands of a few large companies should give us all pause.

We are already seeing people’s jobs change and mass redundancies follow. And yet there remain gaps, functions that AI can’t replace, and shouldn’t try to. At least not yet.

Inclusive user testing is one of them.

A growing number of products now offer ‘synthetic users’ or AI-generated personas. These are simulated disabled users who will, apparently, tell you how accessible your website is without the inconvenience of recruiting, scheduling or paying real people.

In a time of increasing pressures on resources and budgets, it is understandable why this is appealing. Testing with real participants takes time, care and money. A chatbot that role-plays a screen reader user is available instantly and costs pennies.

But there is a fundamental problem here. An AI persona can only ever reflect the data it was trained on. When it comes to the experiences of disabled people online, that data is profoundly incomplete.

The data simply isn’t there

AI models learn from what has been written down and published on the web. So ask yourself: how well is the lived experience of disabled users actually documented online?

Consider what the training data looks like in practice:

  • The overwhelming majority of web pages fail basic accessibility standards to some degree. WebAIM’s annual survey of the top one million homepages consistently finds detectable WCAG failures on well over 90% of them
  • Much of what is written about disabled users online is written about them, not by them.
  • First-hand accounts of access barriers tend to live in forums, community spaces and private conversations, not in the neatly published prose used to train AI

Many blind screen reader users are highly technically literate. Some document their experiences online, and some contribute directly to the Web Content Accessibility Guidelines. But many others are far less experienced at navigating their devices with a screen reader, and have limited technical confidence. The majority, experienced or not, will never publish an account of their experience anywhere.

An AI persona of ‘a blind user’ is therefore built from a thin layer of personal accounts, mostly from the most technically fluent, mixed with a great deal of what non-disabled people have written about blind users. It will confidently reproduce the textbook version of screen reader use. It cannot reproduce the workarounds, the fatigue, the abandoned baskets, the ‘I just phone them instead’ moments that real testing surfaces every time.

This doesn’t only apply to screen reader users. It applies to users of all assistive technologies, and to disabled and neurodivergent users who don’t use any assistive tech at all, whose experiences are even less likely to be visible in the data.

Compounding exclusion

There is a second, less obvious problem and it is the one that worries us most.

If most of the web is inaccessible, then disabled users are being excluded not just from websites, but from the feedback channels about those websites. Survey platforms with unlabelled form fields. Feedback widgets that can’t be reached by keyboard. Pop-up questionnaires that vanish before a magnification user has found them. CAPTCHA walls. Analytics that quietly discard the sessions of people who gave up.

The result is a loop of compounding exclusion:

  1. Access barriers prevent disabled users from completing journeys
  2. The same barriers prevent them from telling anyone about it
  3. Their absence from the data is read as absence of a problem
  4. AI trained on that data learns a version of ‘the user’ in which disabled people barely feature
  5. Products tested against that AI inherit the same blind spots. The loop tightens

Each turn of the loop makes the underrepresentation look more like the truth.

An availability bias, at scale

Researchers would call this selection bias. We mistake the data we happen to have for the data that exists. AI doesn’t correct this bias. It industrialises it. A synthetic persona will happily generate plausible-sounding feedback all day long, and plausibility is precisely the danger. It feels like evidence. It lets a team tick the ‘inclusive testing’ box while never actually encountering a disabled person.

Real usability testing with real participants is the opposite of plausible. It is frequently surprising. It’s the participant who navigates in a way no persona would predict. It’s the barrier nobody on the team knew existed because nobody on the team had ever needed to know. That surprise is the value. A model trained on incomplete data cannot surprise you with what the data never contained.

Where AI can genuinely help

None of this means AI has no place in inclusive research. Used responsibly, it can support the work, for example:

  • Transcribing and summarising testing sessions
  • Helping draft discussion guides and easier-to-read participant information
  • Spotting patterns across large volumes of genuine user feedback
  • Handling admin so budgets stretch to more real participants, not fewer

Used as a support tool, AI can free up time and money for testing with real people. Used as a substitute for those people, it becomes a way of automating their exclusion.

Nothing about us without us

The disability rights movement gave us the principle ‘nothing about us without us’.

An AI persona is, almost by definition, something about disabled people without disabled people.

So before anyone offers you synthetic disabled users, we recommend asking some careful questions:

  • What data were these personas trained on, and who was missing from it?
  • Can this tool ever tell me something a real participant couldn’t?
  • Can a real participant tell me something this tool never could?
  • Who is being paid for this insight? Were the people whose experiences trained the model paid for theirs?

Simulated users can only ever give you simulated confidence. Real confidence comes from watching a real person, with real access needs, complete a real task, or fail to.

But by replacing disabled people with synthesised versions of themselves, we risk something worse than bad research. We deepen the very exclusion this work exists to undo and we tell disabled people, however quietly, that they do not belong. Or worse, that they need not.