Companies releasing AI mental health tools point to safety evaluations as evidence that those tools are ready to deploy. The problem is that most of those evaluations were never designed to test what happens when real people, in real distress, actually use them. A preprint published in January 2026 makes that gap visible at a scale the field has not seen before.

The study, Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety, does two things. It replicates four published safety evaluations on a set of frontier general-purpose models and one purpose-built mental health system. Then it does something most evaluations never attempt: an ecological audit of more than 20,000 genuine user conversations with that purpose-built system. The contrast between the two is the point.

What the evidence shows

Most published safety evaluations rely on small, simulated prompt sets, often a few hundred items, with limited diversity and no known relationship to how real users actually write. They are useful as standardised stress tests. They are not sufficient to establish that a tool is safe in deployment, particularly for the failure modes that matter most: rare, but clinically serious.

The reason is in the language. People in distress rarely state risk directly. Suicidal thinking and self-harm intent are often expressed in ways that are subtle, ambiguous, indirect, or that surface gradually over the course of a conversation. Many guardrails, by contrast, rely on static refusal templates or generic crisis-response heuristics built to catch explicit statements. When the risk is implicit, those systems can miss it.

The findings are not uniformly bleak, and it would be misleading to present them that way. The purpose-built system in the study, designed with layered safeguards for suicide and non-suicidal self-injury, was substantially less likely than general-purpose models to produce harmful content across suicide, self-harm, eating disorder and substance-use scenarios. The authors' conclusion is not that AI mental health tools are inherently unsafe. It is that domain-specific training combined with system-level safety controls outperforms guardrails alone, and that benchmark performance alone tells you very little about real-world behaviour.

That distinction matters for how the results should be read. The critique applies most forcefully to consumer-facing, unmonitored tools that lean on a generic model and a thin layer of refusal logic, then cite a benchmark pass as evidence of safety. It applies far less to systems with clinical oversight and purpose-built safeguards. Not all AI mental health tools are equivalent, and the evaluation gap is widest precisely where the safety architecture is thinnest.

Two caveats are worth stating plainly. The paper is a preprint and has not yet been peer-reviewed. That is a reason to read it carefully, not a reason to dismiss it: the methodology is transparent and the conversational data is real. And the deployment findings come from a single system, so the specific numbers should not be generalised across the market.

Why the scale makes this urgent

The argument becomes concrete at scale. When a tool is used by hundreds of thousands of people, a failure rate that looks small in a benchmark report translates into a large absolute number of harmful interactions. A one per cent failure rate is reassuring on a slide and unacceptable in a population. Evaluations built on a few hundred prompts simply cannot detect low-base-rate failures at the frequency they occur in the wild.

The regulatory backdrop compounds the problem. The US Food and Drug Administration has not yet authorised a single generative AI mental health device, so there is no settled bar to clear. In the absence of one, 43 US states introduced more than 240 AI-related bills in 2026, many targeting AI chatbots used by minors in mental health contexts. The regulatory gap is real. So is the evidence gap underneath it.

What this means for builders and clinicians

For builders, the lesson is uncomfortable but clarifying: passing a benchmark is not the same as being safe, and presenting it as such erodes the credibility the whole field depends on. The more honest claim is narrower and more useful. It describes how a tool was tested, on what kind of data, against which real-world distribution of user language, and where the known limits sit.

Better evaluation looks less like a one-off prompt set and more like continuous, ecological monitoring of real conversations, with particular attention to indirect and evolving expressions of risk. For clinicians weighing whether to recommend a tool, the question to ask is not whether it has been evaluated, but how, and whether anyone is watching what happens after deployment.

This is the field's most important unsolved problem right now. The distance between what a benchmark measures and what a tool does in someone's hands is where the real safety question lives, and it is the gap we will be watching most closely.

🪺