At a recent nest session, founders, clinicians, and researchers kept arriving at the same wall. The conversation would start somewhere specific, clinical credibility, patient trust, privacy trade-offs, outcome measurement, then stall at the same point. None of it could be resolved inside a single conversation. All of it depended on something conversation alone cannot supply.
That something is context.
The industry has made real progress on the conversation itself. Large language models now sustain exchanges that feel responsive, warm, and clinically informed. But a convincing conversation is not a safe one. The gap between the two is where most of this field's hardest problems live, and better language will not close it.
What clinicians mean by context
In clinical practice, context is not a feature. It is the foundation of the work. A clinician opening a new session draws on a patient's history, prior diagnoses, medications, family circumstances, and the trajectory of the relationship itself, the material clinical judgement is built from, not detail meant to personalise a greeting.
A response appropriate for someone in early-stage anxiety management looks nothing like one appropriate for someone with a documented history of crisis, even if what they say today is identical. The conversation is the surface. Context determines how the surface should be read.
Most current evaluations test that surface alone. They rely on single-session testing rather than longitudinal data, which limits what they can tell us about real-world clinical performance (Wang et al., 2025; Guo et al., 2024). What happens across months, across changes in someone's life, remains largely unmeasured.
Why the fixes are not fixes
The technical tools have genuinely improved. Context windows have grown. Retrieval-Augmented Generation lets systems pull relevant information from external memory rather than a single prompt. Persistent memory carries information across conversations. None of this is equivalent to clinical context.
Larger context windows have a documented blind spot: models tend to weight information at the start and end of a long input, and neglect what sits in the middle (Liu et al., 2023). A 200,000-token window does not guarantee that a detail from session seven is read correctly in session twenty-three.
Retrieval improves factual grounding, but it retrieves text, not understanding. A stored note reading "history of crisis" is a string of words. What it should do to a response, raise the threshold for escalation, change how a disclosure is handled, is an interpretive act the retrieval mechanism does not perform.
Persistent memory carries a different risk. Outdated data has been shown to lower the accuracy of AI-supported decisions (Auf et al., 2025). A description that was accurate six months ago can actively mislead now, and a system with no mechanism for revising what it holds will keep surfacing it as current fact. Researchers building frameworks for longitudinal health AI agents describe what real context-awareness would require: coherence across interactions, continuity of care goals, adaptation as circumstances change, accountability that outlasts a single session (Lin et al., 2026). These are design requirements. They are not something today's systems already meet.
Where this becomes a safety problem
Crisis detection is where the gap turns concrete. Mental health disclosures are rarely direct. The signal is often distributed across sessions, a change in tone, a shift in how someone describes their relationships, a pattern of escalating or withdrawing that only becomes visible over time. A system reading each conversation in isolation cannot see that pattern. It can only respond to the message in front of it.
Even within a single exchange, the evidence is not reassuring. A recent expert-rated evaluation of nine LLMs on therapeutic dialogue found strong performance on structural safety but consistent weakness on affective attunement, a pattern researchers call the cognitive-affective gap (Badawi et al., 2026). If models are inconsistent within one response, there is little reason to assume that inconsistency resolves itself across twenty.
False positives carry their own cost. Systems that escalate too readily teach people to discount alerts, including the accurate ones. Calibrating this needs a longitudinal picture of the person, not a probabilistic read of one message.
Informed is not governed
A distinction that came up repeatedly at the nest session: clinically informed products have had a clinician involved somewhere in development. Clinically governed products have a credentialed professional holding accountability for clinical decisions, including escalation, with authority to override the system.
Context sits at the centre of that distinction. A clinician governing a product needs to know what the system is drawing on, or oversight is nominal. The same logic applies to patient trust: one founder in the network described a beta tester who felt unsafe on learning the founder could read full session transcripts. That reaction is not irrational. It reflects a reasonable expectation that therapeutic disclosures get the same defined scope clinical settings provide.
The EU AI Act classifies health AI as high-risk. Meeting that obligation in mental health means more than paperwork. It means being able to account for what a system knows about a user, and why that is shaping its response.
What building for this actually requires
A few design shifts are worth naming, not as solved problems but as directions: structured longitudinal records that track clinical variables rather than raw transcripts; user-controlled, editable memory, so people can see, correct, and delete what is held about them; transparent context summaries, so a system's working assumptions stay visible rather than hidden; genuine uncertainty handling, so a system without enough context says so rather than answering anyway; and measurement-based care, collecting patient-reported outcomes at regular intervals, over optimising for engagement.
The field has spent considerable energy making mental health AI sound right. The harder work is making it know enough. That is a design problem, a governance problem, and an honesty problem before it is a technical one. Founders need to say plainly what their system does and does not know, what it is and is not equipped to decide, and what clinical scope they are and are not claiming.
A product that handles conversation well without clinical context is not a limited mental health tool. It is a plausible-sounding stranger. That is not a safe default in a field where getting it wrong is measured in people's wellbeing.
🪺
If you are finding this topic difficult on a personal level, support is available globally at findahelpline.com.