A 2020 analysis published in Nature Digital Medicine estimated that only 2% of popular depression smartphone apps had a reasonable evidence base.1 The field has grown considerably since, but the proportion has not kept pace. A 2024 systematic review in JMIR found that despite significant funding directed at mental health app development, very few apps have been deliberately adapted to meet the evidence standards that clinical adoption requires.2 These are not fringe findings. They reflect a structural condition that anyone building seriously in this space already knows but rarely names directly.
The gap is not primarily a quality problem. The builders working on mental health technology are, by and large, thoughtful people who understand the stakes. The gap is a systems problem: evidence generation is slow, expensive, and designed for a world in which interventions do not update every two weeks. The result is a field in which the tools with the most users often have the least evidence, and the tools with the most evidence rarely reach the people who need them.
What the Research Actually Shows
The most recent systematic review and meta-analysis of standalone smartphone apps for mental health, published in The Lancet Digital Health in 2025, found small to moderate effect sizes for depression, anxiety, and sleep problems compared with inactive controls, though it flagged significant methodological concerns.3 Risk of bias was moderate to high across included studies. When publication bias was corrected for, effect sizes for depression dropped from 0.45 to 0.18, and for anxiety from 0.35 to 0.18. The honest reading of this literature is that standalone apps can help, particularly when no evidence-based first-line intervention is available. The evidence base, though, is weaker and narrower than the market would suggest.
The engagement picture compounds the problem. A narrative review by Boucher and Raiker (2024) found that dropout was a reported problem in nearly all included studies of mental health apps across diverse populations.4 Establishing and maintaining user engagement was a pervasive challenge. Crucially, positive subjective reports of usability and satisfaction were insufficient to predict objective engagement: users who said they liked an app often stopped using it. This matters because most evidence trials measure short-term outcomes in motivated participants. Real-world attrition consistently exceeds what clinical trials capture.
Innovation and efficacy alone do not result in adoption in real-world clinical settings. This has been demonstrated repeatedly. The field continues to act as though it has not.
The research-to-practice gap runs in both directions. Riper (2021) noted that bridging the gap between evidence-based digital interventions and routine care has proven to take approximately 20 years, a figure drawn from Rogers' innovation cycle.5 A 2024 qualitative systematic review in BMC Health Services Research identified the most persistent barriers as fragmented commissioning structures, unclear regulatory pathways, and the absence of clinical workflow integration at the point of implementation.6 These are not technical failures. They are organisational and systemic ones.
Why This Pattern Persists
Part of the answer is structural. Clinical trial methodology was designed for pharmaceutical interventions with fixed dosing and stable mechanisms. A smartphone app is an evolving product. Its content, interface, and recommendation logic may change significantly between the time a trial is designed and the time it reports. The assumption of a stable intervention, foundational to the RCT model, does not hold.
Part of the answer is financial. Evidence generation is expensive. Venture-backed products operate on timelines that do not accommodate longitudinal outcome data. The incentive structure rewards deployment over evaluation. This is not a moral failing. It is the predictable consequence of applying startup economics to clinical infrastructure.
Part of the answer is fragmentation. The people who generate evidence (academic researchers) and the people who build products (founders and engineers) rarely work in the same rooms or share the same vocabularies. A 2024 analysis of barriers to digital mental health implementation found that the absence of cross-sector collaboration was among the most consistently cited structural obstacles.6 Innovators working in isolation produce work that does not translate. Researchers working in isolation produce findings that do not reach practice.
A note on proportionality. None of this means that building without a full RCT is irresponsible. Context matters. An app that helps people track their mood, find a therapist, or understand their diagnosis does not carry the same evidence burden as one making clinical claims. The problem arises when the field fails to make these distinctions: when marketing language implies efficacy that the evidence does not support, and when builders cannot easily locate the research that would help them understand what has and has not worked before them.
What This Means for Builders
The treatment gap is real and urgent. More than 70% of people with mental health problems cannot access timely treatment, according to WHO data.7 Digital interventions are not an optional add-on to the mental health system. They are, for the foreseeable future, the primary channel through which millions of people will first encounter mental health support. The stakes of building carelessly are high.
The most practical response is not to demand that every founder run an RCT before releasing a product. It is to build the connective tissue between evidence and practice that currently does not exist at sufficient density: forums where researchers and founders can translate findings into product decisions; shared datasets that allow comparisons across interventions; funding models that reward evaluation as part of the development cycle rather than as an afterthought.
Torous et al. (2025) argue in World Psychiatry that the field needs not just more evidence but better integration between the people generating it and the people using it.8 The Society of Digital Psychiatry has proposed a three-pronged approach: improved education, structured digital navigator programmes, and expanded international collaboration.9 These are institutional responses to what is, at its core, a community problem. The right people are not finding each other. The work that could change things is happening in isolation.
That is the problem nest exists to address. Not by solving the evidence gap directly. That is the work of researchers, regulators, and commissioners. nest's role is to make sure that the people best positioned to close it are in the same conversation.