Digital phenotyping has emerged as a promising approach for supporting mental health assessment by continuously capturing behavioural data from smartphones and wearable devices. Recent systematic reviews suggest that passive sensing can identify behavioural patterns associated with depression, anxiety, bipolar disorder, and relapse risk, while offering opportunities for longitudinal monitoring beyond traditional clinic visits (de Angel et al., 2022; Alam et al., 2026). As interest in precision psychiatry grows, digital phenotyping is increasingly discussed as a tool capable of complementing conventional assessment and facilitating earlier clinical intervention.
Despite these advances, an important conceptual question remains insufficiently addressed: what exactly do digital phenotyping models measure?
What digital phenotyping actually measures
Most digital phenotyping systems analyse behavioural indicators, including mobility, sleep, physical activity, device interaction, and communication patterns, and identify statistical associations between these variables and mental health outcomes. However, these measures represent observable behaviour rather than psychological experience itself. Current evidence supports the use of behavioural markers as indicators of clinical change, but does not suggest that passive sensing alone can capture the subjective, contextual, and interpersonal dimensions that underpin psychiatric assessment (Alam et al., 2026).
Why behaviour is not a diagnosis
This distinction has practical implications for clinical decision-making. Behavioural changes are not specific to one cause. Reduced mobility, disrupted sleep, or decreased social interaction may be associated with depressive symptoms, but they may also reflect physical illness, occupational changes, caregiving responsibilities, socioeconomic circumstances, or personal lifestyle choices. As highlighted by de Angel et al. (2022), behavioural markers should therefore be interpreted within a broader clinical framework rather than considered independent indicators of psychopathology.
A patchwork evidence base
This limitation is reflected in the current evidence base. While systematic reviews consistently report promising links between passive sensing features and mental health outcomes, they also highlight considerable variation in how studies are designed and conducted. Differences in sensor selection, feature extraction, outcome measures, follow-up duration, and validation strategies make direct comparison between studies difficult and currently limit the translation of research findings into routine clinical practice (de Angel et al., 2022; Alam et al., 2026). Similarly, recent reviews examining digital phenotyping for relapse detection conclude that behavioural monitoring has considerable potential for identifying early warning signs, particularly in mood disorders. However, they also emphasise that existing predictive models require broader external validation and prospective evaluation before they can be reliably implemented across different clinical populations and healthcare settings (Dormechele et al., 2026).
Biomarkers, not verdicts
For clinicians, perhaps the most important distinction is that prediction does not necessarily equate to understanding. Machine learning models may identify individuals whose behavioural patterns differ significantly from their usual baseline, but they cannot independently determine the psychological meaning of those changes. Clinical assessment is still needed to determine whether a behavioural change reflects worsening symptoms, an environmental stressor, a medical condition, or a normal life event.
Rather than viewing this as a limitation, it may be more helpful to think of digital phenotyping in the same way as other clinical biomarkers. Laboratory results rarely lead to a diagnosis on their own; instead, they provide additional information that needs to be considered alongside the patient's history, examination, and other tests. Behavioural biomarkers generated through passive sensing may serve a similar function by identifying patients who warrant closer assessment, without replacing the interpretative role of clinicians.
This perspective aligns with recent recommendations for clinical implementation. Rather than promoting fully autonomous decision-making, current literature supports combining digital phenotyping with patient-reported outcomes, structured clinical interviews, real-time symptom monitoring, and clinical judgment. Such multimodal approaches acknowledge that behavioural data provide valuable longitudinal information, while recognising that clinical meaning emerges only when those data are interpreted within the patient's personal, social, and medical context (Astill Wright et al., 2026).
What this means for builders
As digital phenotyping moves beyond research environments, future work should prioritise clinical validation over algorithmic optimisation alone. Improving model transparency, standardising methods, and demonstrating clear clinical usefulness across different populations will be essential before passive sensing technologies can be routinely used in mental healthcare pathways. Equally important will be establishing how behavioural data can be presented in ways that support, rather than complicate, clinical judgement.
Digital phenotyping has already shown that it can capture behavioural patterns at scale, but its clinical value will not depend on technology alone. The main challenge is not measuring behaviour, but understanding what it means. High-frequency behavioural data need to be translated into insights that are useful for clinical decisions, without oversimplifying the complexity of mental health. In this sense, the future of digital phenotyping depends less on improving prediction models and more on understanding how behavioural data can be combined with clinical judgment, patient experience, and context to support, rather than replace, psychological assessment.
🪺