Research

How to tell whether an emotional wellbeing app has evidence behind it

on August 21, 2026
Article cover: How to tell whether an emotional wellbeing app has evidence behind it

Abstract

The commercial listing of an emotional wellbeing app presents types of support of different natures in the same format: a university study, a volume of satisfied users or a basis in cognitive behavioral therapy. A systematic review with meta-analysis published in npj Digital Medicine, which started from 16,971 records and retained 110 studies of 112 self-guided mental health apps evaluated in controlled trials, found that the methodological quality of the studies is not associated with an app’s public availability and that the effect size of available apps, g of 0.33, and that of unavailable apps, g of 0.45, do not differ significantly. Apps with evidence of effect exist, and availability in the store does not distinguish them from the rest. The distinction is possible from outside the development team through five criteria with answers verifiable in the publication of each study: the test population, the size and duration, the comparator, the funding and the variable measured.

Introduction

The supply of emotional wellbeing apps comes with heterogeneous claims of support. One listing may mention a university study, another a large number of satisfied users, and a third a basis in cognitive behavioral therapy. All three claims share a format and describe support of different natures.

Public availability does not order that heterogeneity either: the available synthesis data indicate that the methodological quality of studies bears no relation to whether an app can be downloaded. Verifying the support requires criteria of its own, checkable from outside the development team and without clinical training, because scientific publication standards require the elements those criteria examine to be declared.

Types of support and their hierarchy

The word study covers designs with different scopes.

At one end sits evidence of acceptability or perceived benefit. Users are asked whether they found the app useful, whether they would recommend it and whether they would keep using it. It is the most frequent because it is the cheapest to obtain and the quickest to publish, and it describes the experience of use. An app can obtain high acceptance without changing any of the variables that prompted its download.

At the other end sits the systematic review with meta-analysis. It gathers every published trial of an intervention, weighs its risk of bias and calculates a combined effect. It is the most expensive evidence to produce and the least abundant, and also the only kind that supports an estimate of how much effect to expect on average.

Between the two sits the individual controlled trial, comparing those who used the app with a group that did not. Its design settles most of the questions the evaluation criteria examine.

Certification occupies a separate plane. A product can meet technical requirements, pass an accessibility assessment, comply with the GDPR and state clearly that the replies come from an artificial intelligence, as Article 50 of the European regulation has required since August 2026. None of those requirements reports on the effect. A seal certifies what its procedure examines, and reading it as certification of efficacy constitutes an error of interpretation.

Verifiable evaluation criteria

Evaluating the support behind an app admits five criteria with answers verifiable from outside the development team.

  • Test population. A trial with undergraduates carrying no diagnosis says little about people in their sixties with chronic insomnia. The criterion compares the sample described in the study with the target population.
  • Size and duration. A pilot of twenty participants over two weeks explores without establishing. Duration weighs as much as size, because many effects appear in the first week through novelty and fade by the fourth.
  • Comparator. It is the element of the design that most conditions the conclusion. Comparison against a group receiving no intervention measures the sum of the app’s effect and the effect of the attention received; comparison against a waiting list adds the effect of expectation; comparison against an established alternative intervention answers the pertinent question, that of whether the app adds anything beyond what already existed.
  • Funding. Funding by whoever sells the product does not invalidate a result, and it requires closer attention to the design, the prior registration of the protocol and the publication of negative results.
  • Variable measured. Counting how many times the app was opened, how many consecutive days of use it recorded or how many exercises were completed describes use. A change on a validated symptom scale, with assessment before and after and a comparison group, describes effect. Presenting the first as though it were the second is the sector’s most widespread confusion.

The five criteria have verifiable answers because standards exist that require them to be stated. A well-published trial follows CONSORT and the description of the intervention follows TIDieR, so the study record contains the population, the comparator and the measures.

Public availability and quality of the evidence

A systematic review with meta-analysis published in npj Digital Medicine started from 16,971 records and retained 110 studies evaluating 112 self-guided mental health apps in controlled trials. Of the 81 unique apps identified, 42 were publicly available, around half.

The methodological quality of the studies, measured with the Cochrane risk-of-bias tool, was not associated with an app’s public availability. Nor did effect size distinguish one group from the other: g of 0.33 among available apps against g of 0.45 among unavailable ones, with no significant difference.

From both results it follows that apps with evidence of effect exist and that a share of them cannot be downloaded. Availability in a store follows commercial and maintenance decisions, and it works as a filter bearing no relation to the quality of the study.

A mixed methods study published in JMIR Human Factors in 2026, by Zych and colleagues, points the same way from the opposite end: many publicly available mental health apps do not take precautions such as involving health professionals, referencing the literature or running tests to support their content.

Application of the criteria to the Blind Echo Experience pilot

The criteria admit application to any study, those of the hub itself included. The validation results of the Blind Echo Experience protocol were published with their population, their size — 89 participants and 119 sessions across six months —, their pilot study design and the methodological supervision behind them.

A pilot of that size explores without establishing. Publishing its conditions is what allows it to be placed at the level of evidence it belongs to.

Limitations of the available evidence

The distribution of the types of support across the sector lacks a verified quantification in the sources reviewed: the hierarchy is described without proportion figures for each type of evidence.

The meta-analysis is confined to self-guided mental health apps evaluated in controlled trials. Its results describe associations between availability, methodological quality and effect size; the absence of a significant difference between effect sizes is a comparison between groups and does not establish their equivalence.

The JMIR Human Factors study uses mixed methods and is cited for its qualitative conclusion; the magnitude of the phenomenon it describes remains unquantified in the sources reviewed.

The Blind Echo Experience pilot, with 89 participants and a pilot study design, explores without establishing, and its results do not support a claim of efficacy.

Conclusions

The public availability of an emotional wellbeing app carries no information about the quality of its evidence: the methodological quality of the studies is not associated with availability, and effect size does not differ significantly between available and unavailable apps. Apps with evidence of effect exist. Identifying them is possible from outside the development team through five verifiable criteria — test population, size and duration, comparator, funding and variable measured — because publication standards require each to be declared. A seal or certification attests to what its procedure examines, and evidence of effect requires its own verification. A pilot study, regardless of who signs it, explores without establishing.

The ERL scale

The ERL scale that yeshcube develops responds to a specific gap. Technology Readiness Levels describe a technology’s technical maturity, from basic principle to operation in a working environment, and that analysis leaves out the quality of the evidence, the ethical risks, the experience of use and the fit with context. ERL documents that second maturity across six levels, and no solution is transferred below ERL-3.

The framework is compatible with existing research standards, among them CONSORT for randomized controlled trials, STROBE for observational studies and TIDieR for describing interventions.

Collaborating on the validation of digital solutions

yeshcube develops this line within Allies, its scientific collaboration system, with four partner types and three principles: value for value, traceability and independence. No partner can veto a publication.

Producing the evidence these criteria call for takes designs, samples and time that no single team assembles alone. The work is of interest to research groups in digital health and clinical psychology, to health or education organizations assessing tools before adopting them, and to developers willing to register their protocol before collecting the first data point.

Discover our Allies program →

Fill in our contact form →

References

Frequently asked questions

Do app stores sort by quality of evidence?

No. A systematic review published in npj Digital Medicine found that the methodological quality of studies is not associated with an app being publicly available, and that effect sizes for available and unavailable apps do not differ significantly. The set of what can be downloaded and the set of what has been tested overlap only in part.

Which questions reveal whether an app has evidence behind it?

The population it was tested on and its correspondence with the target population, the number of participants and the duration of the study, the comparator used, the funding and the variable measured. The last distinction carries the most weight: usage counts describe how the app is used, while a change on a validated scale describes its effect.

Does a seal or certification prove that an app works?

It proves what it assesses. A product can meet technical, accessibility and data protection requirements without holding any evidence about its effect. These are different planes of the same evaluation and they answer different questions.

What does the ERL scale add to technology readiness levels?

Technology Readiness Levels describe a technology's technical maturity. The ERL scale documents the maturity of the evidence needed to apply it to people, covering efficacy, safety, usability, acceptance, accessibility and fit with the setting. Nothing is transferred below ERL-3.

Subscribe to our updates

By subscribing, you will receive yeshcube news and content by email. You can unsubscribe at any time. See our privacy policy.

Follow us