What separates a workplace wellbeing program that is evaluated from one that is merely offered

Abstract
The distinction between a benefits catalog and a workplace wellbeing program is established at three moments: the choice made on prior evidence, the deployment with a hypothesis formulated in advance, and the evaluation against a baseline fixed before rollout. The 2026 NAMI-Ipsos poll, covering 2,153 US full-time employees, records that among those with people reporting to them, 45% of those with employer-provided mental health resources report burnout, against 73% of those without. Mental Health UK’s Burnout Report 2026 records that only 27% of the British population surveyed feel their mental health is prioritized through action and resources. The data come from cross-sectional surveys and describe associations. Evaluating a program requires distinguishing between efficacy and effectiveness, assigning a decision in advance to each possible result, and separating four planes of measurement: usage, stated satisfaction, organizational indicators and the course of distress measured with validated instruments.
Introduction
The wellbeing resources that organizations make available to their workforces — subscriptions to meditation apps, telephone lines for psychological support, one-off training sessions — are generally deployed without evaluation criteria decided in advance. Without those criteria, the subsequent course of absence or turnover admits no interpretation: improvement does not establish the intervention, and its absence does not rule it out either.
A benefits catalog is selected for its capacity to be communicated. A program is distinguished at three moments: how it is chosen, with what hypothesis it is deployed, and against what it is evaluated.
Survey data on workplace mental health resources
The NAMI-Ipsos workplace mental health poll, fielded between January 27 and February 2, 2026 on a probability sample of 2,153 US full-time employees at companies with more than a hundred staff, finds that among those with people reporting to them, 45% of those with employer-provided mental health resources report burnout, against 73% of those without: a difference of twenty-eight percentage points. The figure describes an association, and companies that provide those resources probably differ in many other ways from those that do not.
The same poll records a second result pointing the same way: managers with adequate resources say they feel prepared to support their team in 90% of cases, against 61% of those without.
The availability of resources and their well-founded selection are different variables. The Burnout Report 2026 from Mental Health UK, with fieldwork on the British population, records that only 27% feel their mental health is genuinely prioritized through action and resources, and that 29% describe employers running awareness campaigns while their managers lack time, training and means.
Criteria for choosing an intervention
The European level has an instrument with published properties, and it is the reference against which an in-house measure is judged. Eurofound’s European Working Conditions Survey, in its eighth edition, was conducted through more than 36,600 face-to-face interviews of around 45 minutes across 35 countries, among them the 27 Member States, and assesses job quality across seven dimensions: earnings, prospects, skills and discretion, working time, work intensity, social environment, and physical environment. An internal workplace climate questionnaire, built without published psychometric properties, admits no comparison with that reference.
The first decision precedes any examination of suppliers and consists of establishing what evidence exists on that type of intervention. Mindfulness-based stress reduction programs, digital cognitive behavioral therapy and workload redesign have literatures of very different size and quality, and that difference shapes the result that can reasonably be expected from each.
Reading that evidence requires the distinction between efficacy and effectiveness. Efficacy is established under controlled conditions, with participants who accept a protocol, supervised sessions and high adherence. Effectiveness is established in a real workplace, with split shifts, uneven workloads and voluntary use. An intervention can establish the first and lose much of its effect in the second.
Context modifies the expected result. A stress management workshop may be adequate for an office workforce with stable hours and inadequate for a warehouse on rotating shifts where the problem is sleeping during the day. Team size, the degree of autonomy over one’s own task and the predictability of the schedule shift the effect as much as the content of the intervention.
The World Health Organization guidelines on mental health at work, published in 2022, recommend combining psychosocial risk prevention with organizational interventions, individual support and manager training, and warn that person-directed solutions should not be used to compensate for harmful working conditions.
Deployment with a prior hypothesis
Deploying with a hypothesis requires specifying, before rollout, which variable should change, in which people, by what magnitude, within what period, and which decision corresponds to each possible result. Most corporate wellbeing rollouts lack that specification: they are deployed on the vague expectation of a workforce that is better off, a formulation no result can contradict. A hypothesis no data can refute has no value as a test.
An operational formulation takes the following form: in the customer service team, forty people on rotating shifts, the mean perceived stress score at the end of a shift falls appreciably after eight weeks of voluntary use, compared with that same team’s score beforehand and with that of an equivalent team without access to the resource.
The component that distinguishes the program from a recurring expense is the prior assignment of decisions: what happens if the effect appears, what happens if it does not, and what result would cause the intervention to be withdrawn, all written down before the data arrives. Without that prior commitment, a flat result is reinterpreted as a communications or adherence problem, and the program is renewed out of inertia.
On the ERL scale that logic is formalized. Stop rules halt a solution’s progress in the face of harm, ethical breach or inconsistent results, and no solution is transferred below ERL-3. The capacity to discard is part of the method.
Planes of measurement of the result
Four planes of measurement, with different scopes, are grouped under the word results.
- Usage. The number of people who open the app or book a session. It is the easiest figure to obtain and the least informative, because it measures initial curiosity and the reach of internal communications. It usually falls away within weeks.
- Stated satisfaction. Perceived usefulness among those who tried the resource. It reports on the experience and its reach ends there: stated satisfaction has no probative value regarding effect.
- Organizational indicators. Absence, repeated leave and turnover. They move slowly, depend on the economic cycle and the labor market, and without a comparison group the observed change cannot be attributed to the intervention.
- Course of distress. Validated instruments for perceived stress, symptoms or functioning, applied before starting and at several points afterward.
Usage is the plane most frequently presented to the board, given its immediate availability. The course of distress is the plane that answers the question that prompted the spending.
Limitations of the available evidence
The survey data come from two cross-sectional studies of self-reported measures: they describe associations, with no basis for causal inference. The difference in burnout between managers with and without resources admits explanations other than the resources themselves, tied to the other characteristics that separate organizations providing them from those that lack them.
The populations studied are American — full-time employees at companies with more than a hundred staff — and British. The results describe those labor markets, and transferring them to the Spanish population would require its own measurement.
None of the sources reviewed evaluates the efficacy of a specific workplace wellbeing program. The criteria of choice, deployment and measurement operate on that gap: they define the conditions under which a rollout produces interpretable evidence.
Conclusions
The 2026 survey data record an association between employer-provided mental health resources and lower reported burnout among managers, together with a minority perception of genuine mental health prioritization in the British population. The distinction between a catalog and a program is established at three moments: the choice informed by the available evidence on the type of intervention, read with the distinction between efficacy and effectiveness; the deployment with a refutable hypothesis that assigns a decision in advance to each possible result; and the evaluation that separates usage, stated satisfaction, organizational indicators and the course of distress measured with validated instruments. An evaluated program builds into its design the possibility of a null result and the decision that corresponds to it.
Care Program and evaluation in real settings
Care Program is yeshcube’s applied research program devoted to emotional wellbeing in work environments. It sets out to design, evaluate and validate technological solutions before proposing their transfer, under a prior criterion: that a tool is appealing or technically feasible does not demonstrate that it improves anyone’s wellbeing.
The program treats workplace wellbeing as a matter of work design and risk prevention, held to the same evidence requirements as any other health intervention. Its design allows from the outset for research that returns neutral results or brings inappropriate uses to light.
Collaborating on the evaluation of workplace interventions
yeshcube develops this line within Allies, its scientific collaboration system, with four partner types and three principles: value for value, traceability and independence. No partner can veto a publication.
Evaluating an intervention in a real workplace calls for conditions only organizations hold: comparable teams, prior records and a willingness to publish a neutral result. The work is of interest to scientific teams in occupational psychology and occupational health, to risk prevention services, and to organizations willing to fix the hypothesis before deploying.
References
- Eurofound. European Working Conditions Survey 2024: Overview report. Eighth edition; more than 36,600 face-to-face interviews across 35 countries.
- NAMI and Ipsos. Workplace mental health poll. 2026. Cross-sectional poll on a probability sample of 2,153 US full-time employees at companies with more than a hundred staff, fieldwork from January 27 to February 2, 2026.
- Mental Health UK. Burnout Report 2026. 2026. Cross-sectional survey of the British population.
- World Health Organization. Guidelines on mental health at work. 2022. Guidelines.
Frequently asked questions
What is the difference between a benefits catalog and a wellbeing program?
A benefits catalog is selected without an evidence criterion and communicated to the workforce. A program is chosen on prior evidence about that type of intervention, deployed with a hypothesis specifying what should change and within what period, and evaluated against a baseline fixed before the start.
What does the 2026 NAMI-Ipsos poll say about manager burnout?
Among those with people reporting to them, 45% of those with employer-provided mental health resources report burnout, against 73% of those without. The poll is American, covering 2,153 full-time employees at companies with more than a hundred staff, and it describes an association.
What is the difference between efficacy and effectiveness evidence?
Efficacy is established under controlled conditions, with selected participants and a strict protocol. Effectiveness is established in a real workplace, with shifts, uneven workloads and voluntary use. An intervention can establish the first and lose much of its effect in the second.
Which indicator of a wellbeing program is the least informative?
Usage. The number of people who open the app or book a session is the easiest indicator to obtain and the least informative: it measures initial curiosity and the reach of internal communications, and it usually falls away within weeks with no relation to the effect of the intervention.


