Recognition error rate as a health outcome

Abstract
Evaluation of a conversational system treats the recognition error rate as a technical performance metric. Evidence published in 2026 also places it among the outcomes of the intervention. A pilot randomized trial with 48 people with Parkinson’s disease recorded, after eight weeks of using a voice-activated assistant, a significant reduction in positive emotion measured with the Mental Health Continuum-Short Form (β = −1.69; p = 0.028), which did not persist at week 12 (β = −0.96; p = 0.31), and which the authors attribute to technical difficulties and the system’s usability. In post-intervention interviews, participants internalized their failed interaction attempts as a characteristic of their own speech, and described the experience as resembling being ignored by a real person. The distribution of error is not uniform: across eight commercial systems evaluated without prior adaptation, severe dysarthria exceeds a 49 % word error rate in all of them.
Introduction
The performance of a speech recognition system is reported as a word error rate aggregated over an evaluation corpus. The figure describes how often the system transcribes incorrectly, and reducing it is the stated objective of most technical work in the field.
In a product whose purpose is emotional wellbeing, that same event has a second recipient. A person who cannot make themselves understood obtains, alongside an incorrect transcription, a repeated experience of failed interaction. The 2026 literature has begun to measure that second consequence with a controlled design.
What standard evaluation measures
The word error rate compares the system’s transcription against a reference and counts substitutions, insertions, and deletions. Latency measures the time to response. Both are properties of the system, measured over a corpus.
Two things fall outside that measurement. The first is the distribution of error across populations: an aggregate average over mostly typical speech is compatible with unacceptable rates in subgroups. The second is the cumulative effect of failure on the person who experiences it, which belongs to the design of the intervention study rather than to component evaluation.
Failure measured in the person
A two-arm, single-blinded pilot randomized trial published in Parkinsonism and Related Disorders in 2026 assigned 48 people with Parkinson’s disease to an eight-week protocol of using a voice-activated intelligent personal assistant or to usual care. The primary outcome was the 13-item Sense of Coherence scale; secondary outcomes included the Mental Health Continuum-Short Form, the UCLA Three-Item Loneliness Scale, and the System Usability Scale. The primary analysis used generalized estimating equations, with multiple imputation as a sensitivity analysis.
In the emotional well-being domain of the MHC-SF, the intervention group showed a significant reduction in positive emotion at the end of the intervention (β = −1.69; p = 0.028). At the week 12 follow-up the effect was no longer significant (β = −0.96; p = 0.31).
The same trial identified favorable exploratory effect sizes in the meaningfulness domain of the SOC-13 (d = 0.27), in manageability (d = 0.19), and in MHC-SF psychological well-being (d = 0.27). The authors conclude that the study identifies preliminary trends of improvement in meaningfulness and psychological well-being, attribute the decrease in emotional well-being to the reported technical difficulties and the system’s usability, rated as fair, and state that a redesign is required to avoid similar adverse effects.
The seven post-intervention interviews supply the proposed mechanism. Participants internalized their failed interaction attempts as a characteristic of their speech, and by their account the experience resembled being ignored by a real person.
A co-design study published in JMIR Rehabilitation and Assistive Technologies in 2026, with twenty participants including people with Parkinson’s disease, carers, speech and language therapists, design and technology experts, and third-sector staff, had gathered the same complaints from the prior literature: having to repeat themselves to be understood, devices timing out before the person had finished speaking, and being unable to hold a conversation.
Who it fails most
The evaluation without prior adaptation of eight commercial systems on atypical speech, published in December 2025, measured four conventional recognizers (AssemblyAI, Whisper large-v3, Deepgram Nova-3, and Nova-3 Medical) and four based on multimodal models (GPT-4o, GPT-4o Mini, Gemini 2.5 Pro, and Gemini 2.5 Flash). In severe dysarthria, the word error rate exceeds 49 % in all of them. The largest improvement observed across systems, 7.36 percentage points in GPT-4o, occurs on that baseline.
That the figure is a property of the design rather than of the speech is indicated by the performance reachable with adaptation. A single-case study published in Frontiers in Rehabilitation Sciences in 2026 trained a recognizer on 1,120 utterances from one speaker with severe dysarthria and global aphasia across thirteen target words, and obtained 72.65 % accuracy on an independent set of 936 utterances, above the 56.75 % mean of twelve rehabilitation professionals familiar with her. The closed vocabulary and the single-case design confine the result to a proof of concept, and with that caution they establish that the margin exists.
The applicable transparency framework
The transparency obligations of article 50 of the European AI Act apply from August 2, 2026, and the European Commission published its guidelines on July 20, 2026. Among those obligations is informing natural persons who are exposed to emotion recognition or biometric categorization systems. What a conversational wellbeing system must communicate about its operation was covered in the analysis of the Regulation published by the hub.
The obligation governs information about exposure to the system. Recognition performance by speech condition falls outside it, and its declaration today rests on the judgment of whoever publishes the product.
What a system that reports its performance declares
Four elements make up a verifiable declaration.
The error rate broken down by speech condition, with a reference to the population over which each band was measured, rather than an average that pools them.
Behavior on repeated failure: what the system does when it fails to understand for a third consecutive time, and whether that exit route leads to a channel that does not depend on voice.
A consultable record of the dialogue, allowing the person to check what was transcribed and to separate the system’s failure from attribution to their own speech, which is the mechanism described in the trial interviews.
Measurement of the effect on the person, and not only of component performance, when the product is offered with a wellbeing purpose.
Limitations of the available evidence
The Parkinson’s trial is a pilot with 48 participants, designed to estimate feasibility and signal, without statistical power to confirm effects. Its principal finding on emotional well-being is one result among several secondary outcomes, and the design itself warns against reading it as confirmatory.
Attribution of the effect to technical difficulty comes from the authors and rests on interviews with seven participants. It is a mechanistic hypothesis consistent with the data, without experimental contrast isolating it from other explanations.
The trial population has Parkinson’s disease, with speech alteration as a frequent characteristic. Transferring the finding to a general population with typical speech has no basis in these sources.
The evaluation of the eight systems was carried out without speaker adaptation and over dysarthric speech corpora. It describes the starting point of commercial systems rather than performance after personalized training.
Conclusions
In a conversational product with a wellbeing purpose, a recognition failure has a recipient beyond the technical log. The only controlled trial available on the question documents a significant reduction in emotional well-being at the end of eight weeks of use, not sustained over the following four weeks, alongside trends of improvement in other domains, and its authors call for a redesign that avoids similar adverse effects. The distribution of error across speech conditions makes the aggregate figure an incomplete description of the system. Disaggregated declaration of performance and the availability of a route that does not depend on voice are the two consequences the evidence supports.
Open research lines: Noor Program
The Somia architecture classifies before responding through a layered instruction system, and its evaluation incorporates the accessibility dimension of the ERL scale.
Noor Program is the line devoted to conversational voice capabilities in underrepresented languages. The question that defines it, which speech falls outside the reach of a recognition system, is the one these data raise for atypical speech, and its evaluation dimensions include prosody, latency, and stability. The program is open and its work has not begun.
Collaborating on the measurement of recognition error
Noor Program is developed within Allies, the scientific collaboration system of yeshcube, with four partner types and three principles: value for value, traceability, and independence. No partner can veto a publication.
Measuring the error rate by speech condition requires labeled corpora held by clinical, university, and advocacy organizations. The line is relevant to speech therapy and neurology services following people with speech alterations, to research teams working on atypical speech recognition, and to organizations willing to measure the effect of failed interaction on the person who experiences it.
References
- Lau, T. K. and Leung, A. Y. M. “Efficacy of the voice-activated intelligent personal assistant (VIPA) intervention on psychosocial well-being among people with Parkinson’s disease: A pilot randomized controlled trial”. Parkinsonism and Related Disorders, 2026, 149, 108392. Pilot randomized trial, 48 participants, eight weeks.
- JMIR Rehabilitation and Assistive Technologies. “Voice-Assisted Technology for People With Parkinson’s Disease Experiencing Speech and Voice Difficulties: Co-Designing Solutions Using Design Thinking”. 2026. Twenty participants across two co-design workshops.
- Alsayegh, A. and Masood, T. “Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models”. arXiv:2512.17474, December 2025. Eight commercial systems.
- Frontiers in Rehabilitation Sciences. “A speaker-dependent Voice-Input Voice-Output Communication Aid for severe dysarthria and global aphasia”. 2026. Single case, thirteen target words.
- European Commission. Guidelines on transparency obligations for providers and deployers of certain AI systems. July 20, 2026. Article 50, applicable from August 2, 2026.
- World Health Organization. Towards responsible AI for mental health and well-being. March 20, 2026.
Frequently asked questions
Why is the recognition error rate a health outcome?
In a pilot randomized trial with 48 people with Parkinson's disease, eight weeks of using a voice-activated assistant were associated with a significant reduction in emotional well-being measured with the MHC-SF (β = −1.69; p = 0.028), which the authors attribute to technical difficulties and the system's usability. In interviews, participants internalized their failed attempts as a characteristic of their own speech.
Did that decline persist at follow-up?
No. At week 12 the coefficient moved to β = −0.96 with p = 0.31, without statistical significance. The same trial identified exploratory trends of improvement in the meaningfulness domain of the SOC-13 (d = 0.27) and in MHC-SF psychological well-being (d = 0.27).
What does the European AI Act require providers to disclose?
Article 50 imposes transparency obligations, among them informing natural persons exposed to emotion recognition or biometric categorization systems. The European Commission guidelines were published on July 20, 2026, and the article 50 obligations apply from August 2, 2026.
What should a voice system declare about its performance?
The error rate broken down by the speech conditions that degrade it, rather than an aggregate figure. In an evaluation of eight commercial systems on atypical speech, severe dysarthria exceeds a 49 % word error rate in all of them, a fact that a general average conceals.


