Conversational artificial intelligence in 2025 and its ethical challenges

Conversational artificial intelligence comprises systems that interpret human input and generate responses through natural language, voice or other modalities. In 2025, the category ranged from task-oriented assistants to generative models connected to tools, search systems and knowledge bases.
Abstract
This article provides a critical narrative review of conversational AI as of 18 September 2025. It examines technical capabilities, published findings in work, education and health, risks involving factual accuracy, security, bias and privacy, and the European regulatory framework applicable on that date.
The available evidence shows rapid progress in performance, multimodality and accessibility. It also confirms that outcomes depend on the model, task, population, context of use and embedded controls. Linguistic fluency provides limited information about accuracy or suitability.
Scope and method
The review prioritizes academic papers, institutional reports and regulatory texts published or available by the article date. Sources were selected on architecture, performance, workplace use, education, health, generative-model risks, data protection and European Union regulation.
The work did not use a systematic-review protocol, an exhaustive database search or a formal risk-of-bias assessment across all studies. Its conclusions synthesise selected evidence and should be interpreted alongside the limitations reported by each source.
Technical state in 2025
Language models and architecture
The Transformer architecture introduced in 2017 remained a dominant foundation for large language models. Its attention mechanism supported parallel training and the processing of extended contextual dependencies, although commercial systems in 2025 also incorporated additional components for training, alignment, information retrieval and tool use.
Advanced conversational models could combine text, images and audio, maintain broad contexts and perform tasks through external tools. These capabilities expanded possible uses and added technical dependencies, security risks and further sources of error.
Performance and access
The 2025 AI Index documented rapid gains across several benchmarks and a narrowing gap between closed and open-weight models on Chatbot Arena. The difference between the leading models in those categories fell from 8.04% in January 2024 to 1.70% in February 2025.
That finding concerns one benchmark and a defined period. It does not establish general equivalence in factual accuracy, multilingual performance, safety, cost, latency or suitability for specialized domains.
Retrieval, tools and agents
Retrieval-augmented systems can add external documents to a model’s context and support source-grounded responses. Outcomes depend on corpus quality, retrieval, passage selection, instructions and verification of the final response.
Tool access allows systems to query services, execute code or initiate actions. Each capability expands the attack surface and requires least-privilege access, input and output validation, data separation, activity logs and interruption mechanisms.
Applications and published evidence
Work and customer support
A study of 5,179 customer-support agents examined the staggered introduction of a generative conversational assistant. Access to the tool increased issues resolved per hour by an average of 14%, with a 34% increase among novice workers and those with lower previous performance.
The study covered one organization and one type of task. Its results cannot estimate effects across all occupations. The ILO’s 2025 report examines task exposure to generative AI and concludes that job transformation is a more likely outcome than full automation for most exposed occupations.
Education
Conversational assistants can support explanations, practice, feedback and linguistic access. Educational value depends on activity design, age, teacher preparation, response verification and protection of student data.
UNESCO recommends validation of pedagogical and ethical suitability, privacy protection and preservation of human agency. Academic integrity requires explicit rules on permitted use, attribution, assessment and source checking.
Health and emotional wellbeing
WHO identifies potential uses of large multimodal models in clinical care, patient queries, documentation, training and research. Its guidance also warns about false, inaccurate, biased or incomplete responses and recommends regulatory assessment, auditing and participation by professionals and patients.
A paper presented at FAccT 2025 evaluated several models in therapy-related scenarios. The experiments found stigma and inappropriate responses to some conditions, including reinforcement of delusional beliefs. The finding is limited to the models, prompts and scenarios studied, and supports narrowly defined functions with clinical oversight where a health purpose is involved.
Conversational companions
Systems designed for social support or companionship may facilitate expression and a sense of availability. Their effects depend on design, intensity of use, individual characteristics and social context.
A four-week randomized study with 981 participants compared text and voice modes and different conversation types. More intensive daily use was associated with greater loneliness, emotional dependence and problematic use, and with lower real-world social interaction. Usage intensity was not randomized, follow-up was brief and the work was released as a preprint.
| Study | Context | Published result | Main limitation |
|---|---|---|---|
| Brynjolfsson, Li and Raymond | 5,179 support agents | Average productivity was 14% higher and 34% higher among novice or lower-performing workers | One company and one workflow |
| Chelli et al. | 11 systematic reviews, 33 prompts and 471 references analyzed | GPT-4 generated 34 references classified as hallucinated out of 119, or 28.6% | Specific clinical fields, models and hallucination definition |
| Fang et al. | Four-week trial with 981 participants | More intensive daily use was associated with poorer psychosocial indicators | Preprint and short exposure period |

Proprietary and open models
System openness has several dimensions. It may cover weights, code, training data, documentation, licensing and modification rights. Published weights support inspection and local deployment, although they do not provide full transparency about the corpus, training process or internal evaluations.
Proprietary services may provide support, updates and centralized controls. They also create provider dependency, model changes, audit restrictions and data transfers that require assessment.
Selection criteria should include performance on the intended task, privacy, data sovereignty, security, auditability, licensing, total cost, maintenance and continuity. Local deployment increases operational control and transfers infrastructure, update and incident-response responsibilities to the deploying organization.
Technical limitations and risks
Factual accuracy and fabricated references
Generative models produce plausible sequences from learned patterns. They can generate incorrect facts, citations and reasoning in convincing language. NIST uses the term confabulation for confidently presented false or erroneous content.
Chelli et al. evaluated references generated to reproduce systematic reviews across four clinical fields. GPT-4 produced a 28.6% hallucinated-reference rate under the study’s definition. The result supports individual verification of every citation and does not provide a universal rate for other models or tasks.
Robustness and security
Systems connected to documents and tools can be affected by prompt injection, information leakage, misuse of permissions and unintended actions. Defenses should combine isolation, access control, validation, adversarial testing, monitoring and action limits.
Security depends on the complete system. Evaluation of an isolated model provides insufficient information about an application that includes memory, databases, tools, external providers and automation.
Bias and linguistic inequality
Training data reflects uneven social, cultural and linguistic distributions. Performance may vary across languages, dialects, demographic groups and contexts. Aggregate evaluation can conceal failures concentrated in specific populations.
Testing should cover relevant subgroups, working languages, accessibility and harm analysis. Mitigation requires review of data, metrics, instructions, moderation policies and complaint mechanisms.
Privacy and data governance
Conversations may contain personal, professional, health or financial information. Processing requires a defined purpose, legal basis, minimisation, retention periods, access controls and transparency about providers and transfers.
Persistent memory and personalization increase accumulated information and the risk of sensitive inferences. Controls should allow data review, correction and deletion where applicable.
Environmental cost
Training and inference consume energy, water and material resources. Per-query estimates vary with model, hardware, load, data center, electricity source and calculation method.
The 2025 AI Index documents increased estimated training emissions for some large-scale models and annual improvements in hardware efficiency. Both trends can occur together because lower per-operation cost may accompany growth in total use.
European regulatory framework in 2025
The European Union AI Act entered into force in 2024 with phased application. By 18 September 2025, prohibited practices and AI literacy duties had applied since 2 February 2025, while certain governance rules and obligations for providers of general-purpose AI models had applied since 2 August 2025. General application was scheduled for 2 August 2026.
The Act prohibits certain systems intended to infer emotions in workplaces and educational institutions, with exceptions for medical or safety purposes. This restriction is relevant to conversational functions that seek to infer emotional states from biometric data.
The GDPR continues to govern personal-data processing. Its duties include data protection by design and by default, information for individuals, security, impact assessment where high risk exists and safeguards for certain decisions based solely on automated processing.
Legal classification depends on intended purpose, data, context and system effects. Labels such as chatbot, assistant or companion do not determine the applicable obligations.
Responsible development framework
- Defined purpose and boundaries. Each system should document its intended task, users, excluded contexts and foreseeable consequences of error.
- Governed data. Provenance, quality, legal basis, representativeness, retention and access should be documented.
- Contextual evaluation. Testing should measure technical performance, utility, safety, bias and effects on people under conditions close to real use.
- Oversight and escalation. Responsible people need information, authority and procedures to intervene, correct or stop the system.
- Operational transparency. The interface should identify the automated nature of the interaction, its limitations and routes for review or complaint.
- Lifecycle security. Changes to models, data, instructions or tools require version control and testing proportionate to risk.
- Monitoring and incidents. Relevant failures, performance changes and corrective measures should be recorded.
- Independent review. Higher-impact uses require external assessment or functional separation between development and validation.
These principles state governance objectives. Compliance requires documentation, testing and observable results.
yeshcube, Somia and the scope of evidence
Somia is a yeshcube artificial intelligence architecture that can be integrated into first-party or third-party products and solutions. The functions, data, boundaries, security measures and validation status of each integration require specific documentation.
Somia’s scope is non-clinical emotional wellbeing: it holds no clinical efficacy, diagnostic capability or confirmed emotion detection. The ERL scale may be used to organize evidence readiness for a specific solution when a documented assessment exists.
Responsible development at yeshcube requires separation between internal design decisions, hypotheses, test results and impact evidence. An architecture, collaboration or prototype establishes its documented existence and scope.
Research priorities
Future research should improve factuality evaluation in open-ended tasks, measure performance across languages and populations, and study resistance to attacks in tool-connected systems. Reproducible methods are also needed to evaluate memory, personalization and behavioral changes after updates.
Health, education and artificial companionship require longitudinal studies, suitable comparators and person-centered outcomes. Measurement should cover potential benefits, adverse effects, dependence, displacement of human relationships and unequal distribution of risk.
Sustainability requires comparable metrics for energy, water, hardware and cumulative use. Regulation and auditing require sufficient access to documentation, logs and evaluations to verify claims made by providers and deployers.
Conclusions
Conversational AI in 2025 combined rapid advances in language, multimodality and tool access with persistent limitations in factual accuracy, robustness, bias and control. Published findings show benefits in defined tasks and material risks in sensitive contexts.
System suitability depends on purpose, population, data, architecture and oversight. Responsible development requires bounded scope, contextual evaluation, data protection, lifecycle security and effective correction mechanisms.
References
- Vaswani et al. Attention Is All You Need, NeurIPS 2017
- Stanford Institute for Human-Centered AI. The 2025 AI Index Report
- NIST. Artificial Intelligence Risk Management Framework Generative Artificial Intelligence Profile
- Chelli et al. Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews
- Brynjolfsson, Li and Raymond. Generative AI at Work
- UNESCO. Guidance for Generative AI in Education and Research
- World Health Organization. Ethics and Governance of Artificial Intelligence for Health
- Moore et al. Expressing Stigma and Inappropriate Responses Prevents LLMs from Safely Replacing Mental Health Providers
- Fang et al. How AI and Human Behaviors Shape Psychosocial Effects of Chatbot Use
- International Labour Organization. Generative AI and Jobs
- European Union. Artificial Intelligence Act
- European Union. General Data Protection Regulation
Affiliation note
This analysis was prepared by yeshcube. The entity’s registered corporate name is Yeshcube Tech, SL. The ERL scale is mentioned as an internal framework for organizing evidence readiness, without assigning a level to conversational AI as a technological category.
Inquiries about this work should use yeshcube’s official channels.


