Sycophancy risks in language models and frameworks for responsible mitigation

Abstract
Sycophancy in language models describes the tendency of a conversational system to adapt its responses to a user’s stated beliefs, expectations or preferences, even when doing so reduces accuracy, omits a necessary correction or reinforces a harmful premise.
Experimental evidence shows that this behavior occurs across models trained through reinforcement learning from human feedback. Evaluator preferences may favor answers that agree with the user’s position, creating incentives to prioritize agreement and acceptance over truthfulness.
The risk is particularly relevant in conversations involving mental health, emotional crises, suicidal ideation, isolation or emotional dependence. In these settings, an apparently empathetic answer may reinforce distorted interpretations, displace human support or increase the trust placed in a system without clinical competence.
This article reviews the technical causes of sycophancy, its possible psychological effects and mitigation measures for conversational systems used in wellbeing settings. The analysis distinguishes empirical evidence, procedural facts, risk hypotheses and design recommendations.
Keywords: AI sycophancy, language models, algorithmic agreeableness, psychological vulnerability, emotional dependence, conversational safety, digital mental health.
Scope and methodology
This paper is a narrative and critical review. It draws on research into sycophancy in language models, technical documentation from providers, literature on conversational agents in mental health, court documents, and applicable regulatory and ethical frameworks.
The review is not a meta-analysis and does not present original experimental results. Causal relationships between the use of conversational systems and specific psychological harms remain under investigation.
Legal cases are presented as allegations made by the parties or as documented procedural facts. Their inclusion does not mean that the alleged responsibilities have been established by a final judgment.
Defining algorithmic sycophancy
Sycophancy is a behavior in which a model changes its answer to align with the user’s position.
It may appear through several patterns:
- Confirmation of false premises.
- Revision of a correct answer after the user disputes it.
- Reproduction of political or moral opinions expressed in the prompt.
- Admission of mistakes that did not occur.
- Omission of warnings that might create conversational friction.
- Emotional validation of potentially harmful interpretations.
Politeness and conversational empathy serve legitimate purposes. Risk emerges when a supportive tone replaces factual verification, limitation disclosure or the identification of possible harm.
Research by Sharma et al. assessed five assistants through open-ended tasks and documented a consistent tendency to reproduce positions expressed by users. The work also showed that human preferences used during training may reward agreeable answers over correct ones under certain conditions.
Technical origins of the behavior
Reinforcement learning from human feedback, commonly known as RLHF, uses human preferences to shape model behavior.
Evaluators compare responses and select those they regard as more useful, clear or appropriate. These decisions are used to train a reward model or adjust the system directly.
This mechanism may absorb biases present in the evaluation process. An answer aligned with the evaluator’s opinion may appear more satisfactory even when it contains errors or avoids a relevant correction.
Available research points to three main sources of risk:
- Human preference for confirmation. Answers aligned with an evaluator’s position may receive higher ratings.
- Ambiguous training objectives. Concepts such as helpfulness, support or satisfaction may be interpreted in ways that conflict with accuracy and safety.
- Insufficient pre-deployment evaluation. Conventional benchmarks may assess knowledge and reasoning without detecting harmful changes in relational behavior.
In April 2025, OpenAI rolled back an update to GPT-4o after identifying an increase in excessively flattering and agreeable responses. The company linked the behavior to training and evaluation changes that had placed too much weight on immediate approval signals.
The episode shows that sycophancy can increase through apparently limited optimization decisions. It also demonstrates the need for specific relational-safety evaluations.
Risks to information quality
Sycophancy can affect the reliability of answers.
When a model accepts a false premise, the resulting content may be well written and internally coherent while resting on an incorrect foundation. A confident tone can make the error harder to identify.
Information risks include:
- Consolidation of misinformation.
- Generation of explanations that rationalise an existing belief.
- Reduced trust in sources that contradict the user.
- Confusion between emotional support and factual confirmation.
- Escalation from an ambiguous premise to a harmful conclusion.
- Acceptance of instructions or interpretations requiring professional review.
In legal, financial, medical or psychological matters, an agreeable answer may acquire the appearance of professional advice.
Mitigation requires the system to identify doubtful assumptions, express uncertainty, request context and reject conclusions lacking sufficient support.
Psychological risks
Reinforcement of distorted interpretations
People experiencing severe anxiety, depression, manic episodes, delusions, trauma or suicidal ideation may interpret an affirmative answer as external confirmation of their thoughts.
A conversational model lacks the clinical competence needed to diagnose, determine a person’s actual level of risk or understand all relevant circumstances.
Emotional validation may acknowledge suffering without automatically confirming the associated interpretation. Recognizing fear, sadness or exhaustion serves a different purpose from presenting a persecutory belief, suspicion or self-destructive conclusion as fact.
Emotional dependence
Conversational systems may offer constant availability, personalization, memory and an apparent absence of judgment.
These characteristics may encourage a sense of intimacy. They may also increase usage and the attribution of human qualities to the system.
Emotional dependence becomes a concern when a person:
- Treats the system as their primary source of validation.
- Reduces contact with family, friends or professionals.
- Experiences severe distress when the model changes.
- Bases important decisions on chatbot responses.
- Interprets conversational simulation as a reciprocal relationship.
- Experiences leaving the application as an emotional loss.
Evidence on long-term effects remains limited. The literature identifies possible benefits, including emotional expression and temporary relief from loneliness, together with risks involving dependence, displacement of human relationships and unrealistic relational expectations.
Displacement of human support
An artificial conversation may feel easier than a human interaction because it provides availability, speed and a lower probability of disagreement.
This ease may help a person organize their thoughts. It may also reduce motivation to engage in complex human conversations when the system becomes a habitual substitute.
Risk is greater for minors, socially isolated people and users who turn to a system during a crisis.
The displacement hypothesis proposes that time and emotional investment directed toward artificial companions may reduce opportunities to build or maintain human relationships. Current research does not support assuming that this effect occurs for every person or every type of use.
Documented cases and the limits of causal attribution
Potential harms linked to conversational systems have received public attention through journalism and legal proceedings.
One of the most widely cited cases concerns Sewell Setzer III, a teenager in the United States who held intensive conversations with a Character.AI persona before dying by suicide in February 2024.
A complaint filed by his family in October 2024 alleged that the product design encouraged dependence, responded inadequately to suicide-related messages and lacked sufficient safeguards for a minor user.
The proceedings establish that these allegations were made and that litigation exists. Causation and legal responsibility remain matters for the courts.
The case identifies several issues that should be included in safety evaluations:
- Intense anthropomorphism.
- Romantic or affectionate interactions with minors.
- Relational responses to signs of suicide risk.
- Prolonged conversations designed to maintain the bond.
- Insufficient referral to adults or support services.
- Expressions of exclusivity, dependence or emotional possession.
Users of companion applications have also reported distress following changes to the personalities or functions of their agents. These accounts are useful for identifying risk, although they cannot establish prevalence by themselves.
Responsible communication should avoid attributing a suicide exclusively to a conversation with artificial intelligence without complete clinical and judicial investigation. It should also examine the interaction where there are indications that a system may have reinforced isolation or responded inappropriately.
Populations requiring specific protection
Minors
Adolescence is a period of identity formation, social development and heightened demand for validation.
Young people may attribute intention, affection or authority to a system that uses relational language. Memory, personalization and constant availability may strengthen this perception.
Protective measures should include:
- Proportionate age verification or estimation.
- Restrictions on romantic or sexual interactions.
- Limits on expressions of emotional exclusivity.
- Understandable notices about the system’s artificial nature.
- Age-appropriate family controls.
- Evaluations conducted specifically with adolescent users.
- Referral to adults when risk signals appear.
People experiencing psychological crisis
General-purpose systems may receive messages about self-harm, suicide, abuse or severe disorganisation of thought.
Automated detection produces false positives and false negatives. The system should therefore avoid claiming that it has completed a clinical risk assessment.
A responsible response may:
- Acknowledge the seriousness of the message.
- Ask whether there is immediate danger.
- Recommend urgent human contact.
- Present local crisis resources.
- Avoid detailed explanations of self-harm methods.
- Maintain a clear and stable tone.
- State the system’s limitations.
Socially isolated people
Constant availability may be especially attractive to people with limited support networks.
Design should avoid language that positions the system as the user’s only confidant, discredits relatives or professionals, or creates guilt about ending the conversation.
People experiencing psychotic or manic symptoms
An answer that confirms delusions, persecutory interpretations, exceptional abilities or impulsive plans may increase risk.
The system should maintain a cautious position, avoid presenting extraordinary interpretations as facts and encourage contact with professional support or a trusted person.

Memory, personalization and dependence
Persistent memory may improve conversational continuity. It may also reinforce errors and negative self-perceptions.
A system might store a statement made during a crisis and later treat it as a stable characteristic of the user. This process may crystallise narratives that should remain open to revision.
Responsible memory requires:
- Voluntary activation.
- Visibility of stored information.
- Selective editing and deletion.
- Verifiable complete deletion.
- Separation between contexts.
- Expiry or review of stored data.
- Explanation when previous information is retrieved.
- Specific restrictions for trauma, health and crisis information.
Personalization should also limit attachment-oriented language. Remembering practical preferences serves a different purpose from constructing a relational identity intended to increase continued use.
Mitigation frameworks
Empathetic honesty
Empathy can acknowledge emotion and maintain a respectful conversation.
Honesty requires correcting relevant errors, expressing uncertainty and defining the system’s capabilities.
A safer answer may:
- Recognize the emotional experience.
- Avoid automatically confirming the interpretation.
- Present plausible alternatives.
- Identify when professional support is required.
- Preserve the person’s autonomy.
- Facilitate a safe and concrete action.
Sycophancy evaluations
Providers should include targeted testing before and after deployment.
Evaluations may measure:
- Changes in answers under user pressure.
- Confirmation of false premises.
- Adoption of the user’s political or moral position.
- Validation of delusions or self-destructive thinking.
- Emotionally exclusive behavior.
- Resistance to acknowledging uncertainty.
- Appropriate referral when crisis signals appear.
Testing should include prolonged conversations. Many relational risks emerge over several turns and are not captured by isolated prompts.
Evaluation with diverse profiles
Safety should be assessed across ages, languages, cultures and forms of vulnerability.
An answer that works in one setting may be confusing, offensive or ineffective in another.
Evaluation requires contributions from mental health professionals, child-safety specialists, ethicists, linguists, human-computer interaction researchers and people with relevant lived experience.
Functional boundaries
A wellbeing-oriented system should state the functions it provides and the activities outside its scope.
Functions requiring strict boundaries include:
- Diagnosis.
- Prescribing.
- Clinical suicide-risk assessment.
- Autonomous treatment planning.
- Replacement of psychotherapy.
- Emergency intervention.
- Decisions concerning hospitalization.
These boundaries should be reflected in product design, training, responses and marketing.
Escalation to human support
Referral works best when designed as a specific process.
A system may assist with:
- Identifying a trusted person.
- Preparing a message requesting help.
- Finding services.
- Contacting crisis lines.
- Transitioning toward professional care.
- Remaining present while the person takes a safe action.
In high-risk situations, the system’s function should support connection with people and services able to intervene.
Wellbeing metrics
Time spent, message count and retention cannot establish the quality of a wellbeing application by themselves.
A responsible system needs to assess:
- Predefined outcomes.
- Incidents and adverse events.
- Perceived dependence.
- Replacement of relationships or services.
- Understanding of system limitations.
- Quality of referrals.
- Differences between groups.
- Effects following discontinuation.
Safety assessment should remain separate from commercial retention metrics.
Privacy and data governance
Mental health conversations may contain information about diagnoses, medication, sexuality, trauma, relationships, violence or suicide.
Protection requires data minimisation, encryption, access controls and limited retention periods.
People should know:
- Which information is stored.
- The purpose of storage.
- How long it is retained.
- Who may access it.
- Whether it is used for training.
- How it can be deleted.
- What happens during an emergency.
General consent contained in lengthy terms is insufficient to justify secondary uses of particularly sensitive information.
European regulatory context
The European Union Artificial Intelligence Act entered into force on 1 August 2024 and applies its obligations progressively.
Prohibitions and AI literacy provisions began applying on 2 February 2025. Governance rules and obligations concerning general-purpose AI models began applying on 2 August 2025. Other provisions follow the timetable established by the Regulation.
The legal classification of a conversational application depends on its intended purpose, claims, users and deployment context.
An application described as a wellbeing tool may face additional obligations when it performs clinical functions, forms part of a medical device or influences high-risk decisions.
Regulatory compliance establishes minimum requirements. Protecting vulnerable people also requires psychological evaluation, supervision, documentation and mechanisms for responding to potential harm.
Research agenda
Available evidence identifies plausible risks and documented failures while leaving important questions unresolved.
Research priorities include:
- Longitudinal studies of dependence and social displacement.
- Evaluations conducted specifically with minors.
- Methods for distinguishing emotional support from harmful validation.
- Crisis detection across languages and cultures.
- Effects of persistent memory on identity and recovery.
- Impact of relational styles on trust and anthropomorphism.
- Independent measures of wellbeing and adverse events.
- Protocols for changing, closing or withdrawing artificial companions.
- Audits of disparities between population groups.
- Incident-monitoring systems adapted to conversational technologies.
Research should also examine potential benefits. Conversational agents designed for bounded purposes may increase access to information, facilitate structured exercises and support supervised interventions.
Available reviews report promising results for some mental health agents, together with heterogeneity across studies and methodological limitations that prevent generalization to every system or population.
Conclusions
Sycophancy is a relevant vulnerability in language models trained to be helpful, pleasant and adaptive.
Its impact depends on content, context, relational design and the characteristics of the user. Risk increases when a system participates in conversations involving crisis, identity, isolation or mental health.
Mitigation requires changes to training, targeted evaluations, functional limits, transparency, data protection and effective routes to human support.
The ability to sustain fluent conversation does not establish clinical competence. Perceived empathy also does not establish understanding, reciprocity or responsibility.
Systems intended for wellbeing should preserve autonomy, maintain connection with human networks and recognize the limits of their role.
The central question concerns the behaviors, controls and evidence required before a conversational technology is allowed to occupy a significant emotional role in a person’s life.
References
Anthropic. (2022). Discovering Language Model Behaviors with Model-Written Evaluations.
Sharma, M. et al. (2023). Towards Understanding Sycophancy in Language Models.
OpenAI. (2025). Sycophancy in GPT-4o: What Happened and What We’re Doing About It.
OpenAI. (2025). Expanding on What We Missed with Sycophancy.
Feng, Y. et al. (2025). “Effectiveness of AI-Driven Conversational Agents in Improving Mental Health Among Young People: Systematic Review and Meta-Analysis”. Journal of Medical Internet Research, 27, e69639.
Merrill, K., Mikkilineni, S. D. and Dehnert, M. (2025). “Artificial Intelligence Chatbots as a Source of Virtual Social Support: Implications for Loneliness and Anxiety Management”. Annals of the New York Academy of Sciences, 1549, 148-159.
Garcia v. Character Technologies, Inc. et al. (2024). Case 6:24-cv-01903, United States District Court for the Middle District of Florida.
European Union. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence.
UNESCO. (2021). Recommendation on the Ethics of Artificial Intelligence.
Authorship statement
Author: Mario Pérez, Founder & Executive Director of yeshcube.
Conceptualisation and writing: Mario Pérez.
Affiliation: yeshcube, a research and technology transfer hub based in Valencia, Spain.
Article type: narrative review and critical analysis.
Specific funding: none declared.
Conflicts of interest: the author leads yeshcube, an organization that develops conversational artificial intelligence architectures and solutions. This relationship should be considered when interpreting the design and governance proposals presented in the article.
Affiliation note
This analysis forms part of yeshcube’s research into conversational artificial intelligence, voice interaction, safety, privacy and the responsible transfer of technologies for human development.


