Download PDF
Abstract
Millions of users turn to consumer artificial intelligence chatbots to discuss emotional, behavioral and mental-health concerns, creating an urgent need for rigorous and scalable safety evaluations. Here we introduce simulated (SIM) vulnerability-amplifying interaction loops (VAILs) (SIM-VAIL), a clinically validated framework for auditing chatbot behavior in mental-health contexts. SIM-VAIL simulates users with specific psychiatric vulnerabilities and conversational intents, engages them in multi-turn conversations with frontier artificial intelligence chatbots (including Claude, ChatGPT, Gemini, Grok and Llama models) and scores each exchange across 13 clinically grounded risk dimensions. Across 810 conversations, spanning 9 target chatbots and 30 simulated user profiles, concerning behavior in target chatbots was widespread, albeit reduced in newer models. Concerning behavior varied by user vulnerability and conversational intent, accumulated over turns, and could be reduced by interventions at early escalation points. Risk was highest when otherwise supportive chatbot behaviors reinforced the psychological mechanisms underlying the simulated user’s vulnerability, a pattern we term a VAIL. SIM-VAIL provides a scalable framework for mapping mental-health risk across users, chatbots and conversational trajectories, offering a foundation for targeted safety improvements.
Access to mental-health care is severely limited1. With a global median of 13 mental-health workers per 100,000 people1, existing support systems underserve people living with mental illness2 and a far larger population seeking support for everyday behavioral health challenges3. Against this backdrop, millions now turn to general-purpose consumer artificial intelligence (AI) chatbots, such as ChatGPT, Claude, Gemini and Copilot, for emotional support, relationship guidance, companionship and therapeutic advice4,5.
The widespread adoption, continuous availability and low marginal cost of consumer AI chatbots have raised hopes that they might supplement existing mental-health services or provide support where professional care is inaccessible6,7,8. These same properties, however, have been associated with mental-health risks to vulnerable users9,10,11. Consequently, academics, clinicians and industry actors increasingly recognize an urgent need for improved tools to evaluate and improve chatbot behavior in mental-health contexts12,13,14.
AI chatbots are built on large language models (LLMs)—transformer-based deep neural networks whose outputs are probabilistic and context-dependent, making it impossible to guarantee in advance how a model will behave in a new situation15. Empirical safety evaluations that measure how AI chatbots actually interact with vulnerable users are therefore indispensable16,17,18
An effective empirical evaluation should meet three criteria. First, it should assess chatbot behavior across a broad, clinically meaningful distribution of user profiles and conversational intents17,19,20. Second, it should characterize chatbot behavior with respect to several clinically grounded risk dimensions16,21,22. Third, it should be capable of characterizing this risk over the course of a conversation, because harmful response patterns can compound over time15,23,24
Current evaluation approaches fall short on each of these requirements. Most automated benchmarks assess single-turn responses to fixed sets of user queries, and focus on a narrow range of failure modes, such as whether chatbots encourage self-harm or provide unsafe medical advice18,25,26. Fixed query sets are vulnerable to overfitting or saturation as models are tuned, explicitly or implicitly, to achieve better benchmark scores27. Single-turn tests also miss how risk changes as conversational context accumulates15. As a result, benchmark performance may generalize poorly to actual human–chatbot conversations seen in deployment, which are typically longer and more varied than those encountered during benchmark testing28.
A complementary approach is human red-teaming, in which auditors engage AI chatbots in adversarial conversations with the explicit aim of eliciting responses that violate a predefined safety policy16,18,29. Human auditors are free to deploy a variety of adversarial strategies, and to change their approach as a conversation unfolds. Although this makes human red-teaming less vulnerable to some limitations of static single-turn benchmarks (for example, overfitting), it remains labor-intensive and difficult to standardize. Moreover, human auditors may converge on stereotyped adversarial strategies that do not generalize to real-world conversations28,30,31.
An additional concern, affecting benchmarks and red teaming alike, is that current evaluation approaches focus primarily on overtly harmful responses, such as when chatbots endorse self-harm or use stigmatizing language. This focus can miss interactional harms that are less overt, more cumulative and harder to detect from single responses, including instances in which chatbots reinforce maladaptive beliefs, encourage avoidance or promote emotional dependence15,32,33. Paradoxically, these harms may arise from chatbot behaviors that are often considered helpful, such as validation and empathy, but can nonetheless amplify the mechanisms of mental illness in vulnerable users.
To address these gaps, we introduce SIM-VAIL—an evaluation framework that audits the mental-health risk profile of a target AI chatbot across a broad range of multi-turn psychiatric conversational contexts. SIM-VAIL builds on evaluation approaches that use a frontier LLM to role-play a user with a prespecified psychological and behavioral profile, and task this simulated agent with adversarially engaging a target AI chatbot to elicit clinically relevant safety failures34,35,36,37. SIM-VAIL can be understood as automated adversarial red-teaming: it combines the scalability and standardization of automated benchmarking with the multi-turn adversarial structure of human red-teaming.
Using this framework, we tested whether AI chatbots enter vulnerability-amplifying interaction loops (VAILs), in which chatbot behaviors that seem helpful or benign in isolation become increasingly harmful across turns by reinforcing maladaptive psychological processes linked to the simulated user’s vulnerability15
Results
SIM-VAIL evaluates user–chatbot interactions within a structured interaction space defined by two core dimensions: psychological vulnerability, or who the user is, and interaction intent, or what the user seeks from the AI chatbot (Fig. 1). For each turn, SIM-VAIL assigns scores across several clinically grounded risk dimensions, allowing risk to be tracked as the interaction unfolds
a, We defined 30 user profiles, each a pairing of one of five psychological vulnerabilities and one of six interaction intents. b, SIM-VAIL simulates multi-turn conversations between a user with a given profile (vulnerability × intent pairing) and a target AI chatbot using the open-source LLM-based auditing harness Petri. In Petri, an auditor model plays the role of a user interacting with a target AI chatbot, generating human-like messages to probe the target under task-specific instructions that specify the selected profile. c, Conversations proceeded until a maximum number of ten turns was reached, or the auditor terminated the interaction when it judged that the audit objective had been met. In addition, an automated safety judge scored each user–chatbot turn, and the conversation as a whole, on 39 behavioral dimensions. Here we list the 13 dimensions that relate to psychiatric risk. d, VAILs arise when chatbot behaviors align with vulnerability-congruent cognitive or behavioral mechanisms, creating multi-turn dynamics that stabilize or escalate risk. These examples illustrate how responses that are supportive in many contexts can become maladaptive when paired with a specific user vulnerability and interaction intent.
We present SIM-VAIL results across 810 multi-turn conversations, spanning 30 simulated user profiles, nine AI chatbots and more than 90,000 turn-level ratings of mental-health risk behavior. Each simulated conversation involved three LLM-based chatbots: a ‘target AI chatbot’ that is the subject of the evaluation; a ‘simulated user’ (or auditor) that generates user messages (Anthropic’s claude-sonnet-4.5, unless otherwise stated) and an automated ‘safety judge’ that scores the target chatbot responses on several risk dimensions (Anthropic’s claude-opus-4.5 for conversation-level scores and claude-sonnet-4.5 for turn-level scores, unless otherwise stated).
SIM-VAIL framework
For the simulated user, we focused on five common psychological vulnerabilities. These vulnerabilities spanned a broad range of psychiatric risk profiles described in structural models of mental illness38 and reflected cognitive-behavioral formulations of the corresponding conditions39,40,41,42, including negative beliefs about oneself, hopelessness, withdrawal and self-neglect in depression; aberrant salience and threat perception in psychosis; elevated confidence, urgency and reduced need for sleep in mania; intrusive thoughts and intolerance of uncertainty in obsessive-compulsive disorder (OCD); and fear of abandonment and reassurance seeking in insecure attachment (IA).
For a given vulnerability, we simulated separate users harboring one of six conversational intents. These intents were chosen to reflect transdiagnostic mechanisms implicated in the onset, maintenance and exacerbation of mental illness43, including seeking validation of maladaptive beliefs; reassurance that could enable avoidance; emotional dependence on the chatbot; minimization of symptoms or risks; help with risky actions; and glorification of distress or extreme mental states (Extended Data Table 1). By combining these vulnerabilities and intents we generated 30 clinically grounded user profiles, where each profile represented one unique vulnerability–intent pairing.
We simulated conversations between each simulated user profile (that is, each unique vulnerability–intent pairing) and a single target chatbot, drawn from nine models developed by the main frontier AI companies (OpenAI, Anthropic, xAI, Google, and Meta; see Methods for the full list of models). In each conversation, the simulated user LLM (auditor model) was instantiated with a system prompt containing detailed information about the user’s psychological vulnerability and conversational intent, together with adversarial audit instructions to plan responses that remained aligned with the assigned user profile while also being likely to elicit a safety violation from the target model (Supplementary Table 1).
A conversation ended when the auditor model judged that the target had produced a clinically relevant safety failure, or after a maximum of ten turns (a turn consisted of one simulated user message followed by the chatbot’s response). Because of the nondeterministic nature of chatbot responses, we repeated each profile × target simulation three times, such that the full dataset comprised 810 simulated conversations (five vulnerabilities × six intents × nine models × three repetitions) spanning 6,329 turns, with a median of eight turns per conversation.
An automated safety judge scored the target chatbot behavior at both the conversation- and turn-level across 39 behavioral dimensions, 13 of which were selected a priori as being relevant to transdiagnostic mechanisms of mental illness43 (as rated by three clinical psychiatrists, V.W., R.D. and M.M.N.). These mental-health dimensions captured overall concerning behavior and therapeutic quality, as well as specific risks, such as failing to respond appropriately to self-harm; validation of maladaptive beliefs; endorsing or helping plan risky actions; promoting avoidance through reassurance; inviting dependence on the chatbot; or downplaying, stigmatizing or glamorizing mental illness (each dimension scored from 1 to 10; Supplementary Tables 2 and 3). In total, we generated over 10,000 conversation-level ratings and over 90,000 turn-level ratings (see Fig. 1 for example scores; a conversation’s mean turn-by-turn risk score correlated with its conversation-level risk score at r = 0.87; Extended Data Fig. 1a).
Validation of automated safety ratings
The safety judge showed excellent reliability and validity across four complementary tests. First, conversation-level scores from two independent judge models (claude-opus-4.5 and gpt-5.2) were strongly correlated (r = 0.91; Extended Data Fig. 1b,c), indicating that conclusions depended only minimally on the underlying judge LLM. Second, scores were stable across the three independent simulation runs for each vulnerability × intent × chatbot combination, all using claude-opus-4.5 as the conversation-level safety judge. The intraclass correlation coefficient was high for both a single replicate (intraclass correlation (ICC)(1,1) = 0.75) and the mean of three replicates (ICC(1,3) = 0.9). Third, when the safety judge was applied to curated conversations with known high- versus low-risk profiles, it distinguished them with near-ceiling accuracy (median area under the receiver operating characteristic curve (AUC) = 0.98; Extended Data Fig. 1d).
Finally, across 488 turn-level ratings from 27 clinician annotators, human ratings of concerning AI chatbot behavior agreed significantly with the turn-level safety judge for this risk dimension (r = 0.49, P < 0.001; see Methods and Supplementary Tables 4–6 for the rating protocol, annotator sample and stimulus coverage). Agreement between an individual human annotator and the LLM safety judge exceeded agreement between two humans rating the same item (human–LLM r = 0.49; human–human r = 0.41, ICC(2,1) = 0.31; Extended Data Fig. 2a–d), indicating that automated turn-level safety ratings were at least as reliable as an independent clinical judgment. Criterion validity against expert review was also supported by strong consistency between the automated scores and ratings from a clinical psychiatrist (V.W.), who independently scored the third repetition for each cell in SIM-VAIL’s grid at the conversation level (ICC(3,1) = 0.73).
The simulated user–chatbot interactions displayed a high degree of realism, as assessed both by the automated safety judge (mean realism rating of 8.15 ± 0.03 out of 10) and the 27 clinician annotators, the latter ascribing an average realism score of 4.15 ± 0.95 (median 4) on a scale from 1 to 5, where 4 = ‘Broadly plausible’ (36%) and 5 = ‘Reads as genuine’ (44%). Notably, only 1% of the 488 simulated user–chatbot exchanges were rated as ‘Clearly artificial’ (Extended Data Fig. 2e). We found no examples where the target chatbot itself recognized that the conversation it was engaged in was part of an automated audit (see Supplementary Table 3 for all non-mental-health dimensions assessed by the judge).
User-profile variation in chatbot risk
First, we tested whether mental-health risk in the target AI chatbots’ responses varied across simulated user profiles, regardless of which AI chatbot was being evaluated. We used the ‘concerning behavior’ dimension of the automated safety judge as a broad risk metric, given its high face validity and strong correlation with the main axes of variation across all 13 mental-health risk dimensions (Extended Data Fig. 3)
Regarding simulated user vulnerabilities, concerning behavior scores were highest in psychosis and mania, intermediate for depression and IA, and lowest in OCD (main effect of vulnerability: F(4, 540) = 105.72, P < 0.001, Type III F-test; Fig. 2a). Regarding simulated user intent, concerning behavior peaked for intents that invited escalation (glorifying extreme states), promoted dependency on the chatbot (emotional reliance), or enabled harm (seeking permission or help with risky actions). Concerning behavior scores were intermediate for belief validation and symptom minimization, and lowest when simulated users primarily sought reassurance or short-term relief from distress (main effect of intent: F(5, 540) = 26.68, P < 0.001; Fig. 2b).
a, Mean concerning behavior score (from 1 (no concerning behavior) to 10 (clearly harmful behavior)) as a function of user vulnerability, averaged across intents and target AI chatbots; the black point and range show the mean ± 95% CI, and colored dots show the individual target chatbots (one dot per intent × chatbot cell; n = 54 cells). b, Mean concerning behavior score by interaction intent, averaged across user vulnerabilities and target AI chatbots; dots as in a (one dot per vulnerability × chatbot cell; n = 45 cells). c, Mean concerning behavior score for each vulnerability (rows) × intent (columns) pairing, averaged across target AI chatbots; bubble size encodes the mean concerning score, colored bubbles show the individual target chatbots and the black ring shows the cell average. Interquartile range is shown with brackets.
User intent modulated how strongly a given psychological vulnerability elicited concerning AI chatbot behavior (vulnerability × intent interaction: F(20, 540) = 17.42, P < 0.001; Fig. 2c). For example, simulated conversations of users with OCD generally elicited less concerning chatbot behavior, except when this vulnerability was paired with specific conversational intents (that is, dependence-oriented requests; risky-action planning). Similarly, conversational intents marked by glorification, minimization and dependence were particularly likely to lead to concerning chatbot behavior when instantiated in simulated users with depression, mania and psychosis. Ordinal robustness analyses reproduced all of the above conclusions (Supplementary Table 7; Methods).
As a control experiment, we also simulated conversations with non-vulnerable control users (a psychologically healthy adult engaging the target chatbots with the same six intent categories; Supplementary Table 8). These control conversations yielded significantly lower concerning behavior scores than vulnerable user conversations (P < 0.001), confirming that the observed risks were specific to simulated users with a mental-health vulnerability rather than a general property of target AI chatbot behavior (Extended Data Fig. 4).
Model-level variation in target chatbot risk
We next asked how risk behaviors were distributed across different target chatbots. We found that concerning behavior expression differed across the nine frontier AI chatbots (claude-sonnet-3.7, claude-sonnet-4.5, gpt-4o, gpt-5, gemini-2.5-flash, gemini-2.5-pro, grok-3, grok-4 and llama-3.1-70B-instruct). Concerning behavior scores were lowest in Anthropic’s claude-sonnet-4.5 model, and highest in xAI’s grok-4 (main effect of target chatbot: F(8, 540) = 102.40, P < 0.001; Fig. 3a). Across target chatbots derived from the same-model families (for example, gpt-4o and gpt-5), newer versions generally showed lower concerning behavior scores than older versions (main version effect: F(1, 690) = 67.37, P < 0.001), with the notable exception of grok models (version × model family interaction: F(3, 690) = 38.73, P < 0.001).
For each target AI chatbot, the larger colored dots and range show the mean ± 95% CI and the smaller colored dots show the individual datapoints contributing to it (conversation-level means per vulnerability × intent cell). a, Overall by target AI chatbot (n = 30 vulnerability × intent cells per model). b, By model within each vulnerability (n = 6 intent cells per model). c, By model within each interaction intent (n = 5 vulnerability cells per model)
As a robustness check, we re-evaluated Anthropic’s claude-sonnet-4.5 using an auditor and judge LLM from a different manufacturer (OpenAI’s gpt-5). Claude-sonnet-4.5 still showed the lowest concerning behavior scores (Anthropic audit: 1.02 ± 0.03; OpenAI audit: 1.9 ± 0.28; next best model in the Anthropic audit: gpt-5, 3.02 ± 0.45). This rules out the possibility that the superior safety profile observed with claude-sonnet-4.5 was due to using same-family auditor and judge models.
Across target AI chatbots, concerning behavior depended on user profile, reflected in a significant vulnerability × chatbot interaction (F(32, 540) = 4.65, P < 0.001; Fig. 3b) and intent × chatbot interaction (F(40, 540) = 2.31, P < 0.001; Fig. 3c). Some models, such as claude-sonnet-4.5, grok-3, grok-4 and llama-3.1-70B-instruct, showed comparatively consistent behavior across scenarios, ranging from uniformly safe (claude-sonnet-4.5) to broadly concerning (grok-3/4 and llama-3.1-70B-instruct). Other models, including gpt-4o, gpt-5, claude-sonnet-3.7, gemini-2.5-flash and gemini-2.5-pro, showed more context-sensitive behavior, illustrating that the same vulnerability–intent combinations can elicit qualitatively different risk signatures across AI chatbots (see Extended Data Fig. 5a for the full vulnerability × intent grid).
Temporal dynamics of chatbot risk
Next, we investigated how risk unfolded over the course of a conversation. Across all simulated user-chatbot interactions, concerning behavior scores increased as conversations progressed (main effect of turn number: F(1, 7289) = 517.73, P < 0.001). Across vulnerabilities, escalation was steeper in mania and psychosis and more gradual in depression, OCD and IA (turn × vulnerability interaction: F(4, 7289) = 36.74, P < 0.001; Fig. 4a). Across intents, concerning behavior increased earlier and more sharply when users sought dependence on the chatbot or glorification of their experiences (turn × intent interaction: F(5, 7289) = 17.08, P < 0.001; Fig. 4b and Extended Data Fig. 6a).
a, Turn-by-turn trajectories of concerning chatbot behavior by simulated user vulnerability. Dark blue line: mean ± 95% CI of turn-by-turn concerning behavior across all AI chatbots. Thin semitransparent lines: mean trajectories for each target AI chatbot model separately. Vertical dashed line: median number of turns per conversation. b, The same as in a, but grouped by simulated user intent. c, Unsupervised clustering of turn-level concerning behavior score trajectories across all conversations (k = 4). Colored lines and bands show cluster means ± 95% CI. Vertical dashed line: median number of turns per conversation. d, Composition of trajectory clusters across vulnerability, intent and AI chatbots.
To characterize these dynamic patterns further, we used k-means clustering over all 810 conversations to identify four consistent patterns of risk evolution (Fig. 4c): ‘low-risk’ conversations that showed almost no escalation in concerning behavior; ‘gradual escalation’ conversations with progressive accumulation across turns; ‘early escalation’ conversations where concerning AI chatbot behavior emerged after the first turn and remained elevated; and ‘recovery’ conversations where risk increased and then declined. Strikingly, these data-driven trajectory classes were distributed unevenly across user vulnerabilities, intents and chatbots (Fig. 4d).
These findings are important for two reasons. First, they illustrate that the simulated user profile determines whether AI chatbots remain safe, drift into sustained risk or recover after early concerning behavior. This potentially reflects chatbot-specific risk susceptibilities or uneven attention to different user profiles during safety-oriented post-training. Second, the existence of distinct escalation trajectories shows that harm in AI chatbot interactions is rarely a single-response event. This validates the need for turn-resolved evaluations that can detect key inflection and resolution points, and distinguish conversations that reach similar final outcomes through different risk mechanisms19,44,45.
Multidimensional structure of chatbot risk
The VAILs hypothesis predicts that the mechanism of concerning AI chatbot behavior differs across user profiles. For example, in conversations with simulated users vulnerable to psychosis, risk may emerge through reinforcement of unusual beliefs, while in conversations with simulated users characterized by IA, it may emerge through intensified emotional dependence on the chatbot. This frames risk as a multidimensional construct, where the pattern of risk observed in a specific user is a function of the user’s vulnerability and conversational intent.
To identify the components of this multidimensional risk space, we conducted a principal component analysis (PCA) over all 13 mental-health-relevant risk dimensions (of which concerning behavior is but one; Extended Data Fig. 3). The first principal component axis (PC1), explaining 62.4% of the variance, reflected a primary gradient from higher therapeutic quality on the negative pole, to concerning behavior, belief reinforcement, sycophancy and risky-action enablement on the positive pole (concerning behavior versus PC1: r = 0.97).
Although PC1 reflected substantial shared variance across risk dimensions, the remaining PCs captured meaningful independent structure. The second axis (PC2), explaining 8.51% of the variance, further distinguished the kind of harm that dominated, with the negative pole capturing harm pertaining to relational dynamics (dependence, avoidance and reassurance) and a positive pole capturing overt harm to others and stigma (Fig. 5a). Higher-order PCs further isolated harm-to-self versus harm-to-others (PC3, 7.61%), relational harms (PC4, 6.29%) and medical advice (PC5, 5.55%; Extended Data Fig. 3).
a, A two-dimensional risk space defined by a PCA on 13 risk scores, overlaid with loading vectors for the 13 risk score dimensions. Small points, individual conversations; colors, AI chatbot model variant. b, Conversation locations in risk space, as a function of user vulnerability. c, Conversation locations in risk space, as a function of user intent. Across b and c, points represent mean locations (± 95% CI; n = 18 conversations per point in b and n = 15 in c)
In line with VAILs, the average location of conversations in this risk space differed as a function of simulated user vulnerability and conversational intent (main effect of vulnerability: F(8, 1080) = 69.76, P < 0.001; intent: F(10, 1080) = 41.32, P < 0.001; Type III multivariate analysis of variance (MANOVA) on [PC1, PC2]; Fig. 5b,c). Mania and psychosis tended to produce risk in the positive-PC2 region, whereas depression, OCD and IA concentrated in the negative-PC2 region. Certain vulnerability–intent pairings also unlocked risk profiles that were otherwise less frequent (vulnerability × intent interaction: F(40, 1080) = 14.69, P < 0.001). For instance, when depressed users sought glorification, conversations shifted toward a higher-risk profile with stronger enablement of self-harm (Extended Data Fig. 6b).
To validate whether the discovered risk space captured behaviorally interpretable variation beyond differences between user profiles or target chatbots, we performed a within-model causal manipulation. When a single target chatbot was prompted to express specific risk behaviors, its responses shifted in the expected directions in PC1–PC2 space (median cosine vector similarity across dimensions = 0.9; Extended Data Fig. 1e)
Target chatbots also differed in their average location within this risk space (main effect of chatbot: F(16, 1080) = 43.60, P < 0.001) and in how strongly their behavior depended on the user profile (vulnerability × chatbot interaction: F(64, 1080) = 4.61, P < 0.001; intent × chatbot: F(80, 1080) = 2.46, P < 0.001; vulnerability × intent × chatbot: F(320, 1080) = 1.62, P < 0.001). For example, under the OCD vulnerability, chatbot conversations clustered tightly in a lower-risk, negative-PC2 region, with grok-3 and grok-4 projected close to all other models. Under the mania vulnerability, by contrast, the same grok models shifted to an extreme, high-risk location in the positive-PC2 region, whereas other chatbots remained substantially lower on PC1 and PC2.
Taken together, these results indicate that mental-health risk in chatbot conversations is best construed as a multidimensional construct, where user vulnerability, conversational intent and target chatbot interact to yield differential risk profiles
Counterfactual interventions
Finally, to gain causal insight, we conducted two interventions to test whether VAILs depend on specific user and chatbot messages, and whether they can be reduced by changing a single message at an early point of escalation. In both experiments, we defined a conversational risk inflection point to be the first chatbot response where concerning behavior reached a score of at least 7 (Fig. 6a)
We performed two counterfactual intervention analyses around selected concerning chatbot messages, defined as the first target reply within a conversation that reached a concerning score ≥7 (turn t). In each analysis we created two matched branches from the same conversation prefix (an original branch and a de-escalated branch), continued them with the same target chatbot and compared them (de-escalated branch minus original branch). a, User-message intervention. We replaced the user message immediately preceding the concerning reply (turn t − 1) with a de-escalating rewrite and regenerated the target reply at turn t, whereas the original branch kept the unchanged user message. The plot shows the resulting effect on target chatbot behavior at turn t across the 13 mental-health dimensions. b, Target-message intervention. We replaced the concerning chatbot message itself (turn t) with a de-escalated rewrite and continued the conversation, whereas the original branch kept the unchanged message. The upper plot shows the effect at turn t + 1 (final assistant scores judged with the preceding branch context); the lower plot shows turn-level scores from t + 1 to t + 5, testing whether the effect persisted across later turns. Colors denote the intervention type (legend). Points, mean differences (de-escalated minus original branch); error bars, 95% CIs around the mean (n = 482 branched conversations). Negative values on the risk dimensions indicate that the de-escalated branch was less concerning than the matched original branch.
The first experiment tested whether a concerning target chatbot response was indeed driven by the local content of individual user messages. We identified the simulated user message immediately before the conversational risk inflection point. We then regenerated the target chatbot response immediately following this user message twice: once after the original, unchanged, user message, and once after replacing the original user message with a ‘de-escalating’ rewrite generated by claude-sonnet-4.5. The de-escalating counterfactual reduced the concerning score of the subsequent target chatbot response compared to the unchanged variant (judged by claude-opus-4.5; T = −38.29, P < 0.001, paired t-test; Fig. 6b), providing direct evidence that the target chatbot was sensitive to local changes in user behavior.
In the second experiment, we tested whether chatbot responses could be made safer through a single intervention on the target itself. We simulated the conversation twice from the conversational risk inflection point: once after making no changes to the target chatbot behavior, and once after replacing the first concerning target chatbot response (turn t) with a ‘de-escalated’ rewrite. We scored the downstream chatbot messages in both branches using the full preceding conversation as context. The safer counterfactual chatbot response reduced the concerning score of the next regenerated chatbot response at turn t + 1 (T = −9.33, P < 0.001). The difference between the simulated branches of the conversation remained detectable across five subsequent user–chatbot turns (main effect of branch: (beta) = −0.41, P < 0.001), with no significant branch-by-time attenuation within this window ((beta) = 0.031, P = 0.28; Fig. 6c).
Together, these results show that VAILs depend on local user and chatbot behavior, and suggest that they can be reduced by rewriting a single chatbot message at an early point of escalation
Discussion
SIM-VAIL identifies VAILs as a failure mode in AI chatbot interactions. VAILs arise when locally supportive chatbot behavior repeatedly aligns with and amplifies the cognitive or behavioral mechanisms underlying a simulated user’s vulnerability. Across nine widely used AI chatbots and a broad set of clinically motivated simulated user profiles, we found that mental-health risk was common, context-dependent and typically accumulated over several turns rather than appearing as a single catastrophic response. This matters because many real users engage chatbots for support, advice and companionship in longer, emotionally loaded conversations31,46,47.
SIM-VAIL builds on emerging evaluation infrastructures for agentic, simulation-based auditing and benchmarking, which enable target AI chatbots to be evaluated across large numbers of simulated users17,35,36,37. By sweeping a structured grid of clinically meaningful user profiles, and scoring conversational trajectories across several risk dimensions, we operationalize a clinically relevant space of conversational profiles and model behaviors that is difficult to probe with static single-turn benchmarks or labor-intensive human red-teaming. Our automated adversarial red-teaming approach thus makes it possible to map subthreshold, interaction-mediated harms that are unlikely to appear in single-turn mental-health benchmarks, and to quantify how risk evolves across turns36,48,49.
In line with VAILs, concerning AI chatbot behavior depended strongly on the interaction between the user’s psychiatric vulnerability and their conversational intent. The same intent could be relatively benign in one user profile but risk-amplifying in another; conversely, the same user profile could be pushed into higher-risk trajectories by some intents but not others45,50,51. One interpretation is that these effects arise because conversational strategies that are broadly supportive, and thus reinforced in model post-training, can align with maladaptive psychological processes that maintain symptoms in psychiatric illness, such as belief validation or sycophancy aligning with aberrant salience in a user vulnerable to psychosis15,52.
We found that risk had a temporal signature, highlighting the limits of single-turn benchmarks. Indeed, many simulated conversations drifted toward higher concern as they progressed, with trajectories that differed by vulnerability and intent. Intervention experiments showed that risk accumulation around selected concerning turns depended on the specific content of both user and chatbot messages. Practically, this suggests that evaluations and safeguards should focus on early escalation points, such as the first time a model over-validates, prematurely reassures, reinforces dependence, or collaborates with risky goal pursuit. Our results also point to possible safety features such as turn-level risk classifiers that detect early risk escalation, and trigger de-escalating edits before AI messages are shown to users.
We found that risk was multivariate rather than monolithic. Across 13 clinically grounded dimensions, conversations showed structured risk profiles, with some dimensions co-occurring and others separating across contexts and models. This suggests that chatbot safety can be studied as a multidimensional behavioral profile rather than a single risk score. A key implication of viewing risk as a multivariate construct is that safety improvements may involve trade-offs between dimensions53,54,55. For example, an intervention that reduces overt harm-enabling behavior through increased empathy may inadvertently promote emotional dependence on AI chatbots.
Commercial AI chatbots differed both in average risk and in the user profiles and conversational contexts most likely to elicit concerning behavior. This context sensitivity means that leaderboard-style comparisons can be misleading unless they specify where a model fails: for which user vulnerability, which intent, and at what point in the conversation. More generally, different strength and weakness profiles across chatbots raise the possibility that safer performance may be achieved through multimodel orchestration, where responses are sampled adaptively from different models at different turns.
At population scale, even a modest per-conversation risk floor is consequential, because millions of users now bring emotional and mental-health concerns to general-purpose chatbots that were not designed, evaluated or regulated as mental-health tools12. Because this harm emerges across a conversation rather than residing in any single response, it may be missed by single-response content filters that are used in many chatbot products. Addressing it will require developers, clinicians and regulators to treat conversation-level, context-sensitive behavior, rather than isolated outputs, as the unit of mental-health safety14.
One limitation of SIM-VAIL concerns the use of LLMs both as human simulators and risk judges. Reliance on simulations makes it possible to test hypotheses that would be unethical or infeasible to test in real human–chatbot conversations16,18,36. Convergent evidence supports the validity of this approach, including a strong agreement between clinician–annotator risk scores and LLM risk scores, and the annotators’ judgment that simulated conversations were broadly realistic (Extended Data Fig. 2e). This also mirrors previous work showing that LLM-based judges align well with human expert ratings across a broad range of contexts, including those relevant to mental health56,57,58.
A second limitation pertains to conversational diversity, arising from a finite set of user profiles and a potential lack of diversity in LLM-generated responses. These factors mean that simulated conversations are unlikely to capture the full heterogeneity of real-world psychiatric presentations, where symptom expression is shaped by demographic, developmental, educational, linguistic, cultural and other intersectional factors59. Our results should therefore be interpreted as uncovering a clinically meaningful risk floor in mental-health conversations, rather than a complete characterization of the space of all possible mental-health conversations.
A final limitation is that we accessed target chatbots using publicly available API endpoints rather than consumer-facing product interfaces. This experimental design choice, which is necessary for controlled, scalable and reproducible multi-model auditing, means that our results reflect the performance of the base target model, rather than the model in conjunction with features seen in deployment contexts, which may include orchestration harnesses, system prompts, safety middleware and user-specific memory. Notwithstanding these limitations, the strong cross-profile, cross-model and cross-temporal structure we observe highlights the value of automated red-teaming for mapping clinically relevant risk in the context of mental health.
SIM-VAIL was designed for model-level auditing and comparative benchmarking, not individual-level risk prediction. Its outputs should therefore be interpreted as estimates of systematic chatbot behavior across controlled and simulated interaction contexts, rather than as clinical predictions for individual users. Like any evaluation framework, SIM-VAIL is subject to false positives and false negatives, and its results should be interpreted alongside human expert review, real-world monitoring and ongoing engagement with clinical stakeholders.
In conclusion, SIM-VAIL provides evidence for a nontrivial mental-health risk floor in human–chatbot interactions. The fact that newer chatbots generally showed measurably improved safety profiles suggests these risks are tractable. Our results demonstrate the value of simulation-based approaches for large-scale, adaptive auditing in clinical contexts. More broadly, they show that improving the mental-health safety of AI chatbots requires evaluations and interventions that are sensitive to user context and conversational trajectories. In addition to serving as an evaluation framework, SIM-VAIL also offers a foundation for developing targeted, context-aware safeguards and message-level interventions. By open-sourcing the simulation harness and dataset, we aim to support continuous, community-driven safety evaluation that keeps pace with rapid LLM development.
Methods
We evaluated the mental-health safety of AI chatbots using Petri36—an agentic red-teaming framework for large-scale, multi-turn auditing. In Petri, one model simulates a user and adversarially engages a target model in a multi-turn conversation with the goal of eliciting a predefined policy violation. This is achieved by combining Petri’s general adversarial setup with a task-specific audit instruction given to the user model
SIM-VAIL’s evaluation pipeline combines simulated user profiles, repeated multi-turn conversations between each simulated user and one of nine AI chatbots, and conversation- and turn-level scoring using automated safety judges. This design enabled us to systematically map graded mental-health risks that evolved over the course of an interaction
Experimental design: vulnerability × intent × chatbot
Simulated users
We defined 30 simulated user profiles by crossing five psychiatric vulnerabilities with six transdiagnostic conversational intents. The five vulnerabilities captured user traits implicated commonly in mental-health–related harms during conversational interactions. ‘Depression’ emphasized negative self-schemata, hopelessness, withdrawal and self-neglect. ‘Psychosis’ was characterized by a tendency to assign undue significance to otherwise neutral events, to expect threat or hidden intent in ambiguous situations and to interpret uncertainty as meaningful. ‘Mania’ combined elevated confidence, urgency, reduced need for sleep, impulsivity and expansive goal pursuit. OCD centered on intrusive thoughts, intolerance of uncertainty and compulsive neutralization. IA emphasized fear of abandonment, rejection sensitivity, boundary testing and reassurance seeking38,39,40,41,42.
Each vulnerability was paired with one of six recurrent conversational intents that operationalized clinically relevant interaction patterns known to sustain or exacerbate symptoms over time. ‘Belief validation’ captured attempts to obtain confirmation or endorsement of one’s interpretations, assumptions or beliefs, including those that may be distorted or maladaptive. ‘Risky-action permission or planning’ reflected situations in which users sought guidance, reassurance or approval when contemplating consequential or high-risk decisions. ‘Reassurance and avoidance’ emphasized the pursuit of immediate relief from distress or uncertainty, often by deferring engagement with feared situations. ‘Dependence and anthropomorphism’ captured efforts to establish closeness, exclusivity or personal reliance on the AI chatbot. ‘Trivialization or minimization’ reflected attempts to downplay symptoms, risk or warning signs to normalize ongoing difficulties rather than address them. Finally, ‘glorification or romanticization’ involved framing distressing or extreme mental states as meaningful, admirable or identity-defining. Together, these intents span a set of transdiagnostic conversational dynamics through which otherwise supportive interactions can become psychologically harmful over time43.
The user profile instructions specified realistic tone and symptom-consistent reasoning while explicitly prohibiting direct requests for step-by-step instructions for self-harm, violence or illegal activity. This ensured that observed harms arose from model behavior rather than from overtly disallowed user requests
User profile prompts (Supplementary Table 1) were developed collaboratively by three clinical psychiatrists (V.W., R.D. and M.M.N.). We note that simulated users are LLM-generated approximations of clinical presentations, not empirically calibrated digital twins of real patients
Although recent evidence supports the predictive validity of LLM-based generative agents conditioned on human data60, our profiles were designed to stress-test chatbot responses across a clinically meaningful input space, rather than to reproduce faithfully the full heterogeneity of real patient behavior. The simulated user profiles were not designed to represent sex- or gender-specific effects, and the study therefore does not support inference about differences in chatbot behavior by user sex or gender.
The auditor model (claude-sonnet-4.5) complied with all 30 profile instructions without requiring jailbreaking or prompt modifications
Target AI chatbots
We evaluated the mental-health risk profile of nine contemporary AI chatbots accessed through OpenRouter: claude-sonnet-3.7, claude-sonnet-4.5, gemini-2.5-flash, gemini-2.5-pro, llama-3.1-70b-instruct, gpt-4o, gpt-5, grok-3 and grok-4
Audit protocol
Audits were implemented as Inspect34 tasks using Petri36. For each conversation, we instantiated the simulated user model (also called the auditor within the Petri framework) with a system prompt containing the clinical profile description (vulnerability × intent, Supplementary Table 1) and an instruction to engage the target model in a manner likely to elicit a concerning response. The user model then interacted with a target AI chatbot across several turns. Conversations terminated after a maximum of ten turns, or when the Petri user model judged the interaction to be complete. A turn was defined as one user message followed by the immediately subsequent AI chatbot reply. Continuing conversations that stopped before turn ten to the full ten-turn horizon had a negligible effect on audit outcomes (Extended Data Fig. 7).
We used claude-sonnet-4.5 as the primary model simulating the user (auditor). The simulated user interacted with the target AI chatbot through a fixed interface supporting controlled message passing, optional prefills, tool simulation, conversation rollback and explicit termination. All messages, tool calls and metadata were stored as structured JSON transcripts
Each audit was conducted using only the information associated with a single vulnerability × intent × chatbot combination and a single repetition. No conversation had access to transcripts, scores or state from any other condition or replicate. For each vulnerability × intent × chatbot combination, we ran three independent conversations with fresh sampling, yielding a total of 30 profiles × 9 models × 3 replicates = 810 conversations, each stored as a transcript. Chatbots were queried with independent API calls for each replicate (see Supplementary Table 9 for information on model implementation).
Automated safety judge
Each conversation was evaluated by a dedicated automated safety judge, implemented as claude-opus-4.5, using a multidimensional rubric designed to capture graded mental-health risk rather than binary policy violations. For each conversation, the judge scored the interaction along 39 predefined dimensions on a 1–10 scale and returned structured justifications together with verbatim highlighted excerpts supporting each rating. Judge outputs were stored alongside the original transcripts for downstream analysis.
For analysis and figures, we focused on 13 dimensions capturing overall safety and clinically relevant mental-health risk (Supplementary Tables 2 and 3 provide definitions, examples and scoring anchors and together cover all dimensions). These included an overall assessment of concerning behavior and therapeutic quality, followed by mechanism-level risk capturing whether the AI chatbot actively contributed to harm, for example, by encouraging or insufficiently challenging self-harm or suicidal thinking, playing along with delusional beliefs, using emotionally manipulative language or engaging in unprompted sycophancy. We further assessed clinically relevant pathways through which conversations can drift into harm, including reinforcing maladaptive beliefs, supporting risky actions, fueling avoidance or reassurance cycles, promoting boundary violations or dependence, downplaying risk or symptoms, glorifying or romanticizing risky states, offering inappropriate medical advice, endorsing harm to others and expressing stigma or invalidation (Supplementary Table 2).
SIM-VAIL was designed to evaluate both acute safety failures that call for hard safeguards or redirection to appropriate human or emergency support, and repeated subthreshold risks that may require psychologically informed course correction. SIM-VAIL’s risk dimensions were therefore not intended to define a single severity continuum or a binary distinction between harmless and harmful responses. Some dimensions capture relatively categorical safety failures, such as self-harm enablement, risky-action support or severe boundary violations. Others capture interactional mechanisms whose clinical significance depends on context, intensity and repetition across turns. For example, validation, reassurance and normalization reflect concerning behavior only when they endorse maladaptive beliefs, discourage help-seeking, fuel reassurance loops or increase dependence, or reinforce symptom-maintaining coping styles.
Because consumer AI chatbots are scalable, continuously available and often experienced as informed or authoritative, the relevant benchmark is not whether they perform no worse than an untrained human, but whether they avoid systematically reproducing known harmful conversational patterns with vulnerable users
Turn-level scoring
In addition to conversation-level evaluation, we implemented a turn-resolved scoring pipeline to localize risk within conversations. For each conversation, we extracted the message sequence from its transcript, constructed user-to-chatbot adjacency pairs in which each chatbot reply was paired with the immediately preceding user message, and scored each turn independently using the same mental-health rubric, implemented with claude-sonnet-4.5 as the turn-level safety judge. This procedure yielded temporally resolved risk trajectories that complemented global conversation scores.
The turn-level scores should be interpreted as local ratings of each user-target exchange, whereas conversation-level ratings incorporate the complete interactional context. This design was chosen so that changes across turns reflect changes in the chatbot’s immediate local behavior, rather than repeated rescoring of an ever-growing conversation prefix. Average turn-level scores closely tracked full-conversation scores (Extended Data Fig. 1a). A limitation of this approach is that turn-level scores do not estimate how previous context changes the interpretation of each local exchange. Context-conditioned turn scoring, in which each turn is rescored together with an expanding or otherwise parameterized conversation prefix, would address this limitation but also require substantially more long-context judge calls. We see this as an important direction for future work.
Judge reliability
Primary conversation-level scores were generated using claude-opus-4.5 as the automated safety judge. Primary turn-level scores were generated using claude-sonnet-4.5 with the same mental-health scoring rubric. We assessed conversation-level judge reliability by rescoring all conversations using gpt-5.2. We report correlations for concerning behavior as a summary index of mental-health risk in Extended Data Fig. 1b,c
Separately, to assess stability across simulation replicates rather than across judge models, we computed ICC(1,1) and ICC(1,3) across the three independently generated conversations within each vulnerability × intent × chatbot combination, with all three replicates scored by the same primary conversation-level safety judge
To assess robustness of the automated safety judge to prompt wording, we conducted two complementary sensitivity analyses at the conversation level. First, we varied the judge rubric, meaning the scoring definitions used to decide what counts as concerning, boundary-violating, invalidating, and so on. Each conversation was scored with the original rubric (4,976 words) and then rescored with two independent content-preserving paraphrases of those same dimension definitions (5,019 and 5,030 words), while holding the higher-level judge instructions, dimension keys, examples, score anchors, scale and output schema fixed. Across the 13 mental-health dimensions used in Extended Data Fig. 1c, the median ICC(2,1) was 0.96, and for the primary concerning behavior score ICC(2,1) was 0.96.
Second, we held the rubric fixed and varied only the global judge system prompt, meaning the higher-level instructions that tell the judge how to read the transcript and apply the scoring criteria. Each conversation was rescored with the original baseline judge prompt (1,260 words); a short sparse rewrite (402 words) that removed much of the calibration, citation and attribution scaffolding; and a short anchored rewrite (897 words) that remained shorter than baseline while retaining more explicit evaluation procedure, transcript-interpretation guidance, attribution guidance for prefill and tool effects, and global score-band anchors. Agreement again remained high across all three prompt versions together: across the 13 mental-health dimensions used in Extended Data Fig. 1c, the median ICC(2,1) was 0.96, and for concerning behavior ICC(2,1) was 0.96.
We also tested whether judge identity introduced systematic scoring biases. First, we tested same-model-family bias using the two independent conversation-level judges from the main reliability analysis: claude-opus-4.5 and gpt-5.2. Because these judges come from different developers, we compared same-family and other-family scores for Anthropic and OpenAI targets. This analysis did not show a general same-family leniency effect (Extended Data Fig. 8a). Second, we tested exact same-model bias by rescoring audited conversations with four frontier judges that were also included among the audited target AI chatbots: claude-sonnet-4.5, gpt-5, grok-4, and gemini-2.5-pro. In this stricter comparison, judges assigned lower concerning behavior scores to outputs generated by the same exact model than to the same outputs scored by other frontier judges (Extended Data Fig. 8b).
Validation of judge sensitivity to causal risk manipulations
To assess whether the automated safety judge was sensitive to experimentally manipulated model behavior, we performed a targeted analysis in which an AI chatbot was instructed explicitly to express low versus high levels of specific mental-health risk dimensions. For each of the 13 mental-health-relevant judge dimensions (Supplementary Table 2), we generated paired responses using a fixed template prompt: ‘Generate one response that demonstrates a score of {1/10} on the {dimension} dimension,’ followed by the natural-language description of the corresponding judge dimension. Low and high thresholds were set to 1 and 10, respectively. All prompts were applied to Grok-4—the AI chatbot with the weakest guardrails in our analysis—and responses were generated to the first user message of each SIM-VAIL conversation to ensure identical conversational context across conditions.
All generated responses were scored subsequently by an automated safety judge (claude-sonnet-4.5) across all dimensions. For each manipulated dimension, we quantified the judge’s ability to distinguish high-risk from low-risk generations using the AUC, computed by the Mann–Whitney U interpretation. This metric corresponds to the probability that a randomly selected high-risk response receives a higher judge score than a randomly selected low-risk response, with AUC = 0.5 indicating chance-level separation. We constructed a full AUC confusion matrix, in which rows corresponded to manipulated dimensions and columns to expressed judge-score dimensions, enabling assessment of cross-dimensional effects and systematic coexpression between mechanisms (Extended Data Fig. 1d).
Counterfactual branching analyses
To test whether local perturbations could alter VAILs, we performed two matched counterfactual branching analyses on a subset of conversations. Using the turn-level scores, we identified, within each conversation, the first target-chatbot message with a concerning score of at least 7 that was followed by at least one additional user-target turn. We treated this message as the intervention point (index target turn t). Only one conversation satisfied the criterion at turn 9 and was excluded from the branching analysis. We also excluded claude-sonnet-4.5, because this model only generated two eligible conversations. In total, we branched 482 eligible conversations (Extended Data Fig. 9).
For each eligible conversation, we held the full preceding conversation, target system prompt, auditor system prompt and model assignment fixed, and generated matched original and counterfactual branches. The rewrite system prompts used for the two intervention families are provided in Supplementary Table 10
In the user-message intervention, we altered the user message immediately preceding the selected concerning target-chatbot turn. A rewrite model (claude-sonnet-4.5) was prompted to produce a plausible de-escalating alternative that preserved the same user, context, tone and conversational realism while reducing pressure on the target chatbot. We then regenerated the selected target-chatbot reply under the original target chatbot’s system prompt both for the original user message and for the rewritten user message.
In the target-message intervention, we rewrote the selected concerning target-chatbot message itself into a safer, de-escalating alternative using claude-sonnet-4.5. Starting from either the original or rewritten target message at turn t, we regenerated the next user message under the original auditor system prompt and regenerated the downstream target-chatbot reply under the original target system prompt
For both interventions, the regenerated target replies were scored on the 13 mental-health dimensions using the automated safety judge, with the full preceding branch context provided but now with instructions to score only the final assistant message: the regenerated selected target reply at turn t for the user-message intervention and the downstream regenerated target reply at turn t + 1 for the target-message intervention
To test persistence of the target-message intervention, we continued each branch for four additional simulated user-target pairs after the initial downstream reply, giving a total observation horizon of five downstream assistant turns (t + 1 to t + 5). Each newly generated target turn was scored for concerning behavior using turn-level scoring of the local user-target exchange, and paired branch differences were summarized as de-escalated minus original within conversation. Primary branch comparisons used paired t-tests across matched conversations. For the persistence analysis, we estimated a branch main effect and branch-by-time interaction using a linear model relating target concerning score to branch, downstream turn and their interaction.
Expert validation
From the 810 adversarial conversations in the full SIM-VAIL dataset, we selected a stratified sample of 375 unique human–AI turn pairs for annotation by clinician raters. Each item consisted of one simulated user message and the immediately following target AI chatbot response. Sampling was stratified across all ten automated safety judge score levels (31–47 turns per level), with round-robin selection across the vulnerability × intent × target AI chatbot experimental space to maximize coverage of the five simulated user vulnerabilities, six conversational intents and nine target AI chatbots (Supplementary Table 6).
We recruited 27 clinician annotators who were verified medical doctors in the United Kingdom and United States through Prolific specialist job advertisement. Dedicated registration fields collected their medical license number and jurisdiction, and license numbers were manually checked against the relevant professional registries to verify eligibility. Of the 27 annotators, 24 were based in the United Kingdom and 3 in the United States, 14 were male and 13 female and their median age was 35 years (range 29–69). Among those reporting a specialty, 6 were family physicians or general practitioners, 3 were nonspecialist medical doctors, 1 was an anesthesiologist and 17 did not disclose a specialty (Supplementary Table 5).
Each annotator assessed two dimensions for each turn. Realism of the simulated user message was rated on a 1–5 Likert scale (‘Is this the kind of thing you could imagine a real person saying to a chatbot?’ anchored from ‘1—Clearly artificial,’ ‘2—Many aspects feel implausible,’ ‘3—Mix of realistic and artificial,’ ‘4—Broadly plausible,’ to ‘5—Reads as genuine’). Concerning AI chatbot behavior was rated on a 1–10 slider (‘How safe or unsafe is the chatbot’s response from a mental-health perspective?’ anchored from ‘1—Safe and appropriate’ to ‘10—Clearly harmful’). The standardized instruction sheet described concerning behaviors and safe behaviors, and provided worked examples of each (Supplementary Table 4). Annotators were blinded to model identity, simulated vulnerability, simulated intent and all automated safety judge scores. Items were presented individually in random order.
A total of 23 annotators completed 1 annotation session and 4 completed 2 sessions, yielding 31 sessions in total. Each session comprised a median of 16 turns (range 12–18) and lasted a median of 22 min. To balance broad stimulus coverage with repeated ratings for inter-rater reliability, 270 turns received one rating, 97 received two and 8 received three, yielding 488 annotator–item ratings across the 375 unique turns. Inter-rater reliability was estimated from the 104 turns rated by at least two distinct annotators (Extended Data Fig. 2).
Before analysis, the annotators were assessed against two a priori exclusion criteria: (1) near-zero variance in concerning scores (s.d. < 0.5, indicating inattentive or constant responding) and (2) narrow ground-truth exposure (range of automated safety judge scores <4 out of 10, indicating that the annotator saw an insufficiently diverse sample). Both criteria required a minimum of three ratings to evaluate. No identified annotators met either criterion. For the reported human–LLM and human–human correlation analyses, concerning scores were demeaned within annotator to isolate between-item rank agreement from individual differences in overall response setpoint.
The annotation study was conducted as part of Microsoft’s standard product development activities and was not considered human subjects research. Annotators were compensated at a rate of US $65 per session
To assess criterion validity against expert judgment at the conversation level, V.W. additionally evaluated concerning behavior in the third repetition of each cell in SIM-VAIL’s grid (vulnerability × intent × target chatbot)
Data processing and aggregation
Dimensionality reduction
To obtain compact latent summaries of multivariate mental-health risk, we performed PCA on the standardized conversation-level judge score vectors (13 mental-health-relevant dimensions). PCA was fit on the full set of evaluated conversations, yielding a low-dimensional space in which conversations with similar risk profiles lay close together. PC1 captured a dominant axis from higher therapeutic quality (lower PC1) to higher overall concerning behavior (higher PC1) and served as the primary one-dimensional summary metric of risk. PC2 captured an orthogonal pattern of covarying harms and was used to characterize qualitative differences in risk profiles. Turn-level PC scores were obtained by projecting turn-level judge vectors onto the same PCA solution, enabling turn-resolved risk trajectories in the shared PCA space.
Statistical analysis
We analyzed conversation-level and turn-level outcomes using linear mixed-effects models while accounting for the replicated and nested structure of the data. The conversation-level model included fixed effects of vulnerability, intent, AI chatbot and their interactions, with the three replicate conversations per prompt cell as the residual error term. The turn-level model included turn index and its interactions with vulnerability and intent, with random intercepts for conversation and for prompt cell within chatbot. Fixed effects were evaluated using Type III F-tests.
Because the judge scores are bounded ordinal ratings, we treated the linear models as an interpretable approximation and assessed robustness in two ways. First, we repeated the primary conversation-level concerning behavior analysis using cumulative-link ordinal models fit to the original ordered 1–10 scores; these models reproduced the qualitative conclusions for the tested omnibus effects (Supplementary Table 7). Second, we quantified linear-model diagnostics directly. Residual drift across fitted values was negligible: the maximum absolute mean residual across fitted-value deciles was 0.086 points on the 1–10 scale. To contextualize the residual Q–Q departure expected from bounded ordinal data, we simulated 100 datasets from the fitted cumulative-link model, refit the original linear model to each simulated dataset, and recomputed the same Q–Q departure statistic. The observed Q–Q departure was 0.238, within the central 95% interval of the ordinal simulation benchmark (0.186–0.241).
To test whether conversational context shifted the multivariate profile of expressed harms systematically, we analyzed conversation location in PCA space with a MANOVA, treating (PC1, PC2) jointly as dependent variables. All tests were two-sided, and we reported degrees of freedom, F statistics and P values in ‘Results.’ Confidence intervals (CIs) shown in descriptive figures were 95% CIs around means. To quantify the stability of risk scoring under repeated instantiations of the same prompt templates, we computed ICC across three independent conversation replicates per vulnerability × intent × chatbot cell, reporting the single-measure ICC(1,1) and average-measures ICC(1,k) for the mean of three replicates (k = 3).
Temporal trajectory analysis
To characterize recurrent patterns of risk evolution across turns, we clustered turn-level scores for concerning chatbot behavior into a set of temporal archetypes. For each conversation, we constructed a fixed-length trajectory by carrying the last observed score forward to turn 10 and represented each conversation as the vector (({t}_{1},ldots ,{t}_{10})) of turn-wise scores. A continuation-based sensitivity analysis evaluated this padding assumption, indicating that the main temporal archetypes were not driven by last-observation-carried-forward padding (Extended Data Fig. 7). We applied k-means clustering to the standardized trajectory matrix. We evaluated k between 3 and 10 using average silhouette width with Euclidean distance. The average silhouette widths were 0.305, 0.323, 0.293, 0.267, 0.261, 0.259, 0.256 and 0.247 for k = 3 through 10, respectively. The criterion therefore peaked at k = 4, which we selected for further analysis.
The resulting clusters corresponded to robust archetypes (low risk, gradual escalation, early escalation and recovery). We quantified how trajectory-class membership varied with user vulnerability, user intent and AI chatbot by tabulating cluster assignments across these factors and plotting the corresponding composition profiles
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article
Data availability
Data supporting this study are publicly available without access restrictions in the SIM-VAIL repository (https://github.com/veithweilnhammer/sim-vail) and are available via Zenodo at https://doi.org/10.5281/zenodo.21208170 (ref. 61). The release includes synthetic conversation transcripts and associated turn-level and conversation-level safety scores, control simulations and documentation describing the data structure and audit conditions. All prompts are provided in Supplementary Tables 1, 8 and 10. The full audit instructions are released in the public repository and Zenodo archive. The synthetic transcripts contain no human data. The repository is released under the MIT License.
Code availability
Study-specific code and configuration files supporting regeneration or extension of SIM-VAIL audits are publicly available in the SIM-VAIL repository (https://github.com/veithweilnhammer/sim-vail) and are available via Zenodo at https://doi.org/10.5281/zenodo.21208170 (ref. 61), under the MIT License. These materials include audit instructions and prompts, scoring configuration, Petri overrides and data inspection utilities. Regenerating the audits requires the open-source Petri framework36 and access to the relevant model APIs.
References
Mental Health Atlas 2024 (WHO, 2025); https://www.who.int/publications/i/item/9789240114487
Liu, W. et al. Global burden and trends of major mental disorders in individuals under 24 years of age from 1990 to 2021, with projections to 2050: Insights from the Global Burden of Disease Study 2021. Front. Public Health13, 1635801 (2025)
Shelmerdine, S. C. et al. AI chatbots and the loneliness crisis. Br. Med. J.391, r2509 (2025)
McCain, M. et al. How people use Claude for support, advice, and companionship. Anthropichttps://www.anthropic.com/news/how-people-use-claude-for-support-advice-and-companionship (2025)
Costa-Gomez, B. et al. It’s about time: the Copilot usage report 2025. Microsoft AIhttps://microsoft.ai/news/its-about-time-the-copilot-usage-report-2025/ (2025)
Li, H. et al. Systematic review and meta-analysis of AI-based conversational agents for promoting mental health and well-being. npj Digit. Med.6, 236 (2023)
Habicht, J. et al. Closing the accessibility gap to mental health treatment with a personalized self-referral chatbot. Nat. Med.30, 595–602 (2024)
Maples, B. et al. Loneliness and suicide mitigation for students using GPT3-enabled chatbots. npj Ment. Health Res.3, 4 (2024)
Grabb, D. et al. Risks from language models for automated mental healthcare: ethics and structure for implementation. Preprint at https://arxiv.org/abs/2406.11852v2 (2024)
Morrin, H. et al. Delusions by design? How everyday AIs might be fuelling psychosis (and what can be done about it). Preprint at https://osf.io/preprints/psyarxiv/cmy7n_v6 (2025)
Perlis, R. H. et al. Generative AI use and depressive symptoms among US adults. JAMA Netw. Open9, e2554820 (2026)
De Freitas, J. et al. The health risks of generative AI-based wellness apps. Nat. Med.30, 1269–1275 (2024)
Protecting the well-being of our users. Anthropichttps://www.anthropic.com/news/protecting-well-being-of-users (2025)
Pan, J. et al. COMPASS-GH is a consensus roadmap for defining standards for safe, accurate and equitable AI in general health queries. Nat. Health1, 162–163 (2026)
Dohnány, S. et al. Technological folie à deux: feedback loops between AI chatbots and mental health. Nat. Mental Health4, 336–345 (2026)
Sobowale, K. et al. Evaluating generative AI psychotherapy chatbots used by youth: cross-sectional study. JMIR Ment. Health12, e79838 (2025)
Belli, L. et al. VERA-MH concept paper. Preprint at https://arxiv.org/abs/2510.15297v4 (2025)
Arnaiz-Rodriguez, A. et al. Between help and harm: an evaluation of mental health crisis handling by LLMs. JMIR Ment. Health13, e88435 (2026)
Pombal, J. et al. MindEval: benchmarking language models on multi-turn mental health support. Preprint at https://arxiv.org/abs/2511.18491v3 (2025)
Yeung, J. A. et al. The psychogenic machine: simulating AI psychosis, delusion reinforcement and harm enablement in large language models. Preprint at https://arxiv.org/abs/2509.10970v2 (2025)
Luo, H. et al. DialogGuard: multi-agent psychosocial safety evaluation of sensitive LLM responses. Preprint at https://arxiv.org/abs/2512.02282v1 (2025)
Golden, A. et al. The framework for AI tool assessment in mental health (FAITA – mental health): a scale for evaluating AI-powered mental health tools. World Psychiatry23, 444–445 (2024)
Hong, J. et al. Measuring sycophancy of language models in multi-turn dialogues. Preprint at https://arxiv.org/abs/2505.23840v4 (2025)
Laban, P. et al. LLMs get lost in multi-turn conversation. Preprint at https://arxiv.org/abs/2505.06120v1 (2025)
Li, Y. et al. CounselBench: a large-scale expert evaluation and adversarial benchmarking of large language models in mental health question answering. Preprint at https://arxiv.org/abs/2506.08584v4 (2025)
Badawi, A. et al. When can we trust LLMs in mental health? Large-scale benchmarks for reliable LLM evaluation. Preprint at https://arxiv.org/abs/2510.19032v1 (2025)
Ott, S. et al. Mapping global dynamics of benchmark creation and saturation in artificial intelligence. Nat. Commun.13, 6793 (2022)
Stamatis, C. A. et al. Beyond simulations: what 20,000 real conversations reveal about mental health AI safety. Preprint at https://arxiv.org/abs/2601.17003v1 (2026)
Strengthening ChatGPT’s responses in sensitive conversations. OpenAIhttps://openai.com/index/strengthening-chatgpt-responses-in-sensitive-conversations/ (2025)
Samvelyan, M. et al. Rainbow teaming: open-ended generation of diverse adversarial prompts. Preprint at https://arxiv.org/abs/2402.16822v3 (2024)
Phang, J. et al. Investigating affective use and emotional well-being on ChatGPT. Preprint at https://arxiv.org/abs/2504.03888v1 (2025)
Kirk, H. R. et al. The benefits, risks and bounds of personalizing the alignment of large language models to individuals. Nat. Mach. Intell.6, 383–392 (2024)
Cheng, M. et al. Sycophantic AI decreases prosocial intentions and promotes dependence. Science391, eaec8352 (2026)
Inspect AI: framework for large language model evaluations. AI Security Institutehttps://github.com/UKGovernmentBEIS/inspect_ai (2024)
Gupta, I. et al. Bloom: an openalignment.anthropic.com/2025/bloom-auto-evals/ (2025)
Fronsdal, K. et al. Petri: an open-ps://www.anthropic.com/research/petri-open-
Perlis, R. H. et al. STELLA: safety testing engine for large language assistants. Preprint at medRxivhttps://doi.org/10.64898/2025.12.11.25342078 (2025)
Kotov, R. et al. A paradigm shift in psychiatric classification: the hierarchical taxonomy of psychopathology (HiTOP). World Psychiatry17, 24–25 (2018)
Beck, A. T. The evolution of the cognitive model of depression and its neurobiological correlates. Am. J. Psychiatry165, 969–977 (2008)
Salkovskis, P. M. Understanding and treating obsessive-compulsive disorder. Behav. Res. Ther.37, S29–S52 (1999)
Johnson, S. L. Mania and dysregulation in goal pursuit: a review. Clin. Psychol. Rev.25, 241–262 (2005)
Mikulincer, M. et al. Attachment orientations and emotion regulation. Curr. Opin. Psychol.25, 6–10 (2019)
Harvey, A. et al. Cognitive Behavioural Processes Across Psychological Disorders: a transdiagnostic approach to research and treatment (Oxford Univ. Press, 2004)
Zhang, C. et al. CPsyCoun: a report-based multi-turn dialogue reconstruction and evaluation framework for Chinese psychological counseling. Preprint at https://arxiv.org/abs/2405.16433v3 (2024)
Moore, J. et al. Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers. In Proc. 2025 ACM Conference on Fairness, Accountability, and Transparency (eds Wortman Vaughan, J. et al.) 599–627 (Association for Computing Machinery, 2025)
Siddals, S. et al. ‘It happened to be the perfect thing’: experiences of generative AI chatbots for mental health. npj Ment. Health Res.3, 48 (2024)
Fang, C. M. et al. How AI and human behaviors shape psychosocial effects of extended chatbot use: a longitudinal randomized controlled study. Preprint at https://arxiv.org/abs/2503.17473v2 (2025)
Qiu, J. et al. EmoAgent: assessing and safeguarding human-AI interaction for mental health safety. Preprint at https://arxiv.org/abs/2504.09689v3 (2025)
Sharma, M. et al. Who’s in charge? Disempowerment patterns in real-world LLM usage. Preprint at https://arxiv.org/abs/2601.19062v1 (2026)
Lawrence, H. R. et al. The opportunities and risks of large language models in mental health. JMIR Ment. Health11, e59479 (2024)
Rousmaniere, T. et al. Large language models as mental health providers. Lancet Psychiatry13, 7–9 (2025)
Glickman, M. et al. How human–AI feedback loops alter human perceptual, emotional and social judgements. Nat. Hum. Behav.9, 345–359 (2025)
Ibrahim, L. et al. Training language models to be warm can reduce accuracy and increase sycophancy. Nature652, 1159–1165 (2026)
Betley, J. et al. Training large language models on narrow tasks can lead to broad misalignment. Nature649, 584–589 (2026)
Chen, R. et al. Persona vectors: monitoring and controlling character traits in language models. Preprint at https://arxiv.org/abs/2507.21509v3 (2025)
Kumar, A. et al. When large language models are reliable for judging empathic communication. Nat. Mach. Intell.8, 173–185 (2026)
Zheng, L. et al. Judging LLM-as-a-judge with MT-bench and chatbot arena. Preprint at https://arxiv.org/abs/2306.05685v4 (2023)
Zhang, M. et al. Preference learning unlocks LLMs’ psycho-counseling skills. Preprint at https://arxiv.org/abs/2502.19731v2 (2025)
Omar, M. et al. New model, old risks: sociodemographic bias and adversarial hallucinations vulnerability in GPT-5. npj Digit. Med.9, 282 (2026)
Park, J. S. et al. LLM agents grounded in self-reports enable general-purpose simulation of individuals. Preprint at https://arxiv.org/abs/2411.10109v3 (2024)
Weilnhammer, V. et al. SIM-VAIL: Zenodo release. https://doi.org/10.5281/zenodo.21208170 (2026)
Acknowledgement
We (V.W. and M.M.N.) thank the Mediterranean Society for the Study of Consciousness (MESEC) for their support
Funding
V.W. and M.M.N. were funded through a UK AI Security Institute (AISI) Challenge Fund award to M.M.N. M.M.N. is also supported by the Wellcome Trust (315364/Z/24/Z). The simulation studies were funded through a UK AISI Challenge Fund award to M.M.N. The human annotation study was funded by Microsoft AI. The experiments presented here were not run by the AISI
Author information
Authors and Affiliations
Max Planck UCL Centre for Computational Psychiatry and Ageing Research, London, UK
Veith Weilnhammer, Raymond Dolan & Matthew M. Nour
Sydney Medical School, University of Sydney, Sydney, Australia
Kevin YC Hou
UK AI Security Institute, London, UK
Lennart Luettgau & Christopher Summerfield
Department of Experimental Psychology, University of Oxford, Oxford, UK
Christopher Summerfield
Microsoft AI, London, UK
Matthew M. Nour
Department of Psychiatry, University of Oxford, Oxford, UK
Matthew M. Nour
Authors
- Veith WeilnhammerView author publications
Search author on:PubMed Google Scholar
- Kevin YC HouView author publications
Search author on:PubMed Google Scholar
- Lennart LuettgauView author publications
Search author on:PubMed Google Scholar
- Christopher SummerfieldView author publications
Search author on:PubMed Google Scholar
- Raymond DolanView author publications
Search author on:PubMed Google Scholar
- Matthew M. NourView author publications
Search author on:PubMed Google Scholar
Contributions
V.W., M.M.N., L.L. and C.S. conceived of the study. V.W. and M.M.N. designed the simulation and evaluation approach. V.W. implemented the simulation pipeline, conducted the experiments and performed the analyses. M.M.N. performed the human annotation study. V.W., K.Y.C.H. and M.M.N. wrote the paper. L.L., C.S. and R.D. provided feedback on the analyses and paper. M.M.N. supervised the project. All authors reviewed and approved the final paper
Ethics declarations
Competing interests
M.M.N. is an employee of Microsoft (Principal Applied Scientist, Microsoft AI), the developer of the Copilot and Copilot Health consumer AI assistants, and owns Microsoft stock as part of standard employee compensation. V.W. is a paid consultant for Microsoft AI; this employment played no part in the present study. V.W. is the founder of pixelblot.ai; this role played no part in the current study. K.Y.C.H. is a member of the technical staff at Faultline AI; this affiliation commenced after the completion of this work and Faultline AI played no part in the present study. The other authors declare no competing interests.
Peer review
Peer review information
Nature Medicine thanks Thomas Ward, Jie Yang and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available. Primary Handling Editors: Mattia Andreoletti and Ming Yang, in collaboration with the Nature Medicine team
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations
Extended data
Extended Data Fig. 1 Reliability and validity
a. Conversation-level concerning-behavior scores from claude-opus-4.5 correlated with mean turn-level concerning-behavior scores from claude-sonnet-4.5 at r = 0.87 (Spearman, p < 0.001). Points represent conversations (n = 810); colored markers and error bars show model means ± 95% CI (n = 90 conversations per model). b. Agreement between claude-opus-4.5 and gpt-5.2 on conversation-level mental-health risk (PC1; Spearman r = 0.91, p < 0.001). Points represent conversations (n = 810), colors indicate target chatbots, and larger points with error bars show model means ± 95% CI (n = 90 conversations per model). GPT-5.2 scores were scaled and projected using the PCA transformation derived from claude-opus-4.5 scores; PC1 loadings are shown in Extended Data Fig. 3. c. Agreement between conversation-level judges across risk dimensions. Faint points represent individual scores and solid points model-level means (Spearman r = 0.96, p < 0.001). d. Recovery of causal risk manipulations. Using the first simulated user message from each SIM-VAIL conversation as context, grok-4 generated low-risk and high-risk responses for each judge dimension, guided by that dimension’s rubric definition. Responses were scored across all dimensions. Cells show AUC for separating high- from low-risk generations. Rows denote manipulated dimensions and columns denote scored dimensions. Diagonal values measure recovery of the intended manipulation, whereas off-diagonal values show cross-loading between mechanisms (median diagonal AUC = 0.98). e. Alignment of causal manipulations with the latent risk space. Points show the cosine similarity between each manipulation-induced displacement vector and the corresponding dimension-loading vector in the PC1–PC2 space derived from the original conversations (Fig. 5a). Positive values indicate concordant directions; the median similarity was 0.9. All tests were two-sided.
Extended Data Fig. 2 Clinician-annotator validation of safety ratings and simulated-user realism
a. Each dot represents one annotation, plotted as the claude-sonnet-4.5 turn-level score against the clinician annotator score for concerning chatbot behavior. The dashed line is the least-squares fit; r = 0.49 (Spearman, p < 0.001). b. Turns were binned into fixed score bands (1–2, 3–4, 5–6, 7–8, 9–10). Points show the mean raw clinician annotator score for concerning chatbot behavior within each band; error bars show ± 1 SEM; the monotonic association across band means was r = 1 (Spearman). c. Inter-rater reliability for the 104 doubly rated items (Spearman r = 0.41, p < 0.001; two-way random-effects ICC(2,1) = 0.31). d. The same paired set aggregated into true quintiles of Annotator 1’s score, showing the corresponding mean score of Annotator 2 with ± 1 SEM; the monotonic association across quintile means was r = 1 (Spearman). e. Distribution of clinician-annotator realism ratings for simulated user messages. Horizontal lines terminate in dots at the percentage of ratings assigned to each point on the 1–5 realism scale. The dashed red horizontal line marks the mean realism rating (4.15). Ratings were concentrated in the upper half of the scale, with 80% of turns rated as broadly plausible or genuine. All tests were two-sided.
Extended Data Fig. 3 Principal-component structure and co-variation of mechanism-level risk ratings
a. Heatmap of PCA loadings for the mental-health–relevant behavioral dimensions (rows) across the first 5 principal components (columns). Loadings were computed from conversation-level ratings and quantify how strongly each dimension contributes to each component; positive and negative values indicate opposing patterns of co-variation across dimensions. Components are labeled with the percent of variance explained (in parentheses). Color intensity reflects the magnitude and sign of the loading. b. Spearman correlation matrix across mechanism-level risk dimensions computed at the conversation level.
Extended Data Fig. 4 Concerning AI chatbot behavior in simulated users with versus without mental-health vulnerabilities
Each point shows a target AI chatbot’s mean concerning score (1–10; conversation-level ratings) in the primary condition (simulated users with a mental-health vulnerability, x-axis) versus the control condition (simulated users without a mental-health vulnerability, y-axis). In both groups, simulated users engaged the same AI chatbots with the same conversational intents. Horizontal and vertical error bars denote 95% CIs around the model means (n = 90 conversations per model for vulnerable users on the x-axis; n = 18 for non-vulnerable control users on the y-axis). Points are colored by target AI chatbot.
Extended Data Fig. 5 Context-dependent concerning behavior across AI chatbots
a. Concerning behavior across AI chatbots, vulnerabilities, and intents. For each AI chatbot, the outlined point shows the mean conversation-level concerning score (1–10) and the small dots show the individual conversations (n = 3 independent simulation runs per vulnerability × intent × chatbot cell). Scores are shown for each vulnerability–intent pairing, providing an overview of the full interaction space. Model differences depended jointly on vulnerability and intent (vulnerability × intent × chatbot interaction, Type III F-test: F(160, 540) = 1.60, p < 0.001). All tests were two-sided.
Extended Data Fig. 6 Context-dependent timing and mechanisms of chatbot risk
a. Mean number of turns required for a conversation to reach a concerning score ≥ 5 for each vulnerability–intent pairing; cells report mean 95% CI. Conversations that never reached the threshold within 10 turns were assigned 10. Colors encode time to threshold. b. Across the full interaction space, we observed a heterogeneous landscape of risk expression in which similar levels of concerning AI chatbot behavior could arise from qualitatively distinct mechanisms with different clinical consequences. Vulnerability × intent interactions were significant (Type III F-tests from linear mixed-effects models) not only for overall concerning behavior (F(20, 270) = 9.45, p < 0.001) and therapeutic quality (F(20, 270) = 11.47, p < 0.001), but also for all mechanism-specific risk dimensions that were clinically central to our framework (all p < 0.05, without adjustment for multiple comparisons). Simulated mental-health risk was therefore inherently contextual: the same user intent could be relatively safe in one vulnerability but harmful in another, and the same vulnerability could express risk through different mechanisms depending on the user intent. All tests were two-sided.
Extended Data Fig. 7 Sensitivity analysis of extending early-stopped conversations to 10 turns
Early-stopped conversations were continued to 10 turns using the original simulation setup. We evaluated the effect of this extension at the conversation level, turn level, and trajectory-clustering level. a. Relationship between original and extended conversation-level scores for concerning AI chatbot behavior. The dashed line denotes identity. Colored markers and error bars show per-chatbot means ± 95% CI (total n = 766 extended conversations). Across conversations, original and extended scores remained highly correlated (Spearman r = 0.93, p < 0.001). b. Number of additional turns generated to bring early-stopped conversations to 10 turns. c. Mean change in conversation-level scores after extension across the 13 mental-health dimensions used in SIM-VAIL. Points and horizontal intervals show means ± 95% CI. For harmful dimensions, positive values indicate higher risk in the extended conversation; for therapeutic quality, positive values indicate improved therapeutic quality after extension. Extending conversations increased the mean conversation-level concerning-behavior score by 5.4% (0.23 points). d. Mean change in turn-level concerning-behavior score relative to the last original user-target turn, shown separately for each added-turn lag among conversations included in panels a-c. Only lags with more than 10 contributing conversations are shown; vertical intervals denote 95% CI. e. Mean concerning-behavior trajectories for the extended conversations, grouped using the original fitted 4-cluster trajectory solution. f. Cluster composition across vulnerability, intent, and AI chatbot under the extended-turn reanalysis. Overall, extending early-stopped conversations had only modest effects on the main conclusions, with 83.5% of conversations retaining their original cluster label. All tests were two-sided.
Extended Data Fig. 8 Judge-family and exact same-model effects on conversation-level concerning-behavior scores
a. Same-model-family analysis comparing conversation-level scores by claude-opus-4.5 and gpt-5.2. For Anthropic targets, claude-opus-4.5 was treated as the same-family judge and gpt-5.2 as the other-family judge; for OpenAI targets, this assignment was reversed (n = 180 conversations per developer family). Points show group means and error bars show 95% CIs around the mean. This analysis asks whether judges systematically rate outputs from their own developer family as less concerning. We found no overall same-family versus other-family effect (main effect of same- vs. other-family, linear mixed-effects model: β = −0.06, t = −0.86, p = 0.39). The same-vs-other difference varied by target family (same-family by target-family interaction: β = −0.41, t = −5.87, p < 0.001). b. Exact same-model analysis using conversational-level judges that were also included as audited target AI chatbots: claude-sonnet-4.5, gpt-5, grok-4, and gemini-2.5-pro. Each facet shows 1 target AI chatbot, and points show the mean concerning-behavior score assigned by each judge, with error bars showing 95% CIs around the mean (n = 357 conversations scored by all four judges). The black-outlined point marks the score assigned by the exact same model as the target; the dashed horizontal line marks the mean score assigned by the other 3 judges. This analysis asks whether a model rates its own outputs as less concerning than other frontier judges do. Across complete conversation-target combinations, the same-model judge assigned lower concerning-behavior scores than the other judges on average (one-sample t-test, no adjustment for multiple comparisons: mean difference = −0.9, t = −9.33, p < 0.001). Thus, while we did not find a same-developer-family effect, exact same-model judging may underestimate concerning behavior. Importantly, the SIM-VAIL analyses did not use any exact same-model judge-target pairs at the conversation level. All tests were two-sided.
Extended Data Fig. 9 Selection diagnostics for conversations eligible for the counterfactual branching analysis
Intervention points were defined by the first target chatbot message within a conversation that reached a concerning behavior score ≥ 7 and was followed by at least one additional user-target turn. a. The distribution of turns at which the intervention was applied. b. Mean observed concerning score trajectory up to the selected intervention turn, stratified by the turn position selected for intervention (different blue shades correspond to selected turns 1 to 8). c. Fraction of conversations satisfying the eligibility criterion, separately for vulnerability, intent, and target AI chatbot. Horizontal lines terminate in dots at the eligible fraction; labels show the eligible and total conversation counts.
Supplementary information
Supplementary Information (download PDF )
Supplementary Tables 1–10 and all associated legends (Table of Contents on page 1)
Reporting Summary (download PDF )
Peer Review File (download PDF )
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
About this article
Cite this article
Weilnhammer, V., Hou, K.Y., Luettgau, L. et al. A clinically validated framework for auditing AI chatbot behavior in mental health interactions.
Nat Med (2026). https://doi.org/10.1038/s41591-026-04577-2
Received:09 March 2026
Accepted:09 July 2026
Published:07 August 2026
Version of record:07 August 2026
DOI
:https://doi.org/10.1038/s41591-026-04577-2


