News Ethics and Policy Ethics, Regulation, and Responsible Use

Simulated AI linked to less stress when delivering bad news

August 14, 2026 By Matthew Solan 6 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

Healthy participants receiving fictitious bad medical news in simulated artificial intelligence physician encounters experienced lower subjective stress than those in human-physician encounters, while memory retrieval of consultation information differed by format, according to a study published in npj Digital Medicine

"These results suggest that conversational technologies should not be evaluated solely in terms of technical feasibility or user acceptance, but also with regard to their psychological effects on patients," the authors wrote.  

The study compared four formats for receiving a simulated prognosis for Huntington's disease: an in-person human physician, a human physician by video visit, an artificial intelligence (AI)-styled text chatbot, and an AI-styled spoken virtual avatar. 

Both simulated AI conditions used a "Wizard of Oz" design in which participants were told they were interacting with an autonomous AI physician while trained human operators controlled the interactions. For the chatbot, participants exchanged typed messages through a custom interface; the operator could select scripted material or type additional responses. For the avatar condition, an adapted Virtual Human Toolkit presented a humanlike figure that delivered spoken responses controlled by the operator. The same three researchers who portrayed physicians in the human conditions served as operators in the AI-styled conditions. 

The investigators assessed subjective stress, physiological stress measured by salivary cortisol, memory retrieval, and perceived credibility of the simulated consultation. 

The final analysis included 163 healthy adults; 45 were assigned to in-person human consultation, 39 to human video consultation, 39 to the AI-styled chatbot, and 40 to the AI-styled avatar.  

All participants underwent two standardized simulated consultations. They were first told that an initial saliva screening indicated a one-third probability of Huntington's disease and, after a 5- to 10-minute break, were given a fictitious positive genetic test result. Consultation content included Huntington's disease symptoms, cytosine-adenine-guanine (CAG) trinucleotide repeats, age at onset, progression, treatment options, and potential effects on life. 

Subjective stress was assessed immediately after the second consultation with two 0-to-100 visual analog items covering perceived stress and challenge. Saliva was collected at five time points for cortisol measurement.  

When the investigators combined the two human-physician conditions and compared them with the two simulated AI conditions, mean subjective stress was 63.5 in the human group and 52.05 in the AI-styled group on the study’s 0-to-100 composite scale, with a standardized effect size of 0.46.  

In the four individual groups, mean scores were 63.31 for in-person human consultation, 63.72 for human video consultation, 51.08 for the AI-styled chatbot, and 53 for the AI-styled avatar. A one-way analysis of variance identified an overall effect of consultation format, accounting for 5% of the variance, although post hoc comparisons did not identify significant pairwise differences. 

Salivary cortisol declined over time in all conditions. In the four-condition model, the decline was steeper for the AI-styled avatar than for the in-person human condition, although that interaction was no longer statistically significant in a sensitivity analysis incorporating covariates. When the conditions were grouped by agent, cortisol declined more rapidly over time in the combined AI-styled group than in the combined human group. 

Memory testing to assess how well participants retained consultation information included six open-ended, seven closed-ended, and three visual-environment questions, with responses scored as correct or incorrect. Open-ended recall did not differ significantly among the four consultation formats. For closed-format memory, the chatbot scored significantly lower than both the human video and AI-styled avatar conditions. Mean scores were 86% for the AI-styled chatbot, 96% for human video consultation, 94% for the AI-styled avatar, and 93% for in-person human consultation. 

For visual memory, human video consultation outperformed both in-person consultation and the AI-styled chatbot, while the AI-styled avatar also outperformed the chatbot. Mean scores were 92% for human video consultation, 75% for in-person consultation, 85% for the AI-styled avatar, and 64% for the AI-styled chatbot.  

Perceived credibility did not differ significantly across the four consultation formats, although the avatar condition had the lowest mean credibility rating. Higher perceived credibility was associated with higher subjective stress, but the association was reversed with the avatar.  

In exploratory analyses, the investigators also found no significant sex differences in subjective stress, cortisol response, perceived credibility, or memory performance. 

The investigators noted several limitations. They cautioned that the "Wizard of Oz" approach was not intended as a validated approximation of existing AI technology. Because human operators generated the responses, the simulated agents did not reproduce the full constraints or behaviors of real-world AI systems. The four conditions also varied simultaneously in agent type, communication modality, social presence, and embodiment, preventing the investigators from isolating which component produced the observed effects. 

Ecological validity was another limitation. Average credibility ratings were below the midpoint of the scale, and the avatar was perceived as less realistic than the other modalities. The investigators therefore said it remains uncertain whether the stress and memory findings would generalize to encounters with actual physicians or real-world AI systems. Visual differences among conditions—including face size, framing, apparent interpersonal distance, and researchers’ physical appearance—could also have influenced participants’ responses. 

All participants were healthy and knew that they were not undergoing a real medical examination, potentially limiting applicability to patients actually receiving serious news. The investigators also did not formally assess participants’ prior familiarity with Huntington's disease. Because data were collected in 2022, before widespread public familiarity with large language models, participants may also have had different expectations of AI than current populations. 

The investigators called for future studies with greater ecological validity, broader age and sex representation, more tightly standardized visual presentation, and systematic variation in AI autonomy and realism. They also recommended independent replication and studies of less emotionally demanding medical encounters. 

"As healthcare increasingly integrates digital communication tools, it will be essential to understand how interaction modalities relate to emotional burden, cognitive processing, and the effective delivery of sensitive medical information," the authors wrote.  

The study was funded by the Baden-Württemberg Foundation in Stiftung, Germany through its Responsible Artificial Intelligence funding line. The funding agency had no role in study design, data collection, analysis, interpretation, or the decision to submit the study for publication. The authors declared no competing interests. 

AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.

Related Content