News Research Ethics, Regulation, and Responsible Use Research and Evidence Precision Endocrinology

AI agents agree with physicians overall—but not on the best clinical LLMs

July 20, 2026 By Matthew Solan 6 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

More than 400 physicians evaluating large language model-generated responses to real clinical cases reached substantially different rankings depending on their clinical experience and practice environment, while artificial intelligence agents showed broad agreement with physician rankings but did not replicate the nuances of physician judgment, according to a multicenter study published in npj Digital Medicine.

"Automated evaluators may be useful for exploratory benchmarking or preliminary comparison of LLM-generated clinical case interpretations, but final assessment of clinical relevance, safety, and utility should remain anchored in appropriately diverse human clinical evaluation," wrote first author Peilun Shi of the Department of Biomedical Engineering at The Chinese University of Hong Kong in China, and colleagues. 

Current evaluations of medical LLMs frequently rely on medical question-answering, which the researchers said may overestimate model performance and fail to capture how physicians evaluate free-text clinical reasoning. To address that limitation, the investigators conducted a multicenter observational study comparing physician assessments with those generated by artificial intelligence (AI) agents configured to simulate physician characteristics.  

The study recruited 450 physicians from eight medical institutions in China. After prespecified quality screening, 421 physician evaluations were included in the final analysis.  Physicians represented seven clinical specialties and were categorized as junior (1 to 3 years of experience; n = 252), middle-career (4 to 14 years; n = 127), or senior (more than 15 years; n = 42). A parallel cohort of 421 AI agents was created using the CAMEL architecture and configured to mirror physician roles. 

Investigators retrospectively assembled seven real-world, deidentified clinical cases collected between November and December 2024. For each case, five LLMs generated clinical interpretations using identical prompts: 

  • GPT-4o 

  • Claude 3.5 

  • Qwen-Max 

  • DeepSeek-R1 

  • OpenAI o1 

Each physician evaluated only cases within his or her specialty and ranked the quality of each model's interpretation across six predefined domains: comprehensiveness, coherence, accuracy, helpfulness, harmlessness, and structural quality. Rankings were converted into standardized scores (100, 75, 50, 25, and 0). AI agents completed the same evaluations using identical case materials and scoring rules. 

Physician preferences varied according to clinical seniority. Junior physicians consistently gave Qwen-Max the highest average scores, whereas senior physicians more often favored GPT-4o and DeepSeek-R1 across multiple evaluation domains.  

Physician rankings also differed according to practice environment. Physicians in plain-region institutions more frequently favored GPT-4o and Qwen-Max for comprehensiveness and helpfulness, whereas physicians from high-altitude institutions assigned relatively higher rankings to OpenAI o1 across several evaluation domains. 

Although AI agents broadly aligned with physicians regarding the relative ordering of models overall, they displayed different preferences for the highest-ranked systems. Across the six evaluation domains, physicians generally favored GPT-4o and DeepSeek-R1, whereas AI agents consistently assigned the highest scores to OpenAI o1. Claude 3.5 received the lowest ratings from both physicians and AI agents.  

Despite these differences in model preferences, investigators performed a global rank concordance analysis using Kendall's coefficient of concordance. Across all five models and six evaluation domains, AI agents demonstrated broad directional alighment with physician rankings (Kendall's W = 0.798), indicating that both evaluator groups generally ordered model performance similarly despite disagreeing on the highest-ranked systems.  

The investigators also compared how consistently AI agents and physicians ranked the five LLMs across experience levels. AI agents produced nearly identical rankings regardless of whether they were configured to simulate junior, middle-career, or senior physicians. Overall, 23 of 30 ranking pairs (77%) were identical across simulated experience levels, with OpenAI o1 consistently ranked first and DeepSeek-R1 second.  

In contrast, physician rankings varied substantially. Only one of 30 ranking pairs (3%) remained identical across junior, middle-career, and senior physicians. Rankings for Qwen-Max ranged from first to fifth depending on physician experience and evaluation domain, while GPT-4o likewise shifted between first and fifth place.  

Senior physicians generally favored Qwen-Max for comprehensiveness and coherence, whereas junior physicians showed broader variation and more frequently assigned higher rankings to Claude 3.5 and OpenAI o1 than did senior clinicians. According to the investigators, these findings suggest that physician judgment incorporates experience-dependent considerations that current AI evaluators do not reproduce. 

The researchers also compared evaluation time between physicians and AI agents. Senior physicians required the longest evaluation times across specialties, followed by middle-career and junior physicians. Oncology cases required the longest average evaluation time—approximately 135 minutes—whereas ophthalmology cases required the least, averaging about 95 minutes. In contrast, AI agents completed identical evaluation tasks in minimal computational time. The researchers emphasized, however, that these efficiency findings were exploratory and did not alter the primary conclusion regarding agreement between human and AI evaluators. 

The study had several limitations. Although the physician cohort was relatively large, the evaluation included only seven real-world clinical cases, each representing a different specialty. Because physicians evaluated only cases within their own specialty, the effects of specialty and case content could not be separated. All cases and evaluations were conducted in Chinese, limiting generalizability to other languages and clinical documentation systems. The study also compared physician evaluation patterns without determining which physician group produced the most clinically valid assessments. 

In addition, the ranking system measured relative preference rather than absolute differences in model quality, and all AI agents were powered on a single underlying LLM architecture, which may limit generalizability to other agent frameworks. Finally, the findings apply only to evaluation of LLM-generated clinical case interpretations and should not be extrapolated to diagnostic decision support, predictive modeling, treatment recommendation, or other medical AI applications. 

The researchers noted that prospective, task-specific validation would be needed to determine whether AI agents can safely perform preliminary screening of LLM-generated clinical case interpretations without missing clinically important errors. They also recommended that future LLM evaluation studies consider practice environment as a stratification variable and incorporate expert adjudication, predefined reference standards, or outcome-linked assessments to distinguish evaluator preference from evaluation quality.  

No conflicts of interest were reported. 

(Editor’s Note: The researchers noted that the study manuscript will undergo further editing before final publication, and there may be errors present that affect the content. All legal disclaimers apply.) 

 

 

AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.

Related Content