Stress tests expose hidden weaknesses in leading medical AI models
Large frontier artificial intelligence models that achieved high scores on multimodal medical benchmarks exhibited important robustness gaps when evaluated with structured adversarial stress tests, suggesting that benchmark performance alone may not adequately reflect readiness for health care applications, according to a study published in Nature Medicine.
"Our results call for more careful interpretation of benchmark performance in health applications," noted researchers. "Benchmark scores alone should not be treated as sufficient evidence of robustness or readiness-relevant performance."
The study evaluated six leading frontier models—GPT-5, Gemini 2.5 Pro, OpenAI-o3, OpenAI-o4-mini, GPT-4o, and Claude 3.5 Sonnet—using a framework of six clinician-informed stress tests designed to assess robustness beyond conventional benchmark accuracy. The investigators examined multimodal medical reasoning across established benchmarks, including the New England Journal of Medicine Image Challenge, JAMA Clinical Challenge, VQA-RAD, PMC-VQA, and OmniMedVQA, while also profiling benchmark characteristics using clinician-developed scoring rubrics.
The clinician-guided profiling showed that widely used benchmarks measure different combinations of visual interpretation and clinical reasoning. The stress tests assessed model performance under input removal, visual-necessity subset, format perturbation, distractor manipulation, visual substitution, and reasoning signal fidelity. Rather than measuring accuracy alone, the framework also evaluated abstention behavior, reasoning fidelity, and model stability under perturbation.
Models Continued Answering Without Required Images
Removing diagnostic images reduced performance across models on the New England Journal of Medicine benchmark, although accuracy often remained well above chance. GPT-5 declined from 81% with image-plus-text input to 67% with text alone, while Gemini 2.5 Pro decreased from 81% to 67%. On the JAMA benchmark, declines were smaller overall, with GPT-5 falling from 87% to 83%.
To test whether models can still answer questions that explicitly require image information, the researchers curated a 197-item subset of the New England Journal of Medicine benchmark, referred to as the NEJM Visual-required Subset, comprising cases selected based on clinical criteria indicating minimal textual cues and a high dependency on visual features.
With complete inputs, GPT-5 achieved 70% accuracy, Gemini 2.5 Pro 67%, and OpenAI-o3 65%. In the text-only condition, GPT-5 achieved 41% accuracy, Gemini 40%, OpenAI-o3 39%, and OpenAI-o4-mini 38%. Accuracy remained well above the 20% random baseline. The researchers said these findings suggested reliance on nonvisual cues or memorized associations rather than sensitivity to missing or conflicting visual evidence. GPT-4o differed from other models by frequently abstaining, achieving 16% overall accuracy because abstentions were counted as incorrect.
Minor Prompt Changes Altered Performance
Additional stress tests showed that seemingly minor modifications affected model behavior. Randomizing answer-choice order modestly reduced performance under text-only conditions, indicating sensitivity to answer formatting, while replacing incorrect answer options with unrelated distractors progressively reduced text-only accuracy toward chance. Conversely, image-plus-text performance often improved when irrelevant distractors were removed because the correct diagnosis became more visually apparent. Replacing one distractor with "Unknown" increased accuracy in both input conditions, suggesting that models treated "Unknown" as an easily eliminated option rather than expressing uncertainty.
Visual Substitution Exposed Grounding Weaknesses
The researchers also replaced original diagnostic images with clinically plausible alternative images corresponding to different answer choices while leaving the accompanying text unchanged.
Under these conditions, GPT-5 accuracy declined from 84% to 53%, Gemini 2.5 Pro from 76% to 53%, OpenAI o4-mini by 23%, and OpenAI-o3 by 33%. GPT-4o was the only model to improve performance. According to the researchers, these findings suggest that many models did not consistently reinterpret changing visual evidence and instead relied on static image-answer pairings or incomplete visual-text integration.
Reasoning Prompts Did Not Consistently Improve Performance
The researchers also evaluated reasoning fidelity by applying chain-of-thought prompting and manually reviewing generated explanations.
Explicit reasoning prompts did not improve performance on the New England Journal of Medicine benchmark across models. On VQA-RAD, gains were small, ranging from less than 1% among reasoning models to approximately 4% among nonreasoning models. Increasing reasoning effort on OmniMedVQA produced inconsistent effects, with longer reasoning chains sometimes introducing hallucinated details. Manual review identified recurring patterns in which models produced correct answers supported by inaccurate reasoning, propagated initial visual errors through subsequent reasoning, or generated coherent but clinically uninformative explanations.
Benchmark Characteristics Varied
Using clinician-developed rubrics spanning 10 evaluation dimensions, the researchers found that commonly used benchmarks differed substantially in the reasoning and visual demands they place on models.
New England Journal of Medicine cases ranked high in both reasoning and visual demands, whereas JAMA cases required substantial reasoning but were mostly text-solvable. VQA-RAD, PMC-VQA, and MIMIC-CXR were visually dependent but low in inference complexity, while OmniMedVQA ranked low in both dimensions. The researchers said these differences may explain why models perform differently across benchmarks and recommended against treating benchmark scores as interchangeable indicators of readiness.
Study Limitations
The researchers acknowledged several limitations. The analyses evaluated performance on health AI benchmarks rather than prospective clinical workflows, relied heavily on multiple-choice tasks that do not capture open-ended clinical decision-making, examined selected adversarial perturbations rather than the full range of clinical uncertainty, and included only a small private chest radiograph data set for supplementary validation. They also noted that both frontier models and benchmark data sets continue to evolve, requiring repeated evaluation over time.
Yu Gu is currently employed by ByteDance. The other researchers declared no competing interests.
AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.