News Research Thyroid Disease Management Diagnostics & Imaging

DeepSeek-R1 tops ChatGPT-4o in thyroid nodule classification

August 19, 2026 By Matthew Solan 5 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

DeepSeek-R1 showed significantly higher sensitivity and accuracy than ChatGPT-4o when differentiating benign from malignant thyroid nodules using ultrasound text reports, although both large language models (LLMs) had lower overall diagnostic discrimination and accuracy than senior radiologists, according to a multicenter retrospective study published in the Journal of Medical Internet Research.

The analysis included 1,063 thyroid nodule ultrasound reports collected at three medical centers in China between January 2023 and December 2024. Of these, 306 nodules with definitive cytopathologic or histopathologic diagnoses: 182 malignant and 124 benign. 

Investigators submitted the original Chinese-language ultrasound reports to both LLMs using zero-shot, task-specific prompts that restricted responses to predefined categories. DeepSeek-R1 was operated with its “Deep Think” mode enabled, whereas ChatGPT-4o used its standard chat configuration. Neither model underwent additional training on dedicated medical imaging data sets. 

Investigators evaluated the models on three tasks: benign-malignant differentiation, Chinese Thyroid Imaging Reporting and Data System (C-TIRADS) classification, and management recommendations. For each task, every report was submitted to each model five times in separate sessions, and the modal response was used as the final result. Benign-malignant differentiation was evaluated against pathology, whereas C-TIRADS classification and management recommendations were evaluated for agreement with senior radiologists and attending clinicians, respectively.

Benign-malignant differentiation. The models classified the 306 nodules with definitive pathologic diagnoses as benign or malignant based on sonographic descriptions. DeepSeek-R1 achieved 88% sensitivity and 73% accuracy, compared with 69% sensitivity and 64% accuracy for ChatGPT-4o. Negative predictive value (NPV) was 74% with DeepSeek-R1 and 56% with ChatGPT-4o. Specificity was 51% and 57%, respectively, and positive predictive value (PPV) was 72% and 70%.  

In a Bayesian analysis assuming malignancy prevalences of 5% to 10%, PPV was estimated at 8% to 15% for ChatGPT-4o and 9% to 17% for DeepSeek-R1. Adjusted NPVs exceeded 94% for both models. 

Area under the curve (AUC) was 0.718 for DeepSeek-R1 and 0.688 for ChatGPT-4o; the difference was not statistically significant. Senior radiologists achieved a higher AUC of 0.865. However, radiologists had access to ultrasound images and clinical context, whereas the LLMs received text reports alone. 

Output stability also differed for benign-malignant classification. DeepSeek-R1 produced substantially more consistent results across repeated queries than ChatGPT-4o (κ=0.869 vs 0.609). The authors noted that inconsistent responses to identical reports could be particularly concerning in patient-facing applications. 

C-TIRADS classification. The models assigned C-TIRADS categories to all 1,063 nodules based on their sonographic descriptions. DeepSeek-R1 showed greater overall agreement with senior radiologists than ChatGPT-4o, with weighted κ values of 0.770 and 0.688, respectively. Across the three centers, agreement ranged from 0.686 to 0.822 for DeepSeek-R1 and from 0.508 to 0.788 for ChatGPT-4o.

Agreement varied by C-TIRADS category. Both models showed 73% to 97% agreement with radiologists for categories 2 and 5 nodules, but substantially lower agreement for category 4 nodules, ranging from 14% to 46%. For category 1 nodules, DeepSeek-R1 achieved 100% agreement, compared with 30% for ChatGPT-4o. 

Overall stability across repeated C-TIRADS classifications was nearly perfect and similar between models, with Krippendorff α values of 0.864 for DeepSeek-R1 and 0.866 for ChatGPT-4o. Category-specific stability varied, however: DeepSeek-R1 was more stable for category 1 nodules, whereas ChatGPT-4o was more stable for categories 3, 4a, 4b, and 4c.

Management recommendations. The models received the complete reports, including radiologist-assigned C-TIRADS categories, for all 1,063 nodules and selected either follow-up or fine-needle aspiration (FNA). Both models showed moderate, nearly identical agreement with clinicians: κ was 0.608 for ChatGPT-4o and 0.606 for DeepSeek-R1. Raw agreement was approximately 81% for both models.  

The authors reported higher concordance for follow-up recommendations (approximately 86%) than for FNA recommendations (75%). However, they noted that the models did not have access to factors such as age, comorbidities, clinical symptoms, or socioeconomic status that can contribute to clinical decision-making. Both models were also highly stable in their management recommendations, with κ values of 0.853 for DeepSeek-R1 and 0.849 for ChatGPT-4o. 

Overall stability across repeated C-TIRADS classifications was nearly perfect and similar between models, with Krippendorff α values of 0.864 for DeepSeek-R1 and 0.866 for ChatGPT-4o. Category-specific stability varied, however: DeepSeek-R1 was more stable for category 1 nodules, whereas ChatGPT-4o was more stable for categories 3, 4a, 4b, and 4c.

The authors identified several limitations. Only 306 of 1,063 nodules had pathologic confirmation, and this subset was enriched for malignancy, with 60% of nodules being malignant. Because pathologic confirmation was available only for patients who underwent FNA or thyroidectomy for clinical indications, the resulting spectrum bias could inflate PPV and overall accuracy and potentially underestimate NPV. 

All reports and prompts were in Chinese, and patients were recruited within a single province, limiting linguistic and geographic generalizability. The models were also evaluated under different consumer-interface configurations, with DeepSeek-R1 using its “Deep Think” reasoning mode and ChatGPT-4o using its standard chat configuration. The authors noted that DeepSeek-R1 may also have benefited from optimization for Chinese-language content.

Other limitations included use of a single standardized prompt for each task, incorporation of radiologist-derived C-TIRADS classifications into the management task, and restriction of management output to follow-up or FNA.  

The authors concluded that both LLMs showed clinical potential but "remained inferior to senior radiologists, suggesting their role as decision-support tools rather than stand-alone diagnostic systems."  

The authors reported no conflicts of interest. 

AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.

Related Content