News Research Diagnostics & Imaging Research and Evidence Ethics, Regulation, and Responsible Use

LLMs show high accuracy interpreting health checkup results

August 07, 2026 By Matthew Solan 4 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

Large language models accurately classified structured health checkup results into predefined clinical categories when paired with structured prompting strategies, according to a study published in npj Digital Medicine. Performance was consistently strongest for routine laboratory measures and weaker for blood pressure interpretation.

The retrospective study evaluated four large language models (LLMs)—Claude Sonnet 4, Gemini 2.5 Pro, GPT-4o, and LLaMA 3.1-70B—for interpretation of structured health checkup data. The models classified laboratory and examination results into predefined clinical categories using numerical health screening inputs, including blood pressure, body mass index, glucose, lipid values, liver function tests, hemoglobin, serum creatinine, and urine protein. 

The study used data from the Korean National Health Insurance Service, which oversees a nationwide standardized health screening program. After excluding records with missing values, investigators retained approximately 320,000 complete records before drawing a stratified sample of 10,000 patients that preserved the original population distribution. 

The investigators compared five increasingly structured prompting strategies, beginning with a simple zero-shot prompt and progressively adding role-based instructions, few-shot examples, constraint prompting, and chain-of-thought reasoning. Average accuracy across models increased from 0.69 with zero-shot prompting to 0.92 after combining role-based, few-shot, and constraint prompting, reaching 0.95 after adding chain-of-thought prompting. 

Using the final prompting strategy, Claude Sonnet 4 and Gemini 2.5 Pro each achieved an overall mean accuracy of 99%, followed by GPT-4o at 98%, and LLaMA 3.1-70B at 84%. When the investigators repeated the same evaluation three times using identical inputs, Claude Sonnet 4, Gemini 2.5 Pro, and GPT-4o produced nearly identical results, indicating highly reproducible performance across repeated evaluations.  

Performance varied considerably by clinical variable. Across all models, fasting blood glucose, total cholesterol, triglycerides, high-density lipoprotein cholesterol, urine protein, serum creatinine, aspartate aminotransferase, and alanine aminotransferase generally achieved accuracies ranging from 97% to 100%. Blood pressure proved substantially more difficult to classify, with accuracies ranging from 88% to 91% for Claude Sonnet 4, Gemini 2.5 Pro, and GPT-4o, compared with 61% for LLaMA 3.1-70B. The investigators attributed the lower performance to the need to integrate systolic and diastolic measurements across multiple clinical categories. 

Sensitivity and specificity also varied across clinical measures. For blood pressure classification, sensitivity ranged from 58% to 77% across models, whereas specificity remained high at 92% to 98%. Mean specificity across all evaluated health checkup items was 99% for Claude Sonnet 4, 100% for Gemini 2.5 Pro, 99% for GPT-4o, and 93% for LLaMA 3.1-70B.  

Subgroup analyses identified model-specific differences. Claude Sonnet 4 demonstrated higher body mass index classification accuracy in male patients. LLaMA 3.1-70B showed sex-related differences for urine protein, serum creatinine, and gamma-glutamyl transferase. Age-related declines in blood pressure classification accuracy were observed with Claude Sonnet 4 and Gemini 2.5 Pro, while GPT-4o maintained consistent performance across age groups. LLaMA 3.1-70B demonstrated age-related variability for low-density lipoprotein cholesterol and hemoglobin.  

The study had several limitations. The analysis used only structured health screening data from South Korea, potentially limiting generalizability to other populations, languages, and health systems. The models were evaluated only on numerical classification tasks without incorporating patient history, concomitant treatments, or other clinical context. Because the investigators did not retrain the underlying models, embedded model biases remained unaddressed. The investigators also noted that deterministic rule-based systems can achieve perfect accuracy for fixed threshold classification and that ethical considerations, including privacy, interpretability, and patient safety, require further study before clinical use. 

The authors said broader validation in multiethnic and multilingual populations, prospective validation, long-term performance assessment, and supervised clinical integration with human oversight are needed before routine clinical implementation. They added that LLMs could complement existing approaches to health checkup interpretation.  

"Rule-based algorithms remain indispensable for deterministic classification, but LLMs extend beyond categorical outputs by providing patient-friendly explanations that contextualize health information and offer actionable guidance," they wrote.  

No conflicts of interest were reported. 

 

AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.

Related Content