Research Ethics and Policy Research and Evidence Ethics, Regulation, and Responsible Use Precision Endocrinology

Where does AI belong in evidence-based medicine?

July 28, 2026 By Matthew Solan 5 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

Large language models can accelerate literature retrieval and synthesis, but their outputs remain information rather than evidence until physicians and methodological experts evaluate their quality, applicability, and relevance to patient care, according to a perspective published in npj Digital Medicine.

"The LLM era will not weaken EBM (evidence-based medicine); it will test whether the institutions and clinicians who depend on it can hold the line between information and evidence precisely when speed and volume most tempt them to blur it," the authors wrote. 

They noted that medicine already faces an overwhelming volume of biomedical literature, creating increasing delays between publication, evidence synthesis, and implementation in practice. The authors added that LLMs may help clinicians manage information overload by rapidly retrieving, summarizing, and tailoring large volumes of literature, but they also risk increasing the production of poorly supported or insufficiently validated information. 

To clarify the appropriate role of AI, they distinguished four stages of knowledge development: 

  • Data: Raw observations, including laboratory values, imaging findings, and clinical measurements. 

  • Information: Organized and contextualized data, such as diagnostic reports or study findings. 

  • Evidence: Information that has undergone methodological evaluation and quality assessment. 

  • Practice: Clinical action that integrates appraised evidence with physician judgment, patient values, risk tolerance, and the local care context. 

For clinical implementation, the authors proposed three structural safeguards. LLMs should be restricted to tasks such as retrieving, reorganizing, and presenting existing information. Methodological quality assessment, applicability judgments, and contextual synthesis should remain operationally enforced human functions, and the complete human–AI workflow, not an isolated model response, should serve as the auditable unit. 

The perspective also cited studies evaluating LLM performance in evidence-synthesis tasks. Reported capabilities included: 

  • Full-text screening approaching human performance when relevant studies substantially outnumber irrelevant studies. 

  • Structured data extraction exceeding 98% accuracy for prespecified information extracted from supplied text. 

  • Population, Intervention, Comparator, Outcome (PICO) element recognition reaching approximately 80%, with additional improvement using knowledge-guided prompting. 

  • End-to-end systems capable of performing question decomposition, literature retrieval, screening, data extraction, and draft recommendation generation while substantially reducing the person-hours required for systematic evidence synthesis. 

The authors cautioned that these performance measures do not indicate that LLM outputs constitute evidence suitable for clinical decision-making. Instead, they identified several recurring limitations. 

One limitation is that LLM performance declines when evidence-synthesis tasks involve genuinely contested clinical questions. The authors cited studies demonstrating that full-text screening performance deteriorates substantially when inclusion and exclusion decisions are evenly balanced.  

At the appraisal stage, studies cited in the perspective reported domain-level accuracy of 57% to 70% for LLM-generated risk-of-bias assessments compared with expert Cochrane judgments. Although LLMs can reproduce the language used in risk-of-bias and GRADE (Grading of Recommendations Assessment, Development, and Evaluation) assessments, the authors said they do not perform the evaluative reasoning required to establish evidentiary quality.  

They identified three broader failure modes. LLMs may produce information without underlying data, including fabricated or unsupported statements and citations; amplify data without meaningful information by contributing repetitive or minimally novel publications; or provide information without verification by summarizing genuine sources without exposing a clear evidence chain or adequately communicating uncertainty. Together, these shortcomings can obscure the evidential status of AI-generated output, making it difficult for readers to determine whether a response is supported, verified, or sufficiently reliable for clinical use. 

In addition, the perspective discussed retrieval-augmented generation (RAG) systems, which ground responses in curated evidence sources rather than model memory. In one cited evaluation, these systems produced actionable answers for 48% of clinical queries for which evidence existed, compared with under 5% for general-purpose LLMs assessed under equivalent conditions. The authors cautioned that retrieving evidence is not equivalent to generating evidence because the quality, consistency, indirectness, and applicability of the retrieved material still require appraisal. 

Beyond accelerating existing evidence workflows, the authors suggested LLMs may help identify areas where evidence remains incomplete. Rather than presenting definitive answers, AI systems could map evidence gaps, retrieve scattered publications, integrate mechanistic findings to generate plausible hypotheses, and organize information about patient preferences that has not yet been incorporated into formal clinical guidelines.  

The authors also argued that increasing automation will concentrate physician effort on appraisal, adjudication, contextualization, and accountable final judgment rather than initial retrieval and synthesis. Drawing on the "centaur" model of human–machine collaboration, they proposed that LLMs could perform rapid information synthesis while physicians direct the system, recognize unreliable output, and retain responsibility for evidentiary and clinical judgments. 

The perspective had several limitations. The authors noted that the article was a conceptual perspective rather than a systematic review, that evidence on LLM performance continues to evolve rapidly, and that regulatory, legal, reimbursement, and implementation issues were outside its scope.

Pearse A. Keane, MD, disclosed he is a cofounder of Cascader Ltd. and reported consulting, equity ownership, speaker fees, travel support, and advisory board participation with multiple ophthalmology and pharmaceutical companies. Tien Yin Wong reported consulting relationships, patents, and cofoundership of companies developing digital solutions for eye diseases. The remaining authors declared no competing interests.  

(Editor’s Note: The article will undergo further editing before final publication, and there may be errors present that affect the content. All legal disclaimers apply.) 

AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.

Related Content