News Research Ethics, Regulation, and Responsible Use Research and Evidence

Commentary: Safeguards needed for AI-assisted guideline development

September 22, 2026 By Matthew Solan 6 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

Large language models could substantially accelerate evidence synthesis and clinical recommendation drafting, but their use in guideline development requires safeguards addressing evidence verification, uncertainty, human validation, and governance, according to a Matters Arising commentary in npj Digital Medicine. 

“The translational challenge, however, is not speed; it is trust,” wrote the authors Quan Zhang, PhD, with the department of economics, and Yuanyuan Fu, PhD, with the department of quantitative health sciences at the John A. Burns School of Medicine at the University of Hawaii at Manoa in Honolulu.  

Drs. Zhang and Fu proposed safeguards for the use of Quicker, an agentic large language model (LLM) system previously developed by Dubai Li and colleagues that automates an end-to-end workflow for evidence-based clinical recommendation development. Their recommendations build on limitations identified in Quicker’s published evaluation. 

For endocrinologists, these safeguards offer considerations for evaluating how AI-generated evidence synthesis and recommendations could be incorporated into guideline development while preserving human oversight. 

Quicker performs question decomposition, literature searches, study selection, evidence assessment, and recommendation drafting. It combines multistage prompting, an iterative literature-search agent connected to PubMed, and retrieval-augmented generation (RAG), and applies the Grading of Recommendations Assessment, Development and Evaluation (GRADE) framework during evidence assessment. 

In system-level testing reported by Li and colleagues, a single reviewer working with Quicker developed each recommendation in 20 to 40 minutes. Its accompanying Q2CRBench-3 benchmark included 26,198 retrieved records from three guideline-development datasets, with 99.5% excluded during screening.  

Although the findings suggest that evidence synthesis and draft recommendations could be generated more rapidly, Drs. Zhang and Fu highlighted performance limitations that could affect the system's use in guideline development. Record-screening sensitivity fell to approximately 0.79 in one dataset using the basic method. Numerical data extraction achieved approximately 72% precision, and risk-of-bias assessments differed substantially from guideline assessments. The system also overlooked selective reporting, had difficulty assessing bias related to missing data, and may have over-downgraded evidence in some cases. 

Making AI outputs verifiable 

To address these limitations, Drs. Zhang and Fu proposed that AI-generated recommendations be accompanied by verifiable evidence bundles rather than presented only as free text. Each bundle would include citations to the underlying studies; a structured evidence table documenting PICO elements, study design, outcome-specific effect estimates, a risk-of-bias or certainty rating, and key limitations, with links to the passages used for extraction; explicit links between evidence and statements in the rationale; and information defining the scope and limitations of the literature search. 

The authors also recommended distinguishing GRADE certainty for each outcome from the model's confidence in its own extraction and synthesis. Because these measures can diverge, decisions to defer an output should be triggered by low model confidence rather than evidence certainty alone, with the confidence signal itself validated. Deferral thresholds should be explicit and calibrated against clinical consequences. “The goal is calibrated caution, not blanket conservatism,” the authors wrote. Excessive deferral could withhold well-supported recommendations, while evaluation should also track recommendations issued despite weak evidence.  

Accounting for study design 

The authors also called for evidence assessment tailored to study design. Quicker's risk-of-bias assessment was built and validated around randomized controlled trials using the Risk of Bias 2 tool. However, guideline recommendations may also depend on observational and other nonrandomized evidence, particularly for harms, rare conditions, and settings in which randomized trials are unavailable or unethical. Drs. Zhang and Fu noted that Quicker misclassified an open-label pilot randomized controlled trial as an observational study, illustrating the importance of explicitly identifying study design. 

They also recommended using design-appropriate risk-of-bias tools for nonrandomized studies and an internally consistent GRADE approach to rating certainty. They suggested flagging recommendations that depend primarily on observational evidence so randomized and observational findings are not treated interchangeably. 

Keeping humans in the loop 

Human validation should focus on stages at which errors could have the greatest effect, according to the authors. Their proposed protocol includes checking a sample of borderline eligibility decisions, verifying primary outcomes and key effect estimates from high-impact trials, reviewing risk-of-bias judgments for selective reporting and missing data, and confirming that recommendations are consistent with the evidence linked to them. Risk-stratified sampling could preserve efficiency without requiring comprehensive re-review, although the authors cautioned that the resulting time savings would be smaller than the reported 20- to 40-minute time for formulating individual recommendations suggests.  

Drs. Zhang and Fu also called for broader prospective validation before widespread deployment. Although Q2CRBench-3 included multiple diseases, Quicker's system-level evaluation used a single reviewer and randomized controlled trial evidence. They proposed testing the approach with guideline committees across specialties and settings, assessing outcomes such as screening sensitivity, extraction errors for key outcomes, recommendation stability across model versions, and subgroup coverage in the underlying evidence. 

Setting quality and governance standards 

Drs. Zhang and Fu proposed evaluating and governing Quicker as a modular evidence-synthesis pipeline with stage-specific quality thresholds. These could include recall floors for screening, extraction-fidelity checks, and consistency checks between evidence tables and drafted rationales, with fail-closed behavior when thresholds are not met. The authors also recommended continued monitoring and updating after deployment. 

Governance recommendations included labeling AI-assisted drafts with the model and version used, maintaining structured audit logs, and preventing unreviewed AI-generated guideline text from being treated as evidence or reused as training data. 

Drs. Zhang and Fu acknowledged limitations to their proposed safeguards. They did not have access to Quicker, did not empirically test the proposed safeguards, and did not present a worked demonstration. The proposals could also increase reviewer workload, computational overhead, and potential latency. The authors therefore presented the safeguards as hypotheses for implementation and study rather than validated requirements. 

Ultimately, Drs. Zhang and Fu emphasized that certainty of evidence is only one input into a clinical recommendation. The move from evidence to recommendation still requires human judgment about benefits and harms, patient values, resource use, equity, and feasibility. 

“Establishing trustworthy pipelines is therefore necessary but not sufficient for warranted trust and safe uptake, and overtrust in fluent yet unverified outputs remains a distinct risk that governance and training must actively manage,” the authors wrote. 

Drs. Zhang and Fu declared no competing interests. 

AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.

Related Content