Effective medical AI oversight requires practical safeguards
Human oversight of medical artificial intelligence can function as a patient safety mechanism only when physicians have the knowledge, time, authority, and practical controls needed to evaluate and alter artificial intelligence–mediated care, according to a commentary published in npj Digital Medicine.
The authors proposed four operational conditions for meaningful oversight—epistemic capacity, cognitive space, decisional authority, and intervention effectiveness—and emphasized that governance should extend beyond reviewing artificial intelligence (AI) outputs to encompass the workflows in which predictive, generative, semi-autonomous, and agentic systems operate. Agentic systems may initiate actions such as outreach, scheduling, information retrieval, and message routing with limited real-time human input.
The commentary authors did not present original empirical data. Instead, they presented a governance framework that translates regulatory expectations for meaningful oversight of high-risk medical AI into operational questions for health systems.
"The central question is not whether a human remains in the loop, but whether the organization enables human judgment to prevent harm," wrote lead author Davy van de Sande, PhD, of the Department of Adult Intensive Care at Erasmus MC University Medical Center in Rotterdam, the Netherlands, and colleagues.
They contended that oversight should be treated as a system-level capability integrated throughout procurement, deployment, monitoring, and decommissioning, rather than as a responsibility assigned solely to frontline physicians.
The framework defined the following minimum conditions for determining whether a physician can recognize, question, and alter an AI-mediated clinical trajectory before avoidable harm occurs:
1. Epistemic capacity. Physicians should understand a model's intended use, validation population, limitations, uncertainty, failure modes, and how its performance varies across patient subgroups. Model cards may improve transparency, but are insufficient without training, onboarding, technical support, and institutional infrastructure.
2. Cognitive space. Clinical workflows should provide adequate time, attention, and context for independent assessment of AI recommendations. Compressed decision windows, competing alerts, and interfaces that make override slower than acceptance may push physicians toward automatic acceptance.
3. Decisional authority. Institutions should explicitly support physicians who question, override, decline, or escalate AI outputs and should protect justified departures from AI recommendations through policy. Organizational expectations, performance metrics, peer norms, and medicolegal concerns should not implicitly discourage disagreement.
4. Intervention effectiveness. AI-mediated recommendations or actions should be interruptible before harm propagates. Pause, correction, escalation, suspension, and deactivation functions should be usable, auditable, and available before downstream actions become difficult to reverse.
The authors also recommended matching responsibility to practical control over the AI system. Developers and vendors should document intended use, validation populations, known failure modes, update procedures, interface assumptions, and available override, pause, or stop functions. Procurement and governance teams should assess whether local validation is necessary, while clinical departments should evaluate workflow integration, define escalation pathways, and provide model-specific training. Health information technology teams should ensure override and suspension functions are usable and auditable, and quality and safety offices should monitor post-deployment performance and determine when systems require reconfiguration, restriction, suspension, or withdrawal.
To support ongoing governance, the authors suggested that a core oversight dashboard could monitor override frequency, review latency, agreement and disagreement between AI and physicians, subgroup-specific performance drift, patterned discordance across patient groups, and AI-associated adverse events or near misses. The authors also recommended scheduled review of subgroup-specific performance and patterns of discordance to identify inequitable performance or other concerning output patterns that may not be apparent during individual clinical encounters.
They cautioned that these metrics should prompt investigation rather than serve as standalone performance measures because a high override rate could reflect poor calibration, workflow mismatch, low user trust, or appropriate resistance in difficult cases. A low rate could indicate good performance, but it could also reflect automation bias—or undue influence from AI recommendations—time pressure, or weak authority to disagree.
The commentary authors also illustrated the framework using two deployment scenarios. For an inpatient deterioration or sepsis alert, users should understand the model’s intended use, limitations, uncertainty, and local performance; receive the alert when it can be appraised rather than reflexively accepted; retain authority to disagree when the output conflicts with the clinical picture; and have discordant cases or repeated overrides incorporated into safety review.
For generative or semi-agentic patient communication systems, oversight should include traceable generation and routing, clear responsibility for configuration rules, protected physician authority to edit, reject, or escalate content, and practical controls to pause, redirect, or suspend downstream actions.
Summarizing the framework’s central premise, van de Sande and colleagues wrote, “Human oversight of medical AI is meaningful only when the clinical, organizational, and technical environment enables users to understand, question, and interrupt AI-mediated care."
The authors reported no funding and declared no competing interests.
AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.