Research News Ethics, Regulation, and Responsible Use Research and Evidence

Researchers call for new safeguards in LLM research

July 22, 2026 By Matthew Solan 4 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

Public biosignals datasets should better support large language model research through updated informed consent, privacy-preserving data preparation, and institutional governance reforms, according to a policy article published in npj Digital Medicine.  

Researchers from the NIH Clinical Center in Bethesda, Maryland, and the University of Oxford in the United Kingdom reported that current restrictions on public biosignals datasets may hinder research needed to evaluate the performance and safety of large language model (LLM)-based systems before clinical deployment. 

"To reduce risk and protect patients, careful research must be performed on public datasets prior to deployment in downstream private settings," wrote lead author James Anibal of the Center for Interventional Oncology at the NIH Clinical Center.  

The biosignals highlighted included electrocardiography, electroencephalography, photoplethysmography, voice recordings, and wearable sensor data. These data increasingly underpin digital health applications but may also reveal sensitive information about patients’ health, activity, mental state, and daily habits.

The authors argued that publicly accessible biosignal repositories are necessary to responsibly develop, evaluate, and benchmark LLMs intended for clinical use. They noted that public datasets have historically enabled investigators to identify bias, benchmark performance, and improve safety across multiple artificial intelligence (AI) applications.  

However, they added that many existing biosignal repositories operate under data use agreements that restrict interactions with commercial LLMs through application programming interfaces, limiting researchers to locally deployed models or specialized private cloud environments that may not be widely accessible. 

According to the researchers, these restrictions make it difficult to conduct realistic academic studies of commercial LLMs that physicians already use in clinical practice. They cited multiple published surveys demonstrating increasing adoption of AI. For example, one survey found that more than 75% of physicians reported informally using general-purpose LLMs during clinical decision-making, while an international survey reported that 76% of health care professionals had used ChatGPT at least once for professional purposes. 

They further argued that restrictive agreements may be difficult to enforce, noting that researchers may use LLMs during initial data exploration or organization without explicitly reporting those activities in published methods, making such use difficult to verify. 

To address these concerns, the researchers proposed three recommendations for biosignals that are currently available online or under review for public release.  

1. Informed consent processes must explicitly cover LLMs.  

Rather than relying on broad consent language for unspecified future research, patients should receive information about possible LLM applications, biomarkers that could be inferred from the data, privacy risks, irreversible aspects of model training, and uncertainties surrounding future model capabilities before deciding whether to contribute data. Digital consent systems would allow patients to update or withdraw consent over time. 

2. Data preparation should be an institutional research focus.  

Stronger de-identification methods—including automated removal of identifiable information, voice anonymization, and adaptation of emerging LLM-based tools—could improve patient privacy protections while supporting broader research access. Greater institutional emphasis should be placed on privacy-aware preprocessing before public release rather than relying primarily on restrictive agreements after datasets have already been released. 

3. Researchers should not assume liability for organizational data security.  

Current agreements often require researchers to implement "state-of-the-art" security measures without clearly defining those expectations, potentially discouraging academic innovation through uncertain liability. The researchers recommended adopting standardized institutional data management guidelines to allocate responsibility more appropriately and protect researchers. 

The authors acknowledged several important considerations and boundaries to their recommendations. They noted that not all biosignal datasets may be suitable for public release, citing rare disease cohorts with records that may be easier to reidentify and datasets from vulnerable populations for whom reidentification could result in stigma, discrimination, or deterrence from care.  

As a policy article rather than an empirical investigation, the report did not include new experimental data, patient outcomes, LLM performance metrics, or comparative model evaluations. It also did not prospectively assess whether the proposed framework would improve research quality, patient privacy, or clinical outcomes.  

The competing interests statement lists multiple cooperative research and development agreements, equipment and research support, patent licensing arrangements, and intellectual property relationships involving Bradford J. Wood, MD. No competing interests were listed for the other authors.

(Editor’s Note: The publisher stated that this is an unedited article-in-press version that will undergo further editing before final publication and may contain errors that affect the content. All legal disclaimers apply.) 

 

AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.

Related Content