Wearable AI may offer a new window into metabolic risk
In a study published in Nature, Ahmed Metwally, PhD, who leads Google's Metabolic Health AI research program, and his colleagues explored whether AI could use data collected through consumer-wearable data, together with routine clinical information, to predict insulin resistance.
In the remotely conducted WEAR-ME study of 1,165 US adults, AI models combined wearable measures—including activity, sleep, heart rate, and heart rate variability—with demographics and routine blood biomarkers. Using a homeostatic model assessment of insulin resistance (HOMA-IR) threshold of 2.9 to define insulin resistance, the multimodal model achieved an area under the receiver operating characteristic curve (AUROC) of 0.80, with 76% sensitivity and 84% specificity.
The researchers also fine-tuned a wearable foundation model pretrained on 40 million hours of sensor data to analyze time-series wearable data. In an independent validation cohort of 72 participants, adding representations from that model to demographic data, fasting glucose, and a lipid panel increased the AUROC from 0.76 to 0.88. The team also developed a large language model–based agent to place insulin-resistance predictions in the context of broader metabolic health and offer personalized recommendations.
The researchers concluded that this approach could provide a scalable framework for detecting metabolic risk and, with further validation, could support timely lifestyle interventions to prevent progression to type 2 diabetes.
In part two of his conversation with AACE Endocrine AI Editor-in-Chief Johnson Thomas, MD, FSCE, FEAA, Dr. Metwally discussed how wearable-based predictions might be used and the role of remote studies in metabolic research.
(The following interview transcript has been edited for length and clarity.)
Dr. Thomas: The study combined wearable data, demographic information, and laboratory data. The AUROC suggests it could be a fairly good screening tool. So, two questions: How do you envision bringing this tool into the clinic and into mainstream use? And how do you decide who should be flagged for further clinical workup?
Dr. Metwally: Those are amazing questions. I think it’s really about how this could be deployed. Google announced plans to make this a product. You would have it on the watch, and if you have enough wearable and demographic data, you would receive a prediction about insulin resistance.
Your question about who should be flagged is also very important. These are screening tools. They are meant to give people information they can use to consider lifestyle changes. If someone wants to proactively check their blood biomarkers, they can certainly do that. But the main purpose is education: letting people know which direction their health may be trending and what lifestyle changes they can make to change it.
Let’s talk about performance first. You don’t want a screening tool that sends people to the clinic all the time when they don’t understand why. It needs a certain level of specificity, and the message shown to the user needs to be delivered with the right level of urgency. Maybe people need repeated classifications showing the same trend, particularly if they are not changing their lifestyle.
One thing I’ve learned from conducting these studies is that a prediction can be a great way to help people understand their inner state—to know whether they may be on the path to developing a disease before they have symptoms or reach a point at which the condition may be more difficult to reverse.
But one of the most important questions for us at Google, and for anyone working in this area, is: If you have this type of information and know the right thing to do, are you doing is? Are you changing your lifestyle? I think that is a different problem.
Detection is one thing. Knowing what to do with the information is another. I feel that’s the gap in digital health right now. It’s not about the prediction or the knowledge. It’s about helping people change their lifestyle. It’s a problem I don’t personally have a solution for yet. We need more studies, particularly behavioral studies, to know how to encourage people to make changes based on where they are in their health trajectory.
Dr. Thomas: Another fascinating aspect of the study was that it was conducted remotely. I’m used to conducting research studies at my institution: Patients come to us, we do blood tests and other examinations and imaging, and then they go home. But this study was entirely remote, from what I understand. Do you see remote data collection as the future of metabolic research? What can remote studies accomplish that traditional clinic-based studies cannot?
Dr. Metwally: I think remote studies are definitely part of the future, but that doesn’t mean everything will be remote. Basic science and early discovery need to happen in a setting where you see participants and communicate with them. I think those interactions are a major component of identifying where a study could fail and what things need to be automated and streamlined.
I learned a lot from conducting studies at the University of Illinois Chicago and Stanford University through those interactions, because we learned where to optimize the protocols. For example, when we developed the protocol for using CGM at home, we drew on what we had learned about how people used it in the clinic and what could go wrong. Then we put checks in place for when participants were at home.
With remote studies all over the US, we also learn how things can fail. We have processes to check data completeness and quality. You don’t want to finish a study only to discover that 10% of the data are usable. You want to make sure that at least 90% are useful.
Remote studies can help validate generalization. They can also bring in enough participants to train certain models. For the wearable foundation models we developed, a small sample wouldn’t be enough to train the models. A simpler model might give better performance, but that performance would not be sufficient to share.
So, it depends on the intended outcome of the study. What do we want to test? In the field, some findings have been established and reproduced by multiple labs with smaller sample sizes. That’s when it’s time to go big.
For CGM, I think now is the time for nationwide studies. Multiple people have validated it. I’ve done it at Stanford. I’m sure two or three groups around the world have done similar studies. They may not have used exactly the same models or settings, but you can connect the dots and see similar trends. It works here; it doesn’t work there. Once you see that, it’s time to scale up. Those are the research questions that are good candidates for remote studies.