News Ethics and Policy Diagnostics & Imaging Ethics, Regulation, and Responsible Use

Review examines AI’s role in liver disease and cancer

September 24, 2026 By Matthew Solan 6 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

Artificial intelligence is being applied in liver pathology to support the diagnosis of liver diseases and liver cancer and the assessment and management of transplantation, but challenges remain as these tools move toward clinical use, according to a review published in The Lancet Digital Health.  

“Additional work is needed to evaluate the real-world usefulness, clinical safety, and successful deployment of AI tools in liver pathology,” wrote Robert Goldin, MD, of the Division of Digestive Diseases, at St. Mary’s Hospital, London, and colleagues.

They reviewed developments in artificial intelligence (AI), digital pathology, and computational image analysis in liver diseases, including applications in metabolic dysfunction-associated steatotic liver disease (MASLD), liver cancer, and transplantation. They began searching PubMed and EMBASE on August 1, 2022, using terms related to digital pathology, histopathology, AI, deep learning, and machine learning, supplemented by updated searches, article alerts, and the authors’ reference collections. The reference list was finalized on October 3, 2025. 

AI in MASLD and fibrosis 

Among the AI approaches reviewed, one machine learning–based automated image analysis system for steatotic liver disease was developed in 100 patients and validated in 146. In the development cohort, interclass correlation coefficients between pathologist and machine assessments were 0.97 for steatosis, 0.96 for lobular inflammation, 0.94 for ballooning, and 0.92 for fibrosis. In a subgroup with paired liver biopsies, quantitative analysis was more sensitive than manual semiquantitative scoring for detecting differences. 

Another study used data from three randomized controlled trials involving 3,320 patients to develop and validate a deep convolutional neural network that identified and quantified components of the nonalcoholic steatohepatitis activity score and fibrosis. Agreement between the model and pathologists was within the range observed among expert pathologists. The Deep Learning Treatment Assessment Liver Fibrosis score also identified antifibrotic treatment effects not detected by routine histologic assessment and correlated with clinical progression. The review additionally cited a 2025 study showing noninferiority of an AI tool compared with manual scoring of MASH features. 

AI has also been combined with newer tissue-imaging techniques to characterize fibrosis. The second harmonic generation B-index, which was developed using AI, is a continuous measure derived from 14 collagen properties. An index greater than 1.76 had 99% diagnostic accuracy and 88% sensitivity for identifying bridging fibrosis.  

In another application, adding an AI image-analysis module to the assessment of 120 liver biopsy samples improved agreement in fibrosis assessment from 89% to 93%. Machine learning has also been used to distinguish fibrosis patterns. In a study of 152 whole-slide images from patients with MASH, areas under the receiver operating characteristic curves were 79% for detection of periportal fibrosis, 83% for pericellular fibrosis, and 86% for portal fibrosis, and exceeded 90% for detection of normal fibrosis, bridging fibrosis, and the presence of nodules or cirrhosis. 

AI beyond metabolic liver disease 

AI applications in liver cancer include tumor detection and classification, prognosis, treatment-response prediction, and prediction of molecular status. In a prospective observational study cited in the review, neural network models improved pathologists’ accuracy in distinguishing hepatocellular carcinoma from cholangiocarcinoma. Less experienced pathologists benefited more from AI assistance but were also more prone to errors when AI predictions were incorrect, which the authors presented as an example of automation bias and overreliance on medical AI. However, they noted that the study used small tumor micrographs rather than whole-slide images, and cases were prescreened to exclude more complex examples. 

A deep learning model developed to predict hepatocellular carcinoma within 7 years after a steatotic liver biopsy achieved accuracies ranging from 81% to 82%, with areas under the curve ranging from 0.80 to 0.84. Other deep learning approaches have been used to predict survival after hepatocellular carcinoma resection and assess microvascular invasion. 

AI applied to hematoxylin and eosin–stained sections has also been investigated as a biomarker for sensitivity to treatment. The authors cited a multicenter study in which an AI model applied to hepatocellular carcinoma biopsies acted as a biomarker for progression-free survival among patients treated with atezolizumab plus bevacizumab. 

In liver transplantation, deep learning has been studied for steatosis detection and tissue segmentation. One model achieved 97% accuracy for steatosis detection in liver transplant samples, although the review authors noted that evaluation in larger datasets is needed. Another deep learning model segmented portal tract regions with an accuracy of 0.89 in liver tissue from patients who had received a liver transplant. 

Foundation and vision-language models represent another emerging AI approach. However, the authors noted that datasets used to train foundation models often contain relatively few liver disease images and even fewer images of non-neoplastic liver disease, highlighting the need for further evaluation of their accuracy and effectiveness in liver pathology. 

Barriers to clinical use 

The authors identified several barriers to translating AI from research into clinical liver pathology. Models require large, diverse, and representative datasets to be robust in clinical settings and avoid bias that could adversely affect real-world accuracy, but high-quality, accessible datasets remain limited, particularly for non-neoplastic liver disease. Scanner vendors, viewing platforms, laboratory practices, staining variation, and image quality can also affect model performance. 

AI systems may also perform worse on real-world data than during experimental evaluation, and tissue contamination artifacts have been shown to reduce diagnostic performance. The authors cautioned that AI can introduce errors into the diagnostic pathway and called for model evaluation, professional oversight, and monitoring before clinical adoption. They also pointed to methodological flaws and bias in published AI studies and highlighted the development of AI-specific reporting guidelines intended to improve research quality. 

Two of the authors reporting being funded through the UKRI-supported NPIC program; industry partners did not fund or contribute to the review. Other authors disclosed relevant unpaid roles, patents, employment, or consultancy relationships.

 

AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.

Related Content