ORIGINAL RESEARCH article
Front. Cell Dev. Biol.
Sec. Cancer Cell Biology
Machine Learning Models for Predicting Liver Cancer: A Real-World Cohort Study in China
- CF
Cexiong Fu 1
- FL
Fang Li 1
- SL
Shibing Li 1
- RL
Ruifeng Lin 2
- ZY
Zhao Yuan 3
1. Hainan General Hospital, Haikou, China
2. Southern Medical University, Guangzhou, China
3. The Second Affiliated Hospital of Guilin Medical University, Guilin, China
Select one of your emails
You have multiple emails registered with Frontiers:
Notify me on publication
Please enter your email address:
If you already have an account, please login
You don't have a Frontiers account ? You can register here
Abstract
Background: Liver cancer has high incidence and mortality worldwide, and timely identification is important for improving prognosis. However, prediction models based on routinely available clinical data remain insufficiently evaluated in hospitalized real-world populations. This study aimed to develop and interpret a machine-learning model for liver cancer prediction using multidimensional clinical and laboratory features. Methods: This retrospective real-world cohort included 9,284 hospitalized patients from Hainan General Hospital between January 2020 and December 2023. Patients were divided into a training set (n=6,426) and an internal validation set (n=2,858). Seventy-four demographic, clinical, and laboratory variables were collected. Feature selection was performed using least absolute shrinkage and selection operator regression in the training set. Six models were developed: logistic regression, random forest, support vector machine, k-nearest neighbors, extreme gradient boosting (XGBoost), and Elastic Net. Discrimination was assessed primarily using the area under the receiver operating characteristic curve (AUC). Decision curve analysis evaluated clinical net benefit, and SHapley Additive exPlanations (SHAP) were used for model interpretation. Results: Liver cancer accounted for 48.7% of the training cohort and 48.8% of the validation cohort. LASSO retained 57 nonzero model terms at the minimum-error penalty. XGBoost achieved the highest AUC among all models, with an AUC of 0.909 (95% CI, 0.903–0.915) in the training set and 0.802 (95% CI, 0.788–0.817) in the validation set. At a threshold of 0.5, validation accuracy, sensitivity, specificity, and F1 score were 0.730, 0.758, 0.698, and 0.753, respectively. Decision curve analysis showed that XGBoost provided the greatest net clinical benefit across a broad range of threshold probabilities. SHAP analysis identified serum sialic acid, the aspartate aminotransferase-to-alanine aminotransferase ratio, alkaline phosphatase, basophil percentage, and monocyte-to-lymphocyte ratio as the leading predictors. Conclusions: XGBoost showed good discrimination and interpretability for liver cancer prediction in a hospitalized real-world cohort using routinely available clinical and laboratory data. The findings support the feasibility of leveraging large-scale inpatient laboratory data for risk-stratification model development. Serum sialic acid, the aspartate aminotransferase-to-alanine aminotransferase ratio, and alkaline phosphatase were important predictors. Prospective validation in outpatient, first-visit, high-risk, and external populations is required before broader clinical or screening application.
Summary
Keywords
auxiliary diagnosis, liver cancer, machine learning, predictive model, XGBoost
Received
01 June 2026
Accepted
13 July 2026
Copyright
© 2026 Fu, Li, Li, Lin and Yuan. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Fang Li
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.