Peer-reviewed

Machine-learning model shows promise for liver fibrosis screening

A model tested in 6,158 people previously infected with Schistosoma japonicum outperformed standard scores, but it still lacks testing outside one Chinese province.

A machine-learning model identified liver fibrosis more accurately than standard scoring tools in people who had previously been infected with Schistosoma japonicum, according to a study in Jiangsu Province, China. On data set aside for testing, the model’s area under the curve, or AUC, was 0.854. AUC is a measure of how well a test separates people with and without a condition, with higher values indicating better discrimination.

The finding comes from a cross-sectional analysis of 2021–2022 follow-up data from a prospective Jiangsu cohort. The 6,158 participants had documented previous infection, had completed standardized praziquantel therapy and agreed to physical and ultrasound examinations. The study’s stated aim was to identify fibrosis in this group using epidemiological and laboratory data.

A clear gap in validation performance

XGBoost, which had the highest reported AUC among the machine-learning approaches tested, recorded a validation AUC of 0.854, with a 95% confidence interval from 0.826 to 0.881. The other machine-learning models scored between 0.609 and 0.776. The traditional FIB-4 and APRI indices recorded AUCs of 0.514 and 0.516. Comparisons between XGBoost and both traditional indices were statistically significant, with P < 0.001.

At the selected threshold, the validation model had an accuracy of 0.818, meaning it classified about eight in 10 cases correctly. Its sensitivity, the share of people with fibrosis that it identified, was 0.853, while its specificity, the share without fibrosis that it correctly ruled out, was 0.736. The positive predictive value, the share of positive predictions that were correct, was 0.884. The study also reported intervals around these estimates.

The model’s AUC was higher in the training data, at 0.928, than in the held-out validation data, at 0.854. The validation set was kept separate for the final assessment, giving the researchers a new group on which to check the model after it had been developed.

Built from study records

Liver fibrosis was assessed by ultrasound and graded from 0 to III. The researchers treated Grade 0 as no fibrosis and Grades I through III as fibrosis. Radiologists were blinded to the other information, and a third hepatologist settled disagreements.

The cohort had a median age of 64 years, with an interquartile range of 59 to 68 years. Men made up 3,518 participants, or 57.1% of the group, and 4,314 participants, or 70.1%, were classified as having liver fibrosis. The data were divided into training and validation sets in a 7:3 ratio.

The researchers first used LASSO logistic regression, a method that narrows a large list of possible inputs, to identify 18 candidate predictors. They then checked whether some variables were too closely related to one another. Five machine-learning models were evaluated, with parameter tuning and tenfold cross-validation carried out on the training data. Synthetic minority over-sampling was used inside that process to address imbalance between outcome groups.

What the model was using

An analysis designed to show which inputs most influenced the model placed alcohol intake first, followed by gamma-glutamyl transferase and triglycerides. Their mean SHAP values were 0.401, 0.295 and 0.292, respectively. SHAP values describe how much a feature contributes to a model’s output; they do not show that any of these factors causes fibrosis.

A decision-curve analysis, which estimates whether using a prediction tool could offer more benefit across a range of risk thresholds, found positive net benefit for XGBoost up to approximately a 95% threshold. The traditional indices had limited benefit beyond 25%. This is a modeled measure of clinical utility, not evidence of an observed improvement in patients’ health.

A promising result with a narrow test

The study does not establish that XGBoost reduces fibrosis, improves treatment or changes outcomes. Its cross-sectional design can identify associations but cannot establish cause and effect. The authors also note that the sample came from a single province, that retrospective epidemiological surveys may introduce recall bias and that liver status before the original infection was unavailable, so pre-existing fibrosis could not be fully excluded.

The single-province sample leaves the model’s wider usefulness unresolved. The study did not report external validation, so it remains unclear whether the result would hold in other endemic provinces or healthcare settings. The findings therefore describe diagnostic discrimination in this cohort rather than evidence that the algorithm is ready for routine bedside use.

The paper is a peer-reviewed, open-access research article. The study reports support from the National Natural Science Foundation of China, the Jiangsu Provincial Health Commission and the Wuxi Municipal Health Commission.

Paper data and sources

Original title: Development and validation of a machine learning-based model for identifying liver fibrosis in individuals with prior Schistosoma japonicum infection: a step toward precision management.
Authors: Tao Wang, Yi Jiang, JianFeng Zhang et al.
Journal/Repository: Infectious diseases of poverty
Status: Peer-reviewed
First online: 2026-08-20
DOI: 10.1186/s40249-026-01485-y
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.