Peer-reviewed

Study reports AI model separates malignant and benign laryngeal lesions

The retrospective evaluation found strong internal performance and attention aligned with physician-marked regions, but it did not test clinical benefit.

An artificial intelligence model reported a mean area under the curve (AUC) of 0.975 when it distinguished malignant from benign laryngeal lesions in a retrospective evaluation of white light endoscopy images. Its reported accuracy was 0.915, sensitivity was 0.876 and specificity was 0.945. AUC is a summary measure of how well a model separates two groups, while sensitivity and specificity describe its performance for the malignant and benign categories. The reported 95% confidence intervals ranged from 0.959 to 0.991 for AUC, 0.883 to 0.947 for accuracy, 0.803 to 0.949 for sensitivity and 0.910 to 0.980 for specificity.

The result came from IMIL-Net, a model designed to assess a patient-level set of endoscopy images rather than a single preselected picture. It used a Swin Transformer encoder and gated attention pooling to produce a patient-level diagnosis while also assigning an attention score to each image. The study's rationale was that clinical assessment works across multiple images, whereas the existing AI approach described in the supplied text often relies on single, pre-selected images and can be difficult to interpret.

A model built around the full examination

The retrospective cohort contained 611 patients with malignant or benign laryngeal lesions who underwent white light endoscopy examinations. The model was evaluated with five-fold cross-validation and compared with four baseline models. Cross-validation means the available data are repeatedly divided into parts used for fitting and checking the model, giving an internal estimate of performance. That design helps test the model within the assembled cohort, but it is not the same as testing it prospectively in routine care or on an independent cohort.

Strong numbers inside the study

Within that comparison, IMIL-Net was reported as superior to all the listed baselines. The non-MIL architectures had AUCs ranging from 0.888 to 0.897, while logistic regression using clinical data alone had an AUC of 0.898. These figures show that the proposed model scored better than the selected computational comparators in this evaluation. They do not establish better performance than physicians, because a clinician benchmark was not reported in the supplied analysis.

A promising explanation check, with clear limits

The researchers also examined whether the model's attention scores landed on areas that physicians had marked as relevant. In an interpretability-validation subset of 30 cases, the physician-annotated region-of-interest images had a median attention score of 20.94%, compared with 15.32% for non-ROI images. The image counts were 63 and 78 respectively, and the reported p-value was below one millionth. This is evidence of an association between the model's weighting and the marked regions in that subset, not proof that the scores are clinically valid explanations or that they improve diagnosis or biopsy decisions.

That distinction matters because the attention check was small and targeted. The 30-case subset, including 63 ROI images and 78 non-ROI images, was an interpretability test rather than a measure of patient benefit. The study did not report changes in clinician decisions, repeat-biopsy rates, patient burden or survival. Nor did the supplied analysis report prospective clinical deployment or external validation across institutions, endoscopy systems or patient populations.

What the result does not answer

Other details needed to judge how broadly the result applies were also absent from the supplied analysis. It did not report the malignant and benign class counts, patient characteristics, inclusion criteria or the number of images per patient. The identities of the four baseline models and detailed performance results beyond their reported AUC values were not provided either. Those gaps make it harder to judge cohort balance, inclusion decisions and whether image or patient-level data leakage was prevented.

The authors characterize IMIL-Net as a high-accuracy, interpretable, clinically aligned and trustworthy solution for integration into otolaryngology workflow. The supplied evidence supports a narrower news finding: the model showed strong reported discrimination in an internal retrospective evaluation, and its attention scores were higher in physician-marked regions. It does not demonstrate clinical effectiveness, improved health outcomes or readiness for routine use. The article's publication record lists it as a version-of-record journal article published on 20 August 2026, and the authors declared no conflicts of interest.

Paper data and sources

Original title: An interpretable multi-instance learning method for accurate differentiation of malignant and benign laryngeal lesions in laryngoscopy.
Authors: Biao Xu, Miao Zhang, Shuai Jiang et al.
Journal/Repository: European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery
Status: Peer-reviewed
First online: 2026-08-20
DOI: 10.1007/s00405-026-10540-1
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.