Preprint

AI System Shows Strong Agreement in Lung Cancer Pathology Tests

Preprint: LUCAID combined nine analysis tools and matched an expert-panel reference in 93% of prospective clinical decisions, but routine clinical benefit remains unestablished.

An artificial-intelligence system designed to analyze lung-cancer tissue matched a jointly agreed expert reference in 93% of prospective clinical decisions, according to a preprint. Individual pathologists in the same assessment had overall agreement rates ranging from 68.3% to 81.1%. That agreement is not evidence of improved patient outcomes or clinical benefit, and the system's influence on routine care remains unestablished.

The system, called LUCAID, is an agentic pathology pipeline. It directs nine AI modules through tasks including slide quality control, tumor detection and segmentation, tumor subtyping, analysis of the tumor microenvironment, tumor-cellularity estimates, immunohistochemistry cell phenotyping, biomarker scoring and structured report generation.

A test of decisions pathologists make

The prospective part of the study included 70 consecutive lung-cancer cases from the nNGM cohort. Five experienced pathologists and LUCAID completed five actionable tasks, producing 328 assessments. A joint expert panel set the reference after the cases were reassessed following a washout period of at least three months.

LUCAID reached 89.9% concordance for PD-L1, 92.5% for MET, 95.2% for membranous TROP-2 and 88.7% for cytoplasmic TROP-2. The corresponding ranges for individual pathologists were 70.1% to 76.8%, 56.7% to 75.8%, 48.4% to 93.5% and 50.0% to 74.2%. At least one deviation from the panel reference occurred in 67 of the 70 cases, and pathologists disagreed with one another in 163 of 328 assessments.

Strong technical agreement across the pipeline

The broader evaluation combined a multicentre development cohort of 1,620 cases, a discovery cohort of 1,001 patients, 105,227 expert annotations and 115 cases used for molecular validation. The development material covered ten staining modalities, allowing the system to be tested across several forms of tissue imaging.

Across its analysis modules, the study reported F1 scores, the benchmark used to compare their classifications with expert annotations, from 0.82 to 0.95. The reported range indicated high agreement in these benchmarked tasks, although the supplied analysis did not give confidence intervals for the summary.

One test compared estimates of tumor cellularity, the estimated proportion of tumor cells in a sample, against KRAS variant allele frequency, a molecular reference. In the 115 KRAS-mutated cases, the approach based on the area occupied by cell nuclei had a reported correlation of 0.46 and a mean absolute error of 17.1 percentage points. That error was 30.0% lower than routine pathologist estimates and 21.2% lower than LUCAID's cell-count approach. The cell-count method had a correlation of 0.39 and an error of 21.7 points, while routine assessment had a correlation of 0.20 and an error of 24.42 points.

Tissue mapping and generated reports

LUCAID's estimates of carcinoma, necrosis and stroma, the connective tissue around a tumor, closely matched joint pathologist assessment. The pooled Pearson correlation was 0.996 and the Spearman correlation was 0.96. Mean absolute errors were 1.6 percentage points for carcinoma, 1.9 for necrosis and 2.9 for stroma.

In the 1,001-patient discovery group with follow-up data, higher levels of stromal tumor-infiltrating lymphocytes and closer carcinoma-to-plasma-cell adjacency were associated with better survival. Higher neutrophil-to-lymphocyte ratios and greater distance between endothelial cells and lymphocytes were associated with worse outcomes. These were observational associations, so they do not establish that the measured tissue features caused the survival differences or would predict treatment response.

The researchers also examined ten generated reports containing 1,365 atomic statements. All 134 statements that referred to module outputs were grounded in those outputs, and 1,194 of 1,231 interpretive statements, or 97.0%, were judged correct. Of 159 citations, 131, or 82.4%, fully or partly supported the statements they accompanied. Reviewers found no omissions across 70 section-level checks, but three contradictions appeared in two reports. None of the 111 errors rated for potential harm was judged severe.

What the preprint does not settle

The authors describe the report assessment as a proof of principle because it involved only ten cases. They also say the system's influence on routine clinical work remains unestablished and that the general-purpose agent could be refined further. The study focused on lung cancer, so the results do not validate LUCAID for other tumor types.

The molecular comparison has its own limits: KRAS variant allele frequency can be affected by copy-number alterations, allelic imbalance, subclonality, tumor heterogeneity and technical factors. The tumor-microenvironment findings likewise need independent biological and clinical validation. Race, ethnicity and socioeconomic status were unavailable because they were not collected in routine documentation, and no formal power calculation or preregistration was reported.

The work is a preprint posted on arXiv as version 2 dated Aug. 31, 2026. It does not show improved survival, treatment response, diagnostic accuracy or workflow efficiency in ordinary practice. Larger studies across institutions, scanners, staining protocols and tumor types will be needed to test whether the system changes care.

Funding and disclosures

The acknowledgements list support from the BIH Charité Digital Clinician Scientist Program, Charité and BIH programmes, Berlin's PROFIT development grant, the German Ministry for Education and Research, the German Research Foundation and Korean government, MSIT and IITP grants. The disclosures say several authors have ties to Aignostics, including co-founder, advisory, board, chief technology officer, employee and scientific-advisory-board roles; the remaining authors reported no conflict of interest. The study received Charité ethics approval, followed the Declaration of Helsinki and obtained written informed consent.

Paper data and sources

Original title: LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology
Authors: Marie-Lisa Eich, Kai Standvoss, Timo Milbich et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.