Peer-reviewed

Three Splicing Models Show Blind Spots Behind Strong Accuracy

A peer-reviewed modeling study found that CpG and stop-codon patterns can skew predictions, while weak RNA-structure sensitivity may miss variant effects.

Strong scores, uneven reality

Three deep-learning models for RNA splicing generally tracked measured exon inclusion, the readout of whether a target exon is retained, but the study also found recurring errors tied to sequence features. After calibration to each assay, Pearson correlations were 0.83 to 0.86 for the Liao library and 0.81 to 0.88 for FAS exon 6. Agreement was weaker for MFASS, with correlations of 0.37 to 0.44.

To look inside those predictions, the researchers trained interpretable surrogate models to imitate the original models. This approach, called interpretable distillation, uses the surrogate's predictive logic to generate hypotheses that can be checked against experimental data. On held-out synthetic sequences, the surrogates matched original predictions with Pearson correlations of 0.94 to 0.99. For SpliceAI, 98.4% of surrogate predictions fell within plus or minus 15% of the original output.

Those surrogates suggested that exon recognition was represented mainly by additive combinations of sequence motifs. The features included patterns resembling known splicing regulatory elements, a liver-specific element in Pangolin and possible uncharacterized elements. Several did not match known RNA-binding protein preferences and were not functionally validated.

The hidden signals in the sequence

One recurring signal was CpG content, the frequency of a short sequence pattern called CpG. CpG composition was associated with model error even after controlling for measured exon inclusion, and higher CpG content inflated predictions. In the matched methylation and RNA-sequencing comparison, genomic CpG methylation was not associated with exon inclusion.

The bias also appeared in a genomic ranking test. Excluding false-positive predictions in intronic CpG islands raised SpliceAI donor top-k accuracy from 93.8% to 94.5%, an 11.3% reduction in error. The reported association between CpG composition and model error had a p-value below 5 × 10−3.

Stop-codon count produced another systematic signal. Across all three models, it was associated with exon-inclusion prediction error, with a reported p-value below 1 × 10−20. Sequences with five or more stop codons received deflated predictions, and SpliceAI and Pangolin incorrectly classified the CFTR G542X variant as splice-disrupting.

A harder problem: RNA structure

RNA structure exposed a separate weakness. Motif features explained 88% to 97% of the variance in model predictions, while predictions were inflated for synthetic sequences with more thermodynamically stable predicted structures. The structure-error association had a reported p-value below 1 × 10−20.

In matched structure-perturbing sequence sets, all three models missed the observed direction of exon-inclusion changes. The models also failed to predict any splicing effect for a BLTP1 variant, even though exon skipping was observed in patient cells and minigene assays.

The authors interpret CpG enrichment and stop-codon depletion as reference-genome proxies for exon identity, and weak RNA-structure sensitivity as a blind spot for atypical sequences and variants. But these analyses are associational. They do not show that CpG or stop codons directly cause the biological splicing changes, or that the identified features are the causal mechanism of human splicing.

What the analysis can and cannot answer

Distillation used approximately 1 million filtered synthetic sequences from a three-exon reporter design. The middle exon had a 70-nucleotide variable region with fixed splice sites, and bases were sampled with equal probability. For experimental checks, the researchers used 239,722 Liao-library sequences, 7,918 FAS exon 6 sequences, including 5,976 unique sequences, and 16,469 MFASS sequences, all with measured exon-inclusion levels.

Those choices set clear limits. The distilled networks were trained and evaluated in a fixed three-exon context with a 70-nucleotide variable middle region, so the study did not directly test every sequence context. The assay comparisons were made after calibration, which supports agreement with measured inclusion but does not by itself establish the biological mechanism behind the models' scores. MFASS used fluorescence rather than direct RNA sequencing, and the structure analysis for synthetic sequences relied on predicted folds.

For users of these scores, the message is straightforward: strong average agreement with assays does not rule out systematic errors in unusual sequences, and a model can miss a variant effect that depends on RNA folding. The findings do not establish clinical diagnostic performance or patient outcomes, and they do not show that every splicing model or every variant has the same failure modes. The authors suggest broader experimental sequence data, multi-calibration and structural information as possible remedies, but those approaches were not tested here.

Paper data and sources

Original title: Interpretable distillation reveals that deep learning splicing models suffer from pervasive confounders and blind spots.
Authors: Simon Liu, Wenjing Zhang, Oded Regev
Journal/Repository: Genome biology
Status: Peer-reviewed
First online: 2026-08-21
DOI: 10.1186/s13059-026-04124-9
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.