A preprint's NAPE-L raster model produced a mixed set of benchmark results: it tied SSLAM on AudioSet-2M, came close to SSLAM on AudioSet-20K and ESC-50, and scored higher than the strongest listed baseline on IEMOCAP.
In the Large configuration, NAPE scored 50.2 mean average precision (mAP) on AudioSet-2M, matching SSLAM. It scored 40.5 mAP against 40.9 for SSLAM on AudioSet-20K and 96.0% top-1 accuracy against SSLAM's 96.2% on ESC-50. On IEMOCAP, it reached 68.0% versus 64.5% for the strongest listed baseline, a 3.5-percentage-point difference.
Predicting the next patch
NAPE is self-supervised: a causal Transformer predicts the next log-mel spectrogram patch embedding from preceding patches. A log-mel spectrogram is a time-frequency map of sound, and causal masking blocks later patches from the prediction.
The retained configuration uses Conv2d patch embedding, raster or diagonal scanning, a three-layer SimSiam-style predictor, patch embeddings as targets, negative cosine loss and target stop-gradient.
Unlabeled pre-training used 1,964,222 AudioSet clips from the unbalanced split and 20,961 from the balanced split; the paper reports that 18,900 evaluation clips were processed.
Tests of the training design
The downstream evaluation covered AudioSet-2M and AudioSet-20K, ESC-50, Speech Commands V1 and V2, and IEMOCAP. AudioSet was scored by mAP, while the single-label tasks used top-1 accuracy; ESC-50 and IEMOCAP results were means across five-fold cross-validation.
The ablations without prediction shift or target stop-gradient were reported as divergent during pre-training. The no-causal-mask version did not diverge, but its reported changes were -7.8 mAP on AudioSet-2M, -14.3 points on AudioSet-20K and -25.4 points on ESC-50.
Adding any predictor outperformed direct encoder prediction in the ablation. The SimSiam-style variant performed best on the audio benchmarks and IEMOCAP and was on par for keyword spotting.
Raster, diagonal and zigzag scanning performed comparably, while time-major scanning consistently underperformed.
Comparisons across model sizes
The scaling experiment used Small, Base and Large configurations with approximately 19 million, 85 million and 303 million parameters. Across all tested benchmarks, NAPE-L improved over NAPE-B and NAPE-B over Small; raster NAPE outperformed Audio-MAE at every scale. At Small scale, the reported gains over Audio-MAE were 2.6 mAP on AudioSet-2M and 4.1 mAP on AudioSet-20K.
In linear-probing tests, which keep the encoder fixed, the strongest layers were the second, sixth and 11th in the Small, Base and Large models, respectively. From Small to Large, probe gains were 3.9 mAP on AudioSet-2M and 3.7 accuracy points on ESC-50.
A qualitative signal
In a qualitative analysis of 500 held-out AudioSet clips, predicted-to-true patch-embedding cosine similarity was close to 1.0 almost everywhere. The attention and embedding maps showed structured temporal-spectral and acoustically coherent patterns.
What the comparisons leave open
These figures are benchmark comparisons within the reported settings, not proof that NAPE is universally better. The manuscript is a preprint, and baseline results came from reported comparisons; confidence intervals, hypothesis tests and seed-to-seed variability summaries were not reported. Independent matched training protocols and broader audio and speech tasks would be needed to show how widely the result generalizes.
Paper data and sources
Original title: Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners
Authors: Umberto Cappellazzo, Xubo Liu, Stavros Petridis, Maja Pantic
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text