Preprint

Neutrino AI model maps hidden signals and ranks events by error

Preprint: The one-million-event PolarBERT analysis identified detector signals and tested an uncertainty head for ranking likely angular error, with calibration limits.

An artificial-intelligence model used to reconstruct neutrino direction contained a broad set of detector-related signals, while a separate uncertainty head ranked events by likely angular error. In a tested one-million-event sample, the uncertainty-head ranking gave median angular errors of 3.15 degrees at 20% efficiency and 11.24 degrees at 50%, compared with 20.2 and 37.72 degrees for total-charge selection. That is a within-study comparison, not evidence that the identified physical features caused the difference, and the uncertainty output was not globally calibrated.

A map of detector signals

Researchers studied PolarBERT, an eight-block, BERT-like transformer pretrained on IceCube pulse sequences and fine-tuned for neutrino direction reconstruction. They used a one-million-event sample divided into five disjoint sets for training, candidate discovery, independent concept validation and intervention testing, with reconstruction and fidelity validation also kept separate. Candidate features were judged on three checks: read-out quality, nuisance selectivity and intervention relevance.

The model's frozen internal representation contained what the analysis describes as a physical atlas of detector quality, auxiliary activity, event brightness and detector depth. A bright-clean read-out separated bright-clean events from strongly auxiliary, dim events with a separation score, AUROC, of 0.911 and remained informative under all matched controls. A layer-2 depth read-out reached AUROC 0.986 and remained essentially unchanged under matched controls, while the final-layer dictionary contained no validated depth coordinate.

Reading a signal is not the same as using it

The layer-dependent result matters for interpretation: a missing final-layer coordinate does not show that the corresponding information is absent from the model, because depth was still readable in layer 2. The direction-head intervention gave a similar result. Zeroing the bright-clean latent shifted angular error by +0.06 degrees, close to the +0.05-degree shift for unrelated latents matched by firing rate; the null replicated across dictionary draws. That evidence applies to the sparse read-outs tested, not to every part of the representation.

Researchers also examined a leading sensitivity axis for the direction head. It differed by 5.6 degrees between the concept-development and concept-test splits and accounted for 98% of aggregate sensitivity on the latter. The result was a dataset-level summary rather than a universal event-level bottleneck, and the axis had no simple physical interpretation.

A second task: estimating error

The analysis then tested a different use for the representation: estimating reconstruction error. The uncertainty head was a one-hidden-layer MLP. On an independent validation split, it reached a Spearman rank correlation of 0.60 with true error and an AUROC of 0.925 when distinguishing events in the highest and lowest true-error quartiles. The score provided an ordering of reconstruction quality, but it was not a perfect stand-in for each event's actual error.

The intervention tests reported shifts in the study's M1 score for clean-side latent sets. Removing the bright-clean latent, a clean pair, a 194-latent clean family and a dense clean core produced M1 shifts of +0.31, +0.32, +0.78 and +0.37, respectively, with the effects holding across dictionary seeds. By contrast, removing the auxiliary latent or the broader auxiliary family did not meet the intervention criteria for changing the uncertainty prediction. The tests could not distinguish redundancy with clean-side information from auxiliary information distributed outside the tested family.

Ranking events, with a calibration caveat

At 20% selection efficiency, meaning the rule kept one in five events, the MLP uncertainty ranking had a median angular error of 3.15 degrees, versus 20.2 degrees for total-charge selection. At 50% efficiency, the corresponding figures were 11.24 degrees for the MLP and 37.72 degrees for total-charge selection. A linear uncertainty head was nearly identical to the MLP: 3.16 degrees at 20% efficiency and 11.32 degrees at 50%.

The ranking did not amount to a globally calibrated event-level error estimate. Predicted and true error had an event-level Spearman correlation of 0.615. Calibration was good in the low-error region but overconfident at intermediate predicted errors. Within this analysis, the output was useful for ordering events, but it still requires further calibration for global event-level use.

How far the evidence goes

The evidence is specific to the tested PolarBERT model, sparse-feature dictionaries, model-level interventions and held-out IceCube events. It does not establish performance on real detector data, other models or other tasks. Nor does the absence of a validated final-layer depth coordinate show that depth information is absent from the model; the layer-2 read-out shows that the result can depend on where the representation is examined. The front matter identifies it as an arXiv preprint dated 26 Aug 2026.

Part of the work was supported by the European Union's Horizon Europe programme under Marie Sklodowska-Curie grant agreement 101168829.

Paper data and sources

Original title: Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders
Authors: Raphaël Bonnet-Guerrini, Johann Ioannou-Nikolaides, Inar Timiryasov, Vincenzo Piuri
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.