No AI summary available for this article.
Why It Matters
We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head barely draws on this atlas. Motivated by this underused information, we train an uncertainty head on th...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2608.26090v1 · Indexed 19 days ago