Harmonic loss, which scores logits by inverse Euclidean distance to class prototypes with a scale-invariant normalization, yields more interpretable class centers, reduced grokking, and faster convergence across MLPs, transformers, MNIST, and a small GPT-2.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Harmonic Loss Trains Interpretable AI Models
Harmonic loss, which scores logits by inverse Euclidean distance to class prototypes with a scale-invariant normalization, yields more interpretable class centers, reduced grokking, and faster convergence across MLPs, transformers, MNIST, and a small GPT-2.