Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

SAFR: Neuron Redistribution for Interpretability

cs.LG · 2025-01-23 · conditional · novelty 4.0

SAFR regularizes a transformer so that important tokens become monosemantic (one neuron per meaning) and correlated tokens share neurons, and it uses the accuracy drop after deleting high-capacity tokens as its interpretability score.

citing papers explorer

Showing 1 of 1 citing paper.

  • SAFR: Neuron Redistribution for Interpretability cs.LG · 2025-01-23 · conditional · none · ref 7

    SAFR regularizes a transformer so that important tokens become monosemantic (one neuron per meaning) and correlated tokens share neurons, and it uses the accuracy drop after deleting high-capacity tokens as its interpretability score.