Pith. sign in

REVIEW 11 cited by

Interpretability as Compression: Reconsidering SAE Explanations of Neural Activations with MDL-SAEs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.11179 v1 pith:QARUC5BN submitted 2024-10-15 cs.LG cs.AIcs.ITmath.IT

classification cs.LGcs.AIcs.ITmath.IT
keywords saesexplanationsactivationsfeaturesframeworkneuralsparsityargue
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sparse Autoencoders (SAEs) have emerged as a useful tool for interpreting the internal representations of neural networks. However, naively optimising SAEs for reconstruction loss and sparsity results in a preference for SAEs that are extremely wide and sparse. We present an information-theoretic framework for interpreting SAEs as lossy compression algorithms for communicating explanations of neural activations. We appeal to the Minimal Description Length (MDL) principle to motivate explanations of activations which are both accurate and concise. We further argue that interpretable SAEs require an additional property, "independent additivity": features should be able to be understood separately. We demonstrate an example of applying our MDL-inspired framework by training SAEs on MNIST handwritten digits and find that SAE features representing significant line segments are optimal, as opposed to SAEs with features for memorised digits from the dataset or small digit fragments. We argue that using MDL rather than sparsity may avoid potential pitfalls with naively maximising sparsity such as undesirable feature splitting and that this framework naturally suggests new hierarchical SAE architectures which provide more concise explanations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features

    cs.LG 2026-05 accept novelty 8.0 of 10

    Many distinct SAE features share identical explanations, with the average annotation resolving only 70% of feature identity in a large annotated dataset.

  2. Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    Block-sparse featurizers recover visual concepts as two- to four-dimensional manifolds and describe activations more compactly than direction-based methods via minimum-description-length comparison.

  3. From Mechanistic to Compositional Interpretability

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Compositional interpretability defines explanations as commuting syntactic-semantic mapping pairs grounded in compositionality and minimum description length, with compressive refinement and a parsimony theorem guaran...

  4. From Mechanistic to Compositional Interpretability

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    The paper introduces compositional interpretability as a category-theoretic framework that casts mechanistic explanations as commuting syntactic-semantic mappings optimized under faithfulness and complexity constraint...

  5. Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression?

    cs.LG 2026-07 accept novelty 6.0 of 10

    A new SAE objective penalizes disagreement between ridge prediction operators, preserving more linear readouts at equal reconstruction error.

  6. The Rate-Distortion-Polysemanticity Tradeoff in SAEs

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    SAEs exhibit a rate-distortion-polysemanticity tradeoff where monosemanticity increases rate and distortion, with optimal polysemanticity set by feature co-occurrence probabilities in the data.

  7. Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Tree SAE learns hierarchical feature structures by combining activation coverage with a new reconstruction condition, outperforming prior SAEs on hierarchical pair detection while matching state-of-the-art benchmark p...

  8. Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Tree SAE learns hierarchical feature pairs in sparse autoencoders by combining activation coverage with a new reconstruction condition, outperforming prior methods on hierarchy detection while remaining competitive on...

  9. FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Training sparse autoencoders on a language model's own generated text can improve seed stability and downstream probing relative to training on web text.

  10. Stable and Steerable Sparse Autoencoders with Weight Regularization

    stat.ML 2026-03 conditional novelty 5.0 of 10

    L2 weight regularization in TopK SAEs increases cross-seed feature overlap and roughly doubles measured steering success on Pythia-70M, at the cost of collapsing most latents to zero.

  11. Evaluating SAE interpretability without explanations

    cs.LG 2025-07 conditional novelty 5.0 of 10

    SAE latent interpretability can be scored directly from activation examples via intruder detection and embedding clustering, with LLM scores correlating strongly with human scores.

Pith tools