Pith. sign in

REVIEW 1 cited by

Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.18895 v1 pith:6INO655O submitted 2024-11-28 cs.LG cs.CL

classification cs.LGcs.CL
keywords shiftsparseannotatorautoencoderseffectivelyhumanintroducemetric
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sparse Autoencoders (SAEs) are an interpretability technique aimed at decomposing neural network activations into interpretable units. However, a major bottleneck for SAE development has been the lack of high-quality performance metrics, with prior work largely relying on unsupervised proxies. In this work, we introduce a family of evaluations based on SHIFT, a downstream task from Marks et al. (Sparse Feature Circuits, 2024) in which spurious cues are removed from a classifier by ablating SAE features judged to be task-irrelevant by a human annotator. We adapt SHIFT into an automated metric of SAE quality; this involves replacing the human annotator with an LLM. Additionally, we introduce the Targeted Probe Perturbation (TPP) metric that quantifies an SAE's ability to disentangle similar concepts, effectively scaling SHIFT to a wider range of datasets. We apply both SHIFT and TPP to multiple open-source models, demonstrating that these metrics effectively differentiate between various SAE training hyperparameters and architectures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Discovering Chunks in Neural Embeddings for Interpretability

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Recurring 'chunks' in neural embeddings can be extracted, predict input patterns, and be perturbed to steer a model's outputs.

Pith tools