Pith. sign in

REVIEW 2 cited by

Tokenized SAEs: Disentangling SAE Reconstructions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.17332 v1 pith:MYMZXOWG submitted 2025-02-24 cs.LG

classification cs.LG
keywords reconstructionfeaturescorrespondinterestingsaessparseachievedauto-encoders
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Sparse auto-encoders (SAEs) have become a prevalent tool for interpreting language models' inner workings. However, it is unknown how tightly SAE features correspond to computationally important directions in the model. This work empirically shows that many RES-JB SAE features predominantly correspond to simple input statistics. We hypothesize this is caused by a large class imbalance in training data combined with a lack of complex error signals. To reduce this behavior, we propose a method that disentangles token reconstruction from feature reconstruction. This improvement is achieved by introducing a per-token bias, which provides an enhanced baseline for interesting reconstruction. As a result, significantly more interesting features and improved reconstruction in sparse regimes are learned.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Truncated Jump Sampling stops the ODE early and outputs the algebraically decoded x0 estimate, reducing NFEs by 20-70% across six model families with no retraining.

  2. How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding

    cs.CL 2025-07 reject novelty 6.0 of 10

    The paper reports that chain-of-thought features extracted by sparse autoencoders and transferred through activation patching improve answer confidence in Pythia-2.8B but not in Pythia-70M, implying a scale threshold ...

Pith tools