Pith. sign in

REVIEW 31 cited by

BatchTopK Sparse Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.06410 v1 pith:DJMEJBMP submitted 2024-12-09 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords batchtopksaeslatentsactivationsnumbersamplesparsetopk
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sparse autoencoders (SAEs) have emerged as a powerful tool for interpreting language model activations by decomposing them into sparse, interpretable features. A popular approach is the TopK SAE, that uses a fixed number of the most active latents per sample to reconstruct the model activations. We introduce BatchTopK SAEs, a training method that improves upon TopK SAEs by relaxing the top-k constraint to the batch-level, allowing for a variable number of latents to be active per sample. As a result, BatchTopK adaptively allocates more or fewer latents depending on the sample, improving reconstruction without sacrificing average sparsity. We show that BatchTopK SAEs consistently outperform TopK SAEs in reconstructing activations from GPT-2 Small and Gemma 2 2B, and achieve comparable performance to state-of-the-art JumpReLU SAEs. However, an advantage of BatchTopK is that the average number of latents can be directly specified, rather than approximately tuned through a costly hyperparameter sweep. We provide code for training and evaluating BatchTopK SAEs at https://github.com/bartbussmann/BatchTopK

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What do Reward Models Memorize?

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Counterfactual memorization maps show RMs misallocate capacity to easy pairs, memorize dataset artifacts, and overgeneralize length/compliance on unseen pairs.

  2. Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    Sparsity regularizers applied before Top-k selection in SAEs improve monosemanticity and make reconstruction robust to inference-time k across vision models and datasets.

  3. Reference Feature Atlases for Mechanistic Auditing of Language Models

    cs.AI 2026-06 conditional novelty 7.0 of 10

    A reference feature atlas attaches a new language model with a linear decoder and surfaces panel-uncovered structure via a separate residual dictionary.

  4. Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations

    cs.LG 2026-05 conditional novelty 7.0 of 10

    SA-GSAE with Bi-Jump-ReLU enables one latent to encode both polarities of anticorrelated features, Pareto-dominating or matching full-width gated SAEs while reducing dead latents by up to 500x on some LLM hookpoints.

  5. Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A two-level mixture-of-experts sparse autoencoder models parent and child concepts together, improving reconstruction and reducing feature redundancy on Gemma 2-2B activations compared to flat top-k SAEs.

  6. Sparse Autoencoders Do Not Find Canonical Units of Analysis

    cs.LG 2025-02 conditional novelty 7.0 of 10

    Across different dictionary sizes, SAE latents are neither complete nor atomic, so SAEs do not learn a canonical set of features.

  7. Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition

    cs.LG 2025-01 conditional novelty 7.0 of 10

    Attribution-based Parameter Decomposition splits a network's parameters into faithful, minimal, and simple components and recovers ground-truth mechanisms in toy models of superposition and compressed computation.

  8. LaPrune: Controllable Differentiable Sparsity at Million Scale

    cs.LG 2026-08 accept novelty 6.0 of 10

    A differentiable top-k mask layer that enforces an exact selection budget and uses a normalized hardness parameter to interpolate from equal-weight masks to hard binary masks, with saturation theory and million-scale results.

  9. ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels

    cs.CL 2026-08 conditional novelty 6.0 of 10

    ChronoLens uses feature-aligned crosscoders to show that historical language change has comparable magnitude across linguistic levels within a language, but divergent timing and direction across five parliamentary languages.

  10. ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

    cs.LG 2026-07 conditional novelty 6.0 of 10

    With matched-scale sparse autoencoders, HuBERT-ECG best preserves its ECG representation while ECG-JEPA best exposes clinical measurements through single features — a leader split that repeats on MIMIC-IV-ECG.

  11. CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A sparse dictionary learned from an ECG foundation model's embeddings recovers interpretable cardiac concepts—PVCs, atrial fibrillation, bundle branch blocks, ST/T-wave segments—and transfers to an external dataset wi...

  12. SADe: Sparse-Atom Support Decontamination for Few-Shot Segmentation with Weak Support Annotations

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A frozen support-only cleaner that uses SAE atom contrast plus dense similarity to strip distractors from weak FSS supports and lifts query mIoU across heterogeneous predictors, especially under expanded boxes.

  13. Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression?

    cs.LG 2026-07 accept novelty 6.0 of 10

    A new SAE objective penalizes disagreement between ridge prediction operators, preserving more linear readouts at equal reconstruction error.

  14. Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders

    cs.CV 2026-04 conditional novelty 6.0 of 10

    Spatio-temporal contrastive SAEs recover temporal coherence lost by hard TopK, improve action probes by +3.9% and retrieval by up to 2.8× R@1, and expose a monosemanticity metric artifact.

  15. Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Coverage of sparse-autoencoder-identified task features predicts post-training performance and can guide synthesis of small, high-impact datasets (2,000 vs. 300,000 samples).

  16. AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders

    cs.SD 2026-02 conditional novelty 6.0 of 10

    SAE features from Whisper and HuBERT are seed-stable, interpretable, and steerable: cutting false speech detections by 70% and correlating with EEG responses to speech.

  17. Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control

    cs.LG 2026-01 unverdicted novelty 6.0 of 10

    Gradient-guided token search against internal persona directions yields gibberish prompts that reduce sycophancy, hallucination, and myopic reward in three tested LLMs.

  18. Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Incorrect L0 makes sparse autoencoders mix correlated features rather than disentangling them, and a decoder projection metric can identify the correct L0.

  19. Teach Old SAEs New Domain Tricks with Boosting

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Training a small secondary sparse autoencoder on the reconstruction error of a pretrained SAE improves domain-specific reconstruction and language-model perplexity without hurting general performance.

  20. Insights into a radiology-specialised multimodal large language model with sparse autoencoders

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Applying Matryoshka sparse autoencoders to a radiology-specialised multimodal LLM reveals a minority of interpretable clinical features, while steering them produces unreliable and often off-target report changes.

  21. Decoding Dense Embeddings: Sparse Autoencoders for Interpreting and Discretizing Dense Retrieval

    cs.IR 2025-05 conditional novelty 6.0 of 10

    Dense retrieval embeddings can be decomposed into interpretable latent concepts that serve both as explanations and as efficient sparse indexing units for retrieval.

  22. Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A position paper proposing feature consistency, measured by PW-MCC, as a core SAE evaluation criterion, with evidence that TopK SAEs achieve high consistency on LLM activations.

  23. Low-Rank Adapting Models for Sparse Autoencoders

    cs.LG 2025-01 conditional novelty 6.0 of 10

    LoRA fine-tuning of the language model around a fixed SAE reduces the SAE-insertion loss gap by 30-55% and matches end-to-end SAEs 2-20x faster.

  24. Transcoders Beat Sparse Autoencoders for Interpretability

    cs.LG 2025-01 conditional novelty 6.0 of 10

    Skip transcoders beat sparse autoencoders on both reconstruction fidelity and automated interpretability scores for transformer MLP layers.

  25. SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders

    cs.LG 2025-01 conditional novelty 6.0 of 10

    SAeUron removes concepts from text-to-image diffusion models by ablating concept-specific sparse autoencoder features during inference, achieving state-of-the-art unlearning on UnlearnCanvas and I2P without weight updates.

  26. Stable and Steerable Sparse Autoencoders with Weight Regularization

    stat.ML 2026-03 conditional novelty 5.0 of 10

    L2 weight regularization in TopK SAEs increases cross-seed feature overlap and roughly doubles measured steering success on Pythia-70M, at the cost of collapsing most latents to zero.

  27. BlueGlass: A Framework for Composite AI Safety

    cs.AI 2025-07 conditional novelty 5.0 of 10

    BlueGlass provides composite AI safety infrastructure; its case studies on object-detection VLMs reveal dataset trade-offs, a decoder-layer phase transition in probe accuracy, and SAE-discovered concepts including spu...

  28. Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy

    cs.LG 2025-05 conditional novelty 5.0 of 10

    HierarchicalTopK trains a single sparse autoencoder that reconstructs transformer activations well at many sparsity levels, matching or beating separate per-level models.

  29. Analyze Feature Flow to Enhance Interpretation and Steering in Language Models

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Cosine similarity between sparse autoencoder features across layers and modules builds flow graphs that explain feature evolution and enable multi-layer steering of language model generation.

  30. Sparsification and Reconstruction from the Perspective of Representation Geometry

    cs.LG 2025-05 reject novelty 4.0 of 10

    Sparse encoding appears to stratify and compress feature representations, but the claimed causal link between cluster separation and reconstruction is not supported.

  31. Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning

    cs.LG 2026-07 accept

    A scoping review surveying circuit analysis, sparse autoencoders, activation steering, and neurosymbolic frameworks for interpreting and controlling Transformer-based neural networks.

Pith tools