Pith. sign in

REVIEW 7 cited by

Sparse Autoencoders for Hypothesis Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.04382 v3 pith:75UUPW26 submitted 2025-02-05 cs.CL cs.AIcs.CY

classification cs.CLcs.AIcs.CY
keywords hypothesaestargetvariabledatadatasetsfeaturesheadlineshypotheses
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We describe HypotheSAEs, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks). HypotheSAEs has three steps: (1) train a sparse autoencoder on text embeddings to produce interpretable features describing the data distribution, (2) select features that predict the target variable, and (3) generate a natural language interpretation of each feature (e.g., "mentions being surprised or shocked") using an LLM. Each interpretation serves as a hypothesis about what predicts the target variable. Compared to baselines, our method better identifies reference hypotheses on synthetic datasets (at least +0.06 in F1) and produces more predictive hypotheses on real datasets (~twice as many significant findings), despite requiring 1-2 orders of magnitude less compute than recent LLM-based methods. HypotheSAEs also produces novel discoveries on two well-studied tasks: explaining partisan differences in Congressional speeches and identifying drivers of engagement with online headlines.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent

    cs.LG 2026-07 conditional novelty 7.0 of 10

    HARP shows that retrieval from an activation database plus linear-probe tools, driven by an LLM agent, matches or exceeds SAEs and activation oracles on four interpretability tasks with zero training.

  2. Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach

    econ.EM 2025-11 unverdicted novelty 7.0 of 10

    A new framework combines AI-derived concept embeddings with high-dimensional selective inference to enable statistically principled, interpretable discovery from unstructured data in empirical economics.

  3. Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A mixed-initiative system where LLMs flag low-confidence annotations, cluster them, and propose codebook rules that human experts review and iterate on.

  4. Aggregated Individual Reporting for Post-Deployment Evaluation

    cs.CY 2025-06 conditional novelty 6.0 of 10

    The authors formalize a mechanism for collecting and aggregating public reports about deployed AI systems, aiming to surface unknown harms and enable accountability.

  5. Correlated Errors in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Large language models from different providers and architectures often make the same errors, and more accurate models are especially likely to share mistakes.

  6. Sparse Autoencoders, Again?

    cs.LG 2025-06 conditional novelty 6.0 of 10

    VAEase gates the VAE decoder input by the encoder's variance, combining sparse-autoencoder adaptive sparsity with a hyperparameter-free loss; a global-minimizer theorem says active latent dimensions recover per-manifo...

  7. BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A pipeline that combines contextual embeddings with two LMs' per-word probabilities and sparse autoencoders to automatically find interpretable slices where one model outperforms another.

Pith tools