A framework that uses LLM sparse autoencoder features as high-dimensional outcome variables, with new k-FWER controlling bootstrap tests and formalized autointerp evaluation, to make interpretable discoveries from unstructured data.
Use sparse autoencoders to discover unknown concepts, not to act on known concepts.arXiv preprint arXiv:2506.23845
5 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Longitudinal analysis of r/ChatGPT posts shows normalization of ChatGPT as an everyday tool alongside rising mental health and emotional attachment discussions after GPT-4o, with PuLSE detecting the latter trend months early.
Introduces conditional hypothesis generation for LLM text analysis incorporating researcher-specified covariates via two econometrics-inspired methods to handle stratum imbalance and sign reversal, outperforming global baselines in synthetic tests and expert evaluations.
ActivationReasoning grounds logical reasoning in LLM latent activations via SAEs to enable structured inference, concept composition, and behavior steering on multi-hop, abstraction, and safety tasks.
citing papers explorer
-
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach
A framework that uses LLM sparse autoencoder features as high-dimensional outcome variables, with new k-FWER controlling bootstrap tests and formalized autointerp evaluation, to make interpretable discoveries from unstructured data.
-
Three Years of r/ChatGPT: Societal Impact Evaluations from Social Media Data
Longitudinal analysis of r/ChatGPT posts shows normalization of ChatGPT as an everyday tool alongside rising mental health and emotional attachment discussions after GPT-4o, with PuLSE detecting the latter trend months early.
-
Conditional Hypothesis Generation for LLM-Based Text Analysis with Researcher-Specified Covariates
Introduces conditional hypothesis generation for LLM text analysis incorporating researcher-specified covariates via two econometrics-inspired methods to handle stratum imbalance and sign reversal, outperforming global baselines in synthetic tests and expert evaluations.
-
ActivationReasoning: Logical Reasoning in Latent Activation Spaces
ActivationReasoning grounds logical reasoning in LLM latent activations via SAEs to enable structured inference, concept composition, and behavior steering on multi-hop, abstraction, and safety tasks.
- Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy