Pith. sign in

Quantifying feature space universality across large language models via sparse autoencoders

7 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.

7 Pith papers citing it
1 external citations · external index

citation-role summary

background 1 method 1

citation-polarity summary

years

2026 7

verdicts

UNVERDICTED 7

representative citing papers

WriteSAE: Sparse Autoencoders for Recurrent State

cs.LG · 2026-05-12 · unverdicted · novelty 8.0 · 4 refs

WriteSAE introduces sparse autoencoders with rank-1 matrix atoms for recurrent state updates, allowing replacement tests that outperform deletion on 92.4% of positions and a formula predicting logit changes with R²=0.98.

Understanding the Mechanism of Altruism in Large Language Models

econ.GN · 2026-04-21 · unverdicted · novelty 6.0

A small set of sparse autoencoder features in LLMs drives shifts between generous and selfish allocations in dictator games, with causal patching and steering confirming their role and generalization to other social games.

Rigorous Interpretation Is a Form of Evaluation

cs.CY · 2026-05-06 · unverdicted · novelty 5.0

Rigorous interpretability can function as a principled form of model evaluation if its claims are falsifiable, reproducible, and predictive.

citing papers explorer

Showing 7 of 7 citing papers.