Training sparse autoencoders on individual chat sequences, instead of concatenated text blocks, improves reconstruction and feature interpretability on instruction-tuned LLMs.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models
Training sparse autoencoders on individual chat sequences, instead of concatenated text blocks, improves reconstruction and feature interpretability on instruction-tuned LLMs.