Interpretability in the Wild: a Circuit for Indirect Object Identification in

Wang, Kevin, Variengien, Alexandre, Conmy, Arthur, Shlegeris, Buck, Steinhardt, Jacob , booktitle=

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

browse 1 citing papers

representative citing papers

Exemplar Partitioning for Mechanistic Interpretability

cs.LG · 2026-05-14 · unverdicted · novelty 7.0 · 2 refs

Exemplar Partitioning creates Voronoi partitions of LLM activation space via leader clustering on streamed activations, yielding comparable, interpretable dictionaries that support interventions and achieve competitive benchmark results with ~1000x less compute than SAEs.

citing papers explorer

Showing 1 of 1 citing paper.

Exemplar Partitioning for Mechanistic Interpretability cs.LG · 2026-05-14 · unverdicted · none · ref 36 · 2 links
Exemplar Partitioning creates Voronoi partitions of LLM activation space via leader clustering on streamed activations, yielding comparable, interpretable dictionaries that support interventions and achieve competitive benchmark results with ~1000x less compute than SAEs.

Interpretability in the Wild: a Circuit for Indirect Object Identification in

fields

years

verdicts

representative citing papers

citing papers explorer