Small changes in data or settings used to find a circuit in a language model often produce very different circuits: under bootstrap resampling, average pairwise overlap of EAP-IG circuits across tasks and models is only 0.561 (Jaccard).
Statistics for experimenters: Design, innovation, and discovery
1 Pith paper cite this work, alongside 94 external citations. Polarity classification is still indexing.
1
Pith paper citing it
94
external citations · OpenAlex
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
Small changes in data or settings used to find a circuit in a language model often produce very different circuits: under bootstrap resampling, average pairwise overlap of EAP-IG circuits across tasks and models is only 0.561 (Jaccard).