Small changes in data or settings used to find a circuit in a language model often produce very different circuits: under bootstrap resampling, average pairwise overlap of EAP-IG circuits across tasks and models is only 0.561 (Jaccard).
Explanation in causal inference: developments in mediation and interaction
1 Pith paper cite this work, alongside 410 external citations. Polarity classification is still indexing.
1
Pith paper citing it
410
external citations · OpenAlex
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
Small changes in data or settings used to find a circuit in a language model often produce very different circuits: under bootstrap resampling, average pairwise overlap of EAP-IG circuits across tasks and models is only 0.561 (Jaccard).