Path patching provides a method to express and quantitatively test hypotheses that neural network behaviors are localized to sets of paths.
Proceedings of the 39th International Conference on Machine Learning , pages =
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2representative citing papers
CIF certifies interventional interpretability metrics as bounded causal means with anytime-valid confidence sequences, including under adaptive sampling, cutting certification cost 10–30× with betting sequences.
citing papers explorer
-
Localizing Model Behavior with Path Patching
Path patching provides a method to express and quantitatively test hypotheses that neural network behaviors are localized to sets of paths.
-
Certified Interventional Fidelity: Anytime-Valid, Adaptive Evaluation of Causal Claims in Mechanistic Interpretability
CIF certifies interventional interpretability metrics as bounded causal means with anytime-valid confidence sequences, including under adaptive sampling, cutting certification cost 10–30× with betting sequences.