Pith. sign in

Finding alignments between interpretable causal variables and distributed neural representations

7 Pith papers cite this work, alongside 9 external citations. Polarity classification is still indexing.

7 Pith papers citing it
9 external citations · external index

citation-role summary

background 1

citation-polarity summary

years

2026 5 2023 2

roles

background 1

polarities

unclear 1

representative citing papers

Localizing Model Behavior with Path Patching

cs.LG · 2023-04-12 · unverdicted · novelty 8.0

Path patching provides a method to express and quantitatively test hypotheses that neural network behaviors are localized to sets of paths.

From Mechanistic to Compositional Interpretability

cs.LG · 2026-05-09 · unverdicted · novelty 7.0

The paper introduces compositional interpretability as a category-theoretic framework that casts mechanistic explanations as commuting syntactic-semantic mappings optimized under faithfulness and complexity constraints derived from minimum description length.

citing papers explorer

Showing 7 of 7 citing papers.