Compositional interpretability defines explanations as commuting syntactic-semantic mapping pairs grounded in compositionality and minimum description length, with compressive refinement and a parsimony theorem guaranteeing concise human-aligned decompositions.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
Bilinear autoencoders decompose neural activations into low-rank quadratic forms to discover interpretable multi-dimensional manifolds, improving reconstruction in language models and challenging linear representation assumptions.
A stability evaluation framework probes parametric projections with Gaussian perturbations around anchor points to quantify mean displacement, bias, nearest-anchor error, and visualize local deformations on MNIST and Fashion-MNIST.
citing papers explorer
-
From Mechanistic to Compositional Interpretability
Compositional interpretability defines explanations as commuting syntactic-semantic mapping pairs grounded in compositionality and minimum description length, with compressive refinement and a parsimony theorem guaranteeing concise human-aligned decompositions.
-
Bilinear autoencoders find interpretable manifolds
Bilinear autoencoders decompose neural activations into low-rank quadratic forms to discover interpretable multi-dimensional manifolds, improving reconstruction in language models and challenging linear representation assumptions.
-
Local Neighborhood Instability in Parametric Projections: Quantitative and Visual Analysis
A stability evaluation framework probes parametric projections with Gaussian perturbations around anchor points to quantify mean displacement, bias, nearest-anchor error, and visualize local deformations on MNIST and Fashion-MNIST.