Pith. sign in

REVIEW 2 cited by

Identifying Sparsely Active Circuits Through Local Loss Landscape Decomposition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.00194 v1 pith:V3CQYRPM submitted 2025-03-31 cs.LG cs.AI

classification cs.LGcs.AI
keywords circuitsdecompositionlosssubnetworksactivationlandscapelocalmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Much of mechanistic interpretability has focused on understanding the activation spaces of large neural networks. However, activation space-based approaches reveal little about the underlying circuitry used to compute features. To better understand the circuits employed by models, we introduce a new decomposition method called Local Loss Landscape Decomposition (L3D). L3D identifies a set of low-rank subnetworks: directions in parameter space of which a subset can reconstruct the gradient of the loss between any sample's output and a reference output vector. We design a series of progressively more challenging toy models with well-defined subnetworks and show that L3D can nearly perfectly recover the associated subnetworks. Additionally, we investigate the extent to which perturbing the model in the direction of a given subnetwork affects only the relevant subset of samples. Finally, we apply L3D to a real-world transformer model and a convolutional neural network, demonstrating its potential to identify interpretable and relevant circuits in parameter space.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Individual Parameters in Weight-Sparse Transformers Appear Interpretable

    cs.LG 2026-07 conditional novelty 6.5 of 10

    An automated LLM pipeline finds that 12–31% of nonzero weights in weight-sparse transformers admit short, held-out-validated descriptions of when they matter, far above dense controls.

  2. Stochastic Parameter Decomposition

    cs.LG 2025-06 conditional novelty 6.0 of 10

    SPD uses stochastic masking and a learned causal importance function to decompose neural network parameters into sparsely active rank-one subcomponents, recovering ground-truth mechanisms in toy models where APD struggled.

Pith tools