Pith. sign in

REVIEW 2 cited by

Mechanistic Permutability: Match Features Across Layers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.07656 v3 pith:NV54W7FP submitted 2024-10-10 cs.LG

classification cs.LG
keywords layersfeaturesacrossfeaturemechanisticneuralaligningapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Understanding how features evolve across layers in deep neural networks is a fundamental challenge in mechanistic interpretability, particularly due to polysemanticity and feature superposition. While Sparse Autoencoders (SAEs) have been used to extract interpretable features from individual layers, aligning these features across layers has remained an open problem. In this paper, we introduce SAE Match, a novel, data-free method for aligning SAE features across different layers of a neural network. Our approach involves matching features by minimizing the mean squared error between the folded parameters of SAEs, a technique that incorporates activation thresholds into the encoder and decoder weights to account for differences in feature scales. Through extensive experiments on the Gemma 2 language model, we demonstrate that our method effectively captures feature evolution across layers, improving feature matching quality. We also show that features persist over several layers and that our approach can approximate hidden states across layers. Our work advances the understanding of feature dynamics in neural networks and provides a new tool for mechanistic interpretability studies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Birth of Knowledge: Emergent Features across Time, Space, and Scale in Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Sparse autoencoder probes of Pythia models show concept activations jump at roughly 410M parameters and during mid-training, while early-layer features re-emerge at the output layer.

  2. Cross-Layer Discrete Concept Discovery for Interpreting Language Models

    cs.LG 2025-06 reject novelty 5.0 of 10

    CLVQ-VAE maps lower-layer transformer activations to higher-layer ones through a discrete codebook, yielding concept vectors evaluated with probe ablation and human annotation.

Pith tools