Pith. sign in

REVIEW 8 cited by

Mathematical Models of Computation in Superposition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.05451 v1 pith:BXO32ZJO submitted 2024-08-10 cs.LG cs.AI

classification cs.LGcs.AI
keywords superpositionemphcomputationconstructfeaturestaskwhenwork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Superposition -- when a neural network represents more ``features'' than it has dimensions -- seems to pose a serious challenge to mechanistically interpreting current AI systems. Existing theory work studies \emph{representational} superposition, where superposition is only used when passing information through bottlenecks. In this work, we present mathematical models of \emph{computation} in superposition, where superposition is actively helpful for efficiently accomplishing the task. We first construct a task of efficiently emulating a circuit that takes the AND of the $\binom{m}{2}$ pairs of each of $m$ features. We construct a 1-layer MLP that uses superposition to perform this task up to $\varepsilon$-error, where the network only requires $\tilde{O}(m^{\frac{2}{3}})$ neurons, even when the input features are \emph{themselves in superposition}. We generalize this construction to arbitrary sparse boolean circuits of low depth, and then construct ``error correction'' layers that allow deep fully-connected networks of width $d$ to emulate circuits of width $\tilde{O}(d^{1.5})$ and \emph{any} polynomial depth. We conclude by providing some potential applications of our work for interpreting neural networks that implement computation in superposition.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Emergent Misalignment Recruits a Pre-existing Persona Subspace

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Fine-tuning on narrow bad data recruits a low-rank persona subspace already present in a frozen instruction-tuned model; holding that subspace out of activations prevents broad misalignment, and injecting it into the ...

  2. Training, Reading, and Editing Legible Transformers

    cs.LG 2026-07 conditional novelty 6.5 of 10

    A variance-floor objective plus learned operator gates produce an end-to-end legible transformer whose crisp units are 50–184× more local to edit and can be reshaped from fan-out to fan-in circuits without quality loss.

  3. Compressed Computation under $L^4$ Loss is likely Computation in Superposition

    cs.LG 2026-07 accept novelty 6.5 of 10

    Training a 50-neuron ReLU network under L4 loss elicits sparse binary codewords over neurons that compute 100 sparse ReLUs in superposition, recovered by a three-scalar ansatz.

  4. Scaling Interpretable Transformers with Parity Bottleneck Layers

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A transformer with fixed algebraic parity-hash bottlenecks has features that are native to the computation and beat post-hoc SAE features on steering, causal edits, and absorption, at a 6-9x training-token tax.

  5. Stochastic Parameter Decomposition

    cs.LG 2025-06 conditional novelty 6.0 of 10

    SPD uses stochastic masking and a learned causal importance function to decompose neural network parameters into sparsely active rank-one subcomponents, recovering ground-truth mechanisms in toy models where APD struggled.

  6. Expand Neurons, Not Parameters

    cs.LG 2025-10 reject novelty 5.0 of 10

    Fixed Parameter Expansion — duplicating neurons and partitioning their incoming weights into disjoint sparse sub-neurons at constant non-zero parameter count — reduces measured feature interference and improves classi...

  7. Linear Spatial World Models Emerge in Large Language Models

    cs.AI 2025-06 reject novelty 5.0 of 10

    Spatial relation words in LLaMA and Qwen models form antipodal, roughly orthogonal directions in a low-dimensional subspace, and steering along these directions changes the model's output.

  8. Feature learning is decoupled from generalization in high capacity neural networks

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Current feature learning measures quantify the magnitude of representation change, which the authors argue is decoupled from the generalization benefit that neural networks show over their neural tangent kernel.

Pith tools