Pith. sign in

REVIEW 4 cited by

Interpreting Neural Networks through the Polytope Lens

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.12312 v1 pith:QIQ35BLA submitted 2022-11-22 cs.LG cs.AI

classification cs.LGcs.AI
keywords neuralpolytopedirectionslensnetworkneuronsactivationcombinations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mechanistic interpretability aims to explain what a neural network has learned at a nuts-and-bolts level. What are the fundamental primitives of neural network representations? Previous mechanistic descriptions have used individual neurons or their linear combinations to understand the representations a network has learned. But there are clues that neurons and their linear combinations are not the correct fundamental units of description: directions cannot describe how neural networks use nonlinearities to structure their representations. Moreover, many instances of individual neurons and their combinations are polysemantic (i.e. they have multiple unrelated meanings). Polysemanticity makes interpreting the network in terms of neurons or directions challenging since we can no longer assign a specific feature to a neural unit. In order to find a basic unit of description that does not suffer from these problems, we zoom in beyond just directions to study the way that piecewise linear activation functions (such as ReLU) partition the activation space into numerous discrete polytopes. We call this perspective the polytope lens. The polytope lens makes concrete predictions about the behavior of neural networks, which we evaluate through experiments on both convolutional image classifiers and language models. Specifically, we show that polytopes can be used to identify monosemantic regions of activation space (while directions are not in general monosemantic) and that the density of polytope boundaries reflect semantic boundaries. We also outline a vision for what mechanistic interpretability might look like through the polytope lens.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training, Reading, and Editing Legible Transformers

    cs.LG 2026-07 conditional novelty 6.5 of 10

    A variance-floor objective plus learned operator gates produce an end-to-end legible transformer whose crisp units are 50–184× more local to edit and can be reshaped from fan-out to fan-in circuits without quality loss.

  2. Patches of Nonlinearity: Instruction Vectors in Large Language Models

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Instruction-following in LLMs is mediated by localized, linearly separable 'instruction vectors' that behave superadditively and appear to select task-specific circuits in later layers.

  3. Even Faster Hyperbolic Random Forests: A Beltrami-Klein Wrapper Approach

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Fast-HyperDT reexpresses HyperDT as pre- and post-processing around standard Euclidean trees, making hyperbolic random forests practical.

  4. Learning Safety Constraints for Large Language Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A polytope learned in LLM representation space can detect unsafe regions and steer outputs back to safety at inference time, reducing jailbreak success across several models.

Pith tools