Pith. sign in

REVIEW 1 cited by

Wasserstein Distances, Neuronal Entanglement, and Sparsity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15756 v4 pith:FNUQ3N7V submitted 2024-05-24 cs.LG cs.AI

classification cs.LGcs.AI
keywords neuronswassersteindisentanglingmixtureoutputaccuracydistanceentanglement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Disentangling polysemantic neurons is at the core of many current approaches to interpretability of large language models. Here we attempt to study how disentanglement can be used to understand performance, particularly under weight sparsity, a leading post-training optimization technique. We suggest a novel measure for estimating neuronal entanglement: the Wasserstein distance of a neuron's output distribution to a Gaussian. Moreover, we show the existence of a small number of highly entangled "Wasserstein Neurons" in each linear layer of an LLM, characterized by their highly non-Gaussian output distributions, their role in mapping similar inputs to dissimilar outputs, and their significant impact on model accuracy. To study these phenomena, we propose a new experimental framework for disentangling polysemantic neurons. Our framework separates each layer's inputs to create a mixture of experts where each neuron's output is computed by a mixture of neurons of lower Wasserstein distance, each better at maintaining accuracy when sparsified without retraining. We provide strong evidence that this is because the mixture of sparse experts is effectively disentangling the input-output relationship of individual neurons, in particular the difficult Wasserstein neurons.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SwiftPrune: Hessian-Free Weight Pruning for Large Language Models

    cs.LG 2025-01 conditional novelty 6.0 of 10

    SwiftPrune prunes LLMs in seconds using a Hessian-free importance metric plus an EWMA threshold, with an O(n) pruning pass and strong 2:4 structured sparsity results.

Pith tools