Pith. sign in

REVIEW 3 cited by

Confidence Regulation Neurons in Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.16254 v2 pith:CPOHFUVY submitted 2024-06-24 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords neuronsentropyfrequencymodelstokencomponentsconfidencedistribution
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite their widespread use, the mechanisms by which large language models (LLMs) represent and regulate uncertainty in next-token predictions remain largely unexplored. This study investigates two critical components believed to influence this uncertainty: the recently discovered entropy neurons and a new set of components that we term token frequency neurons. Entropy neurons are characterized by an unusually high weight norm and influence the final layer normalization (LayerNorm) scale to effectively scale down the logits. Our work shows that entropy neurons operate by writing onto an unembedding null space, allowing them to impact the residual stream norm with minimal direct effect on the logits themselves. We observe the presence of entropy neurons across a range of models, up to 7 billion parameters. On the other hand, token frequency neurons, which we discover and describe here for the first time, boost or suppress each token's logit proportionally to its log frequency, thereby shifting the output distribution towards or away from the unigram distribution. Finally, we present a detailed case study where entropy neurons actively manage confidence in the setting of induction, i.e. detecting and continuing repeated subsequences.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Understanding Gated Neurons in Transformers from Their Input-Output Functionality

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Across 12 language models, neurons in early-middle layers tend to add the direction they detect (enrichment), while later layers tend to reduce it (depletion), based on input-output weight cosine similarity.

  2. Sparse Neuron Ablation Triggers Catastrophic Collapse of the Language Core in Large Vision-Language Models

    cs.AI 2025-11 conditional novelty 5.0 of 10

    Ablating just four neurons in LLaVA-1.5-7b's language-model down-projection layer triggers complete output collapse, with critical neurons concentrated in the language backbone.

  3. Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

    cs.LG 2025-07 conditional novelty 5.0 of 10

    LayerNorm can be removed from all GPT-2 models by fine-tuning with a linear replacement, losing only a small amount of validation accuracy on filtered data.

Pith tools