Attention Layers Add Into Low-Dimensional Residual Subspaces , url =

URL https: //transformer-circuits · 2024 · arXiv 2508.16929

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

read on arXiv browse 3 citing papers

citation-role summary

background 1

citation-polarity summary

background 1

representative citing papers

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers

cs.LG · 2026-05-29 · unverdicted · novelty 6.0

Contribution Weights combine attention, value magnitude, and directional alignment to measure token influence more faithfully than attention alone, and show attention sinks actively suppress information via a convex sink-rate to output-norm relationship.

When and How Long? The Readout-Mediator Angle in Temporal Reasoning

cs.LG · 2026-05-27 · unverdicted · novelty 6.0

Linear probes recover day-of-year from LM activations for temporal reasoning but are orthogonal to the model's causal 4D subspace identified by DAS, with the angle matching the Haar-uniform random null, replicated across scales and families.

RankUp: Towards High-rank Representations for Large Scale Advertising Recommender Systems

cs.IR · 2026-04-20 · unverdicted · novelty 6.0 · 2 refs

RankUp raises effective rank of representations in deep MetaFormer recommenders via randomized splitting and multi-embeddings, delivering 2-5% GMV gains in production deployments at Weixin.

citing papers explorer

Showing 2 of 2 citing papers after filters.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cs.LG · 2026-05-29 · unverdicted · none · ref 41
Contribution Weights combine attention, value magnitude, and directional alignment to measure token influence more faithfully than attention alone, and show attention sinks actively suppress information via a convex sink-rate to output-norm relationship.
When and How Long? The Readout-Mediator Angle in Temporal Reasoning cs.LG · 2026-05-27 · unverdicted · none · ref 11
Linear probes recover day-of-year from LM activations for temporal reasoning but are orthogonal to the model's causal 4D subspace identified by DAS, with the angle matching the Haar-uniform random null, replicated across scales and families.

Attention Layers Add Into Low-Dimensional Residual Subspaces , url =

citation-role summary

citation-polarity summary

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer