LoopUS converts pretrained LLMs into looped latent refinement models via block decomposition, selective gating, random deep supervision, and confidence-based early exiting to improve reasoning performance.
Gated delta networks: Improving mamba2 with delta rule
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.LG 3years
2026 3verdicts
UNVERDICTED 3roles
background 3polarities
background 3representative citing papers
MinMax RNCs are recurrent networks over the min-max semiring that achieve regular language expressivity, log-depth parallel scan, uniformly bounded states, and non-vanishing state gradients while showing competitive empirical performance.
Kaczmarz Linear Attention replaces the empirical coefficient in Gated DeltaNet with a key-norm-normalized step size derived from the online regression objective, yielding lower perplexity and better needle-in-haystack performance.
citing papers explorer
-
LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models
LoopUS converts pretrained LLMs into looped latent refinement models via block decomposition, selective gating, random deep supervision, and confidence-based early exiting to improve reasoning performance.
-
MinMax Recurrent Neural Cascades
MinMax RNCs are recurrent networks over the min-max semiring that achieve regular language expressivity, log-depth parallel scan, uniformly bounded states, and non-vanishing state gradients while showing competitive empirical performance.
-
Kaczmarz Linear Attention
Kaczmarz Linear Attention replaces the empirical coefficient in Gated DeltaNet with a key-norm-normalized step size derived from the online regression objective, yielding lower perplexity and better needle-in-haystack performance.