Pith. sign in

hub

Understanding warmup-stable-decay learning rates: A river valley loss landscape perspective

14 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.

14 Pith papers citing it
1 external citations · external index

hub tools

citation-role summary

background 2

citation-polarity summary

years

2026 13 2025 1

roles

background 2

polarities

background 1 unclear 1

representative citing papers

Layer Collapse in Diffusion Language Models

cs.LG · 2026-05-07 · unverdicted · novelty 7.0 · 2 refs

Diffusion language models develop early-layer collapse around an indispensable super-outlier due to overtraining, resulting in higher compressibility and reversed optimal sparsity patterns versus autoregressive models.

NITP: Next Implicit Token Prediction for LLM Pre-training

cs.CL · 2026-05-24 · conditional · novelty 6.0 · 2 refs

NITP augments next-token prediction with cosine alignment to stop-gradient shallow-layer features of the next token, improving geometry and downstream scores at ~2% extra training FLOPs.

Efficient Pre-Training with Token Superposition

cs.CL · 2026-05-07 · unverdicted · novelty 5.0 · 2 refs

Token-Superposition Training combines multiple tokens into bags for multi-hot cross-entropy pre-training followed by a recovery phase, yielding up to 2.5x reduction in training time at 10B scale under equal-loss conditions.

Scaling Latent Reasoning via Looped Language Models

cs.CL · 2025-10-29 · conditional · novelty 5.0

A 1.4B and a 2.6B looped (weight-tied, recurrent-depth) language model trained on 7.7T tokens match or exceed several 4B–8B transformer baselines on selected reasoning benchmarks.

citing papers explorer

Showing 14 of 14 citing papers.