Pith. sign in

The curse of depth in large language models.arXiv preprint arXiv:2502.05795

11 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.

11 Pith papers citing it
1 external citations · external index

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 8 cs.AI 3

years

2026 11

roles

background 1

polarities

background 1

representative citing papers

AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization

cs.LG · 2026-06-03 · unverdicted · novelty 6.0

AlphaQ performs calibration-free mixed-precision quantization of MoE models by allocating higher bits to experts whose weight spectra exhibit stronger heavy-tailed structure according to HT-SR theory, outperforming calibration-based methods and reaching near full-precision accuracy at 3.5 average bi

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm

cs.LG · 2026-02-08 · unverdicted · novelty 6.0

SiameseNorm is a two-stream architecture that reconciles Pre-Norm and Post-Norm in Transformers by coupling streams via shared residual blocks, yielding performance gains with maintained stability on language, vision, and diffusion models.

citing papers explorer

Showing 11 of 11 citing papers.