Pith. sign in

Canonical reference

Deepnet: Scaling transformers to 1,000 layers.arXiv preprint arXiv:2203.00555,

Canonical reference. 80% of citing Pith papers cite this work as background.

15 Pith papers citing it
54 external citations · external index
Background 80% of classified citations

citation-role summary

background 5

citation-polarity summary

roles

background 5

polarities

background 4 unclear 1

representative citing papers

Delta Attention Residuals

cs.LG · 2026-05-13 · unverdicted · novelty 7.0

Delta Attention Residuals attend over per-sublayer deltas instead of cumulative hidden states, producing higher-contrast attention weights and 1.7-8.2% validation perplexity gains over standard and attention residuals across 220M-7.6B models.

HAARES Half-Split Residual Basis Routing for Deep Transformers

cs.LG · 2026-06-04 · unverdicted · novelty 5.0

HAARES is a lightweight residual basis router that augments block summaries with an RMS-matched half-split detail vector and reports consistent gains over Block AttnRes in 48-layer 201M models across language modeling benchmarks.

Attention Residuals

cs.CL · 2026-03-16 · unverdicted · novelty 5.0

Attention Residuals replaces fixed residual summation with input-dependent softmax attention over preceding layers, and a blocked variant is shown to improve uniformity and downstream performance in a 48B-parameter model pre-trained on 1.4T tokens.

A Survey of Large Language Models

cs.CL · 2023-03-31 · accept · novelty 3.0

This survey reviews the background, key techniques, and evaluation methods for large language models, emphasizing emergent abilities that appear at large scales.

citing papers explorer

Showing 15 of 15 citing papers.