Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2203.00555.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:11.857239Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
54
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 2d3a80d7-d25f-4d34-b7a9-5dae2e110da1 · inbound
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness DeepNet: Scaling Transformers to 1,000 Layers
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f30cc71f-b14a-40bb-b558-4b59c1b0fced · inbound
Language Is Not All You Need: Aligning Perception with Language Models DeepNet: Scaling Transformers to 1,000 Layers
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4809d27e-dd9a-4897-b631-e2f5347ecf70 · inbound
A Survey of Large Language Models DeepNet: Scaling Transformers to 1,000 Layers
Reference 282
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1298970c-a84f-4672-b77e-b7ccfaab61fa · inbound
H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models DeepNet: Scaling Transformers to 1,000 Layers
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1b15f8d7-71c9-4b3f-876f-084f39a99204 · inbound
A Comprehensive Overview of Large Language Models DeepNet: Scaling Transformers to 1,000 Layers
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6f5ab099-0aec-47a4-b763-4e01a88edbb7 · inbound
Retentive Network: A Successor to Transformer for Large Language Models DeepNet: Scaling Transformers to 1,000 Layers
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a1019b78-1466-4b3d-8e4a-39007fabba98 · inbound
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free DeepNet: Scaling Transformers to 1,000 Layers
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5041f34-d033-4787-b3bf-16da13f0a5b6 · inbound
Taming Transformer Without Using Learning Rate Warmup DeepNet: Scaling Transformers to 1,000 Layers
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3fabdea-9f73-48c5-8c82-7cf8c357bba6 · inbound
The Algorithm Is Not the Behavior: Learned Priors Override Look-Ahead in a Chess-Playing Neural Network DeepNet: Scaling Transformers to 1,000 Layers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4979db94-aa97-440e-a0f4-16e9ca7455cf · inbound
Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers DeepNet: Scaling Transformers to 1,000 Layers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f5b1146-dfed-49fb-8341-7c72123b16ed · inbound
Attention Residuals DeepNet: Scaling Transformers to 1,000 Layers
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ae21a8cb-1144-436d-9275-5c41827c84ea · inbound
When Does Sparsity Mitigate the Curse of Depth in LLMs DeepNet: Scaling Transformers to 1,000 Layers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb4ec1ad-925a-4957-bf1b-2ff4e44b2321 · inbound
Delta Attention Residuals DeepNet: Scaling Transformers to 1,000 Layers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a55527ee-9603-4700-980d-b86a5cc646a0 · inbound
Prognostic Value of Lung Ultrasound Biomarkers for Readmission Risk in Congestive Heart Failure: A Pilot Data-Driven Analysis DeepNet: Scaling Transformers to 1,000 Layers
Reference 274
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bf866434-9986-41cb-9eb1-a100cf4345bc · inbound
HAARES Half-Split Residual Basis Routing for Deep Transformers DeepNet: Scaling Transformers to 1,000 Layers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 25bf54da-2c13-4d58-83eb-668bb44d8c1b · inbound
Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models DeepNet: Scaling Transformers to 1,000 Layers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 29b3ba39-91ba-46e7-8dac-bc818def2ea9 · inbound
CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry DeepNet: Scaling Transformers to 1,000 Layers
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d1beed07-4ff1-4138-b2c3-1e679ea7d257 · inbound
Review Residuals: Update-Conditioned Residual Gating for Transformers DeepNet: Scaling Transformers to 1,000 Layers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a5f73626-9260-4004-bab3-cfcd8552fb0b · inbound
AutoNorm: Understanding Adaptive Normalization in Transformers through Differentiable Gating DeepNet: Scaling Transformers to 1,000 Layers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a76f9c-c9e8-4c07-9061-6ae57e9dfcc8 · inbound
Dynamic Parameterization Is Not Dynamic Inference DeepNet: Scaling Transformers to 1,000 Layers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5b95787-2b6c-41e1-8e86-3381df5de39a · inbound
Multi-Head Attention Residuals DeepNet: Scaling Transformers to 1,000 Layers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b4514cd-824d-47b6-a1ff-8b93512acde5 · inbound
Multi-Head Attention Residuals DeepNet: Scaling Transformers to 1,000 Layers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.