Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:12.948856Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2505.21910.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:12.948856Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 257220fa-7306-4f3c-b984-ee6914197076 · outbound
Taming Transformer Without Using Learning Rate Warmup Rezero is all you need: Fast convergence at large depth
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 627bff35-0fc5-4536-92e4-04b298512953 · outbound
Taming Transformer Without Using Learning Rate Warmup Language models are few-shot learners
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17532fd4-7a2f-4153-97f8-e259ec0ba405 · outbound
Taming Transformer Without Using Learning Rate Warmup Palm: Scaling language modeling with pathways
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ade93d-5437-4261-bc70-9188d76b02f4 · outbound
Taming Transformer Without Using Learning Rate Warmup The Road Less Scheduled
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f28c3483-64e4-4317-a86e-c5d6bc0d941b · outbound
Taming Transformer Without Using Learning Rate Warmup Scaling vision transformers to 22 billion parameters
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0321042-1743-4bb4-b921-325db581498a · outbound
Taming Transformer Without Using Learning Rate Warmup Imagenet: A large-scale hierarchical image database
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e99ca12-4aa5-42fd-855b-6fc2a6055561 · outbound
Taming Transformer Without Using Learning Rate Warmup Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7372e20c-c47f-4841-82bb-5da079e2573d · outbound
Taming Transformer Without Using Learning Rate Warmup An image is worth 16x16 words: Transformers for image recognition at scale
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 854b518f-d178-43ed-95c7-47b460d6c760 · outbound
Taming Transformer Without Using Learning Rate Warmup The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d753c21-bfb2-4229-9cd0-22251710c090 · outbound
Taming Transformer Without Using Learning Rate Warmup Adaptive subgradient methods for online learning and stochastic optimization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08543823-8ff1-4e2a-8b78-f5b61eade4d0 · outbound
Taming Transformer Without Using Learning Rate Warmup Openwebtext corpus
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4272df13-9591-4304-bd64-5ceac5594f86 · outbound
Taming Transformer Without Using Learning Rate Warmup Matrix computations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c820b89-57cd-4d03-bd41-95f95040d735 · outbound
Taming Transformer Without Using Learning Rate Warmup Kronecker products and matrix calculus with applications
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f33d8bf-4ad0-4313-90a8-9e874efbb050 · outbound
Taming Transformer Without Using Learning Rate Warmup Flatten transformer: Vision transformer using focused linear attention
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fe4f361-0ce1-4460-8fee-f92622469b58 · outbound
Taming Transformer Without Using Learning Rate Warmup Query-key normalization for transformers
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49aee7c9-c154-4248-94e0-cb2289f9f928 · outbound
Taming Transformer Without Using Learning Rate Warmup Topics in matrix analysis, 1991
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ffb6ea69-f6dd-4d8a-a5f6-2dc73da0c39e · outbound
Taming Transformer Without Using Learning Rate Warmup Matrix analysis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f477f785-e69f-4b32-b923-bdf9488594cd · outbound
Taming Transformer Without Using Learning Rate Warmup Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 169cfeff-2ab4-421b-a40d-5cbd51a47a97 · outbound
Taming Transformer Without Using Learning Rate Warmup The lipschitz constant of self-attention
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4d968143-41b1-4785-b4cf-c4f63278e786 · outbound
Taming Transformer Without Using Learning Rate Warmup Adam: A Method for Stochastic Optimization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd387253-ef1e-484f-82d0-1ba700eba51c · outbound
Taming Transformer Without Using Learning Rate Warmup Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6944f4a5-f7d5-4d2e-bb92-c4b0edd843f4 · outbound
Taming Transformer Without Using Learning Rate Warmup Analyzing & reducing the need for learning rate warmup in gpt training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0933c51d-06cf-4d25-ad53-f083e3f78eac · outbound
Taming Transformer Without Using Learning Rate Warmup Backpropagation applied to handwritten zip code recognition
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c57049ec-059f-4631-93b5-4436c488d358 · outbound
Taming Transformer Without Using Learning Rate Warmup Gradient-based learning applied to document recognition
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb3209bf-7f47-48f9-ab17-046a1f130195 · outbound
Taming Transformer Without Using Learning Rate Warmup Efficient backprop
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79db7c08-6028-4904-8bae-c0774fd91aa3 · outbound
Taming Transformer Without Using Learning Rate Warmup Understanding the difficulty of training transformers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2732e95-40cc-46f9-8159-874e53d30c51 · outbound
Taming Transformer Without Using Learning Rate Warmup Swin transformer: Hierarchical vision transformer using shifted windows
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd256782-b797-43e7-9712-d932a9feac97 · outbound
Taming Transformer Without Using Learning Rate Warmup SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f59261d-ae14-444f-ad00-38a8c63336dc · outbound
Taming Transformer Without Using Learning Rate Warmup Fixing weight decay regularization in adam
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd65d404-4628-46fa-a0c1-358733015b3b · outbound
Taming Transformer Without Using Learning Rate Warmup Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 598dda0a-3158-4bbf-b619-47ff234103bf · outbound
Taming Transformer Without Using Learning Rate Warmup Scalable diffusion models with transformers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f31e6bdf-8d8d-4933-90cb-1cea43146c53 · outbound
Taming Transformer Without Using Learning Rate Warmup The matrix cookbook
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d456acfd-fcd2-4553-857e-3f24fc2b9c91 · outbound
Taming Transformer Without Using Learning Rate Warmup Lipsformer: Introducing lipschitz continuity to vision transformers
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b97858a9-b3d1-466c-989f-5c2b8ab1aff8 · outbound
Taming Transformer Without Using Learning Rate Warmup Understanding Optimization of Deep Learning via Jacobian Matrix and Lipschitz Constant
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb65b2dd-5629-4c47-bf8a-3b19f731b856 · outbound
Taming Transformer Without Using Learning Rate Warmup Improving language understanding by generative pre-training
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa51fde8-5b81-49ae-86f2-e89ce4479b27 · outbound
Taming Transformer Without Using Learning Rate Warmup Language models are unsupervised multitask learners
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 023b964b-3eb8-430f-b40e-14da3b647713 · outbound
Taming Transformer Without Using Learning Rate Warmup Learning transferable visual models from natural language supervision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7c0e253-242e-48df-aecf-9a40aa387c4f · outbound
Taming Transformer Without Using Learning Rate Warmup Zero-shot text-to-image generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f679c03-3b1a-42c7-850d-ed29c4038639 · outbound
Taming Transformer Without Using Learning Rate Warmup A stochastic approximation method
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3f62cb-71d6-4f6f-a084-f403cb1c1f0d · outbound
Taming Transformer Without Using Learning Rate Warmup Learning representations by back-propagating errors
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16719b8a-8d27-4e6f-8747-3abe37ca8fdc · outbound
Taming Transformer Without Using Learning Rate Warmup Cyclical learning rates for training neural networks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c2748a7-0246-48b5-81c1-211a72517b3a · outbound
Taming Transformer Without Using Learning Rate Warmup Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 795361a3-6a79-46e7-ab88-0175b600b518 · outbound
Taming Transformer Without Using Learning Rate Warmup JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffe7fa62-6a19-4c8a-bdf5-0d1005c938a4 · outbound
Taming Transformer Without Using Learning Rate Warmup Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15ca0245-aad0-4545-a921-8d27e25009d2 · outbound
Taming Transformer Without Using Learning Rate Warmup Attention is all you need
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61e6a18f-7c0a-40fe-9936-df9204792b23 · outbound
Taming Transformer Without Using Learning Rate Warmup High-dimensional probability: An introduction with applications in data science, volume 47
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5041f34-d033-4787-b3bf-16da13f0a5b6 · outbound
Taming Transformer Without Using Learning Rate Warmup DeepNet: Scaling Transformers to 1,000 Layers
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f2774c-3042-4510-900d-333208208d21 · outbound
Taming Transformer Without Using Learning Rate Warmup Learning Deep Transformer Models for Machine Translation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3150a55-86c8-429c-8061-24f545beced3 · outbound
Taming Transformer Without Using Learning Rate Warmup Pytorch image models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46f3d844-ff54-4cfb-bdc5-1da34e717569 · outbound
Taming Transformer Without Using Learning Rate Warmup High-dimensional data analysis with low-dimensional models: Principles, computation, and applications
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a1d6b40-b659-4b7d-858c-64b2334feedc · outbound
Taming Transformer Without Using Learning Rate Warmup On layer normalization in the transformer architecture
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10660810-ac57-42c8-835a-697d9671cbea · outbound
Taming Transformer Without Using Learning Rate Warmup Stabilizing transformer training by preventing attention entropy collapse
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc1460b6-f112-4c9c-b2ed-ea07f901c0d4 · outbound
Taming Transformer Without Using Learning Rate Warmup Root mean square layer normalization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42508aa3-13a1-450e-8e89-2b0aae900f3f · outbound
Taming Transformer Without Using Learning Rate Warmup write newline
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a02c986-71ee-4c54-8c00-217fe0d1b11c · outbound
Taming Transformer Without Using Learning Rate Warmup @esa (Ref
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55fd245d-a651-4a37-866a-fd997f6cb4e7 · outbound
Taming Transformer Without Using Learning Rate Warmup Unresolved cited work
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5978d486-dbf6-47f8-a283-ca2b4fba5aaf · outbound
Taming Transformer Without Using Learning Rate Warmup Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.