Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:47:35.336584Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2504.18590.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:47:35.336584Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b728c857-212f-4087-bdef-ac34a6a7c28e · outbound
A multilevel approach to accelerate the training of Transformers Avelin and K
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9d6f7dbd-ab6d-4e3e-b06c-73b6ac5c6935 · outbound
A multilevel approach to accelerate the training of Transformers N-ODE Transformer: A Depth-Adaptive Variant of the Transformer Using Neural Ordinary Differential Equations
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a400bd7-f962-48c1-9774-c890f73bfbdb · outbound
A multilevel approach to accelerate the training of Transformers Brown, B
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 28bf26ff-921c-4538-9a49-1a4261f46b44 · outbound
A multilevel approach to accelerate the training of Transformers Chang, W
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c26b12c7-9df5-4539-b7e7-df7cd162f1fa · outbound
A multilevel approach to accelerate the training of Transformers bert2BERT: Towards Reusable Pretrained Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f457159a-da31-4971-a761-283c595066ab · outbound
A multilevel approach to accelerate the training of Transformers Neural Ordinary Differential Equations
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4707cc99-010e-4f24-b773-94edf14dc961 · outbound
A multilevel approach to accelerate the training of Transformers Net2Net: Accelerating Learning via Knowledge Transfer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0a3eff5-2460-4d3e-a662-09cdeec8b0ad · outbound
A multilevel approach to accelerate the training of Transformers Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 505345ce-d9bd-4783-974c-c3be0b8802be · outbound
A multilevel approach to accelerate the training of Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132c4ec8-514f-4ddc-8667-ad41fefcd172 · outbound
A multilevel approach to accelerate the training of Transformers Gaedke-Merzhäuser, A
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 75745099-7f8c-4d52-a522-0d47aaf49c69 · outbound
A multilevel approach to accelerate the training of Transformers Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1f19f6bf-a763-4515-bbb0-4ecf31ca5fe5 · outbound
A multilevel approach to accelerate the training of Transformers Gratton, V
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 31a7479a-bc73-4642-a71a-028bc7d50c46 · outbound
A multilevel approach to accelerate the training of Transformers On the Transformer Growth for Progressive BERT Training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a0ceb076-e790-41a4-824e-bf8b9b3125c5 · outbound
A multilevel approach to accelerate the training of Transformers Scaling Laws for Neural Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 232b848b-4c0f-4b07-9fd7-30af0633d1b8 · outbound
A multilevel approach to accelerate the training of Transformers Kopaniˇcáková and R
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3cb45f2f-e467-426a-8da9-a6932d1bf5cc · outbound
A multilevel approach to accelerate the training of Transformers Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 588b1dc4-8636-4783-be1f-38e77e8d8df4 · outbound
A multilevel approach to accelerate the training of Transformers Lauga, E
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8c2f8830-7220-476b-affd-18c20cc474e3 · outbound
A multilevel approach to accelerate the training of Transformers ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7a61d24e-f828-4e9b-b5b7-3a64fe57e723 · outbound
A multilevel approach to accelerate the training of Transformers Understanding the Difficulty of Training Transformers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de837293-0650-49db-ba50-e0525d447baf · outbound
A multilevel approach to accelerate the training of Transformers Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c73b403a-9811-4f42-9249-20da5fcef414 · outbound
A multilevel approach to accelerate the training of Transformers Exploring Transformers for Large-Scale Speech Recognition
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1ab8717-6967-42c0-9909-0c2039fb4ba6 · outbound
A multilevel approach to accelerate the training of Transformers Penedo, H
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e282aa30-7704-413c-ae16-182c669bb71a · outbound
A multilevel approach to accelerate the training of Transformers Quemener and M
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7795f1e0-724b-4a88-99db-188c588e0399 · outbound
A multilevel approach to accelerate the training of Transformers Radford, J
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c27f562-a2f7-4ae1-a25a-a92ed8a62b65 · outbound
A multilevel approach to accelerate the training of Transformers LLaMA: Open and Efficient Foundation Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f157f21-6be3-440b-ab52-2fbf396242b6 · outbound
A multilevel approach to accelerate the training of Transformers Vaswani, N
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b4feb1a5-ce67-46aa-962c-446a901e70a7 · outbound
A multilevel approach to accelerate the training of Transformers Learning to Grow Pretrained Models for Efficient Transformer Training
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2488d91-bc84-4c69-9bb7-e3e961e3f049 · outbound
A multilevel approach to accelerate the training of Transformers Speeding up Deep Model Training by Sharing Weights and Then Unsharing
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f58139-59c4-4296-9cc3-9bc13904363c · outbound
A multilevel approach to accelerate the training of Transformers Deconstructing What Makes a Good Optimizer for Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 129c71b4-b63a-4541-b10f-b69a28dfa524 · outbound
A multilevel approach to accelerate the training of Transformers A Multi-Level Framework for Accelerating Training Transformer Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.