Pith. sign in

Paper Citation Record · LEDGER

A multilevel approach to accelerate the training of Transformers

As of 17 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2504.18590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18590 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:47:35.336584Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b728c857-212f-4087-bdef-ac34a6a7c28e · outbound

This paper cites Avelin and K.

A multilevel approach to accelerate the training of Transformers Avelin and K

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.864843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.189309Z digest=sha256:86da2a32c90eecebc8a23a72862fa4139ca72aef03056ae9418072cc2f1a9fcf

Observation 9d6f7dbd-ab6d-4e3e-b06c-73b6ac5c6935 · outbound

This paper cites N-ODE Transformer: A Depth-Adaptive Variant of the Transformer Using Neural Ordinary Differential Equations.

A multilevel approach to accelerate the training of Transformers N-ODE Transformer: A Depth-Adaptive Variant of the Transformer Using Neural Ordinary Differential Equations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.194923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.194923Z digest=sha256:a2700738b76d0815d2354b066c0ba39df2c302491e4dad23e96579c2012b760e

Observation 7a400bd7-f962-48c1-9774-c890f73bfbdb · outbound

This paper cites Brown, B.

A multilevel approach to accelerate the training of Transformers Brown, B

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.849873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.200292Z digest=sha256:601d68b51b28f3d478ff009088f9448071eeacafefda8216b55f54f062461396

Observation 28bf26ff-921c-4538-9a49-1a4261f46b44 · outbound

This paper cites Chang, W.

A multilevel approach to accelerate the training of Transformers Chang, W

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.830869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.205590Z digest=sha256:dd4165d9d2cc5b1c82f604996459bdd69e1b109c693dbc5ed2d6be86a0ca9597

Observation c26b12c7-9df5-4539-b7e7-df7cd162f1fa · outbound

This paper cites bert2BERT: Towards Reusable Pretrained Language Models.

A multilevel approach to accelerate the training of Transformers bert2BERT: Towards Reusable Pretrained Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.210795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.210795Z digest=sha256:09a0dd6dd9c7f804c3bd9c6b48c654899e8e424852abd7f79b25e0efc4fda8a2

Observation f457159a-da31-4971-a761-283c595066ab · outbound

This paper cites Neural Ordinary Differential Equations.

A multilevel approach to accelerate the training of Transformers Neural Ordinary Differential Equations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.216045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.216045Z digest=sha256:367fa0becb1363ecd40d24038f218000eb93bb607daf783d89fbebf9ada57d92

Observation 4707cc99-010e-4f24-b773-94edf14dc961 · outbound

This paper cites Net2Net: Accelerating Learning via Knowledge Transfer.

A multilevel approach to accelerate the training of Transformers Net2Net: Accelerating Learning via Knowledge Transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.221598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.221598Z digest=sha256:d2847c7839e3d5666e626782267f2913d53f04a674b61e88b9a88db9d1de163e

Observation c0a3eff5-2460-4d3e-a662-09cdeec8b0ad · outbound

This paper cites an unresolved cited work.

A multilevel approach to accelerate the training of Transformers Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:35.815550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.226671Z digest=sha256:9b75232cd48a0be3a928bc70f636e25e0b190216c39bfc5b3336f871910fcf21

Observation 505345ce-d9bd-4783-974c-c3be0b8802be · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A multilevel approach to accelerate the training of Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.231248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.231248Z digest=sha256:7640fc8af22ff355b8779b071636a3eefba7e3ed415a7881dca9610414558463

Observation 132c4ec8-514f-4ddc-8667-ad41fefcd172 · outbound

This paper cites Gaedke-Merzhäuser, A.

A multilevel approach to accelerate the training of Transformers Gaedke-Merzhäuser, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.799927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.236367Z digest=sha256:80db5c5af78582f07501f42d1e075eb4fc7b4991e28ef575eb27b43a644bb39a

Observation 75745099-7f8c-4d52-a522-0d47aaf49c69 · outbound

This paper cites an unresolved cited work.

A multilevel approach to accelerate the training of Transformers Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:35.783616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.241107Z digest=sha256:63ca7db61de217434e903baa091f86defb6c61202317f90fe49583842f76d571

Observation 1f19f6bf-a763-4515-bbb0-4ecf31ca5fe5 · outbound

This paper cites Gratton, V.

A multilevel approach to accelerate the training of Transformers Gratton, V

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.767006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.245979Z digest=sha256:a9807c1b709711d4863f93b38f957d44b15b16b274270c7384e4037d4ab1b6a4

Observation 31a7479a-bc73-4642-a71a-028bc7d50c46 · outbound

This paper cites On the Transformer Growth for Progressive BERT Training.

A multilevel approach to accelerate the training of Transformers On the Transformer Growth for Progressive BERT Training

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:47:35.546153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.251108Z digest=sha256:d2de3df28c10bcd5927758903aa3c524f23c72719a4087c12a12e7fe57846058

Observation a0ceb076-e790-41a4-824e-bf8b9b3125c5 · outbound

This paper cites Scaling Laws for Neural Language Models.

A multilevel approach to accelerate the training of Transformers Scaling Laws for Neural Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.256820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.256820Z digest=sha256:542b2e4dfa47c59beb1c79e1f01e8b8317bfb8fb6c2b44187704d2b15d66b482

Observation 232b848b-4c0f-4b07-9fd7-30af0633d1b8 · outbound

This paper cites Kopaniˇcáková and R.

A multilevel approach to accelerate the training of Transformers Kopaniˇcáková and R

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.750989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.261885Z digest=sha256:56ffc8fee466f538674ae88f9168579e30f09caa18330040547bae69ab3aa899

Observation 3cb45f2f-e467-426a-8da9-a6932d1bf5cc · outbound

This paper cites an unresolved cited work.

A multilevel approach to accelerate the training of Transformers Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:35.734119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.266782Z digest=sha256:587cbb9b4066f910f4d40f108335807602a43ed378e2daf591108b1e446d32fa

Observation 588b1dc4-8636-4783-be1f-38e77e8d8df4 · outbound

This paper cites Lauga, E.

A multilevel approach to accelerate the training of Transformers Lauga, E

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.718722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.271991Z digest=sha256:492095fe9461d0e630a66437340c9bd06a89548e28efabfef8f97e83bfcc73ad

Observation 8c2f8830-7220-476b-affd-18c20cc474e3 · outbound

This paper cites ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation.

A multilevel approach to accelerate the training of Transformers ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:47:35.506056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.277017Z digest=sha256:8e11e3033af530341c69c6c829751060214f8b8f5f068b278958d36dcba964a2

Observation 7a61d24e-f828-4e9b-b5b7-3a64fe57e723 · outbound

This paper cites Understanding the Difficulty of Training Transformers.

A multilevel approach to accelerate the training of Transformers Understanding the Difficulty of Training Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.282331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.282331Z digest=sha256:97c272a5335decb50d84c1302add153c8c18bd54f750ac7b7264070834404231

Observation de837293-0650-49db-ba50-e0525d447baf · outbound

This paper cites an unresolved cited work.

A multilevel approach to accelerate the training of Transformers Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:35.703031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.287755Z digest=sha256:bbb203dd2879e5619132d151a0938aa2a712a686536aea00a6b21c6b9bc336dd

Observation c73b403a-9811-4f42-9249-20da5fcef414 · outbound

This paper cites Exploring Transformers for Large-Scale Speech Recognition.

A multilevel approach to accelerate the training of Transformers Exploring Transformers for Large-Scale Speech Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.292408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.292408Z digest=sha256:bc24bbb852507ba38ad1ab0331fa5898c9384522f250c1d2df5f638e70c5ba16

Observation c1ab8717-6967-42c0-9909-0c2039fb4ba6 · outbound

This paper cites Penedo, H.

A multilevel approach to accelerate the training of Transformers Penedo, H

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.686816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.297850Z digest=sha256:fd2db28d1578b4cc497a8f71d878821b3a6223846a034d80d77a7cf8b0d976db

Observation e282aa30-7704-413c-ae16-182c669bb71a · outbound

This paper cites Quemener and M.

A multilevel approach to accelerate the training of Transformers Quemener and M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.670791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.302570Z digest=sha256:5468fda93baac077b2a1b00ae09fa611012fb74a8f2b55a679d0df50ed9c8b48

Observation 7795f1e0-724b-4a88-99db-188c588e0399 · outbound

This paper cites Radford, J.

A multilevel approach to accelerate the training of Transformers Radford, J

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.307185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.307185Z digest=sha256:87f780289ea2c3abd1538ebffeba2b2f5d0c901832dde1afd04631886c4866cb

Observation 0c27f562-a2f7-4ae1-a25a-a92ed8a62b65 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A multilevel approach to accelerate the training of Transformers LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.311821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.311821Z digest=sha256:59ab02a9de8bdcded680d7f8135347d9fa6b9a571468472e6eb73c3c27229892

Observation 3f157f21-6be3-440b-ab52-2fbf396242b6 · outbound

This paper cites Vaswani, N.

A multilevel approach to accelerate the training of Transformers Vaswani, N

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.645250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.316770Z digest=sha256:9835bfa31403970d7c10e1b6ab6c3aa2a766709d76425cd7c11d64e6b66eb98b

Observation b4feb1a5-ce67-46aa-962c-446a901e70a7 · outbound

This paper cites Learning to Grow Pretrained Models for Efficient Transformer Training.

A multilevel approach to accelerate the training of Transformers Learning to Grow Pretrained Models for Efficient Transformer Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.321589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.321589Z digest=sha256:f0fe8d215106445c55afa52aeaf8cff6ac9b76dd8ece7b387cb9603e231d3acd

Observation f2488d91-bc84-4c69-9bb7-e3e961e3f049 · outbound

This paper cites Speeding up Deep Model Training by Sharing Weights and Then Unsharing.

A multilevel approach to accelerate the training of Transformers Speeding up Deep Model Training by Sharing Weights and Then Unsharing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.326873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.326873Z digest=sha256:4f229897e8cb576153e538acaf27e6c5b73348ac81db6d29901d785f70c1d234

Observation d6f58139-59c4-4296-9cc3-9bc13904363c · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

A multilevel approach to accelerate the training of Transformers Deconstructing What Makes a Good Optimizer for Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.331803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.331803Z digest=sha256:59d48e7d9e3c217f61aa5d3bb9577a2bbf916fb0b6e2cc07068993be188c3016

Observation 129c71b4-b63a-4541-b10f-b69a28dfa524 · outbound

This paper cites A Multi-Level Framework for Accelerating Training Transformer Models.

A multilevel approach to accelerate the training of Transformers A Multi-Level Framework for Accelerating Training Transformer Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:47:35.381340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:47:35.336584Z digest=sha256:25f357b5bc5c1ff968e8907321477a6734859bbf720a8bec8b543b20da16b8bf

Pith citing papers

No inbound Pith citation observations are available.