Pith. sign in

Paper Citation Record · LEDGER

A multilevel approach to accelerate the training of Transformers

As of 18 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2504.18590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18590 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:47:35.336584Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b728c857-212f-4087-bdef-ac34a6a7c28e · outbound

This paper cites Avelin and K.

A multilevel approach to accelerate the training of Transformers Avelin and K

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.864843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.189309Z digest=sha256:a4e915253efd31fe0ae83d64ef7aa3dd0fb113cb818f0d81654914dca9de02c0

Observation 9d6f7dbd-ab6d-4e3e-b06c-73b6ac5c6935 · outbound

This paper cites N-ODE Transformer: A Depth-Adaptive Variant of the Transformer Using Neural Ordinary Differential Equations.

A multilevel approach to accelerate the training of Transformers N-ODE Transformer: A Depth-Adaptive Variant of the Transformer Using Neural Ordinary Differential Equations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.194923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.194923Z digest=sha256:a2700738b76d0815d2354b066c0ba39df2c302491e4dad23e96579c2012b760e

Observation 7a400bd7-f962-48c1-9774-c890f73bfbdb · outbound

This paper cites Brown, B.

A multilevel approach to accelerate the training of Transformers Brown, B

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.849873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.200292Z digest=sha256:ecb11f247b4f7047118c54fff53d82f4eb7d3def74e6174e2b0ba8c996d937c6

Observation 28bf26ff-921c-4538-9a49-1a4261f46b44 · outbound

This paper cites Chang, W.

A multilevel approach to accelerate the training of Transformers Chang, W

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.830869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.205590Z digest=sha256:25b193f136bac6365ec2d60e3ccfc156351bd1d8d73467228a03c6bbd2a53b21

Observation c26b12c7-9df5-4539-b7e7-df7cd162f1fa · outbound

This paper cites bert2BERT: Towards Reusable Pretrained Language Models.

A multilevel approach to accelerate the training of Transformers bert2BERT: Towards Reusable Pretrained Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.210795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.210795Z digest=sha256:09a0dd6dd9c7f804c3bd9c6b48c654899e8e424852abd7f79b25e0efc4fda8a2

Observation f457159a-da31-4971-a761-283c595066ab · outbound

This paper cites Neural Ordinary Differential Equations.

A multilevel approach to accelerate the training of Transformers Neural Ordinary Differential Equations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.216045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.216045Z digest=sha256:367fa0becb1363ecd40d24038f218000eb93bb607daf783d89fbebf9ada57d92

Observation 4707cc99-010e-4f24-b773-94edf14dc961 · outbound

This paper cites Net2Net: Accelerating Learning via Knowledge Transfer.

A multilevel approach to accelerate the training of Transformers Net2Net: Accelerating Learning via Knowledge Transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.221598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.221598Z digest=sha256:d2847c7839e3d5666e626782267f2913d53f04a674b61e88b9a88db9d1de163e

Observation c0a3eff5-2460-4d3e-a662-09cdeec8b0ad · outbound

This paper cites an unresolved cited work.

A multilevel approach to accelerate the training of Transformers Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:35.815550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.226671Z digest=sha256:01f83bea68f1082412f8da9ebb6b126f8e227eef6a6e5a211cc0122f7b84ce5e

Observation 505345ce-d9bd-4783-974c-c3be0b8802be · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A multilevel approach to accelerate the training of Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.231248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.231248Z digest=sha256:7640fc8af22ff355b8779b071636a3eefba7e3ed415a7881dca9610414558463

Observation 132c4ec8-514f-4ddc-8667-ad41fefcd172 · outbound

This paper cites Gaedke-Merzhäuser, A.

A multilevel approach to accelerate the training of Transformers Gaedke-Merzhäuser, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.799927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.236367Z digest=sha256:53533c9e811b20a1845985fd653617f5a431bc113a21f55a5bb8e67eafd2bd28

Observation 75745099-7f8c-4d52-a522-0d47aaf49c69 · outbound

This paper cites an unresolved cited work.

A multilevel approach to accelerate the training of Transformers Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:35.783616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.241107Z digest=sha256:3ae97dba7b0e99bd8acfcbd403a8de1d95a58d55a90ed6cf747d10b771fb8b0f

Observation 1f19f6bf-a763-4515-bbb0-4ecf31ca5fe5 · outbound

This paper cites Gratton, V.

A multilevel approach to accelerate the training of Transformers Gratton, V

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.767006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.245979Z digest=sha256:318c7d56677bb53b308eeba6c0d24fe17744df366339223942a536976797edd0

Observation 31a7479a-bc73-4642-a71a-028bc7d50c46 · outbound

This paper cites On the Transformer Growth for Progressive BERT Training.

A multilevel approach to accelerate the training of Transformers On the Transformer Growth for Progressive BERT Training

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:47:35.546153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.251108Z digest=sha256:766f0ba12c0fe91926e9e189a2c60ad7863a0ce959c5f21d3892167148c73c1c

Observation a0ceb076-e790-41a4-824e-bf8b9b3125c5 · outbound

This paper cites Scaling Laws for Neural Language Models.

A multilevel approach to accelerate the training of Transformers Scaling Laws for Neural Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.256820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.256820Z digest=sha256:542b2e4dfa47c59beb1c79e1f01e8b8317bfb8fb6c2b44187704d2b15d66b482

Observation 232b848b-4c0f-4b07-9fd7-30af0633d1b8 · outbound

This paper cites Kopaniˇcáková and R.

A multilevel approach to accelerate the training of Transformers Kopaniˇcáková and R

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.750989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.261885Z digest=sha256:43eb83cf8c2c635e073fed58ccd80a8ed67e92e552449c83ed86a838d7cb0e3d

Observation 3cb45f2f-e467-426a-8da9-a6932d1bf5cc · outbound

This paper cites an unresolved cited work.

A multilevel approach to accelerate the training of Transformers Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:35.734119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.266782Z digest=sha256:d687f383e4050e2674e03d15cddf5c676ca891a2a84c3b8529fd901d0b23a7ff

Observation 588b1dc4-8636-4783-be1f-38e77e8d8df4 · outbound

This paper cites Lauga, E.

A multilevel approach to accelerate the training of Transformers Lauga, E

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.718722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.271991Z digest=sha256:9085b859e568452f7012c502139d54a782fcaa8bcebd0b04d73e0a37d2e4697a

Observation 8c2f8830-7220-476b-affd-18c20cc474e3 · outbound

This paper cites ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation.

A multilevel approach to accelerate the training of Transformers ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:47:35.506056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.277017Z digest=sha256:9ea0dfb23f146094bc1973cac2840ba614ae4d8b5f5cce78e46977dd3e39bc93

Observation 7a61d24e-f828-4e9b-b5b7-3a64fe57e723 · outbound

This paper cites Understanding the Difficulty of Training Transformers.

A multilevel approach to accelerate the training of Transformers Understanding the Difficulty of Training Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.282331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.282331Z digest=sha256:97c272a5335decb50d84c1302add153c8c18bd54f750ac7b7264070834404231

Observation de837293-0650-49db-ba50-e0525d447baf · outbound

This paper cites an unresolved cited work.

A multilevel approach to accelerate the training of Transformers Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:35.703031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.287755Z digest=sha256:b4b32b2ececfa7af3fc485c7a106c0c70be79ed29f4231c2f49563cd921ffaa0

Observation c73b403a-9811-4f42-9249-20da5fcef414 · outbound

This paper cites Exploring Transformers for Large-Scale Speech Recognition.

A multilevel approach to accelerate the training of Transformers Exploring Transformers for Large-Scale Speech Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.292408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.292408Z digest=sha256:bc24bbb852507ba38ad1ab0331fa5898c9384522f250c1d2df5f638e70c5ba16

Observation c1ab8717-6967-42c0-9909-0c2039fb4ba6 · outbound

This paper cites Penedo, H.

A multilevel approach to accelerate the training of Transformers Penedo, H

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.686816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.297850Z digest=sha256:53ea67d5bd8edd2519c791c183347944b54b2379c22fd703f69b0d83e85f180d

Observation e282aa30-7704-413c-ae16-182c669bb71a · outbound

This paper cites Quemener and M.

A multilevel approach to accelerate the training of Transformers Quemener and M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.670791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.302570Z digest=sha256:1a78acab8087dd5bf94d5e7a8237283447128737b6279eb2adbaec24e368c2c8

Observation 7795f1e0-724b-4a88-99db-188c588e0399 · outbound

This paper cites Radford, J.

A multilevel approach to accelerate the training of Transformers Radford, J

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.307185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.307185Z digest=sha256:87f780289ea2c3abd1538ebffeba2b2f5d0c901832dde1afd04631886c4866cb

Observation 0c27f562-a2f7-4ae1-a25a-a92ed8a62b65 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A multilevel approach to accelerate the training of Transformers LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.311821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.311821Z digest=sha256:59ab02a9de8bdcded680d7f8135347d9fa6b9a571468472e6eb73c3c27229892

Observation 3f157f21-6be3-440b-ab52-2fbf396242b6 · outbound

This paper cites Vaswani, N.

A multilevel approach to accelerate the training of Transformers Vaswani, N

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.645250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.316770Z digest=sha256:b1735494501ac477c91af7953c806b1a41a582749c408aafdf70cf7448288764

Observation b4feb1a5-ce67-46aa-962c-446a901e70a7 · outbound

This paper cites Learning to Grow Pretrained Models for Efficient Transformer Training.

A multilevel approach to accelerate the training of Transformers Learning to Grow Pretrained Models for Efficient Transformer Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.321589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.321589Z digest=sha256:a95881445db2ae25ff291ed710bea8eafaa02929978d2c9d0147c5d9b55f8d64

Observation f2488d91-bc84-4c69-9bb7-e3e961e3f049 · outbound

This paper cites Speeding up Deep Model Training by Sharing Weights and Then Unsharing.

A multilevel approach to accelerate the training of Transformers Speeding up Deep Model Training by Sharing Weights and Then Unsharing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.326873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.326873Z digest=sha256:4f229897e8cb576153e538acaf27e6c5b73348ac81db6d29901d785f70c1d234

Observation d6f58139-59c4-4296-9cc3-9bc13904363c · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

A multilevel approach to accelerate the training of Transformers Deconstructing What Makes a Good Optimizer for Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.331803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.331803Z digest=sha256:59d48e7d9e3c217f61aa5d3bb9577a2bbf916fb0b6e2cc07068993be188c3016

Observation 129c71b4-b63a-4541-b10f-b69a28dfa524 · outbound

This paper cites A Multi-Level Framework for Accelerating Training Transformer Models.

A multilevel approach to accelerate the training of Transformers A Multi-Level Framework for Accelerating Training Transformer Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:47:35.381340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:47:35.336584Z digest=sha256:674c52bb0c3abf1b3f91a6ddd77512f8c3d40d998498a778f8a860c3dfead80d

Pith citing papers

No inbound Pith citation observations are available.