Pith. sign in

Paper Citation Record · LEDGER

Taming Transformer Without Using Learning Rate Warmup

As of 7 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2505.21910.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21910 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:12.948856Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 257220fa-7306-4f3c-b984-ee6914197076 · outbound

This paper cites Rezero is all you need: Fast convergence at large depth.

Taming Transformer Without Using Learning Rate Warmup Rezero is all you need: Fast convergence at large depth

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:16.871651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:05.383126Z digest=sha256:32fbae15e1675a432ce155e619b4b5654e1b329608c6aa1c6b29745363e6b6f9

Observation 627bff35-0fc5-4536-92e4-04b298512953 · outbound

This paper cites Language models are few-shot learners.

Taming Transformer Without Using Learning Rate Warmup Language models are few-shot learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.480251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.480251Z digest=sha256:9097416ed15938c81f67d9c38b20cc1bb2b35e065efd5b36afe201e936c0910a

Observation 17532fd4-7a2f-4153-97f8-e259ec0ba405 · outbound

This paper cites Palm: Scaling language modeling with pathways.

Taming Transformer Without Using Learning Rate Warmup Palm: Scaling language modeling with pathways

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.525225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.525225Z digest=sha256:56eea45ffb6b418ec4d7ac4dedc20b37eefdd1b1c7802a9466ec81d7ecf4d404

Observation 42ade93d-5437-4261-bc70-9188d76b02f4 · outbound

This paper cites The Road Less Scheduled.

Taming Transformer Without Using Learning Rate Warmup The Road Less Scheduled

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.649809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.649809Z digest=sha256:0dda0110adf506f1fa2b064c0b819077b4a93a823597784076df4607d80139c1

Observation f28c3483-64e4-4317-a86e-c5d6bc0d941b · outbound

This paper cites Scaling vision transformers to 22 billion parameters.

Taming Transformer Without Using Learning Rate Warmup Scaling vision transformers to 22 billion parameters

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.715302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.715302Z digest=sha256:19e1362abbf32cb568ad1206332fe039c6074d9bdda254a08dd707f7f5e96f15

Observation e0321042-1743-4bb4-b921-325db581498a · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Taming Transformer Without Using Learning Rate Warmup Imagenet: A large-scale hierarchical image database

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.764111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.764111Z digest=sha256:232a6b0c9b0c3c94e128683ae21129465b1889096d7110d9ff07fb08f6caf60e

Observation 1e99ca12-4aa5-42fd-855b-6fc2a6055561 · outbound

This paper cites Attention is not all you need: Pure attention loses rank doubly exponentially with depth.

Taming Transformer Without Using Learning Rate Warmup Attention is not all you need: Pure attention loses rank doubly exponentially with depth

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:16.593547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:05.838849Z digest=sha256:0fcee860bac200487138b0e32f55045d8a66c603598ac1363a5b60f0a6471e81

Observation 7372e20c-c47f-4841-82bb-5da079e2573d · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Taming Transformer Without Using Learning Rate Warmup An image is worth 16x16 words: Transformers for image recognition at scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.881498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.881498Z digest=sha256:ccbfa66c447e05f036d2f90fb92fc906cfda9524fb8f5801adc43e7ab69e0e55

Observation 854b518f-d178-43ed-95c7-47b460d6c760 · outbound

This paper cites The Llama 3 Herd of Models.

Taming Transformer Without Using Learning Rate Warmup The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.929214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.929214Z digest=sha256:6b44a2ff943b7aefeaf396d70f1fc5f86f9ca55297e495a36a2b3ff5a47fbaf6

Observation 2d753c21-bfb2-4229-9cd0-22251710c090 · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

Taming Transformer Without Using Learning Rate Warmup Adaptive subgradient methods for online learning and stochastic optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.005636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.005636Z digest=sha256:e5a1b51296796176ebb0f3955b639612462219d4c3a486723b72d04162b62636

Observation 08543823-8ff1-4e2a-8b78-f5b61eade4d0 · outbound

This paper cites Openwebtext corpus.

Taming Transformer Without Using Learning Rate Warmup Openwebtext corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.065481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.065481Z digest=sha256:1e15341a2177f74ed0255a74d160de881ce263b1df8a2a55fd34ec0c4abe1bb6

Observation 4272df13-9591-4304-bd64-5ceac5594f86 · outbound

This paper cites Matrix computations.

Taming Transformer Without Using Learning Rate Warmup Matrix computations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.120967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.120967Z digest=sha256:94b6103f9f820e7ec83849119871c20a9fe1f6b54ba354aade6c7f35ceb51273

Observation 0c820b89-57cd-4d03-bd41-95f95040d735 · outbound

This paper cites Kronecker products and matrix calculus with applications.

Taming Transformer Without Using Learning Rate Warmup Kronecker products and matrix calculus with applications

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:16.410854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.267817Z digest=sha256:d4e94f59500863dd0b278141c6bf25769301f80b10606148ceb0c74388ca43a7

Observation 8f33d8bf-4ad0-4313-90a8-9e874efbb050 · outbound

This paper cites Flatten transformer: Vision transformer using focused linear attention.

Taming Transformer Without Using Learning Rate Warmup Flatten transformer: Vision transformer using focused linear attention

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:16.242762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.354068Z digest=sha256:7f78e05e9079b53287a1848919785cf5e2a7bbd120300cdf1a8b62099fc08743

Observation 0fe4f361-0ce1-4460-8fee-f92622469b58 · outbound

This paper cites Query-key normalization for transformers.

Taming Transformer Without Using Learning Rate Warmup Query-key normalization for transformers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:15.997633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.453396Z digest=sha256:dc00e61a3d9eb468cfb3a703aad919bfa4cddaba76d85445aa7d46056f76e74d

Observation 49aee7c9-c154-4248-94e0-cb2289f9f928 · outbound

This paper cites Topics in matrix analysis, 1991.

Taming Transformer Without Using Learning Rate Warmup Topics in matrix analysis, 1991

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:15.735652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.516079Z digest=sha256:1c908c61839963bd6e05e171ebd122831bdeade5c33e54b8bb41e60a97d0c57e

Observation ffb6ea69-f6dd-4d8a-a5f6-2dc73da0c39e · outbound

This paper cites Matrix analysis.

Taming Transformer Without Using Learning Rate Warmup Matrix analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.586362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.586362Z digest=sha256:4c76e99f84de3c3112e0af9e8998cb5dd75fbf5984179f98e0dc675f490b0780

Observation f477f785-e69f-4b32-b923-bdf9488594cd · outbound

This paper cites an unresolved cited work.

Taming Transformer Without Using Learning Rate Warmup Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:26:15.518844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.722088Z digest=sha256:0012e2370be0a9373ccc812911fa921cb782c82e6356875e98616842aa103322

Observation 169cfeff-2ab4-421b-a40d-5cbd51a47a97 · outbound

This paper cites The lipschitz constant of self-attention.

Taming Transformer Without Using Learning Rate Warmup The lipschitz constant of self-attention

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:15.345563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.837314Z digest=sha256:11b5a4d029f16462b15b21cf95a19d930fa693ce603826d7c1c24cf5f8a7f765

Observation 4d968143-41b1-4785-b4cf-c4f63278e786 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Taming Transformer Without Using Learning Rate Warmup Adam: A Method for Stochastic Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.941052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.941052Z digest=sha256:c93347d21fa20dba6efbbaf420609aa163085eeee285b53c7b4a89d5e0455a76

Observation cd387253-ef1e-484f-82d0-1ba700eba51c · outbound

This paper cites Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks.

Taming Transformer Without Using Learning Rate Warmup Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:07.038954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:07.038954Z digest=sha256:425e66b36bc0b6588e0c5077e38fd3e5510562ec2436d799625ea636b68a5b57

Observation 6944f4a5-f7d5-4d2e-bb92-c4b0edd843f4 · outbound

This paper cites Analyzing & reducing the need for learning rate warmup in gpt training.

Taming Transformer Without Using Learning Rate Warmup Analyzing & reducing the need for learning rate warmup in gpt training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:15.212279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:07.233908Z digest=sha256:da7c55ebb9da86f9df1985bed75447a9f61f9c7f73923b476ad3e3834b3b2297

Observation 0933c51d-06cf-4d25-ad53-f083e3f78eac · outbound

This paper cites Backpropagation applied to handwritten zip code recognition.

Taming Transformer Without Using Learning Rate Warmup Backpropagation applied to handwritten zip code recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.994693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:07.415329Z digest=sha256:aec461a105b9249635c42b91a9715863ac90ec9a254ed3233bc1c8e5139cfa87

Observation c57049ec-059f-4631-93b5-4436c488d358 · outbound

This paper cites Gradient-based learning applied to document recognition.

Taming Transformer Without Using Learning Rate Warmup Gradient-based learning applied to document recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:07.676301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:07.676301Z digest=sha256:fbe65fe973f001792f84c7ae733eb451195f98056159a53da1fab2b836be78ee

Observation fb3209bf-7f47-48f9-ab17-046a1f130195 · outbound

This paper cites Efficient backprop.

Taming Transformer Without Using Learning Rate Warmup Efficient backprop

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.840983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:08.240797Z digest=sha256:6ec26e6a9f60326166f26c982264d356f0f92f872c06fc6b743e199ad40aef5f

Observation 79db7c08-6028-4904-8bae-c0774fd91aa3 · outbound

This paper cites Understanding the difficulty of training transformers.

Taming Transformer Without Using Learning Rate Warmup Understanding the difficulty of training transformers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.661451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:09.058294Z digest=sha256:fe38fc52291ccaf8b1c8bafd00f47594603ec842f457960fa868698a62cbfcaa

Observation f2732e95-40cc-46f9-8159-874e53d30c51 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Taming Transformer Without Using Learning Rate Warmup Swin transformer: Hierarchical vision transformer using shifted windows

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.453992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:09.982728Z digest=sha256:7ddcadf41aaba9c20a396a94858b820d4abd8a0d4fb96c9bd1554828965acd3f

Observation cd256782-b797-43e7-9712-d932a9feac97 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Taming Transformer Without Using Learning Rate Warmup SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.126111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.126111Z digest=sha256:d9b3b3b0f9a2781006bbfb2c017403fdecd3ba3f493e5824c1d466cf23f332b5

Observation 5f59261d-ae14-444f-ad00-38a8c63336dc · outbound

This paper cites Fixing weight decay regularization in adam.

Taming Transformer Without Using Learning Rate Warmup Fixing weight decay regularization in adam

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.237757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:10.206724Z digest=sha256:801e2e02d0110faa439b9e14ca4ad0fbff44354d8dbaa6737cc4fd0df8e4c295

Observation dd65d404-4628-46fa-a0c1-358733015b3b · outbound

This paper cites Signal propagation in transformers: Theoretical perspectives and the role of rank collapse.

Taming Transformer Without Using Learning Rate Warmup Signal propagation in transformers: Theoretical perspectives and the role of rank collapse

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.020113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:10.257907Z digest=sha256:d47a4344446b80c1f996c1bcbe0237a1c1d21efcab8802455875b95ed5a15c7f

Observation 598dda0a-3158-4bbf-b619-47ff234103bf · outbound

This paper cites Scalable diffusion models with transformers.

Taming Transformer Without Using Learning Rate Warmup Scalable diffusion models with transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.332842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.332842Z digest=sha256:36df4a0b9d6ac2e80801a2d07b4e6fba27fd924ea3120c3956a048e8e68059f6

Observation f31e6bdf-8d8d-4933-90cb-1cea43146c53 · outbound

This paper cites The matrix cookbook.

Taming Transformer Without Using Learning Rate Warmup The matrix cookbook

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.441366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.441366Z digest=sha256:5d92b77897533efe8d555e3bb1f2dc0a3a52e32abdc0131ae820c162374e69a7

Observation d456acfd-fcd2-4553-857e-3f24fc2b9c91 · outbound

This paper cites Lipsformer: Introducing lipschitz continuity to vision transformers.

Taming Transformer Without Using Learning Rate Warmup Lipsformer: Introducing lipschitz continuity to vision transformers

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:13.882027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:10.517077Z digest=sha256:1b70574dc5ef32975e400d9d6d77ae426d4d374b57c91a0c11ba27e140028ed5

Observation b97858a9-b3d1-466c-989f-5c2b8ab1aff8 · outbound

This paper cites Understanding Optimization of Deep Learning via Jacobian Matrix and Lipschitz Constant.

Taming Transformer Without Using Learning Rate Warmup Understanding Optimization of Deep Learning via Jacobian Matrix and Lipschitz Constant

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.591325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.591325Z digest=sha256:73292e42b8896a909cf6b9d1d51549b08d4a047634954a715b88b00dc9016c34

Observation fb65b2dd-5629-4c47-bf8a-3b19f731b856 · outbound

This paper cites Improving language understanding by generative pre-training.

Taming Transformer Without Using Learning Rate Warmup Improving language understanding by generative pre-training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.643909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.643909Z digest=sha256:0c02cf195714473ff65715952b595c2e76cf6f8738fdeb12c06b54d979b44841

Observation fa51fde8-5b81-49ae-86f2-e89ce4479b27 · outbound

This paper cites Language models are unsupervised multitask learners.

Taming Transformer Without Using Learning Rate Warmup Language models are unsupervised multitask learners

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.768095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.768095Z digest=sha256:a03e9366933fbe02897bc988220029ba64643aae5908d5ed3ef6e6fc397cfb69

Observation 023b964b-3eb8-430f-b40e-14da3b647713 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Taming Transformer Without Using Learning Rate Warmup Learning transferable visual models from natural language supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.824743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.824743Z digest=sha256:6f5d6d181d02fb15cab45eed45656f53fac8a7a9a245656978f280ee9dd1a666

Observation a7c0e253-242e-48df-aecf-9a40aa387c4f · outbound

This paper cites Zero-shot text-to-image generation.

Taming Transformer Without Using Learning Rate Warmup Zero-shot text-to-image generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.911598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.911598Z digest=sha256:ef48be501fe344f1cb6255a087976695f5b43e3fe7c8e53e0df39582a40c8bfd

Observation 7f679c03-3b1a-42c7-850d-ed29c4038639 · outbound

This paper cites A stochastic approximation method.

Taming Transformer Without Using Learning Rate Warmup A stochastic approximation method

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.015079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.015079Z digest=sha256:9d557263bde05027557a98fe555730e5dd70bdd76d11058ceca3e0134cd3bd3d

Observation 1e3f62cb-71d6-4f6f-a084-f403cb1c1f0d · outbound

This paper cites Learning representations by back-propagating errors.

Taming Transformer Without Using Learning Rate Warmup Learning representations by back-propagating errors

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.102419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.102419Z digest=sha256:a301053a903f8d9ced87f30c8b4a34d618a67d661dd777b52283f8cf1b5da951

Observation 16719b8a-8d27-4e6f-8747-3abe37ca8fdc · outbound

This paper cites Cyclical learning rates for training neural networks.

Taming Transformer Without Using Learning Rate Warmup Cyclical learning rates for training neural networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:13.656492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:11.173579Z digest=sha256:0f5e97e41ed258c51cfd84afb29198111dec2a48699c6a176dfd619997392d32

Observation 9c2748a7-0246-48b5-81c1-211a72517b3a · outbound

This paper cites Scan and snap: Understanding training dynamics and token composition in 1-layer transformer.

Taming Transformer Without Using Learning Rate Warmup Scan and snap: Understanding training dynamics and token composition in 1-layer transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:13.453488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:11.280585Z digest=sha256:aa760d86f08cae95719d00e0002ae54a189f8e114789d8f6a0d409c520e93386

Observation 795361a3-6a79-46e7-ab88-0175b600b518 · outbound

This paper cites JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention.

Taming Transformer Without Using Learning Rate Warmup JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.405603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.405603Z digest=sha256:cbde3ec1d6447ea3e307b39fb1636215e6cf2b823f54ef7765f4c9f55b162a09

Observation ffe7fa62-6a19-4c8a-bdf5-0d1005c938a4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Taming Transformer Without Using Learning Rate Warmup Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.518646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.518646Z digest=sha256:82ecb0429269749e22525c1f33952d8d923e7ebd7e1a65eb1f160f181be6f646

Observation 15ca0245-aad0-4545-a921-8d27e25009d2 · outbound

This paper cites Attention is all you need.

Taming Transformer Without Using Learning Rate Warmup Attention is all you need

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.617287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.617287Z digest=sha256:db31ed76b684d46a45dd49683abafe10113c84ec057125431329d97342dec094

Observation 61e6a18f-7c0a-40fe-9936-df9204792b23 · outbound

This paper cites High-dimensional probability: An introduction with applications in data science, volume 47.

Taming Transformer Without Using Learning Rate Warmup High-dimensional probability: An introduction with applications in data science, volume 47

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.746254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.746254Z digest=sha256:af7f588c69d6bc0d84f760e665f0c6f40ccf5d1daba7269feff053f6188659ed

Observation c5041f34-d033-4787-b3bf-16da13f0a5b6 · outbound

This paper cites DeepNet: Scaling Transformers to 1,000 Layers.

Taming Transformer Without Using Learning Rate Warmup DeepNet: Scaling Transformers to 1,000 Layers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.857239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.857239Z digest=sha256:b4878df9f83da2b3fd66ff34cffcb09fc1b3682dfe4a4950e8d6357c76e6ce8b

Observation 81f2774c-3042-4510-900d-333208208d21 · outbound

This paper cites Learning Deep Transformer Models for Machine Translation.

Taming Transformer Without Using Learning Rate Warmup Learning Deep Transformer Models for Machine Translation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.909056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.909056Z digest=sha256:7f30e47592fb5922f3f7a361dc30be74cc6d533aa929d991f5f64c26c3380d85

Observation a3150a55-86c8-429c-8061-24f545beced3 · outbound

This paper cites Pytorch image models.

Taming Transformer Without Using Learning Rate Warmup Pytorch image models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.056554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.056554Z digest=sha256:5e87c0eb39ccee3d46170e951f0cdfd274630dae73a95151c8765824c2111c5c

Observation 46f3d844-ff54-4cfb-bdc5-1da34e717569 · outbound

This paper cites High-dimensional data analysis with low-dimensional models: Principles, computation, and applications.

Taming Transformer Without Using Learning Rate Warmup High-dimensional data analysis with low-dimensional models: Principles, computation, and applications

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.154560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.154560Z digest=sha256:62e27491b39d47b7f7fcb6c73808865e86fa50644ad5f6a0a9658518bd9831f9

Observation 7a1d6b40-b659-4b7d-858c-64b2334feedc · outbound

This paper cites On layer normalization in the transformer architecture.

Taming Transformer Without Using Learning Rate Warmup On layer normalization in the transformer architecture

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:13.239775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:26:12.285606Z digest=sha256:bff007810ee4a334a0294725729b83b22598f59f188301c3f2fb7d4b7c353d4f

Observation 10660810-ac57-42c8-835a-697d9671cbea · outbound

This paper cites Stabilizing transformer training by preventing attention entropy collapse.

Taming Transformer Without Using Learning Rate Warmup Stabilizing transformer training by preventing attention entropy collapse

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.409813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.409813Z digest=sha256:dfe08a312f2dbaedf5a75cd2df1850c5d17247db1e2f64449bf2db0bb7c9ecf2

Observation cc1460b6-f112-4c9c-b2ed-ea07f901c0d4 · outbound

This paper cites Root mean square layer normalization.

Taming Transformer Without Using Learning Rate Warmup Root mean square layer normalization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.544561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.544561Z digest=sha256:005685498ff4f02fe1b225ece644f952e7b398e2fc2982cbd41f370de6ded305

Observation 42508aa3-13a1-450e-8e89-2b0aae900f3f · outbound

This paper cites write newline.

Taming Transformer Without Using Learning Rate Warmup write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.625867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.625867Z digest=sha256:a8797f26f60fb4529efbd992542ddaa61ba8ebdf7f23a3e988f47f41681dce6f

Observation 0a02c986-71ee-4c54-8c00-217fe0d1b11c · outbound

This paper cites @esa (Ref.

Taming Transformer Without Using Learning Rate Warmup @esa (Ref

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.700048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.700048Z digest=sha256:420cea9e9553f89c9f7c38e0d486f1dacca15ac420c01d949603e6fd8d29aaef

Observation 55fd245d-a651-4a37-866a-fd997f6cb4e7 · outbound

This paper cites an unresolved cited work.

Taming Transformer Without Using Learning Rate Warmup Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.817304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.817304Z digest=sha256:e43fd9820b152a5181e7158b190d2284a761197812c253c01ca124b876a86eea

Observation 5978d486-dbf6-47f8-a283-ca2b4fba5aaf · outbound

This paper cites an unresolved cited work.

Taming Transformer Without Using Learning Rate Warmup Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.948856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.948856Z digest=sha256:ec9b1168192f8437f31188d04121a5dbaa0a051a42f6138cf4e42996da2b10cc

Pith citing papers

No inbound Pith citation observations are available.