Pith. sign in

Paper Citation Record · LEDGER

On the Nonlinearity of Learning Rate Scaling for LLM Training

As of 6 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2606.29158.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.29158 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T08:15:20.191222Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact14
  • verified fuzzy26
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ed80a9e-47b6-467a-a23c-9d2ae7b9bb60 · outbound

This paper cites Tune My Adam, Please!.

On the Nonlinearity of Learning Rate Scaling for LLM Training Tune My Adam, Please!

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:24:26.805017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:df9ef24c2b5b211856dc5186c87b66043b88ebaa71b6228893fbca34f496121b

Observation b476bcbc-729f-4330-94de-9e18dd9d44db · outbound

This paper cites Layer Normalization.

On the Nonlinearity of Learning Rate Scaling for LLM Training Layer Normalization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:24:26.802509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:d1f6ac615d1c27b8c7d747e1ecd04d660481a19a8754e193f4f3e68145202951

Observation 2b49e810-a187-4c01-be95-494c36a627de · outbound

This paper cites Power lines: Scaling laws for weight decay and batch size in llm pre-training.

On the Nonlinearity of Learning Rate Scaling for LLM Training Power lines: Scaling laws for weight decay and batch size in llm pre-training

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.419325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:4efcf669a3634036782b68e652900eed6deb833e6c9bda9a3c898dfb873ec3c6

Observation d04d66e2-b1fe-4124-bc94-a89b80481c78 · outbound

This paper cites Scaling optimal lr across token horizons.

On the Nonlinearity of Learning Rate Scaling for LLM Training Scaling optimal lr across token horizons

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.453162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:c9f5b08e0af5ab4b8e027d9ef96629721ce5369943804ff07a37df2eb5a14bb3

Observation 2942bb31-e203-45f3-807e-d8b56bbdb2f0 · outbound

This paper cites Y., Deiseroth, B., Cruz-Salinas, A.

On the Nonlinearity of Learning Rate Scaling for LLM Training Y., Deiseroth, B., Cruz-Salinas, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.352095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:22cfca713ea14dcd01be47d7e06110a57356e160ae356794f447f6f4a1ae8ead

Observation 85bfd103-e0ed-4bf5-8edf-3ae048ad1712 · outbound

This paper cites Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit.

On the Nonlinearity of Learning Rate Scaling for LLM Training Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.368176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:a0f6a4f6455367e305c6c23393f53ea3b60d93aa8227c0172e777302ea2c3280

Observation 326afa66-a69d-41c6-b494-96c0f9d9240f · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

On the Nonlinearity of Learning Rate Scaling for LLM Training DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T08:24:26.799840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:ee9143b6bb598b4aa5cca729278e90d501f676265f1a80bf191cc093a0f45098

Observation acdbe677-c97d-4750-a1ca-d887f791c183 · outbound

This paper cites Don’t be lazy: Completep enables compute-efficient deep transformers.

On the Nonlinearity of Learning Rate Scaling for LLM Training Don’t be lazy: Completep enables compute-efficient deep transformers

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:24:26.796924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:1f339470a54210cc04b29cd267d93f6106e875f5ed2a47d36aff3457f6b29bf7

Observation fad2912c-901a-4aa5-ac76-44710a4a7a07 · outbound

This paper cites E., Xiao, L., Wortsman, M., Alemi, A.

On the Nonlinearity of Learning Rate Scaling for LLM Training E., Xiao, L., Wortsman, M., Alemi, A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.399756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:295ad07a584e6acc9a7fa1d63039996b5e367a13406a34897e7205cb0c78ac6a

Observation 8cfee3b9-eef0-4808-9210-cc8a42fb94df · outbound

This paper cites Robust layerwise scaling rules by proper weight decay tuning.

On the Nonlinearity of Learning Rate Scaling for LLM Training Robust layerwise scaling rules by proper weight decay tuning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:24:26.782528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:e1e6a858064415fe7cabe5f50a88ba26ebfc5742d4753778becbeeaaecb75be9

Observation d5ba71d2-ef7d-4300-a530-e599984dba5b · outbound

This paper cites Nemotron-flash: Towards latency-optimal hybrid small language models.

On the Nonlinearity of Learning Rate Scaling for LLM Training Nemotron-flash: Towards latency-optimal hybrid small language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.337043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:a29d8d14c820ceeb345d9693504b4baa7a84f62c38217699f790510e2fceddf4

Observation 5e5b1cde-973b-4d68-9a8f-4bef60ff7942 · outbound

This paper cites Norm matters: efficient and accurate normalization schemes in deep networks.Advances in Neural Information Processing Systems, 31, 2018.

On the Nonlinearity of Learning Rate Scaling for LLM Training Norm matters: efficient and accurate normalization schemes in deep networks.Advances in Neural Information Processing Systems, 31, 2018

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.292267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:8576feb1d8363776a398f145c473b851878f03372b9de66a8b88d2a50f8c78f2

Observation b6e36f17-a54f-4e53-a522-bcbc3966accc · outbound

This paper cites Minicpm: Unveiling the potential of small language models with scalable training strategies.

On the Nonlinearity of Learning Rate Scaling for LLM Training Minicpm: Unveiling the potential of small language models with scalable training strategies

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.275959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:5b3c82809f369c80a82e857f7e7ff9ab03bf94bd8991647f77be559189d48e57

Observation b9fc8db1-5b25-45c6-b8e5-5912fbd4607b · outbound

This paper cites H., and Leyton-Brown, K.

On the Nonlinearity of Learning Rate Scaling for LLM Training H., and Leyton-Brown, K

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.307151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:17294f47cbb32601f359ced5fde4f5a8e4856e368cb73ccadfdc8bfce0d9244e

Observation 216b8889-4b67-4bab-a95b-9f06bbd1caa0 · outbound

This paper cites and Szegedy, C.

On the Nonlinearity of Learning Rate Scaling for LLM Training and Szegedy, C

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.321669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:3a217b2e9ab7d5c095695441f755bac90f391267aca276580e97eacc9b1701b2

Observation f1ed5e63-8183-4b39-9fb3-943fbbad2508 · outbound

This paper cites Three Factors Influencing Minima in SGD.

On the Nonlinearity of Learning Rate Scaling for LLM Training Three Factors Influencing Minima in SGD

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:24:26.785633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:24681b4a2bd5a70e5e4cc95565aabf3c1cb6ae810ebeb125bd4f78d11d5ccee6

Observation b2d893dc-ba19-47b2-9599-720bda639dba · outbound

This paper cites Scaling Laws for Neural Language Models.

On the Nonlinearity of Learning Rate Scaling for LLM Training Scaling Laws for Neural Language Models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T08:24:26.788145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:0bacc3d1cd48d7e1ca8d2250ae133803bad80e2522e4cf5590bbf465a5f6170b

Observation 93482c5a-9c43-43d1-8828-1011787364a5 · outbound

This paper cites nanoGPT.https://github.com/karpathy/nanoGPT, 2022.

On the Nonlinearity of Learning Rate Scaling for LLM Training nanoGPT.https://github.com/karpathy/nanoGPT, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.383752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:bed794b8a78cc0a5f9a6ec2f4884abafcdaa0169789753b7268e201ac2cbb6dd

Observation 70994301-6fd4-488d-9b4b-cb94d4fd16c2 · outbound

This paper cites Rotational equilibrium: How weight decay balances learning across neural networks.

On the Nonlinearity of Learning Rate Scaling for LLM Training Rotational equilibrium: How weight decay balances learning across neural networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.437186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:f4eb8414d629027a08d6d8581f7a568634b534adb7d8ea7025158b5024a9e2e2

Observation 1d47b165-16e1-4c2f-b048-95fcb2473e01 · outbound

This paper cites Weight decay may matter more than mup for learning rate transfer in practice.

On the Nonlinearity of Learning Rate Scaling for LLM Training Weight decay may matter more than mup for learning rate transfer in practice

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:24:26.791192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:56ef4086f83d4365f6a1f6f65a1aa627fd12ddc7eb1a9154fe2626ad3bdcf21b

Observation 9b8a732c-2eb5-45f3-af9d-0924c8286b4f · outbound

This paper cites Efficient hyperparameter tuning via trajectory invariance principle.arXiv preprint arXiv:2509.25049, 2025.

On the Nonlinearity of Learning Rate Scaling for LLM Training Efficient hyperparameter tuning via trajectory invariance principle.arXiv preprint arXiv:2509.25049, 2025

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:24:26.772306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:1211f8c6f1a9fdcda81d61a83ab82b8c4ef6048e2ddb6e855422b6f75468dc5d

Observation 23542f01-bb32-401a-9b90-90b1e45102bb · outbound

This paper cites Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining.

On the Nonlinearity of Learning Rate Scaling for LLM Training Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:24:26.775578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:30f5e62142bf179d4bb8b84786126acc8fa49ad44d7cffe8fb941259acfecdb8

Observation 910d5cd0-ae41-426c-805e-22c39419f4f0 · outbound

This paper cites Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate.Advances in Neural Information Processing Systems, 33: 14544–14555, 2020.

On the Nonlinearity of Learning Rate Scaling for LLM Training Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate.Advances in Neural Information Processing Systems, 33: 14544–14555, 2020

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.470216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:99cba340e3d0b8976a7c5bd61b160d3e927189787fe4b0e4deab03b9c70fd0fb

Observation fc444885-f740-4f4d-bc3f-7df9c1d70f0a · outbound

This paper cites Muon is Scalable for LLM Training.

On the Nonlinearity of Learning Rate Scaling for LLM Training Muon is Scalable for LLM Training

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:24:26.786545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:805385e82f8dab6c0faffc3dc1c8f78f053e34df4cc0925ae4c594c0ca51c64e

Observation d7dd65cb-4a23-4d67-83b9-227a7d2daab7 · outbound

This paper cites and Hutter, F.

On the Nonlinearity of Learning Rate Scaling for LLM Training and Hutter, F

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.177327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:8041ea67b08f312868e6cb57195e9291b397b76a75ee9661c534a03f68103537

Observation f04a8424-4645-408b-a2fd-edd22af09e6b · outbound

This paper cites ngpt: Normalized transformer with representation learning on the hypersphere.

On the Nonlinearity of Learning Rate Scaling for LLM Training ngpt: Normalized transformer with representation learning on the hypersphere

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.193385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:fad57ee30411989f75717e897013276930a28fdc3b4d11dd40bba813e2be72fc

Observation fb782221-1cd1-4f28-9fa5-ed58f30af273 · outbound

This paper cites A multi-power law for loss curve prediction across learning rate schedules.

On the Nonlinearity of Learning Rate Scaling for LLM Training A multi-power law for loss curve prediction across learning rate schedules

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.209774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:6fbd9c7f06fa4c268d7dd1819f152138456006e22d899a0fee20986d5c795794

Observation 7336a97c-5837-4d9d-832f-2fce3e899626 · outbound

This paper cites On the sdes and scaling rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022.

On the Nonlinearity of Learning Rate Scaling for LLM Training On the sdes and scaling rules for adaptive gradient algorithms.Advances in Neural Information Processing Systems, 35:7697–7711, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.241692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:a7b004e5b7e8f3546f989ad80033d755766399b2a90ca6375f4ce26bcd1fdd88

Observation 70b8bbfd-6f20-459a-b87b-080c0bc99d32 · outbound

This paper cites G., and Goldblum, M.

On the Nonlinearity of Learning Rate Scaling for LLM Training G., and Goldblum, M

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.161987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:2a62b5d740e48217d8fac390d5a1ec55d949e3db75a81f8a9430316731a83a98

Observation 6c3812b3-5e75-4cf6-8c25-6dd533c7b0c2 · outbound

This paper cites An Empirical Model of Large-Batch Training.

On the Nonlinearity of Learning Rate Scaling for LLM Training An Empirical Model of Large-Batch Training

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:24:26.769946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:c3f164521e645586feb73365ff168a746c10bff73761b3e3fe5aa92d2f6c8eb6

Observation 8d1dc6d3-028a-4b07-aa18-9ce07405935b · outbound

This paper cites Completed hyperparameter transfer across modules, width, depth, batch and duration.

On the Nonlinearity of Learning Rate Scaling for LLM Training Completed hyperparameter transfer across modules, width, depth, batch and duration

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:24:26.772683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:46dc5b2abee9df379dcdde2e27d5fba39dd578c6446604ba3e7c4ab8c720cc5e

Observation f0dc6fca-e4f4-49b3-be1a-e3b4f7e87669 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems, 37:30811–30849, 2024.

On the Nonlinearity of Learning Rate Scaling for LLM Training The fineweb datasets: Decanting the web for the finest text data at scale.Advances in Neural Information Processing Systems, 37:30811–30849, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.258185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:d40cfb77a338aefc19b290a6d7add7e82cc082c57de11158db1803600678dfbe

Observation 70f0f246-979c-4d3a-9506-8d9675bd2dab · outbound

This paper cites Resolving Discrepancies in Compute-Optimal Scaling of Language Models.

On the Nonlinearity of Learning Rate Scaling for LLM Training Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:24:26.779384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:c7997e8110583aec805c20664abad8313ed0a1212178e63da439d89938742e96

Observation a7ae3b72-9161-42ae-9408-24e534c7285f · outbound

This paper cites Language models are un- supervised multitask learners.

On the Nonlinearity of Learning Rate Scaling for LLM Training Language models are un- supervised multitask learners

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.226047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:bd81410a5c065d5a7dd19031c64dabbb8a3bf35719ddf6fecc25d6b87b883216

Observation 87995147-01b3-4849-ad29-b17493b54b62 · outbound

This paper cites an unresolved cited work.

On the Nonlinearity of Learning Rate Scaling for LLM Training Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-07-10T22:27:43.146959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:30de887103004aa341baa292babfa87e2a7e2f19aafe18951c1f7e31dd925f23

Observation 068cd734-0431-4fcc-8a9b-977367f48ca5 · outbound

This paper cites Scaling Law with Learning Rate Annealing.

On the Nonlinearity of Learning Rate Scaling for LLM Training Scaling Law with Learning Rate Annealing

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:24:26.793968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:547fb0a6ff5bdee504366bcab9c7d544d3a10e3e96430ecbb4042361d633fe70

Observation 1b0bb89b-2852-45b5-8e8c-2359a173e3d5 · outbound

This paper cites L2 Regularization versus Batch and Weight Normalization.

On the Nonlinearity of Learning Rate Scaling for LLM Training L2 Regularization versus Batch and Weight Normalization

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:24:26.807479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:1bff2c6e357d80bb864ed000626a1d19a77b452a5e771ce6902324f2657b3445

Observation ec48e0a1-bd5e-4a9c-b52e-9034a6388842 · outbound

This paper cites N., Kaiser, L.

On the Nonlinearity of Learning Rate Scaling for LLM Training N., Kaiser, L

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.033764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:b7e61191528b39216fc8b47a5a524432efa9e8b5e6bc26dea18ad1cf4ff922da

Observation 0b715f71-b032-424c-bf81-edaac5d03c5b · outbound

This paper cites The sharpness disparity principle in transformers for accelerating language model pre-training.

On the Nonlinearity of Learning Rate Scaling for LLM Training The sharpness disparity principle in transformers for accelerating language model pre-training

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.087074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:8f0d04f1cdd1ee3935aec8a5ffe1172ccb0bdfccb3433a2510b7c45e9839987f

Observation 81f308be-1414-4f09-97d9-38b1fe3beed7 · outbound

This paper cites Scaling laws across model architectures: A comparative analysis of dense and M o E models in large language models.

On the Nonlinearity of Learning Rate Scaling for LLM Training Scaling laws across model architectures: A comparative analysis of dense and M o E models in large language models

Reference 40

Resolution
verified exact
doi, observed 2026-06-30T08:24:26.153118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:884ff07050fce462753d3918954a1fe107c46a26d06cef1cab9cf2d405121de2

Observation 14de7d54-e1a2-439f-a978-41d89baa9892 · outbound

This paper cites and Aitchison, L.

On the Nonlinearity of Learning Rate Scaling for LLM Training and Aitchison, L

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.018928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:7f2af3c6b753a2888cb4301153e87d3f562ef7f8f4d5bc853bb1dae0ab4120e4

Observation d5b9adb0-4a9f-418d-bc2e-5ad7c3c40901 · outbound

This paper cites Fantastic pretraining optimizers and where to find them 2.1: Hyperball optimization, 12 2025.

On the Nonlinearity of Learning Rate Scaling for LLM Training Fantastic pretraining optimizers and where to find them 2.1: Hyperball optimization, 12 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.117232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:7b41d747a969d90499be23178e5d7ce076ab92a801c9a70a28adab8ba8029959

Observation ed70c51b-7dd4-42f4-9e0f-e43c287b5ac2 · outbound

This paper cites Controlled llm training on spectral sphere.

On the Nonlinearity of Learning Rate Scaling for LLM Training Controlled llm training on spectral sphere

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:24:26.775149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:44f95afb6ad54b0b544d549584fd7bbf046c3c3b5db359794e1f4585196ad8bc

Observation ac30bcd1-4fd1-4bdd-b1f5-2cddd2272537 · outbound

This paper cites and Hu, E.

On the Nonlinearity of Learning Rate Scaling for LLM Training and Hu, E

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.051503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:1d243a5a631778ffaf1914c2fdc2c637ebba7e05d2aa15f4878fb154bbde9daa

Observation 31917e82-803b-4c49-8f4a-7a98ccdc03d2 · outbound

This paper cites Tuning large neural networks via zero-shot hyperparameter transfer.Advances in Neural Information Processing Systems, 34:17084–17097, 2021.

On the Nonlinearity of Learning Rate Scaling for LLM Training Tuning large neural networks via zero-shot hyperparameter transfer.Advances in Neural Information Processing Systems, 34:17084–17097, 2021

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T22:27:43.069067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:35165084cf58aa4bc0cdb4a070df37f9c8e740fd82ecc7288d58e0e49bb3f7db

Observation 46be5cf6-44e2-4b78-b138-60d16ba37af6 · outbound

This paper cites an unresolved cited work.

On the Nonlinearity of Learning Rate Scaling for LLM Training Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-07-10T22:27:43.132477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:32c5a6ab3230a8eff6141da1752015e94cb8bbe200db630bb2335fb52ddeae81

Observation fbcbe078-632d-4549-a4b2-31dae2eaa7de · outbound

This paper cites arXiv , author =:2602.10300 , file =.

On the Nonlinearity of Learning Rate Scaling for LLM Training arXiv , author =:2602.10300 , file =

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:24:26.784084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:4b16d9adf2d6a5293f379836c3158229f65f8ed748ab1cd74c2e42b23985b8d4

Observation 1e8e661a-dcd3-42cf-8e8d-9ce6fdbf9989 · outbound

This paper cites How to set the learning rate for large-scale pre-training?arXiv preprint arXiv:2601.05049, 2026.

On the Nonlinearity of Learning Rate Scaling for LLM Training How to set the learning rate for large-scale pre-training?arXiv preprint arXiv:2601.05049, 2026

Reference 48

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T08:24:26.758109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:622630363b54cc7a94769ba6004701220d17580af0b94faf97d6b08b54663a0b

Pith citing papers

No inbound Pith citation observations are available.