Pith. sign in

Paper Citation Record · LEDGER

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs

As of 15 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2605.22297.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.22297 v3

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:58:56.074206Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact15
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 819e1380-7647-451f-85fc-da480930b13f · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.584086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:5b1ab03f8b1313dba661c269e5a1ca5ac87911117db20e0a2600a664b15512ab

Observation cbe0943e-ee89-4248-9b27-fbfc0a7c29cd · outbound

This paper cites Don’t be lazy: Completep enables compute-efficient deep transformers.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Don’t be lazy: Completep enables compute-efficient deep transformers

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:34:57.587016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:55a4c013bd18885f69043b20f2bbeeebfda488e43817b89d156289faa7593a4d

Observation bc462715-4679-4150-bef0-3cee690c782f · outbound

This paper cites Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.595450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:15437592ef2b4d022e12c24f53dec11c2381300049fc0fd9ec9bdc0de4ba4eff

Observation f9de646d-499a-4748-9436-1d405e734dfd · outbound

This paper cites Alphadecay: Module-wise weight decay for heavy-tailed balancing in llms.arXiv preprint arXiv:2506.14562.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Alphadecay: Module-wise weight decay for heavy-tailed balancing in llms.arXiv preprint arXiv:2506.14562

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.601205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:9095065cca047c6dbe37f6f464eb35b83e0af1ee0ff1d204e4d0faad50f72732

Observation 67fc3bb7-d4f8-4a94-9474-7e88faef9e28 · outbound

This paper cites Training Compute-Optimal Large Language Models.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Training Compute-Optimal Large Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.633293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:56e49fcc673606bb3d9c6322faef038a622b2362af26419aa37069ac30b8894d

Observation fe5e839b-61a3-46ef-a501-15c2a0fe2df3 · outbound

This paper cites LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.589775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:83809ecd66803f8f37b0a54ec05c04a65987af0f13902222b88cef1ecd41e322

Observation 400b49a7-d530-4f35-8c50-7b2f34679cef · outbound

This paper cites Muon is Scalable for LLM Training.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Muon is Scalable for LLM Training

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.604761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:c5e4fcf70700dd3c3604c4f40a9b45cdb2c4f5fbfa707e4fda1dc91433ae06f9

Observation 15ab6c0a-27ba-4ae0-8781-557c59d4bbb8 · outbound

This paper cites Model Balancing Helps Low-data Training and Fine-tuning.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Model Balancing Helps Low-data Training and Fine-tuning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.627601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:ffaedefda5f2b5f6b8f6a5fb1c86f4648df067290dfd6367691035b614e018eb

Observation 335f0174-667c-4858-9817-ade1ff1f7da0 · outbound

This paper cites Decoupled Weight Decay Regularization.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Decoupled Weight Decay Regularization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.624687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:3d35dc9ff01fc03c921cfbdfe8e8f9fafe17729b0db87064ed148d88bb16da6e

Observation 2c26e020-6f35-494f-bbb9-bcd4e03b191f · outbound

This paper cites Traditional and Heavy-Tailed Self Regularization in Neural Network Models.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Traditional and Heavy-Tailed Self Regularization in Neural Network Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.621695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:a5c02c77e66c3bd54ee1c049cc1b1cfd57767c9a0ce29b102545ae79c1c7c585

Observation cec8cb3f-1977-48a2-b2df-caee9a25e30b · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.630682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:bddc2bc8de3532acff9094b852c3aac712f0e08f6ea309d1c5563f40cee5b525

Observation 1be24f73-fb06-45b3-82d2-97602b6189a3 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T17:34:57.613690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:989c0a48ebe5c44545e08b3af08264647433d4a0640bf2ab32b2cb962b39cd63

Observation 5eb45495-a603-400a-bc8a-6c3873053e2a · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs SocialIQA: Commonsense Reasoning about Social Interactions

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T17:34:56.910283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:6e64f9e63528ae458f619a05fa754dfbe0d1e670a39b2aa72ad9cfc429f0a0b4

Observation 8b49eecf-6cbd-4f43-9330-9c01c3b7f723 · outbound

This paper cites The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.616526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:262c7e1bb5b9d516df21f064c6b2a352b068ecd1a9e5d6e53cfac88d7909248a

Observation 1eb83793-db37-4efa-95da-fff07bd1647f · outbound

This paper cites Feature Learning in Infinite-Width Neural Networks.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Feature Learning in Infinite-Width Neural Networks

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.603774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:b08e1e0f4d701576043f4a6b3c67e88109ae44d73a86f5dafeb7c23969cdc5bd

Observation e9a303f1-aa0d-4451-af4e-933cde503842 · outbound

This paper cites Large Batch Training of Convolutional Networks.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Large Batch Training of Convolutional Networks

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.606238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:4d24ba7419af4bc4d784634d0f490d77129ce503a17c47ad51a40a4e44b613b0

Observation e173ea4f-b4a6-42d8-8117-72b6db0c3b36 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.608787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:33c2cc0ca34cd6a38e56f2bbe13209e7d5355a414857da7c1fd35ffa901d0507

Observation 1112d5bd-7908-4d8c-b9f0-4a1d17511620 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.611239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:183250ea523853d97362ba93f4ca096cc9cfb59e493e94d862601654a215df53

Observation 7e02c1ac-30c9-489b-a4c3-8e5c92189eac · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Adam-mini: Use Fewer Learning Rates To Gain More

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:34:57.598365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:4f25cffd131c1902df431c66fde84dfc8f144866c31d689c3ffe2d4ffcdf50db

Observation 7f1bf2a3-b974-4d96-a975-cf417390f4bb · outbound

This paper cites Details of Experiments This section provides detailed configurations for both pre- training and finetuning experiments.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Details of Experiments This section provides detailed configurations for both pre- training and finetuning experiments

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T08:54:50.066199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:dbc22f9eece33c53f706a7b4fc2c1bbf7e06cfe30ac5bcc9e6d49611dd1240fa

Observation 02c5b434-0d25-44cd-b0b6-d74a456cb6bb · outbound

This paper cites Among these, Linear achieves the best results across all LR settings, showing a notable advantage over other methods.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Among these, Linear achieves the best results across all LR settings, showing a notable advantage over other methods

Reference 21

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T17:34:57.619301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:a8af1dc85b7ee73875c719b2a22c1ce57b1d9dcaef3c02a456a34db3ca4b6646

Pith citing papers

Observation 10f40f6a-933f-426d-ad80-6a7aac84ea9b · inbound

Muse: Representation Geometry of Muon Beyond Normalized Momentum cites this paper.

Muse: Representation Geometry of Muon Beyond Normalized Momentum One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:58:56.074206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:58:56.074206Z digest=sha256:655f74498b1310ad59163c92b3c0f674dc8e6898f62bd28ca8f2cb840f755cc3