Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2605.22297.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T01:58:56.074206Z
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 819e1380-7647-451f-85fc-da480930b13f · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cbe0943e-ee89-4248-9b27-fbfc0a7c29cd · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Don’t be lazy: Completep enables compute-efficient deep transformers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bc462715-4679-4150-bef0-3cee690c782f · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f9de646d-499a-4748-9436-1d405e734dfd · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Alphadecay: Module-wise weight decay for heavy-tailed balancing in llms.arXiv preprint arXiv:2506.14562
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 67fc3bb7-d4f8-4a94-9474-7e88faef9e28 · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Training Compute-Optimal Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fe5e839b-61a3-46ef-a501-15c2a0fe2df3 · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 400b49a7-d530-4f35-8c50-7b2f34679cef · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Muon is Scalable for LLM Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 15ab6c0a-27ba-4ae0-8781-557c59d4bbb8 · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Model Balancing Helps Low-data Training and Fine-tuning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 335f0174-667c-4858-9817-ade1ff1f7da0 · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Decoupled Weight Decay Regularization
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2c26e020-6f35-494f-bbb9-bcd4e03b191f · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Traditional and Heavy-Tailed Self Regularization in Neural Network Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cec8cb3f-1977-48a2-b2df-caee9a25e30b · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1be24f73-fb06-45b3-82d2-97602b6189a3 · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T20:38:10.326889+00:00.
Observation 5eb45495-a603-400a-bc8a-6c3873053e2a · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs SocialIQA: Commonsense Reasoning about Social Interactions
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8b49eecf-6cbd-4f43-9330-9c01c3b7f723 · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1eb83793-db37-4efa-95da-fff07bd1647f · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Feature Learning in Infinite-Width Neural Networks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e9a303f1-aa0d-4451-af4e-933cde503842 · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Large Batch Training of Convolutional Networks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e173ea4f-b4a6-42d8-8117-72b6db0c3b36 · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1112d5bd-7908-4d8c-b9f0-4a1d17511620 · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7e02c1ac-30c9-489b-a4c3-8e5c92189eac · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Adam-mini: Use Fewer Learning Rates To Gain More
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7f1bf2a3-b974-4d96-a975-cf417390f4bb · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Details of Experiments This section provides detailed configurations for both pre- training and finetuning experiments
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 02c5b434-0d25-44cd-b0b6-d74a456cb6bb · outbound
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Among these, Linear achieves the best results across all LR settings, showing a notable advantage over other methods
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 10f40f6a-933f-426d-ad80-6a7aac84ea9b · inbound
Muse: Representation Geometry of Muon Beyond Normalized Momentum One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.