Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:29.278658Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 2 inbound Pith citation observations for arXiv:2506.01049.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:29.278658Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T22:10:49.683444Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T16:41:06.415113Z
92 of 92 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0963deb6-c74b-4b8c-80b6-ca00a44c6488 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 551f9183-6011-4774-a690-29566546a9c1 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f78582-7e3d-46ff-ad52-36f938fb7e84 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0746c1e7-aa97-4e50-b269-2fbc6807cd5e · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f31865e9-46bc-4f4e-83af-6ca0a2eb761f · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff8d07a4-c54e-4f38-97eb-367bb43c0c7f · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52fac5ca-1489-47af-8830-ce9d1cc63cc6 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acd9c9c6-d16c-4306-a6d0-9b5881e8ea8c · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping LLaVA-KD: A Framework of Distilling Multimodal Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2b3a661-29ae-442f-9be4-8049a765a487 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad7b1bb8-15bc-46ae-9929-e2f44bc0ca30 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cd80351-bc8c-47b5-abf1-f6c9500b5ca5 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68aeed3-1427-4dc0-bee3-9e33ee9beed7 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 692be308-d7b0-470f-9cb0-d27506ec7c40 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e73f06e3-5284-43c7-a3a7-e651d2fa2cef · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4038c6b4-e5b0-41c2-acfa-ba023fc04a99 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc5028c-58e4-4e6e-b142-8b2c4c3e1e08 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d09eed9-f3f1-4f1b-b336-868922556f8a · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b3ff932-08e5-4c17-91a2-f03696222cde · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82a835d2-ab35-4e35-8390-b26a6bc7eb62 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 971fbb92-8263-4497-9930-f1671856488c · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c48c3115-6596-43e0-b15e-8798df9e7736 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd64ccbf-f79b-4857-82be-e64792012b92 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRA+: Efficient Low Rank Adaptation of Large Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7292eecc-2457-4f45-a096-18a097d6434f · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 165ad793-8fe0-459a-8b08-99fa783084ed · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d79f9994-2fba-40fe-89ff-79a511e26d96 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRA: Low-Rank Adaptation of Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05cda9ed-9b95-4933-ad63-6ddbaff80b92 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dd4f4fe-d38b-496a-ae49-bc58b853c8da · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 331539f7-7b6a-463e-93b9-846c81e04baf · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbab1350-ed83-40e4-9c6d-b1b7bb62ffb6 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Exploring Low Rank Training of Deep Neural Networks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f00894e-3412-4e42-903f-4742cd1e9907 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c83b9a68-badd-459f-8bb2-6f2024d8bbd8 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Kingma and Jimmy Ba
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae32cc3f-32fa-495d-8410-ca4a74769a87 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping o pf, Yannic Kilcher, Dimitri Von R \
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05593c1e-0953-4188-becd-72abe6222c2d · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 109676e1-f73c-4df9-a5d3-778c96f7cd5f · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c826ac-215f-4559-8d2c-002b3bdd6f5a · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c01ecb-d986-44d9-846c-1c85806a28e3 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88980bb8-475c-4efc-8aa0-b0dffb53b9f5 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 364bbec4-e582-4037-b2b8-ce6cc5bc78cc · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7cdfcf04-d675-4a18-854c-468ef8462bf5 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc732852-44aa-4a48-9b37-8a32aae9c42d · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f23a6cc7-1999-48cb-a9f3-eea0fd23118a · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58996f95-d9c1-4333-a044-f7b8b12fb5d0 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d965385b-4881-4497-a283-25194e44450c · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f77f1f80-dc91-4ecc-b75a-d070ee67d29a · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f249484d-dc3b-4482-b5f3-a6ae44cf3c29 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01d72d3a-a244-4dcc-9acc-ac26c3dc937d · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Muon is Scalable for LLM Training
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c88cdf3a-b3a5-47fd-9c7a-756791a6b415 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd877286-b983-49ef-a3a1-6eeba05ac596 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60d1ffe9-8272-4615-84eb-d64e913ea568 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping DoRA: Weight-Decomposed Low-Rank Adaptation
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc653f9c-4569-4eb6-86cf-d8074085a207 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9786330-16e2-47b7-8b01-43fb2e0337be · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 476339c5-5912-4f3c-bb0e-e5fb722a7511 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b931a06c-dd32-4450-988e-2683564428c6 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16288887-825e-4553-aa28-8fd0e6c42175 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1cd5d36-7d70-46d1-9dfc-ce586ea37d32 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dfa9f3e3-acae-45c6-8047-415b0567a2ef · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping CAME: Confidence-guided Adaptive Memory Efficient Optimization
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28962728-0aa8-47cb-bd0e-ce6920dfe53a · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Visual Perception by Large Language Model's Weights
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 897b4096-df09-470e-8495-9f998b50dd0e · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a9a2b64-3069-4a7c-ae52-e76ccfec22df · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 432c66dc-b5e1-4d55-91c0-e6a82faaa734 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping A Theory on Adam Instability in Large-Scale Machine Learning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21b291a7-59a5-4f8b-8bc7-13a0211d6ec5 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping EDoRA: Efficient Weight-Decomposed Low-Rank Adaptation via Singular Value Decomposition
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c681ca2-5a6a-42e7-b647-7079d06871f0 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b89d3d72-25b6-4554-8823-3b457dc481da · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63d0ecf8-4439-4f80-b8f9-a4d88e9b00a6 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Reddi, Satyen Kale, and Surinder Kumar
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89ed90fc-8074-4108-bc77-7327446f2e31 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10022d71-56be-4848-a29a-443a275c82c6 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping SocialIQA: Commonsense Reasoning about Social Interactions
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35116b00-3bc9-47c6-9fa4-a821ae2ae174 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e966505-a288-4ba3-b062-3fa242178992 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5acd949-8072-4ba1-91a2-67b0f7a43217 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7cbdbc6-d00e-48ee-bb36-3159467da3d8 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35eca4f5-bb6e-41e3-a6af-5695ac55c152 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Sinha and Michael P
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05858b3f-1b4d-414c-91b4-fcfbaf86820a · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c41112a3-43f6-4036-b58a-eefa42ff0149 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93957090-2b1a-4fb8-b488-492be99e021e · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping SOAP: Improving and Stabilizing Shampoo using Adam
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67f2c45a-9ca7-4b76-84dc-c6dcdb8ce9d1 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a56bb22a-4cd7-401b-a945-b401ac0a3a33 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 58806685-fab1-4a62-8eef-0b1953f973ce · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Qwen2.5-Omni Technical Report
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27195656-a9b3-4348-bf49-d5c7179c9221 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40064f04-3bae-490b-bee2-4066438baa9a · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping A Survey on Multimodal Large Language Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0699f0b-c24d-4f2e-b5d5-575773268cff · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 488b4aeb-24ea-4312-9f2b-73aa3153362d · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29cfd7c2-10a0-46a8-859c-59632a011363 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca378888-4d58-4a7c-b368-cd91497d0e09 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a396661-7194-4d09-9c32-792aaac1e40c · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Parameter-Efficient Fine-Tuning for Foundation Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 210d5578-f9c6-4cd4-8a05-e7b376f11d57 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08551d6d-0321-4715-9421-bc9c26773f70 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Adam-mini: Use Fewer Learning Rates To Gain More
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d65b0bc-6e39-43ed-8c87-c595703282c0 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19eaa3e-d997-4475-af1e-df2e6a124e0c · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Deconstructing What Makes a Good Optimizer for Language Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcc700fc-9253-4936-a883-bf2c1addffe4 · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping TinyLLaVA: A Framework of Small-scale Large Multimodal Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ec620a7-60ca-46d6-8a67-437c76a0961b · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping APOLLO: SGD-like Memory, AdamW-level Performance
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4fcc1d0-1de3-4fe9-9690-5d3932797f3c · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Transformers without Normalization
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52272ecf-36fd-481c-87d1-ff47692244fe · outbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 39b37602-eb7a-4980-b4a0-f8e3a6ae7eb7 · inbound
Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio Taming LLMs by Scaling Learning Rates with Gradient Grouping
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1e36740-728e-47b1-8590-2f2aeec73c55 · inbound
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Taming LLMs by Scaling Learning Rates with Gradient Grouping
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.