Pith. sign in

Paper Citation Record · LEDGER

Taming LLMs by Scaling Learning Rates with Gradient Grouping

As of 8 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 2 inbound Pith citation observations for arXiv:2506.01049.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01049 v1

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:29.278658Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T22:10:49.683444Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T16:41:06.415113Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved87
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0963deb6-c74b-4b8c-80b6-ca00a44c6488 · outbound

This paper cites online" 'onlinestring :=.

Taming LLMs by Scaling Learning Rates with Gradient Grouping online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.960786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.960786Z digest=sha256:4c5ce7635f3022b3260e7647e6c1031cfd994ea89a6b30362df16e14e59d68b1

Observation 551f9183-6011-4774-a690-29566546a9c1 · outbound

This paper cites write newline.

Taming LLMs by Scaling Learning Rates with Gradient Grouping write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.965573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.965573Z digest=sha256:63ced1c44588f0d04619db20e2b5ab9fad330c924fa45e19e58b87cd1b61a847

Observation b2f78582-7e3d-46ff-ad52-36f938fb7e84 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.969116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.969116Z digest=sha256:ea81fba78002c312ddc4f219fbc398f720b106a915e8ec95d5b7d8e2cd4ff48f

Observation 0746c1e7-aa97-4e50-b269-2fbc6807cd5e · outbound

This paper cites LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.972955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.972955Z digest=sha256:0c0cc01294c33c97a797bed18514f5bad5f4483ca5c5a22183871abe7a90ab7b

Observation f31865e9-46bc-4f4e-83af-6ca0a2eb761f · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.976631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.976631Z digest=sha256:31a9cf280812657587f8bb2a7038b473f88d1a368c58e3c529a3dabad5959fce

Observation ff8d07a4-c54e-4f38-97eb-367bb43c0c7f · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.980071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.980071Z digest=sha256:c212005e0cd5892e880eb108364fed324e136fc59630822b1d4a95eeeebc1509

Observation 52fac5ca-1489-47af-8830-ce9d1cc63cc6 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.983339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.983339Z digest=sha256:92060f4968e26fa2ac476860d66107280acbbe2df3c9137705c77cedd88ddd28

Observation acd9c9c6-d16c-4306-a6d0-9b5881e8ea8c · outbound

This paper cites LLaVA-KD: A Framework of Distilling Multimodal Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LLaVA-KD: A Framework of Distilling Multimodal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.986563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.986563Z digest=sha256:d2bdcd0a69ea72b505b47bd50ae84b6b241ef8d59c67dd66c615dc16a47f1158

Observation f2b3a661-29ae-442f-9be4-8049a765a487 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.991360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.991360Z digest=sha256:a3e3b94928ca9d7667a320fc73e8bbddc1841d58fd771a005682b8c517da18ff

Observation ad7b1bb8-15bc-46ae-9929-e2f44bc0ca30 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.995330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.995330Z digest=sha256:11d3b5429e8b17f61c134703a77eaa7b13afecdedca3474a6a1b789e3a25cab4

Observation 3cd80351-bc8c-47b5-abf1-f6c9500b5ca5 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.998658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.998658Z digest=sha256:83e71563a4d3744bfe8cf09c7893d37f4b79667553fbb008059a23a2278b2256

Observation e68aeed3-1427-4dc0-bee3-9e33ee9beed7 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.002055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.002055Z digest=sha256:e9cc7119e4629cbeb1582dd1a681ae0e03a504a79e33886c074a69aa2c08a4b4

Observation 692be308-d7b0-470f-9cb0-d27506ec7c40 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Taming LLMs by Scaling Learning Rates with Gradient Grouping BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.005938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.005938Z digest=sha256:5d1801fa7a78fe06d930fae5210d02b2babc4d9c87be3d7f57a5bc0cabe4618c

Observation e73f06e3-5284-43c7-a3a7-e651d2fa2cef · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.010001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.010001Z digest=sha256:f84bdcaa5820f2a2124b8502cc326c3bf4c95a3ad160f7e8f621250ec444f9bd

Observation 4038c6b4-e5b0-41c2-acfa-ba023fc04a99 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Taming LLMs by Scaling Learning Rates with Gradient Grouping InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.013438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.013438Z digest=sha256:bbda13464e6407bedb8bf60f33ae0921db76fd9fdbfc64b49e5a9dd7fb1959c4

Observation 9bc5028c-58e4-4e6e-b142-8b2c4c3e1e08 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.016941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.016941Z digest=sha256:e98045c2644e00291c2dc93fd5b7f41015463ba1e9ee9f57e3cf323d19d22da4

Observation 3d09eed9-f3f1-4f1b-b336-868922556f8a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.020514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.020514Z digest=sha256:c6c578bbb4cd4ccacd71badaf09eda298cc78d33dfdb4f562efa1ef093bcf053

Observation 2b3ff932-08e5-4c17-91a2-f03696222cde · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.024107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.024107Z digest=sha256:0caf90c191ca3a29ce9defe07e149dfc42ec440e521f85a9a83a8ae8544648de

Observation 82a835d2-ab35-4e35-8390-b26a6bc7eb62 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.027322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.027322Z digest=sha256:5ea044ffc56d1499ff7d5fa5d8ed256813ae66ee885f82036bd0259cf44adf45

Observation 971fbb92-8263-4497-9930-f1671856488c · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.364643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.030485Z digest=sha256:790d08870c7f38709f0e184f6da2e6580ab09786e68fa33fd0869d1e95601348

Observation c48c3115-6596-43e0-b15e-8798df9e7736 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.352712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.033741Z digest=sha256:f021b69a2e06b0cea309b73c43074bb5d6c7cb938ff4a2897f58f47715cf465a

Observation fd64ccbf-f79b-4857-82be-e64792012b92 · outbound

This paper cites LoRA+: Efficient Low Rank Adaptation of Large Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRA+: Efficient Low Rank Adaptation of Large Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.037629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.037629Z digest=sha256:3d65641dfc66f941fafc13cdbbb4f8fb10a7cd997f656e6db3113e7d789734d3

Observation 7292eecc-2457-4f45-a096-18a097d6434f · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.339368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.040979Z digest=sha256:59201bd4bd4215c871dd2978fb8d42979f259ef387dd00fbc9b9a797b1da9125

Observation 165ad793-8fe0-459a-8b08-99fa783084ed · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.328100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.044311Z digest=sha256:825ff411b2d90e57cbedc714cd265782eb195332bd1b3bd1c8533d348a347dff

Observation d79f9994-2fba-40fe-89ff-79a511e26d96 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRA: Low-Rank Adaptation of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.047501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.047501Z digest=sha256:127232656f8346b2d7b0d54bd5574b4715ba39686edde3038d167e5f8818ee32

Observation 05cda9ed-9b95-4933-ad63-6ddbaff80b92 · outbound

This paper cites LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.051249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.051249Z digest=sha256:39e33f4e1e6bdcca9bb6d9aad049085e73a2371df273919b86413d301c56cc5e

Observation 9dd4f4fe-d38b-496a-ae49-bc58b853c8da · outbound

This paper cites SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training.

Taming LLMs by Scaling Learning Rates with Gradient Grouping SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.054680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.054680Z digest=sha256:b7b81298647b18008c6a127ce16db9cc9ddf6651c490ca15e95d91ebaa0db459

Observation 331539f7-7b6a-463e-93b9-846c81e04baf · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.058208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.058208Z digest=sha256:e701e5efe1340f1253cc109dcba18d920f58ddaffaf1ece86517c36dd5f4f1d3

Observation cbab1350-ed83-40e4-9c6d-b1b7bb62ffb6 · outbound

This paper cites Exploring Low Rank Training of Deep Neural Networks.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Exploring Low Rank Training of Deep Neural Networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.061428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.061428Z digest=sha256:f7dbfc9ecbd63ac59fd8092c60e5eaa12a6ca92f8ea1743188d6a6ec55b207fd

Observation 0f00894e-3412-4e42-903f-4742cd1e9907 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.307740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.064748Z digest=sha256:42e468b9d950d1f8a93adcc207717cd23bc3f195c04c517859886e68b7861fe0

Observation c83b9a68-badd-459f-8bb2-6f2024d8bbd8 · outbound

This paper cites Kingma and Jimmy Ba.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Kingma and Jimmy Ba

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:30.296503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.067938Z digest=sha256:ec5aae78e9f218a165e9d9af31f83dac2e0ba66de38b2235a4c84e6c300750f0

Observation ae32cc3f-32fa-495d-8410-ca4a74769a87 · outbound

This paper cites o pf, Yannic Kilcher, Dimitri Von R \.

Taming LLMs by Scaling Learning Rates with Gradient Grouping o pf, Yannic Kilcher, Dimitri Von R \

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:30.285335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.071052Z digest=sha256:508ea909b6dacad47cf79c3ade4c5a177f1fe93a38640196d473d6fe0b15cd2a

Observation 05593c1e-0953-4188-becd-72abe6222c2d · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Taming LLMs by Scaling Learning Rates with Gradient Grouping SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.074356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.074356Z digest=sha256:bb2419a78a33cf95135b71dc9d6624489d5381868ce178e831b901386565552d

Observation 109676e1-f73c-4df9-a5d3-778c96f7cd5f · outbound

This paper cites LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.078239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.078239Z digest=sha256:31a5318ccccd14450db71718ddf7172275ede81bb45809e68e1d3fb57b623bca

Observation 10c826ac-215f-4559-8d2c-002b3bdd6f5a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.081480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.081480Z digest=sha256:3bd26b2a88d1a33ead0793f5a805073ec47c0fc231bbc5caf1c80ab1dd321c85

Observation 84c01ecb-d986-44d9-846c-1c85806a28e3 · outbound

This paper cites Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.084798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.084798Z digest=sha256:30fae30dff0717eedc5fa9e0771b01c189741868c25755f029a9ed5737e81211

Observation 88980bb8-475c-4efc-8aa0-b0dffb53b9f5 · outbound

This paper cites Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.088289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.088289Z digest=sha256:926615f409446a0615a45ed74ef77ecac322f8d6c39e2d3e42a0450c50443ec0

Observation 364bbec4-e582-4037-b2b8-ce6cc5bc78cc · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.264637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.091805Z digest=sha256:c27949292ded1a77417edb93d02b14019d68a37a1eb14d1db57f43b75e2a6f32

Observation 7cdfcf04-d675-4a18-854c-468ef8462bf5 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.252122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.095039Z digest=sha256:15b7cfce98ab7b85c6e278a6f8a0f42e42dba514094cbac6c52668580f261a83

Observation dc732852-44aa-4a48-9b37-8a32aae9c42d · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.239884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.098283Z digest=sha256:76e590bebfef62118c6ec231902feff6c60cb0518bb25564b56f718c9de652aa

Observation f23a6cc7-1999-48cb-a9f3-eea0fd23118a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.228139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.101419Z digest=sha256:2f5955f5aea434d153b069a9bb9d3b548b96d78d4ddaac6173a24741b202f8b1

Observation 58996f95-d9c1-4333-a044-f7b8b12fb5d0 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.216422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.104867Z digest=sha256:03d8c4db2f6b73fdfb112ff1564723d5966b047b7704503662ddec9abea67c3c

Observation d965385b-4881-4497-a283-25194e44450c · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.108339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.108339Z digest=sha256:b0282eacb569f6e589ee41e1dc5c64bb8b8822424a8843aa30e371724409d273

Observation f77f1f80-dc91-4ecc-b75a-d070ee67d29a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.111921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.111921Z digest=sha256:14d6007b652c29d2f56f59bda00f5bc333323a650b2c245eaae15558c75e2e9f

Observation f249484d-dc3b-4482-b5f3-a6ae44cf3c29 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.115370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.115370Z digest=sha256:76576bd24e842aaba41f428d21f93eb4990ff7f5e2a92717690513f66f36dcf8

Observation 01d72d3a-a244-4dcc-9acc-ac26c3dc937d · outbound

This paper cites Muon is Scalable for LLM Training.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Muon is Scalable for LLM Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.118553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.118553Z digest=sha256:4b3a6092a06fab1266120ccf63c22976025f2e4b8566cda03dd17e25e5268801

Observation c88cdf3a-b3a5-47fd-9c7a-756791a6b415 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.188577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.122075Z digest=sha256:d3d0a6de95ddd4bf40fbc0b8ba9eb760b76cac1ab6402b94ea3044093e88ec22

Observation bd877286-b983-49ef-a3a1-6eeba05ac596 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.176770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.125346Z digest=sha256:6ce75fba9ef3cc164501f2dd9b06deac6b2c8e51c9b814670c90cf52929f3ae6

Observation 60d1ffe9-8272-4615-84eb-d64e913ea568 · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

Taming LLMs by Scaling Learning Rates with Gradient Grouping DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.129514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.129514Z digest=sha256:c099dca5970a1b4e2ef9bc554c0e1033a726bbe230e49a2a86222b84a7a36a38

Observation cc653f9c-4569-4eb6-86cf-d8074085a207 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.164748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.132995Z digest=sha256:7020b6e335b193a8d4c0115997db9ef6476423d5188a02e511b4db1adc830d4b

Observation e9786330-16e2-47b7-8b01-43fb2e0337be · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.153328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.136203Z digest=sha256:f9fb2ebd277b9d276722897ecc0bc93250701c69d3436eb77e545556312af8f2

Observation 476339c5-5912-4f3c-bb0e-e5fb722a7511 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.141876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.139446Z digest=sha256:0bd145f14754b509911dadfda4bf53db7f7d92660e2211b1817259a3a505a5de

Observation b931a06c-dd32-4450-988e-2683564428c6 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.129665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.142621Z digest=sha256:c508a964b941f2dfe219c4b4d930e7731a3eb5824f1cef9e7b81c4787219a54f

Observation 16288887-825e-4553-aa28-8fd0e6c42175 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.118083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.145820Z digest=sha256:41a351697ed74007562fb71af235b91365c8ccd20e89af3bf6a312b73e719d8f

Observation d1cd5d36-7d70-46d1-9dfc-ce586ea37d32 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.107042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.149567Z digest=sha256:beb69e4d90ce7d2d787456d0904e4cb87ee081de6aa9e2216930a4078f57d029

Observation dfa9f3e3-acae-45c6-8047-415b0567a2ef · outbound

This paper cites CAME: Confidence-guided Adaptive Memory Efficient Optimization.

Taming LLMs by Scaling Learning Rates with Gradient Grouping CAME: Confidence-guided Adaptive Memory Efficient Optimization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.152833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.152833Z digest=sha256:65730284e95ff14620ec94dab9b24761ebca2ce87478ba3ea60370ebe6d4562e

Observation 28962728-0aa8-47cb-bd0e-ce6920dfe53a · outbound

This paper cites Visual Perception by Large Language Model's Weights.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Visual Perception by Large Language Model's Weights

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.156262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.156262Z digest=sha256:a705641488c0d9bc1a642af7099b6339fb2a54281bc88ed5f4a36d121b242bd6

Observation 897b4096-df09-470e-8495-9f998b50dd0e · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.095539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.159763Z digest=sha256:c6173a90f98deca8b818942a9e1c51523f37c54b46fdb5f453da22ecd30491d2

Observation 3a9a2b64-3069-4a7c-ae52-e76ccfec22df · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.163047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.163047Z digest=sha256:a52c64493aff92ee1798b1f969b1d5c5d36e142e573a6e841bbe748f84f32528

Observation 432c66dc-b5e1-4d55-91c0-e6a82faaa734 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

Taming LLMs by Scaling Learning Rates with Gradient Grouping A Theory on Adam Instability in Large-Scale Machine Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.166603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.166603Z digest=sha256:73fc98c117b421ee969c39047a1ad7c4d7be0fc4430181ee16627eb48213ad36

Observation 21b291a7-59a5-4f8b-8bc7-13a0211d6ec5 · outbound

This paper cites EDoRA: Efficient Weight-Decomposed Low-Rank Adaptation via Singular Value Decomposition.

Taming LLMs by Scaling Learning Rates with Gradient Grouping EDoRA: Efficient Weight-Decomposed Low-Rank Adaptation via Singular Value Decomposition

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:57:29.517503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.170521Z digest=sha256:54b76d716078eb6b99f17409817b213df4ad169aa4c539e7b3c783e1a131b528

Observation 7c681ca2-5a6a-42e7-b647-7079d06871f0 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.082626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.174065Z digest=sha256:5f1ec08a55f15773efed0b5675982fa5acab3a3410a8c316209508ec96ac0b93

Observation b89d3d72-25b6-4554-8823-3b457dc481da · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.177371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.177371Z digest=sha256:ab60aa5ad62bba14f580968c37ea6cc6da9a359a7b59607956400bf3ca502c3c

Observation 63d0ecf8-4439-4f80-b8f9-a4d88e9b00a6 · outbound

This paper cites Reddi, Satyen Kale, and Surinder Kumar.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Reddi, Satyen Kale, and Surinder Kumar

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:30.062417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.180610Z digest=sha256:0a7147accfe98a9452b6d9ea10419d3643c80f2427e71ea7a227434f2494ed13

Observation 89ed90fc-8074-4108-bc77-7327446f2e31 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.183773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.183773Z digest=sha256:4abf711c94c6f0c1be43cf028b5d70d2713fd3b608a02eec54937ae966ae818f

Observation 10022d71-56be-4848-a29a-443a275c82c6 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Taming LLMs by Scaling Learning Rates with Gradient Grouping SocialIQA: Commonsense Reasoning about Social Interactions

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.187200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.187200Z digest=sha256:810d1899b87519cdd52a2e0dfdccc82f7e4dfa19550ee93e8ee7eaf8df11e4ac

Observation 35116b00-3bc9-47c6-9fa4-a821ae2ae174 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.041991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.190488Z digest=sha256:e9cfc525862a6b63f5cd654eb0e37d41c639ac2af9d398dc21d032c937dc826e

Observation 2e966505-a288-4ba3-b062-3fa242178992 · outbound

This paper cites Adafactor: Adaptive Learning Rates with Sublinear Memory Cost.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Adafactor: Adaptive Learning Rates with Sublinear Memory Cost

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.193839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.193839Z digest=sha256:6a818447d4f025ccef257103627d9cc2c9e3c963262418fd7a0f21d40c5cc1e3

Observation e5acd949-8072-4ba1-91a2-67b0f7a43217 · outbound

This paper cites LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.197574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.197574Z digest=sha256:90d3539ef0dac5cf9ae4124495f41e11d97d8df075431ca2e8b05004184f32e4

Observation e7cbdbc6-d00e-48ee-bb36-3159467da3d8 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.200848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.200848Z digest=sha256:30c63e8a2c775c12cc620b1f0cbd45731818d3f81ea7a7c9d6eb47eb24cb98bc

Observation 35eca4f5-bb6e-41e3-a6af-5695ac55c152 · outbound

This paper cites Sinha and Michael P.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Sinha and Michael P

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:30.021951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.204136Z digest=sha256:c2ba81782afb2df1ee865c965e4cdbe4357216e0648cf66527f5e026407a9f9d

Observation 05858b3f-1b4d-414c-91b4-fcfbaf86820a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.207378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.207378Z digest=sha256:bca58f8306ea42c9ade51007567590302ead6693b91e2b374e246d720204de44

Observation c41112a3-43f6-4036-b58a-eefa42ff0149 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.210678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.210678Z digest=sha256:9fe6dad0ff49934a38fbd1f527d33e3606549f579e526b5a2acf575af1125775

Observation 93957090-2b1a-4fb8-b488-492be99e021e · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Taming LLMs by Scaling Learning Rates with Gradient Grouping SOAP: Improving and Stabilizing Shampoo using Adam

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.214022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.214022Z digest=sha256:384c42cd11c0b81c811fe04a4a2356b73b8244ce9430ef411d526b7f351c7adf

Observation 67f2c45a-9ca7-4b76-84dc-c6dcdb8ce9d1 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Taming LLMs by Scaling Learning Rates with Gradient Grouping GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.217387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.217387Z digest=sha256:0f6acb397998a24c9ea1eb41626d44ce1cf3c9567386ef6cd9ba260887dca83e

Observation a56bb22a-4cd7-401b-a945-b401ac0a3a33 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.992788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.220927Z digest=sha256:e0423e931a11cfbfdc0523219df7521aa75e5b7dcfb0c3eb2da50b16fd408152

Observation 58806685-fab1-4a62-8eef-0b1953f973ce · outbound

This paper cites Qwen2.5-Omni Technical Report.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Qwen2.5-Omni Technical Report

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.224160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.224160Z digest=sha256:cf3148b07896f8e1d83af0cbc78157e49c2deb76f1fef324e299dca16afdce68

Observation 27195656-a9b3-4348-bf49-d5c7179c9221 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.980801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.227542Z digest=sha256:67ba867832c2fbbeb68deb19d9658c83da683fcbbf8cb907e37db520adbaa0a5

Observation 40064f04-3bae-490b-bee2-4066438baa9a · outbound

This paper cites A Survey on Multimodal Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping A Survey on Multimodal Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.230732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.230732Z digest=sha256:e3fd33ccb977d0693e5d1e3f2512132da3af591768f782886bda28ea1c0d8899

Observation e0699f0b-c24d-4f2e-b5d5-575773268cff · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.967560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.233998Z digest=sha256:bc9099100428545d5033a339c6463d61cb3e4f7aef90a9b47e5ef8ea4409181f

Observation 488b4aeb-24ea-4312-9f2b-73aa3153362d · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.955405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.237565Z digest=sha256:0cebf310a9a8c77e880df8e9a094f55e47f8eb36837e11ef25389628e1b757d7

Observation 29cfd7c2-10a0-46a8-859c-59632a011363 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.943249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.241519Z digest=sha256:3c80c7434914cdc81a88f805f108390fe51842a5b413991d9ac5a70738e820b5

Observation ca378888-4d58-4a7c-b368-cd91497d0e09 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Taming LLMs by Scaling Learning Rates with Gradient Grouping HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.245272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.245272Z digest=sha256:1f263eb2e2674004180fb28d6965b2ab659e8af856be506c289c79826bcc89c8

Observation 9a396661-7194-4d09-9c32-792aaac1e40c · outbound

This paper cites Parameter-Efficient Fine-Tuning for Foundation Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Parameter-Efficient Fine-Tuning for Foundation Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.248704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.248704Z digest=sha256:f8acd335be62037a54979f6fb42e5093a6a726b2d6b59e6b2b79d700360b2d60

Observation 210d5578-f9c6-4cd4-8a05-e7b376f11d57 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.932025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.252216Z digest=sha256:91edee7825d4deb4132d462caaa2124e003e72aadab04c5f9e0589940e6965ca

Observation 08551d6d-0321-4715-9421-bc9c26773f70 · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Adam-mini: Use Fewer Learning Rates To Gain More

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.255600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.255600Z digest=sha256:2cf4183f999da79091caf289e2d8adfffaccfa731d65fcef8703eb2e29b19756

Observation 5d65b0bc-6e39-43ed-8c87-c595703282c0 · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Taming LLMs by Scaling Learning Rates with Gradient Grouping GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.259287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.259287Z digest=sha256:e9bd54e1b923ac5f73c40551b7dc045eab60bba3ae49b46278816bd41e0e49ba

Observation b19eaa3e-d997-4475-af1e-df2e6a124e0c · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Deconstructing What Makes a Good Optimizer for Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.263146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.263146Z digest=sha256:0076f12916135117681c440590bde9bd6e440ffa61af725795ab77e59814fa9c

Observation fcc700fc-9253-4936-a883-bf2c1addffe4 · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.266772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.266772Z digest=sha256:9a04c47cc5df479ecc41c57ea4e1536fa951aae3352312947d7908a5e09de2f9

Observation 4ec620a7-60ca-46d6-8a67-437c76a0961b · outbound

This paper cites APOLLO: SGD-like Memory, AdamW-level Performance.

Taming LLMs by Scaling Learning Rates with Gradient Grouping APOLLO: SGD-like Memory, AdamW-level Performance

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.271257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.271257Z digest=sha256:a957f77dd4dd71d389284cc85608fae2278cc0a11247612874883289dd2aa330

Observation b4fcc1d0-1de3-4fe9-9690-5d3932797f3c · outbound

This paper cites Transformers without Normalization.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Transformers without Normalization

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.274935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.274935Z digest=sha256:3a2f8e5ca3f793eccdf6ab58cf2275cb4ac50408eb5cc903fda6046f224f0b03

Observation 52272ecf-36fd-481c-87d1-ff47692244fe · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.920689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.278658Z digest=sha256:722ed9596932096e11f2209e4948db1e642811606c2d1a65b2291c34f04cd7aa

Pith citing papers

Observation 39b37602-eb7a-4980-b4a0-f8e3a6ae7eb7 · inbound

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio cites this paper.

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio Taming LLMs by Scaling Learning Rates with Gradient Grouping

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:06.419261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T15:39:51.611115Z digest=sha256:cad24edfdc559637b4f108a0d23977cb8118803a958993129d728e7424304bf0

Observation b1e36740-728e-47b1-8590-2f2aeec73c55 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Taming LLMs by Scaling Learning Rates with Gradient Grouping

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:c3744b2a4840b4851865e09eca79b6334a9fcc15b099cc4062b41641b8da60c1