Pith. sign in

Paper Citation Record · LEDGER

Taming LLMs by Scaling Learning Rates with Gradient Grouping

As of 23 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 2 inbound Pith citation observations for arXiv:2506.01049.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01049 v1

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:29.278658Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T22:10:49.683444Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T16:41:06.415113Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved87
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0963deb6-c74b-4b8c-80b6-ca00a44c6488 · outbound

This paper cites online" 'onlinestring :=.

Taming LLMs by Scaling Learning Rates with Gradient Grouping online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.960786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.960786Z digest=sha256:e4a10c600d9747eaf91a70a122086bbdb9cb0c67dd5e5596f3eb212150c436a4

Observation 551f9183-6011-4774-a690-29566546a9c1 · outbound

This paper cites write newline.

Taming LLMs by Scaling Learning Rates with Gradient Grouping write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.965573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.965573Z digest=sha256:e0a278d38e5a353fbedf5c4978b8c7dadfca6fc47672eee8344fad37379b628a

Observation b2f78582-7e3d-46ff-ad52-36f938fb7e84 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.969116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.969116Z digest=sha256:03898da6e959c5c6f7605a1133adb371533a9ce0b1faefffe4451b97a3e78ed6

Observation 0746c1e7-aa97-4e50-b269-2fbc6807cd5e · outbound

This paper cites LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.972955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.972955Z digest=sha256:8268583064d596811a14aa3e04a07d6b124784d045e332c28bad45aabfdc5baa

Observation f31865e9-46bc-4f4e-83af-6ca0a2eb761f · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.976631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.976631Z digest=sha256:49067eb214e0873046c2404fa656e928b4d8e28c8f0dd47674a6cc62ea92bb94

Observation ff8d07a4-c54e-4f38-97eb-367bb43c0c7f · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.980071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.980071Z digest=sha256:96f2d0c9cc08a91cfe31b82ca3d5378e4cb771acf99407a20d795b4cdae7ac1e

Observation 52fac5ca-1489-47af-8830-ce9d1cc63cc6 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.983339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.983339Z digest=sha256:fa1fd555b5e94cbb011d8d120c61f1051fec873da851d51058e94f3b94fd27e1

Observation acd9c9c6-d16c-4306-a6d0-9b5881e8ea8c · outbound

This paper cites LLaVA-KD: A Framework of Distilling Multimodal Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LLaVA-KD: A Framework of Distilling Multimodal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.986563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.986563Z digest=sha256:edc139458186be4e953221134debe29da8a67db9144688e948f03c4da3e6c365

Observation f2b3a661-29ae-442f-9be4-8049a765a487 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.991360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.991360Z digest=sha256:4984741986f8ffe5cde7522fc2bd526a5ec83ae83348fe29330e949c6992055b

Observation ad7b1bb8-15bc-46ae-9929-e2f44bc0ca30 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.995330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.995330Z digest=sha256:0f44f6f68c211890aaeceb3176eb9bce283855e5daf64a1443fe73a71824e6bb

Observation 3cd80351-bc8c-47b5-abf1-f6c9500b5ca5 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.998658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.998658Z digest=sha256:c4b96dfb1d942f5a198d7ab62c512d55cb594399cd5faaeb7ba6723ccdeb3506

Observation e68aeed3-1427-4dc0-bee3-9e33ee9beed7 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.002055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.002055Z digest=sha256:d396bba8feac68db5c199110d1ea8b0ff7ab546702e54ca4a62f090ab785c714

Observation 692be308-d7b0-470f-9cb0-d27506ec7c40 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Taming LLMs by Scaling Learning Rates with Gradient Grouping BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.005938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.005938Z digest=sha256:b748c45fd2358bbae48dc41e2d9c634ee8d199162cf4934b23d0e25909aa7d89

Observation e73f06e3-5284-43c7-a3a7-e651d2fa2cef · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.010001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.010001Z digest=sha256:8425e1b47c4b47640acc4c4596a77628639f7339b5da07ec2d8470ebe4ade340

Observation 4038c6b4-e5b0-41c2-acfa-ba023fc04a99 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Taming LLMs by Scaling Learning Rates with Gradient Grouping InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.013438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.013438Z digest=sha256:30b6a743d04d45b35e3b0f15eebeeed5fa4ea72a6c65b42d748431f4186c980b

Observation 9bc5028c-58e4-4e6e-b142-8b2c4c3e1e08 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.016941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.016941Z digest=sha256:bce9ff30eb9092ac7e5b4d170822c2fd4d04c7a6352b4e938c7bcb28ebcf6c99

Observation 3d09eed9-f3f1-4f1b-b336-868922556f8a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.020514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.020514Z digest=sha256:68d463e8dc82fc59897dc99eb533e6c9e228a7095f2b034b45865af12e02c84e

Observation 2b3ff932-08e5-4c17-91a2-f03696222cde · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.024107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.024107Z digest=sha256:a988eed07373431a885f0aa636993d705cf416c40984c51566ea39eb152f3a0e

Observation 82a835d2-ab35-4e35-8390-b26a6bc7eb62 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.027322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.027322Z digest=sha256:0ad582075296d10bcb4d67f467a37785ddb8831b5f804789314d81f456d34c3a

Observation 971fbb92-8263-4497-9930-f1671856488c · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.364643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.030485Z digest=sha256:432b94d1b769228a31d42d16f190282cb2636cd333fab7dcd516718c9978eec7

Observation c48c3115-6596-43e0-b15e-8798df9e7736 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.352712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.033741Z digest=sha256:dbde56751012c9bd06446484e51e2eda448a4e451ea89d0d29ae70b51632902d

Observation fd64ccbf-f79b-4857-82be-e64792012b92 · outbound

This paper cites LoRA+: Efficient Low Rank Adaptation of Large Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRA+: Efficient Low Rank Adaptation of Large Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.037629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.037629Z digest=sha256:57c1d16034d5c7107717cb3c80b92e19305ca6dd470a5b21a375fb1f7a297118

Observation 7292eecc-2457-4f45-a096-18a097d6434f · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.339368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.040979Z digest=sha256:422bced85a06d8805d25a63eb82129390b7a91c718774a8e8420ddc90c859612

Observation 165ad793-8fe0-459a-8b08-99fa783084ed · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.328100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.044311Z digest=sha256:9c9225a0ca16ae038d06f4c1a4438e0ed166560fa0c7bc139e2060ab2824f1b3

Observation d79f9994-2fba-40fe-89ff-79a511e26d96 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRA: Low-Rank Adaptation of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.047501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.047501Z digest=sha256:e5dacd2d5c664c3aa6553e51b7fb57a9ceeb50303e859740b6148e733fbd3437

Observation 05cda9ed-9b95-4933-ad63-6ddbaff80b92 · outbound

This paper cites LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.051249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.051249Z digest=sha256:0197545c857c020e4db640ff22c528a91e500c7f8d40bb0bfed331f6a7a03b89

Observation 9dd4f4fe-d38b-496a-ae49-bc58b853c8da · outbound

This paper cites SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training.

Taming LLMs by Scaling Learning Rates with Gradient Grouping SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.054680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.054680Z digest=sha256:a594df18602321061ae27259e0ec13df1a78872b254a4f48f924af336b903e51

Observation 331539f7-7b6a-463e-93b9-846c81e04baf · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.058208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.058208Z digest=sha256:f349041fbb425330f59dbfba835bff4549367698f5940978eab047a53fcb3ef5

Observation cbab1350-ed83-40e4-9c6d-b1b7bb62ffb6 · outbound

This paper cites Exploring Low Rank Training of Deep Neural Networks.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Exploring Low Rank Training of Deep Neural Networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.061428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.061428Z digest=sha256:8c577fc040dbf91b142ad350b4652bd95c4959743d2b2e86dcb57c6ef6495bcf

Observation 0f00894e-3412-4e42-903f-4742cd1e9907 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.307740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.064748Z digest=sha256:dc50ba4936ceb4c71c2e21aa3a96d0c53fa572d00757a07069e131e80901926d

Observation c83b9a68-badd-459f-8bb2-6f2024d8bbd8 · outbound

This paper cites Kingma and Jimmy Ba.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Kingma and Jimmy Ba

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:30.296503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.067938Z digest=sha256:d2ed60ccea85aea2567a958983c9ab6d1cfb17c6ab15687159c1fa0472ce7004

Observation ae32cc3f-32fa-495d-8410-ca4a74769a87 · outbound

This paper cites o pf, Yannic Kilcher, Dimitri Von R \.

Taming LLMs by Scaling Learning Rates with Gradient Grouping o pf, Yannic Kilcher, Dimitri Von R \

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:30.285335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.071052Z digest=sha256:dba06372b808172b6d047670748488548ee8f0d4c90575eb354efd57a7cc2faf

Observation 05593c1e-0953-4188-becd-72abe6222c2d · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Taming LLMs by Scaling Learning Rates with Gradient Grouping SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.074356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.074356Z digest=sha256:9b80b6df7b5c3a8254d13b3f857193f87030b537f0ff6a1918bd3658271b0830

Observation 109676e1-f73c-4df9-a5d3-778c96f7cd5f · outbound

This paper cites LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.078239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.078239Z digest=sha256:f8c605dc5e75e0dab38473309d6718878a0c70ab5c4a0c7b685ba5f147392426

Observation 10c826ac-215f-4559-8d2c-002b3bdd6f5a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.081480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.081480Z digest=sha256:a456052a880a14447f7ba6a5a1c939ed32d3f85dad99f6e02223876e4e5b974d

Observation 84c01ecb-d986-44d9-846c-1c85806a28e3 · outbound

This paper cites Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.084798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.084798Z digest=sha256:14571c815056c1aa8db7b646158bfd6f3f493e3d848065aad30a464ac7be1e33

Observation 88980bb8-475c-4efc-8aa0-b0dffb53b9f5 · outbound

This paper cites Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.088289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.088289Z digest=sha256:20d56fe5f35886bf1119a495d5f36301a175615ba996a8b0b7ad37822aa9b169

Observation 364bbec4-e582-4037-b2b8-ce6cc5bc78cc · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.264637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.091805Z digest=sha256:9149f9a0c5cd91abb2bd3a629f43bf7b0d05e9aabe70dfb7dbb5a4412be563a9

Observation 7cdfcf04-d675-4a18-854c-468ef8462bf5 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.252122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.095039Z digest=sha256:4962a1b80afd42c2fbe50394352f2cfd10f3e037659c7954d795314e6e03a033

Observation dc732852-44aa-4a48-9b37-8a32aae9c42d · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.239884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.098283Z digest=sha256:21bfe837f658be1b5629bb223522b1d3b7deadafa2ccd9c202b9ae9d301b3cb0

Observation f23a6cc7-1999-48cb-a9f3-eea0fd23118a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.228139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.101419Z digest=sha256:cd0244760b49d10413fdca9b12dcd30baa1ea05445299902328c4ce61568674d

Observation 58996f95-d9c1-4333-a044-f7b8b12fb5d0 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.216422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.104867Z digest=sha256:ba2a9e45097e1964293e5d51719b9c0f9e5e6a509feddd935613af1ca9e1c884

Observation d965385b-4881-4497-a283-25194e44450c · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.108339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.108339Z digest=sha256:775b8f6c98d5dd09eee927f0b0e6138daa25e9cbfcb21a5312dcc9c510f9123b

Observation f77f1f80-dc91-4ecc-b75a-d070ee67d29a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.111921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.111921Z digest=sha256:3d2c1da984d3e0881f7b62da6e1d33f7a37b8d0efd985feb9a7d70c28a777a0d

Observation f249484d-dc3b-4482-b5f3-a6ae44cf3c29 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.115370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.115370Z digest=sha256:6edc58f3f02c984c93a709776982ac1e93281c530b7ac9366265dbd799086a48

Observation 01d72d3a-a244-4dcc-9acc-ac26c3dc937d · outbound

This paper cites Muon is Scalable for LLM Training.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Muon is Scalable for LLM Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.118553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.118553Z digest=sha256:5bbe43a0f3eb1390e92188a0b70b8ca73d6b04c2cbc4544d9ba9161ccd259a68

Observation c88cdf3a-b3a5-47fd-9c7a-756791a6b415 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.188577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.122075Z digest=sha256:9f9f01cec3032ed3e0047eda969ac37a3b716044251cf1e5e9b15a60e99a3c02

Observation bd877286-b983-49ef-a3a1-6eeba05ac596 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.176770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.125346Z digest=sha256:11b605ab2a833ac82c6ae3d87776e4aadeb0d23b7c708b02a9a1409dd3f3fe8d

Observation 60d1ffe9-8272-4615-84eb-d64e913ea568 · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

Taming LLMs by Scaling Learning Rates with Gradient Grouping DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.129514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.129514Z digest=sha256:cf5f46c2d25abd5296b24cd200b840d034d6b1d7d0d739cf9ac550d55065d45e

Observation cc653f9c-4569-4eb6-86cf-d8074085a207 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.164748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.132995Z digest=sha256:caf15cf4dfbf165712f6ae6c425b2103ff0804532e1c981e64fdef737aabc2d0

Observation e9786330-16e2-47b7-8b01-43fb2e0337be · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.153328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.136203Z digest=sha256:bc192fa9c98bcc85b728e95bf7fe112a888cb3833b1b668b27b126b9c51bd888

Observation 476339c5-5912-4f3c-bb0e-e5fb722a7511 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.141876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.139446Z digest=sha256:6ad5cbc98e3c06562f17fa279485e65515cf5ca4d02a037e6b5872e4d61fd420

Observation b931a06c-dd32-4450-988e-2683564428c6 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.129665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.142621Z digest=sha256:6e670e4a9afd147df697493d008f6897e91537976df4f3984bdf697ec3965bc3

Observation 16288887-825e-4553-aa28-8fd0e6c42175 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.118083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.145820Z digest=sha256:aed5a20221411b44947ced7c76e3e0e6616d3295137eebdac0730feeef7ec2a0

Observation d1cd5d36-7d70-46d1-9dfc-ce586ea37d32 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.107042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.149567Z digest=sha256:c9b1efd43e62714d069c98b9e0f1c8597b1b37ffee8fd1d6ae96502a7513ab9a

Observation dfa9f3e3-acae-45c6-8047-415b0567a2ef · outbound

This paper cites CAME: Confidence-guided Adaptive Memory Efficient Optimization.

Taming LLMs by Scaling Learning Rates with Gradient Grouping CAME: Confidence-guided Adaptive Memory Efficient Optimization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.152833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.152833Z digest=sha256:6fe157015e36c99ba738ff00b92c2f14b26a44602b1d32b68f514a4b77b28bb6

Observation 28962728-0aa8-47cb-bd0e-ce6920dfe53a · outbound

This paper cites Visual Perception by Large Language Model's Weights.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Visual Perception by Large Language Model's Weights

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.156262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.156262Z digest=sha256:1fd2597c62ded05a2982e5d14e0b2391a65af7945f6dd9a84a3ba7c1b36e646d

Observation 897b4096-df09-470e-8495-9f998b50dd0e · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.095539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.159763Z digest=sha256:067cb75f0a76cdf3139649764b2471eea3bcc83420a7e4470434709ed89e9ed0

Observation 3a9a2b64-3069-4a7c-ae52-e76ccfec22df · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.163047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.163047Z digest=sha256:961ab54290ecdcadd7db53ccfec335d39517342e3f7e95d668cf9290a77ca1ad

Observation 432c66dc-b5e1-4d55-91c0-e6a82faaa734 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

Taming LLMs by Scaling Learning Rates with Gradient Grouping A Theory on Adam Instability in Large-Scale Machine Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.166603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.166603Z digest=sha256:fc3873219caaef6d020feb5d518113cb1575dbebe7f277542d6fa5b225a8ef27

Observation 21b291a7-59a5-4f8b-8bc7-13a0211d6ec5 · outbound

This paper cites EDoRA: Efficient Weight-Decomposed Low-Rank Adaptation via Singular Value Decomposition.

Taming LLMs by Scaling Learning Rates with Gradient Grouping EDoRA: Efficient Weight-Decomposed Low-Rank Adaptation via Singular Value Decomposition

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:57:29.517503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.170521Z digest=sha256:3006d8f5f9ed345c41ed164993a045e870fb337ffa8195bcfdc1327b71e5642e

Observation 7c681ca2-5a6a-42e7-b647-7079d06871f0 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.082626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.174065Z digest=sha256:21b3636b20bce925962eab245aa0ced1fa7c23e3026aff211368fad47685b971

Observation b89d3d72-25b6-4554-8823-3b457dc481da · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.177371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.177371Z digest=sha256:840bc44b04d96c49befe3c3b97212d04ab41da73db96bc1b7daf898f8e816b83

Observation 63d0ecf8-4439-4f80-b8f9-a4d88e9b00a6 · outbound

This paper cites Reddi, Satyen Kale, and Surinder Kumar.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Reddi, Satyen Kale, and Surinder Kumar

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:30.062417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.180610Z digest=sha256:d49023cb894a0bc38adb0892669bd8309975fca5df76b29225289567f1ecf213

Observation 89ed90fc-8074-4108-bc77-7327446f2e31 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.183773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.183773Z digest=sha256:c20502dc35dfdc95047a929dc9b443cdbdd047df48dc552dd455166822e47d08

Observation 10022d71-56be-4848-a29a-443a275c82c6 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Taming LLMs by Scaling Learning Rates with Gradient Grouping SocialIQA: Commonsense Reasoning about Social Interactions

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.187200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.187200Z digest=sha256:4ac8cba85db08fd99fff0dda88a6ab4324e07b13b39dd97a1b5ad561a3e8cb79

Observation 35116b00-3bc9-47c6-9fa4-a821ae2ae174 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.041991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.190488Z digest=sha256:ab082b004f83b671f97933489ad189263e8d91bea0bf47e544bddbd6aa8f97c7

Observation 2e966505-a288-4ba3-b062-3fa242178992 · outbound

This paper cites Adafactor: Adaptive Learning Rates with Sublinear Memory Cost.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Adafactor: Adaptive Learning Rates with Sublinear Memory Cost

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.193839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.193839Z digest=sha256:f66d88fb79f68e6737912e30642e40a7a8a25a3db0af666990caccb561f92902

Observation e5acd949-8072-4ba1-91a2-67b0f7a43217 · outbound

This paper cites LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.197574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.197574Z digest=sha256:011a31aa5589d1a19f9a0ddbc845e8384f08249c86f12bbc84c30f6a33886399

Observation e7cbdbc6-d00e-48ee-bb36-3159467da3d8 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.200848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.200848Z digest=sha256:80928fdf0be829b55e2f48b24f29b8651fa411eb8243bc35132db3fdf9e295e1

Observation 35eca4f5-bb6e-41e3-a6af-5695ac55c152 · outbound

This paper cites Sinha and Michael P.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Sinha and Michael P

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:30.021951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.204136Z digest=sha256:5517a42330da420bfff6ed4b9ef3ab99602cddefd393a91ec17bb22bceef4f2a

Observation 05858b3f-1b4d-414c-91b4-fcfbaf86820a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.207378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.207378Z digest=sha256:b5a9e0c1514fd4e097175872d955c99dabaf24989cccf63f4d6866961c4d4f0a

Observation c41112a3-43f6-4036-b58a-eefa42ff0149 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.210678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.210678Z digest=sha256:3c80aaf258c4b56c2f51ae312dd8a45a63c0e66d1a41b8d0eee523addf53199f

Observation 93957090-2b1a-4fb8-b488-492be99e021e · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Taming LLMs by Scaling Learning Rates with Gradient Grouping SOAP: Improving and Stabilizing Shampoo using Adam

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.214022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.214022Z digest=sha256:3a1978e7d1b08d5c6838f614203ff031df7b9a30b799e497b3f8e862f5010c78

Observation 67f2c45a-9ca7-4b76-84dc-c6dcdb8ce9d1 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Taming LLMs by Scaling Learning Rates with Gradient Grouping GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.217387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.217387Z digest=sha256:93779debf0b560c67f672c4c93ebe68d16d243607750a6f4afe518a98f34eb6f

Observation a56bb22a-4cd7-401b-a945-b401ac0a3a33 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.992788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.220927Z digest=sha256:3c5403572983106b4de7c57f0483dc7906b0b50c9ffdfea0ffa31d475b4a6d0f

Observation 58806685-fab1-4a62-8eef-0b1953f973ce · outbound

This paper cites Qwen2.5-Omni Technical Report.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Qwen2.5-Omni Technical Report

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.224160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.224160Z digest=sha256:6ce0b81bbb9680c5c33599614b2fce74a6c9a278564665d7a067d96fb6cae145

Observation 27195656-a9b3-4348-bf49-d5c7179c9221 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.980801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.227542Z digest=sha256:d5863888c196bc954868a886f526e674eb622f7379961f5343d8f162b2b9364f

Observation 40064f04-3bae-490b-bee2-4066438baa9a · outbound

This paper cites A Survey on Multimodal Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping A Survey on Multimodal Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.230732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.230732Z digest=sha256:bc292301ab27c71667d057458805668bee60085357573f4108411815ad822b99

Observation e0699f0b-c24d-4f2e-b5d5-575773268cff · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.967560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.233998Z digest=sha256:e998f790b9633b326da0661cf6c09a7718d0a31808411e55edd8a2369a2ded05

Observation 488b4aeb-24ea-4312-9f2b-73aa3153362d · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.955405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.237565Z digest=sha256:717b1726118d0bc41180e419eec3ddb9303b7c9f54c6b2b5bb38cdcb78962638

Observation 29cfd7c2-10a0-46a8-859c-59632a011363 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.943249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.241519Z digest=sha256:755d3199c97604ed7ded5bf29cd36e6295c0aeab19e3a20b0953b7fe20baa01e

Observation ca378888-4d58-4a7c-b368-cd91497d0e09 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Taming LLMs by Scaling Learning Rates with Gradient Grouping HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.245272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.245272Z digest=sha256:b9a9c6c938762f89f3c30e6d3cbc402cecd80837ffaf117cb8e33cbf2a186eed

Observation 9a396661-7194-4d09-9c32-792aaac1e40c · outbound

This paper cites Parameter-Efficient Fine-Tuning for Foundation Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Parameter-Efficient Fine-Tuning for Foundation Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.248704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.248704Z digest=sha256:440e32be4f4f6693755b9501fd8ee9cd47755a8bb8eeaa9709999285448d2110

Observation 210d5578-f9c6-4cd4-8a05-e7b376f11d57 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.932025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.252216Z digest=sha256:6e12c92be408c178dda9b37c66ce6b38bc73e68831b68ce1e7254d366361b7d3

Observation 08551d6d-0321-4715-9421-bc9c26773f70 · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Adam-mini: Use Fewer Learning Rates To Gain More

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.255600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.255600Z digest=sha256:57856e65839f08bc79fd4827e1bdf187ca503ec51945b4f6ff8f3e97552332b3

Observation 5d65b0bc-6e39-43ed-8c87-c595703282c0 · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Taming LLMs by Scaling Learning Rates with Gradient Grouping GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.259287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.259287Z digest=sha256:329f16f781210eb20a1bc25d32c2e67efb4b522f442b8ddc353d315f0a1f78b5

Observation b19eaa3e-d997-4475-af1e-df2e6a124e0c · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Deconstructing What Makes a Good Optimizer for Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.263146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.263146Z digest=sha256:aa3b1e83007441b92e12d78d093181fe17037b5b4f06a793db4eefa53ee7abea

Observation fcc700fc-9253-4936-a883-bf2c1addffe4 · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.266772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.266772Z digest=sha256:d4c32df1b76426690f60597e3cbfc762c269ba42b987291f98eb9da5dd38e05f

Observation 4ec620a7-60ca-46d6-8a67-437c76a0961b · outbound

This paper cites APOLLO: SGD-like Memory, AdamW-level Performance.

Taming LLMs by Scaling Learning Rates with Gradient Grouping APOLLO: SGD-like Memory, AdamW-level Performance

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.271257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.271257Z digest=sha256:c50d23db173ae352058ea9e89dbfc1b3977fb84ea010fdb35ef1e42c455a68bf

Observation b4fcc1d0-1de3-4fe9-9690-5d3932797f3c · outbound

This paper cites Transformers without Normalization.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Transformers without Normalization

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.274935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.274935Z digest=sha256:7f09b601c33a8e8a29599b0c8899222e698e43da4606d42208ab303051bcf161

Observation 52272ecf-36fd-481c-87d1-ff47692244fe · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.920689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.278658Z digest=sha256:daa021d775e8d770bca480dda52f15b9a906b75a940d3943fd7c248f592ef5c9

Pith citing papers

Observation 39b37602-eb7a-4980-b4a0-f8e3a6ae7eb7 · inbound

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio cites this paper.

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio Taming LLMs by Scaling Learning Rates with Gradient Grouping

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:06.419261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T15:39:51.611115Z digest=sha256:93cf6fac0775e85efa3b882be46900d4ad74fd8eb43aec07bcdcb6f26ea80ed4

Observation b1e36740-728e-47b1-8590-2f2aeec73c55 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Taming LLMs by Scaling Learning Rates with Gradient Grouping

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:7cf5a4c0fb102469a3ac2095039d2ac3729c3bb06e8a6472cbf1105e8810f845