Pith. sign in

Paper Citation Record · LEDGER

Low-rank Momentum Factorization for Memory Efficient Training

As of 7 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2507.08091.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08091 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:35:40.252196Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact3
  • verified fuzzy20
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d091fdb-94ec-4d6e-8c7c-3f6b9ada94e2 · outbound

This paper cites Memory efficient adaptive optimization.

Low-rank Momentum Factorization for Memory Efficient Training Memory efficient adaptive optimization

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.844330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.092595Z digest=sha256:ae4dfb36fdbc15d50b78728b1533edab200614754dc033fd3f84f3e84461fc33

Observation b532b623-71d0-40f4-b010-b66bc957a580 · outbound

This paper cites Lower bounds for non-convex stochastic optimization.

Low-rank Momentum Factorization for Memory Efficient Training Lower bounds for non-convex stochastic optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.095813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.095813Z digest=sha256:f41e46a00aa4df3dcccfb1597670df7f93f76967a276e7f3f2fcf0e6d9955fa6

Observation 3fe12a6e-83c1-43dd-9681-700cffe0ce6e · outbound

This paper cites Modular Duality in Deep Learning.

Low-rank Momentum Factorization for Memory Efficient Training Modular Duality in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.098289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.098289Z digest=sha256:c9fb2163b91d0b6fd00756c96182d68e18b759d1752d6129797f190d41c857a2

Observation ce0bfadf-c388-47e9-b543-633023dd2d42 · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

Low-rank Momentum Factorization for Memory Efficient Training Old Optimizer, New Norm: An Anthology

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.101337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.101337Z digest=sha256:2d16f8adc0e29a4b7f10486aadec916d0b6440c7862ff2bd7fa04d5388aeb017

Observation ca466eb8-9b8a-426e-aba8-48cdb118eb72 · outbound

This paper cites signsgd: Compressed optimisation for non-convex problems.

Low-rank Momentum Factorization for Memory Efficient Training signsgd: Compressed optimisation for non-convex problems

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.832449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.104099Z digest=sha256:03f94a27bb83d42c48696b05fa00017576cb7e7f2f04dc7d6e376c90c1439d3e

Observation bbb69756-4f58-421b-9dca-bf1f25d5c1a4 · outbound

This paper cites Automatic Gradient Descent: Deep Learning without Hyperparameters.

Low-rank Momentum Factorization for Memory Efficient Training Automatic Gradient Descent: Deep Learning without Hyperparameters

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.106492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.106492Z digest=sha256:ad5e6fedab2d13c1f071f206c05234d9226b83a029f55a7467380ca539923655

Observation 01992ade-7bfc-4cf9-bad7-33c5713dd188 · outbound

This paper cites Symbolic discovery of optimization algorithms.

Low-rank Momentum Factorization for Memory Efficient Training Symbolic discovery of optimization algorithms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.825256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.109327Z digest=sha256:3ee592460e12aa41a015502c65df729c1aae9c69ffdce95c02376badfa9c2a4f

Observation 37add5ea-1f3f-4a0b-8e0b-d16bbe160b7b · outbound

This paper cites 8-bit Optimizers via Block-wise Quantization.

Low-rank Momentum Factorization for Memory Efficient Training 8-bit Optimizers via Block-wise Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.111493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.111493Z digest=sha256:540cbc30faf07fd848f872faf3c80443ccdfa37ea75a7877a8086ee0563841ad

Observation 1e8c6ba0-3238-45b3-965b-ea5cae56aedf · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

Low-rank Momentum Factorization for Memory Efficient Training Qlora: Efficient finetuning of quantized llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.114360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.114360Z digest=sha256:b688096feb0e63704504028c1feb1e1185706478bfd00a811fde2892b796836c

Observation 92ea956e-c64b-40a3-8cc9-9336f072b503 · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

Low-rank Momentum Factorization for Memory Efficient Training Adaptive subgradient methods for online learning and stochastic optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.117442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.117442Z digest=sha256:211636dcc8bc8e0f002605e1e6ab7eae699d7712a9fc934b2da5f0f12d68907b

Observation fba15dfa-01fe-458d-bf68-7558f6d01345 · outbound

This paper cites Combining axes preconditioners through kronecker approximation for deep learning.

Low-rank Momentum Factorization for Memory Efficient Training Combining axes preconditioners through kronecker approximation for deep learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.808185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.119531Z digest=sha256:fafa4b358912b01f365c06b2b2eaaaf4048bf45cede17178b36e2b7b7630d97a

Observation d8b1410a-a235-4978-bcc1-8cd76cb2083a · outbound

This paper cites Sketchy: Memory-efficient adaptive regularization with frequent directions.

Low-rank Momentum Factorization for Memory Efficient Training Sketchy: Memory-efficient adaptive regularization with frequent directions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.800933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.121654Z digest=sha256:9cc827e7425b4eb2a442f6151da36a552c00a88d002ad431a2b1aa4af215ea89

Observation f3523b01-4493-4be9-9ec8-8377958da5fd · outbound

This paper cites Fast approximate natural gradient descent in a kronecker factored eigenbasis.

Low-rank Momentum Factorization for Memory Efficient Training Fast approximate natural gradient descent in a kronecker factored eigenbasis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.792974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.123726Z digest=sha256:7d910f3ec0c9c4c4d064ea11bf68432e2bb78665d48a90268ce59aff543352d8

Observation 43c30931-6709-46ec-ba7b-63fdf9549231 · outbound

This paper cites Improving neural network training in low dimensional random bases.

Low-rank Momentum Factorization for Memory Efficient Training Improving neural network training in low dimensional random bases

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.785399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.125769Z digest=sha256:f59e4309ba13fadfb40cf7bd5e76e3fa2210f4316d46dd16a112c2520b178989

Observation 46d2730b-4bb7-4c8c-83c8-1e7073e1420b · outbound

This paper cites A kronecker-factored approximate fisher matrix for convolution layers.

Low-rank Momentum Factorization for Memory Efficient Training A kronecker-factored approximate fisher matrix for convolution layers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.777756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.127897Z digest=sha256:76ea871330c3513c354d8d646ad842a2229bcf511331b08a191b3ace01ae6a69

Observation d2334a67-9cef-41e9-bad2-19cb0492beaa · outbound

This paper cites OLMES: A Standard for Language Model Evaluations.

Low-rank Momentum Factorization for Memory Efficient Training OLMES: A Standard for Language Model Evaluations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.130029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.130029Z digest=sha256:0858390c489fead9565a03730b7884716c1fc55e7424c8b2f44d4e10f8e2c952

Observation cc7931f7-ca09-446e-b00e-37a06f26928d · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

Low-rank Momentum Factorization for Memory Efficient Training Shampoo: Preconditioned stochastic tensor optimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.770861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.132495Z digest=sha256:ada1015e01b916baabf656ec5f93738cfc069b2aa6ecca59a7028dbda9f75265

Observation 71536acc-4a68-4f4f-bfcb-ea96cc205e28 · outbound

This paper cites Gradient Descent Happens in a Tiny Subspace.

Low-rank Momentum Factorization for Memory Efficient Training Gradient Descent Happens in a Tiny Subspace

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.134826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.134826Z digest=sha256:0bf01f87e81f860e8ab7e49e8e6fbb163abfea176dbb563d7c2ded034e2d1506

Observation 7d13c518-a176-44f3-a921-2569a2abe3a7 · outbound

This paper cites Flora: Low-Rank Adapters Are Secretly Gradient Compressors.

Low-rank Momentum Factorization for Memory Efficient Training Flora: Low-Rank Adapters Are Secretly Gradient Compressors

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.137244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.137244Z digest=sha256:b92a95031b2746f71e87b482dc54ee145f4226fd8cf4c9713f4f6a4d418db4f2

Observation bd10f289-2c68-460e-b447-ed4f2b5479b8 · outbound

This paper cites Topics in matrix analysis.

Low-rank Momentum Factorization for Memory Efficient Training Topics in matrix analysis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.763802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.139626Z digest=sha256:3407eaa504ba599ff04e4c9efc0f9a7275ca2c449963c1e2e154341a97f967fe

Observation a823097d-4846-4f24-9830-cd0abd3df947 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

Low-rank Momentum Factorization for Memory Efficient Training Parameter-efficient transfer learning for nlp

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.141663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.141663Z digest=sha256:494be25e810c5b421fd825a33f6c7d6aba0bac540b04b2fc4467e6628848a2c1

Observation 81fac90f-2158-41c7-b5b7-fdaa8185c668 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Low-rank Momentum Factorization for Memory Efficient Training LoRA: Low-Rank Adaptation of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.143830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.143830Z digest=sha256:4f8e9d20c8dc684521c28085094db7ff608d56d3c1e0747b8967cc3c2442a657

Observation 305cc2a1-a26b-4291-9bf4-8fd0fb450d06 · outbound

This paper cites modded-nanogpt: Speedrunning the nanogpt baseline, 2024 a.

Low-rank Momentum Factorization for Memory Efficient Training modded-nanogpt: Speedrunning the nanogpt baseline, 2024 a

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.751685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.146130Z digest=sha256:42c2a166524157d05e27f51d66d43b1eb6c62a4fda94da971cba5a44a430d41d

Observation 86df6ef4-909a-4fda-b1dd-0e25d75367f4 · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024 b.

Low-rank Momentum Factorization for Memory Efficient Training Muon: An optimizer for hidden layers in neural networks, 2024 b

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.743734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.148307Z digest=sha256:c3baa019a390ba9af3e7cb1db6d10779a0f6667d4f487cc1423a8855ebd9c457

Observation 705f936e-4b9c-4cdc-84c2-2819e351c7c1 · outbound

This paper cites A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA.

Low-rank Momentum Factorization for Memory Efficient Training A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.150396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.150396Z digest=sha256:5de5758124925c0c5333a62a0f0edaf385a5815d19c7f993ab06f3e76ce489a0

Observation b05009ba-c43f-40bb-8ca5-f9d0c7d3d80f · outbound

This paper cites Scaling Laws for Neural Language Models.

Low-rank Momentum Factorization for Memory Efficient Training Scaling Laws for Neural Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.153241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.153241Z digest=sha256:b52edb1ab60ae3ea037d50a9b2c8e3a52931551d4eed98f2c9e4fe7958b2d718

Observation 70e53d3a-442f-4abb-aa35-4c9f769be0fb · outbound

This paper cites Accelerating Neural Network Training: An Analysis of the AlgoPerf Competition.

Low-rank Momentum Factorization for Memory Efficient Training Accelerating Neural Network Training: An Analysis of the AlgoPerf Competition

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:35:40.577672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.155578Z digest=sha256:3f93026b5f4cc7db3dbed70f1c6c3598a01482337f28c5f48f64a37fbc477a07

Observation 58235faa-8bb5-44c0-b417-5c000c6c6e7f · outbound

This paper cites Kingma and Jimmy Ba.

Low-rank Momentum Factorization for Memory Efficient Training Kingma and Jimmy Ba

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.735655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.157919Z digest=sha256:ad7c7bc27b92a5b4965e099552b0616d0936d3e1e247cd8ef44c4e5389a7921e

Observation 5c145465-698e-4d14-940f-d61f01f4069e · outbound

This paper cites VeRA: Vector-based Random Matrix Adaptation.

Low-rank Momentum Factorization for Memory Efficient Training VeRA: Vector-based Random Matrix Adaptation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.160096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.160096Z digest=sha256:d389f80db148a3a11cb9cbe3ad4503b023d00542fd08999e48b7e722caab72fe

Observation b72f3667-9544-45d7-a43d-53d61327f79e · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Low-rank Momentum Factorization for Memory Efficient Training Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.162460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.162460Z digest=sha256:f94ef02ea1423f4f66915fc5a2912f94585cf05ae7daec883d2445e04feab2a7

Observation 33ae95ce-59c6-4ccc-a4b7-14e26ed11296 · outbound

This paper cites Scalable optimization in the modular norm.

Low-rank Momentum Factorization for Memory Efficient Training Scalable optimization in the modular norm

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.728448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.165018Z digest=sha256:5b901b25bb9dcc1a0d8f63d8b91d9e5974f0a699a24f0ee61a5eb83fc608df81

Observation 48215550-78b8-43a5-ba07-9e7dadf5e798 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Low-rank Momentum Factorization for Memory Efficient Training The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.167207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.167207Z digest=sha256:58c1839de270865e559f3b30b7cffd5ed7afaf517663670ef6a71276c83f9705

Observation 7b5e13d4-b336-4951-b0b3-e6aac2eafcf5 · outbound

This paper cites Memory efficient optimizers with 4-bit states.

Low-rank Momentum Factorization for Memory Efficient Training Memory efficient optimizers with 4-bit states

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.721169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.169646Z digest=sha256:12841c75ebd8c6944e051a951505fa11a446f8580cdb478c45c6e022a85851b9

Observation cd7943ca-2435-4c8e-9273-49dbd1343839 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Low-rank Momentum Factorization for Memory Efficient Training Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.171944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.171944Z digest=sha256:dfeaf246727b39b632c491de65dc59062236651455ac205765db68c318634c77

Observation 805a5003-36a7-43bb-bef0-dc9b00e5268b · outbound

This paper cites Relora: High-rank training through low-rank updates.

Low-rank Momentum Factorization for Memory Efficient Training Relora: High-rank training through low-rank updates

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.713960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.174393Z digest=sha256:7ecc37bc8714c32544981a2603d0ce74e620ed03e7ffc34610b40417294cf86d

Observation 2c3af3d2-a44b-4111-bd5e-3e79e5bc9429 · outbound

This paper cites On the limited memory bfgs method for large scale optimization.

Low-rank Momentum Factorization for Memory Efficient Training On the limited memory bfgs method for large scale optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.176629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.176629Z digest=sha256:13e8fb3426f2c1bc2e181a01d1df2393fced4ea68a470aefc46772bb9be8682e

Observation 5bd23c7b-b0fa-4202-ba6c-cbb38d1c8ce9 · outbound

This paper cites Muon is Scalable for LLM Training.

Low-rank Momentum Factorization for Memory Efficient Training Muon is Scalable for LLM Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.178866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.178866Z digest=sha256:b1d6da4112ed44f5428684808833bdc9385c22f1e9c02357188a8470fa025a43

Observation 8e34f1bb-6d34-4aaa-b435-50c22f551eb7 · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

Low-rank Momentum Factorization for Memory Efficient Training DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.181226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.181226Z digest=sha256:0cc6e6fc085d51b33f856532769bbbd58be6347407d56c58510b4839e46836ad

Observation bdd016c8-8623-473e-b4f0-d65feca6e5fd · outbound

This paper cites Decoupled Weight Decay Regularization.

Low-rank Momentum Factorization for Memory Efficient Training Decoupled Weight Decay Regularization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.183763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.183763Z digest=sha256:88937cfb43810f8136eddc2aff769eb3870ad8d6b7b4361620b9a354bf7e2cf9

Observation 41f8934f-c0ca-4531-9d2b-174c4e6cfcac · outbound

This paper cites BAdam: A Memory Efficient Full Parameter Optimization Method for Large Language Models.

Low-rank Momentum Factorization for Memory Efficient Training BAdam: A Memory Efficient Full Parameter Optimization Method for Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.186047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.186047Z digest=sha256:3f3b9c77a681c4344f9b7fdedb58ced1a1945e3ffa7b2bf39b9c57c6288225c4

Observation bbe63c38-2c25-47b1-a48b-e8cd765c3073 · outbound

This paper cites CAME: Confidence-guided Adaptive Memory Efficient Optimization.

Low-rank Momentum Factorization for Memory Efficient Training CAME: Confidence-guided Adaptive Memory Efficient Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.188746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.188746Z digest=sha256:bc8b620eb20949f1d0906f48ce98243953abb85523c130b85e3eae64c4971cc1

Observation 5758d07c-1abf-450f-9b2b-0ada3efa269a · outbound

This paper cites AdaLomo: Low-memory Optimization with Adaptive Learning Rate.

Low-rank Momentum Factorization for Memory Efficient Training AdaLomo: Low-memory Optimization with Adaptive Learning Rate

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.191264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.191264Z digest=sha256:05e8046f82bdb4ed5c910615fdbc3ce0d9e0202bb6d9e40ddc1f5f0ecacc55b9

Observation aa8b78af-79f6-4412-8c45-a7c3df309054 · outbound

This paper cites Full Parameter Fine-tuning for Large Language Models with Limited Resources.

Low-rank Momentum Factorization for Memory Efficient Training Full Parameter Fine-tuning for Large Language Models with Limited Resources

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.193728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.193728Z digest=sha256:5104c1c2c64d2adddad1ada3f5979dec023582ccb1b22067630bf2447f3d956c

Observation b33f2a32-f26b-4cab-bc92-e0b243e5deb4 · outbound

This paper cites SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training.

Low-rank Momentum Factorization for Memory Efficient Training SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.196266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.196266Z digest=sha256:8399b79ec3591be5a0f21d30a98169c6b4a56040ad5a96d3c254799626e36831

Observation 1ecbedd0-ba15-4420-b2ea-843301b4e023 · outbound

This paper cites New insights and perspectives on the natural gradient method.

Low-rank Momentum Factorization for Memory Efficient Training New insights and perspectives on the natural gradient method

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.199023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.199023Z digest=sha256:b342b955a08f59148b6f85c6dee7f4747567bbce1306394737372f09416cc3f6

Observation 535a1cff-93fd-4595-b5dc-57ac99e1d34f · outbound

This paper cites Optimizing neural networks with kronecker-factored approximate curvature.

Low-rank Momentum Factorization for Memory Efficient Training Optimizing neural networks with kronecker-factored approximate curvature

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.201548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.201548Z digest=sha256:91a9c0acc0b511b4310b31f2f6d0500fbd1f6d0d1a28212034b86eb36b430a57

Observation d96bc255-68ef-4318-8d27-7152607d0cfa · outbound

This paper cites MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence.

Low-rank Momentum Factorization for Memory Efficient Training MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:35:40.279717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.203626Z digest=sha256:087b28e5150657987d508e8969c0338e9e6e6f840b262eef357d42f7aa060968

Observation 3f9068b4-d42c-48fe-8539-1770cd4965d8 · outbound

This paper cites A New Perspective on Shampoo's Preconditioner.

Low-rank Momentum Factorization for Memory Efficient Training A New Perspective on Shampoo's Preconditioner

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.206232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.206232Z digest=sha256:8ae4eefa47ffbb30aeaef5ab531e0d53af136839a04bc65e0d663f84028ed2d5

Observation 1efd01c2-0826-46a9-bcfa-d986b1cb4467 · outbound

This paper cites Training language models to follow instructions with human feedback.

Low-rank Momentum Factorization for Memory Efficient Training Training language models to follow instructions with human feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.208989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.208989Z digest=sha256:ee01d3cc0531e8ca558b87ea20880b05b5829eb3268cda42d9388973e4a73ede

Observation 6f316196-de61-474c-baa8-f2798e19a7d8 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.

Low-rank Momentum Factorization for Memory Efficient Training The fineweb datasets: Decanting the web for the finest text data at scale

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.688519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.211977Z digest=sha256:b3f909be9f7c1e82c358f71237f07510653bc79fc29428caaafe6bed00929148

Observation af92c5c3-9423-4e42-906b-a077d1a73022 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Low-rank Momentum Factorization for Memory Efficient Training Zero: Memory optimizations toward training trillion parameter models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.214420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.214420Z digest=sha256:cec36e4c9e03e9218f92285804566275ce4512c15cc2e66c99151e55b1dda55c

Observation 9adae0d8-c7be-4909-8e96-53e62d9dc309 · outbound

This paper cites Adarankgrad: Adaptive gradient-rank and moments for memory-efficient llms training and fine-tuning.

Low-rank Momentum Factorization for Memory Efficient Training Adarankgrad: Adaptive gradient-rank and moments for memory-efficient llms training and fine-tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.216796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.216796Z digest=sha256:b2ed3aac66ded42d5b15b21a5b7a01bb44057982e0730c6f3a8b482ebc474a1a

Observation 7e249672-1596-48c3-943e-0e72ef748e99 · outbound

This paper cites Ldadam: Adaptive optimization from low-dimensional gradient statistics.

Low-rank Momentum Factorization for Memory Efficient Training Ldadam: Adaptive optimization from low-dimensional gradient statistics

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.676755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.219217Z digest=sha256:0a0af98fd9e0f165cc8f382a0df8818f45eb9f2026e07831effe95dbb3b8ba90

Observation b23d3601-7f59-4904-8d2e-0039c1fa3197 · outbound

This paper cites Gradient Multi-Normalization for Stateless and Scalable LLM Training.

Low-rank Momentum Factorization for Memory Efficient Training Gradient Multi-Normalization for Stateless and Scalable LLM Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.221296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.221296Z digest=sha256:4f6e80109f97adadbf0a8def7ee9312fc80e71b9466dcbc1b896cf2bc38b08f3

Observation 34e4caf4-1165-4a4f-9045-929d96a595e8 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

Low-rank Momentum Factorization for Memory Efficient Training Adafactor: Adaptive learning rates with sublinear memory cost

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.223758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.223758Z digest=sha256:de36d53b66a915d46ba3a72a39f8fcf488dba6f4cf9ca3925d14bcd804c53d84

Observation 841a71e7-81fb-4fbb-85f4-c201c60929ff · outbound

This paper cites Tieleman and G.

Low-rank Momentum Factorization for Memory Efficient Training Tieleman and G

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.664744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.225792Z digest=sha256:8b1d327c0d219989f10a4ee0f4a9af35b88b3d1709f5fbf2dcdec1e5418b3b8b

Observation 110f4a2b-637f-488c-ae65-3bfc5ac28b6f · outbound

This paper cites Practical low-rank communication compression in decentralized deep learning.

Low-rank Momentum Factorization for Memory Efficient Training Practical low-rank communication compression in decentralized deep learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.657442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.227937Z digest=sha256:c7f38a655003e0702aa90dc314b8d30c9c89251989b9bc6bd70c13d1d6d09513

Observation c37c9c3d-7ce6-4a10-95b5-e88c51857c64 · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Low-rank Momentum Factorization for Memory Efficient Training SOAP: Improving and Stabilizing Shampoo using Adam

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.230139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.230139Z digest=sha256:ed760497c9281ccbc3dcda8f7e34c33c0c6cedfc8fd6d8e1bfbcfc8bf06c636d

Observation 47ef6f09-ff35-46fa-80a5-2adbad33c61f · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Low-rank Momentum Factorization for Memory Efficient Training GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.232517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.232517Z digest=sha256:43c1178ef0be3746124e5d7fe4e96469c0fbdacff52a5897f38f46bb7e66dff6

Observation 70c1579f-dfb6-489b-ae7b-dac080c96305 · outbound

This paper cites How far can camels go? exploring the state of instruction tuning on open resources.

Low-rank Momentum Factorization for Memory Efficient Training How far can camels go? exploring the state of instruction tuning on open resources

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.234934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.234934Z digest=sha256:e3aa859119860bbd65438958c75e384080f0c280f546acfb935bbe55657e762f

Observation 0d750cd9-6d15-4d64-8b9f-3f5302b61e32 · outbound

This paper cites A Spectral Condition for Feature Learning.

Low-rank Momentum Factorization for Memory Efficient Training A Spectral Condition for Feature Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.237394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.237394Z digest=sha256:29de226c0ae07abfb9a6086d7f98c568d0aa09f1b70776ec0c8557239be3798e

Observation 44fa1c87-6636-4c72-b2e2-b83704a4c2f4 · outbound

This paper cites Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models.

Low-rank Momentum Factorization for Memory Efficient Training Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.239930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.239930Z digest=sha256:42db7436b075f549f683bfac57611be6e3efd4cfe3c3efe9501accf5336e9269

Observation f90f451e-8e67-4718-ab89-7507bd97bf7f · outbound

This paper cites AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning.

Low-rank Momentum Factorization for Memory Efficient Training AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.242074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.242074Z digest=sha256:3f8815f4fceae429aa31f6b794013a4950edcfec1fa51969ae9e2447566c7f4a

Observation ce95a6af-a70d-440f-91d7-1eabe494034a · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Low-rank Momentum Factorization for Memory Efficient Training GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.244756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.244756Z digest=sha256:7cd36007226d13a5eb9531348820781ebeca1bcccd4a1b1ad716c41c025186a7

Observation 0b28fa6a-0895-48f6-8cb4-0bdc270963a8 · outbound

This paper cites Adapprox: Adaptive Approximation in Adam Optimization via Randomized Low-Rank Matrices.

Low-rank Momentum Factorization for Memory Efficient Training Adapprox: Adaptive Approximation in Adam Optimization via Randomized Low-Rank Matrices

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:35:40.307682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.247371Z digest=sha256:e1a167597decf4075953e75968281e397e8b366b35afbe89e1505ffeeed90d75

Observation 008ee817-5a97-436a-97d1-a07d2e8019f7 · outbound

This paper cites APOLLO: SGD-like Memory, AdamW-level Performance.

Low-rank Momentum Factorization for Memory Efficient Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.249822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.249822Z digest=sha256:fda6e6aaa9d614dea80b44223c2936d61bc6b35f7861a87f60d4559e1bbbee11

Observation 9b521a52-776b-4ec0-9b32-ea53ebece1f2 · outbound

This paper cites write newline.

Low-rank Momentum Factorization for Memory Efficient Training write newline

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.252196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.252196Z digest=sha256:697dd0ea3c23dc7eacfecbaa3d208f8826068d886dd18cbeb1f211abb41bf92d

Pith citing papers

No inbound Pith citation observations are available.