Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T13:23:55.233840Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 6 inbound Pith citation observations for arXiv:2506.16659.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-22T13:23:55.233840Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T18:42:01.854481Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-06-30T18:45:00.461346Z
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8b6c1845-e2e0-4cb8-ae9e-26953546d9e7 · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Old Optimizer, New Norm: An Anthology
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9bf11316-10ee-492a-8df4-02127bc376cc · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Fira: Can we achieve full-rank training of llms under low-rank constraint?arXiv preprint arXiv:2410.01623, 2024a
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c96ffdcd-e6c1-45c5-b1ce-5b8491a3b64b · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b342fcff-4960-49e5-8f96-e760f28afdb2 · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design GradientStabilizer:Fix the Norm, Not the Gradient
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c4f6eaf-247d-4b6d-9e4c-2e8188b0ff02 · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design When Can You Get Away with Low Memory Adam?
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21c944f3-4c7b-4de5-9933-1a28d8c163d4 · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design The Power of Normalization: Faster Evasion of Saddle Points
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aaaa3d9f-c7df-4133-ba39-87287c04398e · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ac2c9a3-d5ce-4a81-bdec-76d82c467ae7 · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Muon is Scalable for LLM Training
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c7ad440-ac60-41d4-af55-8523aebdfe02 · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design BAdam: A Memory Efficient Full Parameter Optimization Method for Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff737671-9801-4638-be91-b3d68b07fc22 · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 810ec3b8-47cf-4a7f-b507-e13f8881f7e8 · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Muhamed, O
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66a0262e-6d5e-4225-8fdb-2b45289408b4 · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Training Deep Learning Models with Norm-Constrained LMOs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d590ed14-3955-4500-abfa-d640c286bd69 · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design BlockLLM: Memory-Efficient Adaptation of LLMs by Selecting and Optimizing the Right Coordinate Blocks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2058620b-bee9-4c8c-9e54-e6f4bade211c · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Gradient normalization provably benefits nonconvex sgd under heavy-tailed noise
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c732579d-e615-4f62-8939-4f4a3b6dc7df · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94586757-8b76-4e00-8ecc-be0d4d039b79 · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Adam-mini: Use Fewer Learning Rates To Gain More
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1570a9a9-4d81-42fb-af3c-658e9de81ced · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design (Cited on pages 3, 11, and 12.) A Appendix A.1 Details of memory estimation for 1B and 7B models Here we compute the memory estimate for both 1B and 7B LLaMA models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5e9bb58-295f-45be-899d-728c310a9a61 · outbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design 7B model: Pre-last layers include 6.607B parameters and last layer includes 0.131B parameters, which in total leads to 6.738B parameters
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 607922fb-7c6c-4628-9847-25677acbfa58 · inbound
Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d6315e0f-af1f-4b82-a7d6-dbe4598f41f3 · inbound
MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eeca4c05-e80d-49d1-9015-de8493d9bb24 · inbound
Demystifying Manifold Constraints in LLM Pre-training Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d5d3a02-59a2-4547-ba84-af9e45b65520 · inbound
Budget-aware Auto Optimizer Configurator Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2d0252ca-2c12-4380-a92d-b6a9c13c8373 · inbound
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 698ffaf2-a6c9-4a4f-a3ef-35ff3efd7ed4 · inbound
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.