Pith. sign in

Paper Citation Record · LEDGER

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design

As of 7 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 6 inbound Pith citation observations for arXiv:2506.16659.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16659 v3

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T13:23:55.233840Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T18:42:01.854481Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-30T18:45:00.461346Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact12
  • verified fuzzy3
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b6c1845-e2e0-4cb8-ae9e-26953546d9e7 · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Old Optimizer, New Norm: An Anthology

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:24:53.256819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:f1247733f40b9dacd6fe7587e1f379f75c918f2a117febe88aaa5bf56dc862b6

Observation 9bf11316-10ee-492a-8df4-02127bc376cc · outbound

This paper cites Fira: Can we achieve full-rank training of llms under low-rank constraint?arXiv preprint arXiv:2410.01623, 2024a.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Fira: Can we achieve full-rank training of llms under low-rank constraint?arXiv preprint arXiv:2410.01623, 2024a

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:24:53.237262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:558e92a301f377de20eece042dfefb465cfce9585a017d13b2754b2c63535361

Observation c96ffdcd-e6c1-45c5-b1ce-5b8491a3b64b · outbound

This paper cites Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:24:53.193892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:10726155fe70fed6a352279887db2994b9af1f9538cb23dbfd0b55628d32825a

Observation b342fcff-4960-49e5-8f96-e760f28afdb2 · outbound

This paper cites GradientStabilizer:Fix the Norm, Not the Gradient.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design GradientStabilizer:Fix the Norm, Not the Gradient

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-28T02:03:51.577898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:b6ece44c26c6fa9941dbc43e4392d84f573716183f19b28a1f544c034fa5cb9f

Observation 2c4f6eaf-247d-4b6d-9e4c-2e8188b0ff02 · outbound

This paper cites When Can You Get Away with Low Memory Adam?.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design When Can You Get Away with Low Memory Adam?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:24:53.261901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:4d741f9116a88e0f1ace048a4c594db3035ed5f8effd1a413c8d680237bca2e4

Observation 21c944f3-4c7b-4de5-9933-1a28d8c163d4 · outbound

This paper cites The Power of Normalization: Faster Evasion of Saddle Points.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design The Power of Normalization: Faster Evasion of Saddle Points

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T13:24:53.221808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:f497b9a0182a4219ab3265f6606de2c86acb10b2ba368f8f975b393b6a22060e

Observation aaaa3d9f-c7df-4133-ba39-87287c04398e · outbound

This paper cites GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design GaLore$+$: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:24:53.221906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:49f4e06ba12f9086c11ed6696d152f45b5d0f05f1bc7e357c39e977ebfb0bfbd

Observation 8ac2c9a3-d5ce-4a81-bdec-76d82c467ae7 · outbound

This paper cites Muon is Scalable for LLM Training.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Muon is Scalable for LLM Training

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T13:24:53.207881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:abfd6fbb9ae71a444dfd9f125ea671a5de1c2d719530ac051927f16dac9b37d0

Observation 2c7ad440-ac60-41d4-af55-8523aebdfe02 · outbound

This paper cites BAdam: A Memory Efficient Full Parameter Optimization Method for Large Language Models.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design BAdam: A Memory Efficient Full Parameter Optimization Method for Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:24:53.227020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:2eb93577aaf0faf9b7462301211a4e7539bc8a351579cc12213f56858b04687a

Observation ff737671-9801-4638-be91-b3d68b07fc22 · outbound

This paper cites SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:24:53.210558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:a69389e53ecc36e832ce04c25621522b485d8d8fd5445e01509b35220c148178

Observation 810ec3b8-47cf-4a7f-b507-e13f8881f7e8 · outbound

This paper cites Muhamed, O.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Muhamed, O

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T13:24:53.488044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:765ce1289fc55de8735f9ba7eaf908f14a3bd0058c12c5a9e5ba49b9189b324b

Observation 66a0262e-6d5e-4225-8fdb-2b45289408b4 · outbound

This paper cites Training Deep Learning Models with Norm-Constrained LMOs.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Training Deep Learning Models with Norm-Constrained LMOs

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:24:53.266465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:c150d8f02e264b1f1da4c0caee7da49d6aab3349cfb8022c6cf5e5af60684723

Observation d590ed14-3955-4500-abfa-d640c286bd69 · outbound

This paper cites BlockLLM: Memory-Efficient Adaptation of LLMs by Selecting and Optimizing the Right Coordinate Blocks.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design BlockLLM: Memory-Efficient Adaptation of LLMs by Selecting and Optimizing the Right Coordinate Blocks

Reference 13

Resolution
malformed identifier
arxiv_id, observed 2026-05-22T13:24:53.242252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:2d53a10b6ae3c8380cec5caf94401454527febcd275dfbbdeeb59096177abb5c

Observation 2058620b-bee9-4c8c-9e54-e6f4bade211c · outbound

This paper cites Gradient normalization provably benefits nonconvex sgd under heavy-tailed noise.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Gradient normalization provably benefits nonconvex sgd under heavy-tailed noise

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:24:53.232134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:bdc7dbca7b34425859dddf01d36319582e62e163002d896ccf7546e75b5ce5f8

Observation c732579d-e615-4f62-8939-4f4a3b6dc7df · outbound

This paper cites No More Adam: Learning Rate Scaling at Initialization is All You Need.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:24:53.247110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:d259bcca0e0f5f64fcbd1a72b5520c091cca1bed6aa89edbd0fc9943690aa629

Observation 94586757-8b76-4e00-8ecc-be0d4d039b79 · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design Adam-mini: Use Fewer Learning Rates To Gain More

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:24:53.226332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:a07a18efc350c6ce56c88d7e31a7d69b69b4d09caac2e16be789a53b096f5e15

Observation 1570a9a9-4d81-42fb-af3c-658e9de81ced · outbound

This paper cites (Cited on pages 3, 11, and 12.) A Appendix A.1 Details of memory estimation for 1B and 7B models Here we compute the memory estimate for both 1B and 7B LLaMA models.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design (Cited on pages 3, 11, and 12.) A Appendix A.1 Details of memory estimation for 1B and 7B models Here we compute the memory estimate for both 1B and 7B LLaMA models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T13:24:53.480331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:1fbc52413e76db11430d37d52676b76f9e8aa1a8c3bdce6e314211cba598a026

Observation e5e9bb58-295f-45be-899d-728c310a9a61 · outbound

This paper cites 7B model: Pre-last layers include 6.607B parameters and last layer includes 0.131B parameters, which in total leads to 6.738B parameters.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design 7B model: Pre-last layers include 6.607B parameters and last layer includes 0.131B parameters, which in total leads to 6.738B parameters

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T13:24:53.484309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:927dd92e8967f05302a17af0587e8f6648b98b719e443c9a2b98fac2f3945744

Pith citing papers

Observation 607922fb-7c6c-4628-9847-25677acbfa58 · inbound

Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training cites this paper.

Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training Memory-Efficient LLM Pretraining via Minimalist Optimizer Design

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:26.475417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T15:48:43.422716Z digest=sha256:7c02f512917610c429c6437e48ef96b8f2fe0d91b50e79a151976c2b9a0c5170

Observation d6315e0f-af1f-4b82-a7d6-dbe4598f41f3 · inbound

MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration cites this paper.

MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration Memory-Efficient LLM Pretraining via Minimalist Optimizer Design

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T02:03:26.475417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:12:32.579136Z digest=sha256:ad5dab7b9f96d1118fecdf2db45e17df6b22179408e4baa9d3a88c1380265227

Observation eeca4c05-e80d-49d1-9015-de8493d9bb24 · inbound

Demystifying Manifold Constraints in LLM Pre-training cites this paper.

Demystifying Manifold Constraints in LLM Pre-training Memory-Efficient LLM Pretraining via Minimalist Optimizer Design

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T02:03:26.475417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:44:44.438637Z digest=sha256:cd2e0140a997b8c86e88117361c777c75e735d44316ebc551cceba40dff90a97

Observation 9d5d3a02-59a2-4547-ba84-af9e45b65520 · inbound

Budget-aware Auto Optimizer Configurator cites this paper.

Budget-aware Auto Optimizer Configurator Memory-Efficient LLM Pretraining via Minimalist Optimizer Design

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T02:03:26.475417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T17:46:09.673049Z digest=sha256:f235886d04e1b5304f4c9e4c63b18b8d82d13d42bd8b0b24ad3a140708d42df6

Observation 2d0252ca-2c12-4380-a92d-b6a9c13c8373 · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Memory-Efficient LLM Pretraining via Minimalist Optimizer Design

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:26.475417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T09:34:45.186929Z digest=sha256:2e9670b44e29953f642f1095ea63f5316100c9735eaf9956c6d44928bac3868e

Observation 698ffaf2-a6c9-4a4f-a3ef-35ff3efd7ed4 · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Memory-Efficient LLM Pretraining via Minimalist Optimizer Design

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.462706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:42:01.854481Z digest=sha256:447b55f20f48a18732d2a3a0a8e581258d9a767db2c8e1b99b6c57a4996d75b0