Pith. sign in

Paper Citation Record · LEDGER

GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2504.20437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.20437 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:22:33.116839Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T22:25:39.562757Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6271ed04-48b0-4f08-9354-d7415b1682a5 · inbound

Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning cites this paper.

Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T05:22:33.116839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:22:33.116839Z digest=sha256:cd9cffcc5655d81e9e4ec5842b6fa1f53a777d81d5f31ab6e92f0b65b3ce537f

Observation e1196af2-e476-43c0-b6e7-eea51c406164 · inbound

Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking cites this paper.

Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:06:26.598216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:06:26.598216Z digest=sha256:5b3bf06f88f8597a5036ae56681302dea85cc3ec2ba426c6eea61a52390f67cb

Observation 752e5699-365f-4d5a-b44f-e5fe9bffeff9 · inbound

Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters cites this paper.

Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.442424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:01:37.086159Z digest=sha256:cbc2e0691e028ff24c23475a124b70245e146f1be5497c7f61e4bd603893c57d

Observation 603d1909-7060-44cb-bd28-0c799ebbad0e · inbound

Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training cites this paper.

Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:28:59.911059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T20:24:26.681383Z digest=sha256:29956c03476c2eaacdcc97fdef481a4be823a21750e9d82654646785b8b3a1ba

Observation 5526992a-b395-49ef-b7ab-c1c204fee7e7 · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:38:10.965089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T09:34:45.186929Z digest=sha256:41c8163ea98010596c06f0dfa472375ad50e43d0d063ba81f4610b32b59406f8

Observation ead297c7-32fa-4ca7-b7de-28cea3b8fb8d · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.322879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T18:42:01.854481Z digest=sha256:23ec491a0b3075c7260ced0fdf301d0189507854e77385f976a47c21528af09d

Observation 312471be-0631-45bb-83c4-7336807d2ebf · inbound

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training cites this paper.

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:25:39.564479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T22:16:44.668076Z digest=sha256:61f23a2ece99f5b5406579c5048a5ecd7350821510597bc8ce3ccd99592d7e99

Observation ffd69fd4-8e4f-401b-ab51-86250de8f7f2 · inbound

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training cites this paper.

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:27:40.399086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:27:40.399086Z digest=sha256:c6d557f1ae206d04ac23651c4b29bafcbd7ee3839d74316bf164b8ca8da64862