Pith. sign in

Paper Citation Record · LEDGER

GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2504.20437.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.20437 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:22:33.116839Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T22:25:39.562757Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6271ed04-48b0-4f08-9354-d7415b1682a5 · inbound

Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning cites this paper.

Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T05:22:33.116839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:22:33.116839Z digest=sha256:39102bb6ac993efeab3fb887f6857c880eb9853d88837794301149192beefa03

Observation e1196af2-e476-43c0-b6e7-eea51c406164 · inbound

Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking cites this paper.

Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:06:26.598216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:06:26.598216Z digest=sha256:3328d69474252dde4101378338f0851219f3c338d8d445949cc9f6d687e0606c

Observation 752e5699-365f-4d5a-b44f-e5fe9bffeff9 · inbound

Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters cites this paper.

Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.442424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T07:01:37.086159Z digest=sha256:a25a15719e7778e8b3340544569bb39edf5cb5d89c784a37e122973e1539a00a

Observation 603d1909-7060-44cb-bd28-0c799ebbad0e · inbound

Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training cites this paper.

Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:28:59.911059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T20:24:26.681383Z digest=sha256:32a4980d5703ddec0ee0650200b7e563ff7bbd11a0cad79e62e0cbb03a8500c1

Observation 5526992a-b395-49ef-b7ab-c1c204fee7e7 · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:38:10.965089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T09:34:45.186929Z digest=sha256:b1316b275dbc7ae97205570e29b42bc599a2ed1809c360787cecb99e81e35094

Observation ead297c7-32fa-4ca7-b7de-28cea3b8fb8d · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.322879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T18:42:01.854481Z digest=sha256:d1dfc1e54b66df55371fc9f9e364e4a3ba5a2af8300c75ebd391b3bb4191bfee

Observation 312471be-0631-45bb-83c4-7336807d2ebf · inbound

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training cites this paper.

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:25:39.564479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T22:16:44.668076Z digest=sha256:b5993009d9c8b78ccfb99941745541495cbb643336bab18fc7d95b66c8bdae10

Observation ffd69fd4-8e4f-401b-ab51-86250de8f7f2 · inbound

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training cites this paper.

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training GaLore 2: Large-Scale LLM Pre-Training by Gradient Low-Rank Projection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:27:40.399086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:27:40.399086Z digest=sha256:794509a2b9a6855be146810921bbaf821d70d5998bf350a3bfe8e0f4becf6957