Pith. sign in

Paper Citation Record · LEDGER

APOLLO: SGD-like Memory, AdamW-level Performance

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2412.05270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05270 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:19:19.840427Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T22:25:39.550471Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4a38ff1a-5543-4da6-9d7f-a057b390944b · inbound

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training cites this paper.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.483763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.483763Z digest=sha256:7b626b5ef3d6f97d189a75a39467447bd5792a5cb6c8fd2e489f0f1035cf5cea

Observation 5f0b7288-dee8-4114-b32d-914e834b639b · inbound

CE-LoRA: Computation-Efficient LoRA Fine-Tuning for Language Models cites this paper.

CE-LoRA: Computation-Efficient LoRA Fine-Tuning for Language Models APOLLO: SGD-like Memory, AdamW-level Performance

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T15:36:04.609015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:36:04.609015Z digest=sha256:d67bb4473e26b00c9896bdd1c9c5cbaf6a307abff6c4f54f15d123858b88e7bd

Observation a6b8083a-4cd7-47bb-ad53-6e19cd7f3df8 · inbound

Gradient Multi-Normalization for Stateless and Scalable LLM Training cites this paper.

Gradient Multi-Normalization for Stateless and Scalable LLM Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.815685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.815685Z digest=sha256:8e13fb4484a7fb7814e10b3dc1d076115efb51b96842afbb55670abf43eea2c9

Observation 1c9a1073-bf9e-481b-a50b-8c2ce0ac64dc · inbound

Memory-Efficient LLM Training by Various-Grained Low-Rank Projection of Gradients cites this paper.

Memory-Efficient LLM Training by Various-Grained Low-Rank Projection of Gradients APOLLO: SGD-like Memory, AdamW-level Performance

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T04:19:19.840427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:19:19.840427Z digest=sha256:e940d269b17888a5fcfd9b52eba5977e51d49663ca472a3527c5016c259949a0

Observation b25dd6bd-35ff-4847-b463-a31567e7f612 · inbound

Gaussian Weight Sampling for Scalable, Efficient and Stable Pseudo-Quantization Training cites this paper.

Gaussian Weight Sampling for Scalable, Efficient and Stable Pseudo-Quantization Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:33.493241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:33.493241Z digest=sha256:081c50930fc8ff08af22e0a5e28f074f2184744835b09a2d2e6f1508804048db

Observation 655bdb2d-85e9-4505-b312-13a44980bbd8 · inbound

Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking cites this paper.

Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking APOLLO: SGD-like Memory, AdamW-level Performance

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:06:27.700501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:06:27.700501Z digest=sha256:062a77334384236f490809d0d501176db7926bb9f24c9ebcc2650f9e47ae295b

Observation 4ec620a7-60ca-46d6-8a67-437c76a0961b · inbound

Taming LLMs by Scaling Learning Rates with Gradient Grouping cites this paper.

Taming LLMs by Scaling Learning Rates with Gradient Grouping APOLLO: SGD-like Memory, AdamW-level Performance

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.271257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.271257Z digest=sha256:bb75167e3286240cc9b10281d59fe881a18f17b4be7582338f44dad6162a464a

Observation cc89ca25-c48d-4b56-a5f7-8ed01bc56d6d · inbound

E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models cites this paper.

E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models APOLLO: SGD-like Memory, AdamW-level Performance

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:07.981342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:07.981342Z digest=sha256:1a66a3501403e08e781351fb006be03232ad7bd03b4ec78abdf12cee03850d02

Observation 008ee817-5a97-436a-97d1-a07d2e8019f7 · inbound

Low-rank Momentum Factorization for Memory Efficient Training cites this paper.

Low-rank Momentum Factorization for Memory Efficient Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.249822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.249822Z digest=sha256:35c44f3d1546de8aef70ed45059f7069be23be59053e30d42aba378fb6392e2d

Observation 21ef9879-1b50-426c-ae10-2f909a2f3e20 · inbound

LOST: Low-rank and Sparse Pre-training for Large Language Models cites this paper.

LOST: Low-rank and Sparse Pre-training for Large Language Models APOLLO: SGD-like Memory, AdamW-level Performance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:44:20.002334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:44:20.002334Z digest=sha256:1ebe73db69827cc003debeb309a814b552e275966177d80729d3c74826997465

Observation e821d31e-544f-4a05-9313-2182433925af · inbound

CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure cites this paper.

CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure APOLLO: SGD-like Memory, AdamW-level Performance

Reference 66

Resolution
malformed identifier
arxiv_id, observed 2026-05-18T14:52:41.115843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T14:51:30.312509Z digest=sha256:9a4788a229901c496ec5e58e23b2b99ea8c3687b67cf6eab7cf6afaa6b4b9c29

Observation 17c8cba1-b950-42e7-87c1-9f2477ed13d0 · inbound

Geometrically Principled Randomized Optimization for Efficient LLM Training cites this paper.

Geometrically Principled Randomized Optimization for Efficient LLM Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T12:51:27.373729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:51:27.373729Z digest=sha256:1f48e228c77dec7ade59cfd6bf4e2ebb33a45d1529afb25bcd83e18112cf8df8

Observation 8f25b09b-a1c8-4319-a8dd-7557c3b9e99c · inbound

BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models cites this paper.

BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models APOLLO: SGD-like Memory, AdamW-level Performance

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T23:21:21.529366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T23:19:02.358348Z digest=sha256:079f33a43fb085b8fc773f7479d03be39ca5a0f0b3107c6cea36892819e2af65

Observation 0bd458b2-0d94-4159-9439-f18c11a1393f · inbound

SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization cites this paper.

SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization APOLLO: SGD-like Memory, AdamW-level Performance

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:15:54.526152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:37:16.046496Z digest=sha256:9580e83ecf2d84eab558c3cb756d83ff804a493721356d754b3376ab571b1c24

Observation 7ca6f309-152d-4b9e-aac2-f30151b65e20 · inbound

Disposition Distillation at Small Scale: A Three-Arc Negative Result cites this paper.

Disposition Distillation at Small Scale: A Three-Arc Negative Result APOLLO: SGD-like Memory, AdamW-level Performance

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:00.993265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T16:09:21.845152Z digest=sha256:f9dc171d1ca932eca9525551fc0c728f8a3597361200c16c68c57e5e730840dc

Observation ddb8d743-3783-47a7-adc1-b84d8c41579d · inbound

Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers cites this paper.

Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers APOLLO: SGD-like Memory, AdamW-level Performance

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:41:45.424865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T04:01:32.057022Z digest=sha256:e17e4b54d5b1e89c1dc3762dcc0d8ec2d3d6f5f631539c14ba7f99f992e48ce3

Observation fc6b2427-7f56-471d-a466-d846900c002f · inbound

Gefen: Optimized Stochastic Optimizer cites this paper.

Gefen: Optimized Stochastic Optimizer APOLLO: SGD-like Memory, AdamW-level Performance

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T18:00:57.430970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:00:57.430970Z digest=sha256:df162ccc52749a9f7366df3471af575d9c7871ee0eee3bc715f7442baafe5791

Observation a8699242-10a0-4e8a-8ba6-3ddd1afd4b28 · inbound

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training cites this paper.

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T22:25:39.551975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:16:44.668076Z digest=sha256:0a918328896b5f56f38b5c05e30b85a8cf6a5832a8924e410507593f1a6729e4

Observation d2e2757c-9340-43b9-8b47-67ebdd72bca5 · inbound

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training cites this paper.

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:27:40.404285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:27:40.404285Z digest=sha256:98771c72e7b46986f165a8ddc4dcacc119c7b8c36153ea0bfb23c60806f60d77