Pith. sign in

Paper Citation Record · LEDGER

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 6 inbound Pith citation observations for arXiv:2502.07222.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07222 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:31:42.126884Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:06:15.845262Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T14:52:41.305916Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c43fe9cd-617f-4505-ac56-f4d5740570ed · outbound

This paper cites Language Models are Few-Shot Learners.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Language Models are Few-Shot Learners

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.007786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.007786Z digest=sha256:9edd046735ef5f9b2dff58fa00aefa7f07846f6de3701db3fcb2bc54c037adc1

Observation e0be1fcc-0bdd-44f7-b901-b882f8008b00 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.038311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.038311Z digest=sha256:99fc1ec6c3a19720d2f211885fc2172fa3d3157e16e6889ea7b3159bd631c275

Observation 1e1eff25-4cd1-4051-8c2d-52f462d793c4 · outbound

This paper cites Chain of LoRA: Efficient Fine-tuning of Language Models via Residual Learning.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Chain of LoRA: Efficient Fine-tuning of Language Models via Residual Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.043331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.043331Z digest=sha256:82351f05605dc87c098f32165fbdd53adcbbf508ecd86e7e93f4b575f6a142ba

Observation 723f37ca-d918-4659-b148-8b7389db98c5 · outbound

This paper cites Subspace Optimization for Large Language Models with Convergence Guarantees.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Subspace Optimization for Large Language Models with Convergence Guarantees

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.048006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.048006Z digest=sha256:33a7c68473104b2c1382716cb02ca11d41ce95ceb3bb3faf91abc612ee821eb6

Observation c1dd313c-bc00-4608-967b-bcdeee29589f · outbound

This paper cites Flora: Low-Rank Adapters Are Secretly Gradient Compressors.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Flora: Low-Rank Adapters Are Secretly Gradient Compressors

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.052527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.052527Z digest=sha256:d89c0e267730a1e060c61a3d53e007c8e2e11d0cb32e897f1c45c4f7063c2f4b

Observation 3416b57f-a747-44e1-aa4f-53cd5a854338 · outbound

This paper cites Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.057316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.057316Z digest=sha256:11e4a17f7b51a6c68deab0ae64454bcc2f4c93f216fa75b8462185ea224b41c4

Observation 9ba0fd91-1273-4f46-8c25-048b37646563 · outbound

This paper cites LoRA Learns Less and Forgets Less.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models LoRA Learns Less and Forgets Less

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.061884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.061884Z digest=sha256:c0ed166349c7cde90a41300da26fb0a39c51bb7bcdef7b9e0d2f6545b57a6a5e

Observation 962db0d4-9782-4af4-8935-6fcb5d559388 · outbound

This paper cites Memory-Efficient LLM Training with Online Subspace Descent.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Memory-Efficient LLM Training with Online Subspace Descent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.066484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.066484Z digest=sha256:c4cc36fc64d2e4a53891cd317549e58ea4287271fc9eaf6c6fe4326ac848999b

Observation 7257a49f-7d7f-4d3f-bb85-442a9660787d · outbound

This paper cites Fira: Can we achieve full-rank training of llms under low-rank constraint? arXiv preprint arXiv:2410.01623, 2024b.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Fira: Can we achieve full-rank training of llms under low-rank constraint? arXiv preprint arXiv:2410.01623, 2024b

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.070965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.070965Z digest=sha256:052c5c2d7c4b490ee62ea79f33380765bf2efafba9ffe782dd442967db02cf9d

Observation a7221865-1152-4da2-a76c-69f946720f23 · outbound

This paper cites GWT: Scalable Optimizer State Compression for Large Language Model Training.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models GWT: Scalable Optimizer State Compression for Large Language Model Training

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-08T13:31:42.282331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:31:42.080116Z digest=sha256:79884dcbf3a8e502fed781a02607bc1b5ecff3d5bac097fd828722e778be0908

Observation 5512ac83-5cbc-40c7-8f4b-cfb6d32aa5ec · outbound

This paper cites Second-Order Fine-Tuning without Pain for LLMs:A Hessian Informed Zeroth-Order Optimizer.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Second-Order Fine-Tuning without Pain for LLMs:A Hessian Informed Zeroth-Order Optimizer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.089462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.089462Z digest=sha256:55262031e7d74e40f06c14b536e30d4773d6861dec780139405572c8270770c2

Observation 51eb4ee5-6276-4bf5-93c6-39a0e6e05ef5 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Training Deep Nets with Sublinear Memory Cost

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.094129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.094129Z digest=sha256:c7ba8e35c915308713d6d63759220e61278c55b6879089c5adb8f6cf84af09b0

Observation 8999f02c-d660-4eb4-96f3-afa40bc85f4f · outbound

This paper cites {Zero-offload}: Democratizing {billion-scale} model training.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models {Zero-offload}: Democratizing {billion-scale} model training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:31:42.702554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:31:42.103753Z digest=sha256:9b6c4ce2978a46634700991b1e46dba8846f3126fc2f472d63b63759e7f8b5f5

Observation b2700321-27a8-44be-bb10-b64b6f316864 · outbound

This paper cites Unified Convergence Analysis for Adaptive Optimization with Moving Average Estimator.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Unified Convergence Analysis for Adaptive Optimization with Moving Average Estimator

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-08T13:31:42.202894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:31:42.112866Z digest=sha256:c7ea8ca6a74c8fc363d079304fb76c6743ea20c51627d505c6452e4f77a0f144

Observation 0e2b3cf8-96b4-49c7-9de4-517f0a01e3aa · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.117526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.117526Z digest=sha256:f03556fc45122a50e52c3800a8285370712bc596de2fcdd58b6e2e2288ef4180

Observation dc7a6963-77c3-4ea5-9329-d7b9aa275ed5 · outbound

This paper cites We also report the memory overhead and total training time for each method.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models We also report the memory overhead and total training time for each method

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:31:42.672786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:31:42.126884Z digest=sha256:c002a3622a0e328bdfd99d0cb26b8ee31cadb3aa6325cdd76dc407a6e1033bb2

Observation 7d117e1a-e632-4350-889c-e33cbc2f503e · outbound

This paper cites Decoupled Weight Decay Regularization.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Decoupled Weight Decay Regularization

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.028545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.028545Z digest=sha256:8c1cb207b7c5acb0b1f98e0b984d626c4295018452e65a61b82f3b8d26d8eca5

Observation 7dbf2f18-468c-4ae3-9814-34798345d6b9 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models QLoRA: Efficient Finetuning of Quantized LLMs

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.098753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.098753Z digest=sha256:072471e3f5ca3a92abf6715982e353fa9f91237dc8957181cd62f17abf4f3dd3

Observation 81b546d6-cd09-41b1-b43e-8fe64df9ee27 · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.033460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.033460Z digest=sha256:4602b30f3f175c73f6ea1dd832f1d029f4f0969a246a50c60097239b3fd1fb70

Observation a964815e-6c11-4934-8371-7c97acad7c5c · outbound

This paper cites Natural GaLore: Accelerating GaLore for memory-efficient LLM Training and Fine-tuning.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Natural GaLore: Accelerating GaLore for memory-efficient LLM Training and Fine-tuning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.075566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.075566Z digest=sha256:622068f6086f68778749a5b9820e287347af86e534e35352092920e04c5cbfa8

Observation 3347d7e8-4a3d-48c4-92ac-90dd0a3a50dd · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.122265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.122265Z digest=sha256:c4399aea4c73d3d857bba01286435e3eebe08f80345689a9694e1b716d1190b3

Observation d0cd6cf5-8f8b-4b11-99ce-5024066cf95e · outbound

This paper cites GPT-4 Technical Report.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models GPT-4 Technical Report

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.013680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.013680Z digest=sha256:cd95ee90f49083560a8ffa4d48ba6072e6f92c178e11610499fcbf8a278ccee6

Observation cfea90d6-d746-4ff9-877f-9862585cee59 · outbound

This paper cites G10: Enabling an efficient unified gpu memory and storage architecture with smart tensor migrations.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models G10: Enabling an efficient unified gpu memory and storage architecture with smart tensor migrations

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:31:42.688081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:31:42.108417Z digest=sha256:94cdeb36f7dac58336f4a29ee850783736177d75f86196f3153f09c220b49bdf

Observation fe92f3d9-0eb2-4e29-bbda-735c56a10869 · outbound

This paper cites The Llama 3 Herd of Models.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models The Llama 3 Herd of Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.018811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.018811Z digest=sha256:5f9ecae8eb18332aeebecab7ce9ddb1a5e84cc5d95b30185c82cbaae614aa3c1

Observation c68f13ae-ebba-426c-98f6-6fdfa3823743 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Adam: A Method for Stochastic Optimization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.023758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.023758Z digest=sha256:50c605c33f42df1e3f46a6f44eec395760a889dce55e445dde604029105f4da4

Observation 762db2e0-86aa-46b6-9993-bec36b85da3a · outbound

This paper cites Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models.

A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T13:31:42.084774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:31:42.084774Z digest=sha256:1efc6ea9750d2d9552ceaa90a9306c21ad08ddcb77e821ef88a97c7aa7c89070

Pith citing papers

Observation 975a7038-cd34-481d-aa55-a59dd83bf1a5 · inbound

FZOO: Fast Zeroth-Order Optimizer for Fine-Tuning Large Language Models towards Adam-Scale Speed cites this paper.

FZOO: Fast Zeroth-Order Optimizer for Fine-Tuning Large Language Models towards Adam-Scale Speed A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:06:15.845262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:06:15.845262Z digest=sha256:1fdabe0b84e550ac0691405e7fb8b5718a1433e8084e56a4bb72476e96e92c99

Observation d48eac71-acc1-4a57-a145-b1334f1674a9 · inbound

CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure cites this paper.

CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:52:41.310077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:51:30.312509Z digest=sha256:af95075a675f6d7bcc6d3c0a21888b5198af6d270b120965bcf091f27e0e8c53

Observation b9b5e486-9ed5-43bd-9601-4a81e8f559c6 · inbound

BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models cites this paper.

BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:21:21.624728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:19:02.358348Z digest=sha256:5079509485849268dc8072207e23a3f7eee2e52b76a795c9b45dee2d123e1f1e

Observation fe18d896-397a-46a8-b4d8-f4cd285c224a · inbound

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments cites this paper.

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:07.546357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T19:50:50.653184Z digest=sha256:df8f34fe6180fb55bfcef153eae4f00d11a028f8edfda32f5024489b0041238e

Observation ec5d3761-5ce8-44bf-84cd-c91704fedc2a · inbound

BROS: Bias-Corrected Randomized Subspaces for Memory-Efficient Single-Loop Bilevel Optimization cites this paper.

BROS: Bias-Corrected Randomized Subspaces for Memory-Efficient Single-Loop Bilevel Optimization A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:56:26.672855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:47:04.868735Z digest=sha256:b387ff7d0074c5f4d009ce4d4071c8b0335ab476f230ae234986944a7e705a5d

Observation 15be2455-b720-4779-9d82-bceece556a22 · inbound

BROS: Bias-Corrected Randomized Subspaces for Memory-Efficient Single-Loop Bilevel Optimization cites this paper.

BROS: Bias-Corrected Randomized Subspaces for Memory-Efficient Single-Loop Bilevel Optimization A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:22:23.334369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:20:42.781613Z digest=sha256:6f9e4d03d748322ddf407d155e6bdf7e2a4b69850908645f4429649c0cf7caa9