Pith. sign in

Paper Citation Record · LEDGER

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension

As of 16 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 3 inbound Pith citation observations for arXiv:2502.07752.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07752 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:46:47.840408Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:27:40.373666Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T22:25:39.545387Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact4
  • verified fuzzy10
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5bfff0f7-79e4-4529-936b-889f94da49b4 · outbound

This paper cites Scalable Second Order Optimization for Deep Learning.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Scalable Second Order Optimization for Deep Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.441001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.441001Z digest=sha256:a8a11b0461493809e63104574690c1b14e505f771965b49290337adffbfe899c

Observation 7c25bd3b-2481-4d21-8029-2ef4fa51b299 · outbound

This paper cites M1 0 0 0 M2 0 0 0 M3 # Example of Diagv(·) This will stack the vector element into a pure diagonal matrix: Diagv([a11, a22, a33]T ) =.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension M1 0 0 0 M2 0 0 0 M3 # Example of Diagv(·) This will stack the vector element into a pure diagonal matrix: Diagv([a11, a22, a33]T ) =

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:46:48.475556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.802256Z digest=sha256:586313e0f6884a5d40c3f64848054b29d620c059bfb8c042d4a8183024b0797a

Observation d0ed36c7-5554-4911-9910-5b40f9a75dab · outbound

This paper cites From the Theorem D.1, the iterative procedure for Q can be simply obtained by taking the diagonals of M: Q = Diag E GSGT ∥S∥2 F.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension From the Theorem D.1, the iterative procedure for Q can be simply obtained by taking the diagonals of M: Q = Diag E GSGT ∥S∥2 F

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:46:48.401760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.822095Z digest=sha256:f1700e1094a076af17f08a65589a7491b060d741ac67e9139561b5b06542dd64

Observation 41d59989-3507-4b07-a1f4-167591649010 · outbound

This paper cites Effects of last layer One crucial setup difference during evaluation for low-rank methods is whether the last layer is trained by full-rank Adam or not.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Effects of last layer One crucial setup difference during evaluation for low-rank methods is whether the last layer is trained by full-rank Adam or not

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:46:48.351953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.840408Z digest=sha256:f67a2108f3e3689ee8e409134f360331631327c4d2a2d95479f32442d6d812f9

Observation 6d6d2a75-85dc-41f1-b244-cfc84dcdd9f9 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension LoRA: Low-Rank Adaptation of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.457800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.457800Z digest=sha256:496cdf26362bb2a352b40b3b7463695e35aa04003fd3920130c1bfcc674c290d

Observation 3f0ec22f-5055-46d2-a09d-78f7a9910efd · outbound

This paper cites FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.479823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.479823Z digest=sha256:a122eb00b390b5cf1d81d40c1f9c56e5ab25edc90cf4f8dffac3f3d8ed35c676

Observation 27822886-f0b7-4bb6-b64a-620ff692c6cc · outbound

This paper cites For 8-bit optimizers, we assume weights are stored in BF16, but optimizer states use FP8.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension For 8-bit optimizers, we assume weights are stored in BF16, but optimizer states use FP8

Reference 9

Resolution
verified exact
raw_fallback, observed 2026-08-08T11:46:47.934780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.836779Z digest=sha256:f2254afe275f2b10aaea840819a41abe897acfeaf36f4a294daf1162ecd6bbea

Observation 1847dc01-59e5-4b35-9f69-8e0a19be6459 · outbound

This paper cites an unresolved cited work.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T11:46:48.364266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.833300Z digest=sha256:4a8c8b6adf5e3b83409d9c0463e7d40382ae10c57fdb3bf0c14366465d45dd23

Observation af1dfbbe-8afb-4c20-93c2-cca4ee01e73c · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.747483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.747483Z digest=sha256:14e2b237feb726e6c940ac81dabf892ba0674052f2c52f20790c314a9dcfa0b7

Observation 34d89f11-b35a-49a3-a208-31cb801b1f64 · outbound

This paper cites A New Perspective on Shampoo's Preconditioner.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension A New Perspective on Shampoo's Preconditioner

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.755604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.755604Z digest=sha256:3d50dd631cbbd1dec138797935cee109c3b5574515dfb3cd2b058fbcf827ef86

Observation 07dcae0d-99ef-421d-9a27-abaf2e1bb770 · outbound

This paper cites Curvature-Informed SGD via General Purpose Lie-Group Preconditioners.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Curvature-Informed SGD via General Purpose Lie-Group Preconditioners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.759576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.759576Z digest=sha256:550e597aa89fdaa073ea687ca3886b407e1720e90b7ce356c4c39cb855d46a4c

Observation 1ee16ce6-d40b-4be5-b189-fa3c6654258f · outbound

This paper cites Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuning.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.763401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.763401Z digest=sha256:c7025bcf0ee6baa0e8761c421cdd514ac94d930147fbf800b8672368e03107c6

Observation 87834fae-442a-46ee-884f-0294812c3003 · outbound

This paper cites Connection of diagonal hessian estimates to natural gradients in stochastic optimization.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Connection of diagonal hessian estimates to natural gradients in stochastic optimization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:46:48.486121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.767173Z digest=sha256:2757fc0d5fe8a7f8d1f8240b60e449a55fa1a4237abd81ee933a720de1a793bb

Observation b3c55a12-7f0f-436f-bc2d-a67a591c09dd · outbound

This paper cites Understanding Self-supervised Learning with Dual Deep Networks.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Understanding Self-supervised Learning with Dual Deep Networks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.771026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.771026Z digest=sha256:3de5988a0ea91b647e8b58327760076f578d4c980a4e1ece2c9e00d0870e8981

Observation 6bd0fe42-d56c-4794-83c5-42f239b0fc19 · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension SOAP: Improving and Stabilizing Shampoo using Adam

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.778814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.778814Z digest=sha256:9a724ffae7acead81ee2bdb1493eb4d23a812ab099f04ac8a73033213ce586f1

Observation 938f875f-cad8-42ce-916b-ecc6ec5e8f30 · outbound

This paper cites No More Adam: Learning Rate Scaling at Initialization is All You Need.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.782800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.782800Z digest=sha256:986a36f5810a179a09f420ec00538c3c944c917970611109c96ca5e334a66375

Observation 3b7c6b6a-edde-44c1-ba10-87bc84748d90 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.790437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.790437Z digest=sha256:8fc045cbf1f641d92ef88ced1187ebfec34539e79263d7f5a9b54e2d6dbb1630

Observation 87ec5990-80e3-4c54-abd5-0c85ea4bf764 · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Adam-mini: Use Fewer Learning Rates To Gain More

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.795211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.795211Z digest=sha256:1b677bf47e8cc3d0246193dd5ff142b0c088fda461e74d453304881af3dd7a55

Observation 41446892-5841-44ff-806f-c9e190c9a2f9 · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.799002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.799002Z digest=sha256:d734591c5624bbd4ee4da0195156487f4075e68d2e190c816e85e9cde91c110e

Observation 681d9b23-720e-479d-9aa9-2a74eff61afc · outbound

This paper cites for t = 1,.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension for t = 1,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:46:48.464545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.805888Z digest=sha256:795a614a984081a4190460169051590d55c38e14c505adacc13ccc8aa6117f65

Observation 9b1a17ee-3f19-456a-a908-df93999b7bc1 · outbound

This paper cites Note that our paper assumes Vec(·) is stacking columns of matrix whereas Gupta et al.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Note that our paper assumes Vec(·) is stacking columns of matrix whereas Gupta et al

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:46:48.452711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.808909Z digest=sha256:6bcb0d0fd67772abc542042b203e54d456e1071564aa26249e7433fd0ff22a47

Observation 2da2c5cc-9156-4851-b16c-d659a268811b · outbound

This paper cites an unresolved cited work.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T11:46:48.439953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.811783Z digest=sha256:83cc93911765f5f52545e56761a1d712f727c845e802302ebe6cd234d53cc865

Observation 49c3e7fc-8115-4ce3-b3b4-7f35fb270d6c · outbound

This paper cites In practice, SW AN proposes to compute the(GGT )− 1 2 using Newton-Schulz iterations.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension In practice, SW AN proposes to compute the(GGT )− 1 2 using Newton-Schulz iterations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:46:48.427956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.814682Z digest=sha256:4b026da4b0b5b2f9ac889160343f89caa72d33f3bf0678af6bc20833c449169e

Observation 2e1d9700-dd5f-4773-8de0-7500648dca74 · outbound

This paper cites At the same time, GaLore [Zhao et al., 2024a] popularizes the use of low-rank optimizers, which demonstrates on-par performance compared to full-rank Adam training.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension At the same time, GaLore [Zhao et al., 2024a] popularizes the use of low-rank optimizers, which demonstrates on-par performance compared to full-rank Adam training

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:46:48.415008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.818309Z digest=sha256:cf5834e9bf2833fe986eb14aee809ec9d4792323ab6dc0b89fc7f31b0a1a6efa

Observation d3e11205-1153-40ae-b572-72d1e534efd7 · outbound

This paper cites Then, the normalization step of Lamb and Lars can be viewed as a 1-sample approximation to FIM* under the structure considered in Sec.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Then, the normalization step of Lamb and Lars can be viewed as a 1-sample approximation to FIM* under the structure considered in Sec

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:46:48.389263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.825916Z digest=sha256:e416c90bf4df7082b3577c9239997eace17d1fa90ec4704fe9490343fa145736

Observation 262c1a26-ccc7-45a0-80ed-caa8d5b905e6 · outbound

This paper cites U T U T c G ⊙2# =[U , Uc] U T G U T c G vuutE.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension U T U T c G ⊙2# =[U , Uc] U T G U T c G vuutE

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:46:48.376523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.829466Z digest=sha256:26633748099dd89c095018b13a6d98762a63c67f3c19109c640c7f176130c714

Observation 26634398-10a7-4e11-8a13-13b53a4aaa42 · outbound

This paper cites Large Batch Training of Convolutional Networks.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Large Batch Training of Convolutional Networks

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.786590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.786590Z digest=sha256:1f72ae5abcac83599a7b500004b66616f829c6b20670c2b195f5bf5feb1961a0

Observation 8d0ec684-065f-49f1-905a-3d8e8d99d594 · outbound

This paper cites Efficient Approximations of the Fisher Matrix in Neural Networks using Kronecker Product Singular Value Decomposition.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Efficient Approximations of the Fisher Matrix in Neural Networks using Kronecker Product Singular Value Decomposition

Reference 2014

Resolution
verified exact
local_arxiv, observed 2026-08-08T11:46:48.234529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.587382Z digest=sha256:4775867ebbe5e1e928040a3eda985fd27a1467d08c9c2e6c56d97de50073965a

Observation 860c6b52-a841-4eaf-9efe-f4947bd814be · outbound

This paper cites SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.751693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.751693Z digest=sha256:d50f3cb55bfb8f82ba9ab9e3228661858f873f3d9823c5a05241e5255ae760da

Observation 91c9ad27-332d-4c82-812c-716d28f52cc3 · outbound

This paper cites Preconditioner on Matrix Lie Group for SGD.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Preconditioner on Matrix Lie Group for SGD

Reference 2017

Resolution
verified exact
local_arxiv, observed 2026-08-08T11:46:48.218281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.629321Z digest=sha256:d4918637b745e122493594125f21e0f851bc8da4ce4bcb95b41b941ebb8bd4ed

Observation df762821-163a-4888-8ce3-1a2c866041c0 · outbound

This paper cites Fira: Can we achieve full-rank training of llms under low-rank constraint? arXiv preprint arXiv:2410.01623, 2024a.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Fira: Can we achieve full-rank training of llms under low-rank constraint? arXiv preprint arXiv:2410.01623, 2024a

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.446203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.446203Z digest=sha256:e24e17b890ec815da2d8f2401d4ab8da141116073d78ee7280a37fb3c8605373

Observation f55c5c32-f0e7-4f6e-9262-f79e92191c9b · outbound

This paper cites On Empirical Comparisons of Optimizers for Deep Learning.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension On Empirical Comparisons of Optimizers for Deep Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.449567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.449567Z digest=sha256:c5fa2d1807b15a37443de5bee412a0fbc4a5bc2c17f39ae02c52a84af30fd714

Observation 1cb85550-a248-4f7a-b11c-f9226b101e92 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension LLaMA: Open and Efficient Foundation Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.775109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.775109Z digest=sha256:eca545617e23668c4bb4108c36c146604f3e45ca6836a2219220424a5e2b4067

Observation d031e4fd-ea87-48a9-a5c7-62fcc633626b · outbound

This paper cites The Llama 3 Herd of Models.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension The Llama 3 Herd of Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.453851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.453851Z digest=sha256:a8616e63a8eccd69bcd9e1c16ed1a17064042eb05c849acf7583c43e55d8fc57

Observation 98d71768-a389-41a1-9b93-9794e8b2545b · outbound

This paper cites Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective

Reference 2023

Resolution
verified exact
local_arxiv, observed 2026-08-08T11:46:48.200943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T11:46:47.689130Z digest=sha256:d43d8c5b5b534b902c40088460afb7d9dc0ed37d6b27a119736a5a104aeba363

Observation dc51d578-f541-4b7b-8309-3d1ace46c9da · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension Adam: A Method for Stochastic Optimization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.554330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.554330Z digest=sha256:7333944e350d2160285a5c632655d76b4729787d273fdb32a9f21d1dfa7d9bee

Pith citing papers

Observation 91f4ecc9-1ff4-4058-9bf0-0ff00be0822d · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:0a07066014f2ccfc2b1eb0b1effea27c304cb4f0e17fa4b3eda8a4a859315b74

Observation 88a4455d-ebb9-4db0-a1a4-4857364a2afe · inbound

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training cites this paper.

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T22:25:39.546940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-08T22:16:44.668076Z digest=sha256:f0775101d40272cd6367c84ccf03605efef97198d9a3867e0e5650f22ea7838e

Observation cf227b65-8840-4ae3-a742-e515ac419a16 · inbound

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training cites this paper.

No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T08:27:40.373666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:27:40.373666Z digest=sha256:92b6d5981d78498013aa0ff3930781d4d3a9d6b10b38f596f3f3001afe274926