Pith. sign in

Paper Citation Record · LEDGER

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training

As of 19 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2412.19616.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19616 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:15:00.504663Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:25:51.616737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T19:01:10.343530Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3427dafa-22d5-4dbe-9646-a1008616c640 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.195051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.195051Z digest=sha256:f791379662cbed51135a2b9d46e6476dc91c86dd8e868025e8911e5809eb891e

Observation 87582f04-ea93-4143-8bfa-4bc6190cd8f3 · outbound

This paper cites write newline.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.200908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.200908Z digest=sha256:82c7a16cfee8d2051d5b90ffa3ee6dbdb1e5d8726f1555d2339b06cc209195a4

Observation bd526c06-7f09-47be-879a-792597b11c37 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.207607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.207607Z digest=sha256:3e7fcd441a1c65785572011a405b90fa6aa77516746436c3ebf68fa5f2ddd393

Observation 9b084c1d-a145-42f2-a154-bafc3944b22a · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.462533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.212482Z digest=sha256:476d2b9689689d414bd7ddca85319f815c14d62fb7fe7c72d24e48bf7bba60a8

Observation d42873f8-c90b-4aa7-b418-5626e636d748 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.445217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.218821Z digest=sha256:c5d3a63117ef87dd75dfe8275ce09445fa5452aaad46b273a214b1991c6045fc

Observation f070a56c-a574-4a79-ad70-a9b8569394d8 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Training Deep Nets with Sublinear Memory Cost

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.224035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.224035Z digest=sha256:5028372c246bc8e83a9497a450196bcbb7d11fdeed0683db6d4ab3927c1b648a

Observation 4baa2e8c-0c31-4e57-9e07-04f1f799a94e · outbound

This paper cites Fast low-rank estimation by projected gradient descent: General statistical and algorithmic guarantees.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Fast low-rank estimation by projected gradient descent: General statistical and algorithmic guarantees

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.229457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.229457Z digest=sha256:2fc51464f944b572dbca939cfdfd30cf3a446f6d53a99effd280ea09587e23f3

Observation 7640b6f2-e1ce-428e-9341-5263c77915d1 · outbound

This paper cites 8-bit Optimizers via Block-wise Quantization.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training 8-bit Optimizers via Block-wise Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.234867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.234867Z digest=sha256:9d0c6a745de978fff872f73492e3ca100f4d307f2a2e58de43357a18b54f1245

Observation 68826332-96d6-498d-b75d-67ead10f9895 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.241175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.241175Z digest=sha256:5fe45f4af448e9408ebfbd1af430bbfb94d71ceacd033dee27c1913db6e7869b

Observation 4bed34df-db91-4be1-b79a-6dd30eff3dc0 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.246576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.246576Z digest=sha256:b2bbc0bea55bdad9513c7505e5a9a4eb6a08dade99a23632ed8ea0edb22b2c39

Observation 4235ad2b-d503-4767-84dc-9a00623d53a7 · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.251637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.251637Z digest=sha256:73ac60af5679112216777eab14fd86a39f2e19e4cf8ff7ccdc0eaff4497f37a4

Observation 307401e0-2f3c-4d55-bed0-2f160a1713b2 · outbound

This paper cites K.; Roy, D.; and Carbin, M.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training K.; Roy, D.; and Carbin, M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:15:01.418427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.256937Z digest=sha256:f01f5266fbff0eee8b216c4ad123edf609f16dca55d1d1d5eafb1376815eddb9

Observation 24a5fade-4b96-4745-8942-32b530b0a522 · outbound

This paper cites N.; Ren, M.; Urtasun, R.; and Grosse, R.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training N.; Ren, M.; Urtasun, R.; and Grosse, R

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:15:01.401913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.262687Z digest=sha256:22ea833db35ec28c788b807410403447d71f9f32c1f58dcd73411e7a17dc729f

Observation e942dbec-3523-4edf-b613-6ed144056c1d · outbound

This paper cites N.; Ren, M.; Urtasun, R.; and Grosse, R.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training N.; Ren, M.; Urtasun, R.; and Grosse, R

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:15:01.386380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.267598Z digest=sha256:fbccecf4abc096e0000181298bb5beb48574ff41b13f4ede909ac03a5472b68c

Observation 6132dab5-739d-416e-b566-5c284c4e3042 · outbound

This paper cites Gradient Descent Happens in a Tiny Subspace.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Gradient Descent Happens in a Tiny Subspace

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.271903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.271903Z digest=sha256:a6ec131d738595fe144ce77a74c7bc01d086df7474a7eae572ee4f0f6c11b522

Observation 5205b8f8-5a72-4753-a7aa-b4927a4cb012 · outbound

This paper cites WARP: Word-level Adversarial ReProgramming.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training WARP: Word-level Adversarial ReProgramming

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.276399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.276399Z digest=sha256:5d3160d89bf102e62ac635c221ecb4bf50a45045170f30ce7be6729790f97245

Observation 2d1c1057-972a-4ead-9425-a13a643a79c8 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.282431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.282431Z digest=sha256:f2cd0eb229d09aa84533b75d1789f6796a9d4e5a2b4af66d7de1952bd297fe16

Observation 4168b098-7944-412c-be82-10844964ae2e · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Distilling the Knowledge in a Neural Network

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.287268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.287268Z digest=sha256:6338732ad01476aac9113ddbe67d4b76243db272c9d05ae057516f7731567aed

Observation fa6d330f-c52d-4c5c-91e3-2bf43a08eb6c · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.292775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.292775Z digest=sha256:f4cd7817f7f808803b6b5d5795bc71bca25d529e825e59cbf5bab022b590bf51

Observation 78772e50-0441-4ffa-92ed-6a3627115396 · outbound

This paper cites J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:15:01.348671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.297666Z digest=sha256:bcc6a3b3c166b4eca7ac2ec9c98a9b31bf2745f2069a97b2905a1aee5078adb8

Observation fabdeda1-4538-42dd-bf7f-d96e777cc12b · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Adam: A Method for Stochastic Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.302574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.302574Z digest=sha256:b1f52f4affa75c3fcc3ff9c9c1fe9fda964e10f877bd688da29a512cdc0792f0

Observation 20694219-570b-4f87-af11-587298110cb1 · outbound

This paper cites Reformer: The Efficient Transformer.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Reformer: The Efficient Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.308254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.308254Z digest=sha256:cdc151b04ff3b458b5c6d2614c216422548c56fcd32266c8ddd2b1157236e16a

Observation a7928e46-dd3b-4203-9cb4-31dd1af1925d · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.331544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.313723Z digest=sha256:c3267dd5b60f04f20493439d55d519f97c41bc22f37fc05e1b165e34b2764b78

Observation 26e960b6-3ead-40a0-bbe2-7bcee29779bb · outbound

This paper cites J.; Blankevoort, T.; and Asano, Y.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training J.; Blankevoort, T.; and Asano, Y

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:15:01.314771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.318731Z digest=sha256:fbcf2e7b30c2323ad5f261038a3e915aa2d09ea06dd429d51399cb4e5d014483

Observation ca890914-af7e-4993-8224-80cda051e354 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.297393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.323938Z digest=sha256:e514f375927c0837c3eb1ab4be99bb865b7b4b29b951e1e11c325fecaf5e623c

Observation d9110df9-ad4f-4915-bf96-666ad9c581c8 · outbound

This paper cites How many degrees of freedom do we need to train deep networks: a loss landscape perspective.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training How many degrees of freedom do we need to train deep networks: a loss landscape perspective

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.329382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.329382Z digest=sha256:297535b2e738626b82e75735dbfaa7f0a10537b158ab71fd47d6857a745793fc

Observation 30d1dd21-73df-4f7a-86c3-23b41b6bd2f1 · outbound

This paper cites S.; Yvon, F.; Gall \'e , M.; et al.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training S.; Yvon, F.; Gall \'e , M.; et al

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:15:01.279830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.335097Z digest=sha256:61ab86fb2e2c88f39c005e56ce140400611e053a7c6b2881c7e105ed376b4840

Observation 5719f4dc-b8e6-46be-839d-5cacea2eab23 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.263198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.340116Z digest=sha256:108de7f98fbad47b4c0090e4895b75e5be34aac0886d69d3287859114c82fa84

Observation b3454d66-b87f-4dc6-ad6e-ac1a2b8c915c · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.345066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.345066Z digest=sha256:1996bb2ba412dee9c59de60f63418fafac9f3fb15d2feeb55f0ca57a8df61c4b

Observation 963cd4db-ea28-4ba6-beae-0fe00e536335 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.246268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.350600Z digest=sha256:8b2fb6fde61e812e7aee36d57874cbef35ad46ab281db147a5ec1aff7e7b42a8

Observation caac4706-4262-454f-bedb-c5d91a80238a · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.356289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.356289Z digest=sha256:9f38eb9b25b2f045fa8d9419f635d01eb17ceacb69a24ccc68734c88f4d66890

Observation 85bcc0d9-c4a8-4783-aff6-45c4fdb8b13b · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.229471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.362613Z digest=sha256:916e7cbcae74e7c5465fd2fb3a788512300640f558a354da013c712109a7e507

Observation d432afda-fda1-47f2-b05f-9dd5d7ef9f64 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.213084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.368267Z digest=sha256:9cfc3764a5e0ec4532d654295d60708fbe45189359936ea6f9e2a9857eff0507

Observation 7eaac570-b915-4e45-9e1e-ea7b3a00e5aa · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.373419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.373419Z digest=sha256:732d8282734c5ea8df57d5bddbe7501b9fe18a3d1f78dd141eb105931e664db2

Observation 4cbba197-ae2b-4cfe-a858-895a08b9adbd · outbound

This paper cites F.; Cheng, K.-T.; and Chen, M.-H.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training F.; Cheng, K.-T.; and Chen, M.-H

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:15:01.186856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.378809Z digest=sha256:87a6d4f1e44779ff7e513dd8aa100f31dd3590b8bddd8ea3b0858ec50a31c85d

Observation c83febbc-78a1-4548-8894-cd45578a14c4 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.170936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.383449Z digest=sha256:15f44190985d6f0a373f07fe877e4a8b2a0efbf15fa599f9752cdfb151c2f6a4

Observation 520703aa-3def-4430-a34a-7a202b81d09a · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.153742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.388011Z digest=sha256:93e9cc7003c15117e78b498e7e5ef4be3d59cf12dd70f5f58089ad1e3544a5f9

Observation b82d4a7e-6a29-486e-ba49-760cd46625f3 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.392562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.392562Z digest=sha256:0e53ea5f08a88c6fd7c825f07f388cd57a88efdaf419a96f1fd9e58792c5413c

Observation 085a736c-45ee-4123-a1e9-2479a82df7fc · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.397581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.397581Z digest=sha256:afd0b0eca1688c55855760cfdeed3bf1e63cc8133247d64d0665929107b4426d

Observation 30cc5fd2-21bb-4715-8f60-c997d7edf5a6 · outbound

This paper cites Full Parameter Fine-tuning for Large Language Models with Limited Resources.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Full Parameter Fine-tuning for Large Language Models with Limited Resources

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.401994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.401994Z digest=sha256:8b8b1d25d92403c824357845daac26eb7397808ce80b18afe836e9fb49786e40

Observation fce756cb-eb76-48ee-b90e-759f839ebd8d · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.125570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.406970Z digest=sha256:4bdde291330c359ad05ae1b401137921ef51c3b24acb9ad1a823c93905813f69

Observation 48d6fb39-5dae-42a7-9cf1-11c426896565 · outbound

This paper cites AdapterFusion: Non-Destructive Task Composition for Transfer Learning.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training AdapterFusion: Non-Destructive Task Composition for Transfer Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.411551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.411551Z digest=sha256:f53814edd6c5aee3c18306fe91f3ee175cb1a3be0bc8b3c92c50a52be8f8e932

Observation 12b6b328-2023-4be5-b4fd-c8611e5dc047 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.109748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.415977Z digest=sha256:3be39aa2f9d41a8dd341d1c609d5e9eaf66cd27941a5ebd0e566e6a3e32ad199

Observation 39439011-6017-44a1-a179-7d6837f9b635 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.092755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.420394Z digest=sha256:dd5123ef77e68e7559e8186f959d3a043e325323dff772439b316412704c6cb5

Observation 00a77363-0931-4fc7-b716-bb6d353c2fd2 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.075683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.424752Z digest=sha256:c5a5b6ed1e9cf094f336a833fb934cdf7799ac6ad93d30c20b75577f8bf27215

Observation 87bd92aa-76dd-462c-9cac-8a035a314955 · outbound

This paper cites AdapterDrop: On the Efficiency of Adapters in Transformers.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training AdapterDrop: On the Efficiency of Adapters in Transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.430704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.430704Z digest=sha256:9777a79f81fc400cc1276c7bbb7d9f22517a5df26b52943ba49c13f2d75afb50

Observation 83dd164f-6802-416d-9339-bf77ec22b640 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.059747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.435623Z digest=sha256:7354f0edb3718c21ba008904e1e662073fb2711ad7ce217a87eb708a0280fa2f

Observation bcbbb98e-f4ee-4838-9edd-d500c56b193e · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.440055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.440055Z digest=sha256:40f76eb9c56949cd29332c19c88c9ff7669bd01d1c8e34fd3f8d5bb5ae057d05

Observation b1c0ac83-783a-462c-889c-babe5f0a0333 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.042963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.445598Z digest=sha256:41ff5095321308f8a51a3623727d92a514e2e6d935b65ba6677643b6ebd36da5

Observation 095cb70b-5d8a-48d6-80cc-877d6e3d3140 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.023603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.450133Z digest=sha256:6be920348edc80f0f692f3f6e413d46e814714371aad3a839e26c34881093ebc

Observation eb8aa5ed-74a0-4030-9fd7-33bba2000dd8 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:01.005644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.455246Z digest=sha256:21cb42d8d1ccb394c5540b730b3e8213a0f71a5404149cffef07de7f452f3b85

Observation f743604f-12b8-420b-a41c-0cf55b0450bf · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:00.988975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.460841Z digest=sha256:64e45983c4b44ec6ec98cafb3d7231c7c4c9bb9c4c97eb776d92274605cc3c56

Observation 65e4d4e5-6b03-4ed9-8bb9-29c977de77cd · outbound

This paper cites Q.; Garcia, X.; Wei, J.; Wang, X.; Chung, H.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Q.; Garcia, X.; Wei, J.; Wang, X.; Chung, H

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:15:00.971681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.465785Z digest=sha256:fd7807042a01b2e4e6edf0445235dc4284818d5ee848b5fa889ad7932785b79b

Observation 638cc6ad-01c2-4f94-a770-334d44e06e4b · outbound

This paper cites Understanding Self-supervised Learning with Dual Deep Networks.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Understanding Self-supervised Learning with Dual Deep Networks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.470740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.470740Z digest=sha256:c090a05362d9ce573440202cd3dd664ad30fb63cf8b5717ca0e315ee079d5df0

Observation 5e058069-8e84-4926-aa68-0e9f392f6b36 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training LLaMA: Open and Efficient Foundation Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.475904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.475904Z digest=sha256:69841c043ade87c6b1e9360d09c7e878c3fa51899c22e030a27dfe7b17794e30

Observation dce234f8-0a1e-48b7-8551-f10a770e5a52 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.480905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.480905Z digest=sha256:684c3a3fe4a8e5cc781d8672dd51bbee9c3d190ce3e2bc489b21af4041b1188c

Observation cec5e46e-4392-4d0a-be61-32a7461e9491 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:00.954812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.486271Z digest=sha256:66a87808ac1445e8720dc7fb7c7a51d6fcde3b8ee2049fb457c6a7c7d41f9ba6

Observation 013cfe90-6c4c-4089-ac13-29e17f8e0b88 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.490930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.490930Z digest=sha256:6db9b3f6ef755a4dc5fa710a8520429768d605fe19541622d8bcbd457d80a137

Observation 076a740c-a580-477a-a141-adfe131e1097 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training OPT: Open Pre-trained Transformer Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T00:15:00.495448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:15:00.495448Z digest=sha256:6f5ee4800a524835ac02b5fc8329e52be82807eae335dbfcf97a6829417bfc20

Observation 426e1664-c24f-4453-8ae7-70384f40d797 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:00.926948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.500331Z digest=sha256:e586d078374f73705764d45dfcdf375ee47617eaabaa7f9e1fca6f41d773d405

Observation 5c03486c-852f-411f-ae57-06d3590d81a3 · outbound

This paper cites an unresolved cited work.

Gradient Weight-normalized Low-rank Projection for Efficient LLM Training Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:15:00.910098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-11T00:15:00.504663Z digest=sha256:365dacdbfdca901fd54995b074aeba5621cc36037dfd22dc00685c814ef918ee

Pith citing papers

Observation 059d4fc8-321e-480a-b3f2-1ea11ba969c7 · inbound

SSH: Sparse Spectrum Adaptation via Discrete Hartley Transformation cites this paper.

SSH: Sparse Spectrum Adaptation via Discrete Hartley Transformation Gradient Weight-normalized Low-rank Projection for Efficient LLM Training

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:01:10.348560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-08T19:01:09.456667Z digest=sha256:a5963ad78f6558ea69fbad2c5dc025f764d77116c9fd4f1dd0ebcb746beaac6e

Observation d7d4a625-365a-454f-a19d-b137a97ca165 · inbound

Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data cites this paper.

Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data Gradient Weight-normalized Low-rank Projection for Efficient LLM Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:25:51.616737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:25:51.616737Z digest=sha256:1c98c242249a9a56c5fd69c06813ad5a54315ded4a1e4cf79c4d4beab78447e4