Pith. sign in

Paper Citation Record · LEDGER

Efficient Pre-Training of LLMs through Truncated SVD Layers

As of 8 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2605.28573.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.28573 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T13:39:53.206470Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 803ca0db-dfff-419d-a183-a4fb9dd8a4d4 · outbound

This paper cites Princeton University Press, 2008.

Efficient Pre-Training of LLMs through Truncated SVD Layers Princeton University Press, 2008

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:c1fb2e6021615c66087f7278642e336b049c0046c6175f18896f808fba0e6bb5

Observation eb999300-a5e4-4e24-8c8d-c2217c04322f · outbound

This paper cites GQA: Training generalized multi-query transformer models from multi-head checkpoints.

Efficient Pre-Training of LLMs through Truncated SVD Layers GQA: Training generalized multi-query transformer models from multi-head checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:26b9128b897062e2a6b0ded6017bbd5a51d226ea0dad46bab04c170a455bf8ad

Observation 99b47e68-0876-4c50-be0d-b28ea703478d · outbound

This paper cites Unitary evolution recurrent neural networks, 2016.

Efficient Pre-Training of LLMs through Truncated SVD Layers Unitary evolution recurrent neural networks, 2016

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:1556e77c3b5d903fd26e2e83abd1e9478034390d0402a7c4528a15e98db0d3ab

Observation cc58cfc9-086b-4e95-a96e-02da620e13d7 · outbound

This paper cites Can we gain more from orthogonality regularizations in training deep cnns?, 2018.

Efficient Pre-Training of LLMs through Truncated SVD Layers Can we gain more from orthogonality regularizations in training deep cnns?, 2018

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:fd25c895dec2ab7d8310745a8fde6d28c447f925dcd680b3d57d22eee47ac9aa

Observation 33d73fa0-0af7-48be-a9c3-3c6ae314c3eb · outbound

This paper cites Cambridge University Press, 2023.

Efficient Pre-Training of LLMs through Truncated SVD Layers Cambridge University Press, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:31822acd97319118072069f16739b65bca511419c5f912d5ef0f20120bdce359

Observation ba6164c2-f849-4d16-aa4f-a427d06134b3 · outbound

This paper cites an unresolved cited work.

Efficient Pre-Training of LLMs through Truncated SVD Layers Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:02e13362312936892d3d214f9f8af5fc6fcf0a4a6bb8aa72509a3a5005e8ec40

Observation 4a6ca715-2088-4ab3-a9d6-8df46e5a36c8 · outbound

This paper cites an unresolved cited work.

Efficient Pre-Training of LLMs through Truncated SVD Layers Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:2075cc68d480c9043b09f74d3ac3f1915cdc36bc9dd4f0b48c1e9f3e5a741683

Observation 611fa5f3-172d-4acb-a452-19547ff7e3f0 · outbound

This paper cites Reducing overfitting in deep networks by decorrelating representations, 2016.

Efficient Pre-Training of LLMs through Truncated SVD Layers Reducing overfitting in deep networks by decorrelating representations, 2016

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:b3a69844ca154bf6029bfc801c00be7c4908af210d84a783b9f2d4b4d2aff542

Observation 8d66f8d9-5628-406e-a03b-55ab48ecf332 · outbound

This paper cites IOS Press, October 2024.

Efficient Pre-Training of LLMs through Truncated SVD Layers IOS Press, October 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:d0db8e00f7be1bc997041b370c684201c74898c44f80a3ccc6fc175d84475b2e

Observation b3cefa28-6997-4de1-979c-4220b5cc62b4 · outbound

This paper cites Documenting large webtext corpora: A case study on the colossal clean crawled corpus, 2021.

Efficient Pre-Training of LLMs through Truncated SVD Layers Documenting large webtext corpora: A case study on the colossal clean crawled corpus, 2021

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:293c3fe1dc95207927ef4b65e3e49a6755469fc624e4816de64ac19650987ae3

Observation a2551fa4-5e2b-4dd5-8c8a-9dd6b84f3f37 · outbound

This paper cites The Llama 3 Herd of Models.

Efficient Pre-Training of LLMs through Truncated SVD Layers The Llama 3 Herd of Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.727141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:e8688003aa944e9b20f358957d82883c9b224ae18700d8f439c5cb9eb46c40a0

Observation 71be59f2-fb3a-44fb-8de5-47c194ac368b · outbound

This paper cites The approximation of one matrix by another of lower rank.

Efficient Pre-Training of LLMs through Truncated SVD Layers The approximation of one matrix by another of lower rank

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:36dfcd738d7d2baa1ea86fc448cc873dbc4629dc21a21e25eecc4365bdac6162

Observation 4888d7ca-8ad8-431e-938a-937d7a7f5710 · outbound

This paper cites Arias, and Steven T.

Efficient Pre-Training of LLMs through Truncated SVD Layers Arias, and Steven T

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:23fd0e51b4ca4e26b2ac8c351637045f62b2bd0c2768b7e1025743b1cdfaf8b0

Observation fcd53eef-4ceb-4e08-9487-bd379e41b4e2 · outbound

This paper cites Full-rank no more: Low-rank weight training for modern speech recognition models, 2024.

Efficient Pre-Training of LLMs through Truncated SVD Layers Full-rank no more: Low-rank weight training for modern speech recognition models, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:08122d79e30d162406cc845ef15a540d8f8da62a6ff46e5bdb48136ba9afb8d6

Observation 813dba24-c7c6-4b33-8aeb-6b60d7b52e69 · outbound

This paper cites Deep feedforward networks.Deep learning, 1:161–217, 2016.

Efficient Pre-Training of LLMs through Truncated SVD Layers Deep feedforward networks.Deep learning, 1:161–217, 2016

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:ec56ffd7ad0dc5d739d0400a776ffa9fb5daf52833eb7fe04cb4cd4b541c9fa3

Observation 46786ee7-0e73-4c67-aad8-4297ab996ab7 · outbound

This paper cites From gpt to llama: Tracing the growth of large language models.Theoretical and Natural Science, 142:144–155, 11 2025.

Efficient Pre-Training of LLMs through Truncated SVD Layers From gpt to llama: Tracing the growth of large language models.Theoretical and Natural Science, 142:144–155, 11 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:c4257ce59ad9bc19dbd1b6e6b5efb1fb8330e95069d77f32eb06a292ee601f79

Observation c9514802-bbac-4a32-9529-7845407417e0 · outbound

This paper cites Sltrain: a sparse plus low-rank approach for parameter and memory efficient pretraining, 2024.

Efficient Pre-Training of LLMs through Truncated SVD Layers Sltrain: a sparse plus low-rank approach for parameter and memory efficient pretraining, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:8a3aa22c69fba48d108f5304525b9c73438f524c07204deaaa2849db82c25899

Observation 8de2ee8c-5454-41af-a9e5-0f048087185b · outbound

This paper cites Rae, Oriol Vinyals, and Laurent Sifre.

Efficient Pre-Training of LLMs through Truncated SVD Layers Rae, Oriol Vinyals, and Laurent Sifre

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:f8b8aa6663a7121b823c991c537acbe1ba1c35d9794dec993fa6b4cd125cc256

Observation 0da1ecbc-3590-4db9-b3ba-f30a79bdd00a · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Efficient Pre-Training of LLMs through Truncated SVD Layers Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:481d09de18064a222cc54801991b478f2a2156d55aac8e225481917079363fd6

Observation 880813a3-cbb5-4af9-a3a7-62de6314e3d9 · outbound

This paper cites an unresolved cited work.

Efficient Pre-Training of LLMs through Truncated SVD Layers Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:c4b3d6b4788910e60d7119a13b471cced64e74bae6f3078ffb7093e9810aeb17

Observation 71870e05-89b7-4237-a8a1-79d0fc10ab72 · outbound

This paper cites Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei.

Efficient Pre-Training of LLMs through Truncated SVD Layers Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:b709991679b28aa853aba2d1d63499b7f1915d41bc4b14a7fde6b383ce10e56b

Observation 9df9d7a1-5758-4e84-9db4-2f7221b41631 · outbound

This paper cites Initialization and regular- ization of factorized neural layers.

Efficient Pre-Training of LLMs through Truncated SVD Layers Initialization and regular- ization of factorized neural layers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:b47761327c81f71f817ad7cf8143594815bbfa33247ac1b065b931cc45a24a74

Observation c958fc24-c032-4584-bc87-a468cc2d1af2 · outbound

This paper cites Initialization and regular- ization of factorized neural layers, 2022.

Efficient Pre-Training of LLMs through Truncated SVD Layers Initialization and regular- ization of factorized neural layers, 2022

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:5178d641a7edb748bf094ee6bdc4381635ad1cc582d4327133edc0ffd727b897

Observation 56d8b64a-c1f2-4d7a-8ec5-512369ef81fc · outbound

This paper cites Lost: Low-rank and sparse pre-training for large language models, 2025.

Efficient Pre-Training of LLMs through Truncated SVD Layers Lost: Low-rank and sparse pre-training for large language models, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:2d4cbee6f6fafd9c129c143cac3d89fbc44c0a240a8a1de6edc8c4c518a15354

Observation ea4f1aee-409b-404d-b0c6-989b44cf5229 · outbound

This paper cites Relora: High- rank training through low-rank updates, 2023.

Efficient Pre-Training of LLMs through Truncated SVD Layers Relora: High- rank training through low-rank updates, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:2b8f6eabdc3d50f094f7267d79f2628fefe7cc516c5a561b0619b262ab7821b7

Observation ffb52c38-215f-4810-af53-aa8909c67714 · outbound

This paper cites Cola: Compute-efficient pre-training of llms via low-rank activation, 2025.

Efficient Pre-Training of LLMs through Truncated SVD Layers Cola: Compute-efficient pre-training of llms via low-rank activation, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:e84bf5c4c449f46025110e5b19bd461fab594ca432388c785402a0aabaf27d07

Observation 583e94d8-c514-4a0b-8f1d-d813c34878f7 · outbound

This paper cites Large language models: A survey, 2025.

Efficient Pre-Training of LLMs through Truncated SVD Layers Large language models: A survey, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:75bd1d408606c2523a868f372adc5154371e6cd0ea4862cf3019785619eeff6c

Observation 22064a8e-f43f-43f4-a44b-3d622b4420e8 · outbound

This paper cites Parameter and memory efficient pretraining via low-rank riemannian optimization.

Efficient Pre-Training of LLMs through Truncated SVD Layers Parameter and memory efficient pretraining via low-rank riemannian optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:4b34086bf12e5b8f35a516f3f81e3026207873ed07e9a30ca58d59f490358805

Observation 73748c87-332c-4023-8d3f-228568de0478 · outbound

This paper cites Olmo 3.

Efficient Pre-Training of LLMs through Truncated SVD Layers Olmo 3

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.744541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:7a4a17548a702207a84ad576efebf71f542462927cc06d14afa6db09526c6a2d

Observation 05b2e6af-98d9-4ff0-9121-53322f2be3f7 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Efficient Pre-Training of LLMs through Truncated SVD Layers Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.731874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:b27360bc1f41c96370b22e4282b06c780a5faf0c53ed10de8e1ecdc4b92bb4a9

Observation 151ff332-23a2-4de5-96af-322218e385bc · outbound

This paper cites Robust low-rank training via approximate orthonormal constraints, 2023.

Efficient Pre-Training of LLMs through Truncated SVD Layers Robust low-rank training via approximate orthonormal constraints, 2023

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:e64392b10ff747d509c1f51b83e1852497f4d75d40dee13e439a58b2011774ca

Observation 0b38f402-96f7-4b3e-96f6-1da574865883 · outbound

This paper cites Low-rank lottery tickets: finding efficient low-rank neural networks via matrix differential equations, 2022.

Efficient Pre-Training of LLMs through Truncated SVD Layers Low-rank lottery tickets: finding efficient low-rank neural networks via matrix differential equations, 2022

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:c092d5d2c3f9c00ea9a6d98b99fee3526625767edd3729a41ad32cf328ede1f6

Observation c7d54399-7dba-4ffe-b6c2-0192317e9479 · outbound

This paper cites Compact: Compressed activations for memory-efficient llm training, 2025.

Efficient Pre-Training of LLMs through Truncated SVD Layers Compact: Compressed activations for memory-efficient llm training, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:5da830ee6ec1dbe2e43d15d7567d82ed7013e056db8d4ab47027ef73dab206ef

Observation 63391c11-29bf-44e0-a151-124d2792ff23 · outbound

This paper cites Cuttlefish: Low-Rank Model Training without All the Tuning.

Efficient Pre-Training of LLMs through Truncated SVD Layers Cuttlefish: Low-Rank Model Training without All the Tuning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.736773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:21cf225adaf3fb409ae99451cc1f73eeb1163570f2f8673be0ff4b9cb1cefd44

Observation 60b0ab53-a955-47f1-85d7-e9ed03adf229 · outbound

This paper cites Dynamic rank adjustment for accurate and efficient neural network training.arXiv preprint arXiv:2508.08625, 2025.

Efficient Pre-Training of LLMs through Truncated SVD Layers Dynamic rank adjustment for accurate and efficient neural network training.arXiv preprint arXiv:2508.08625, 2025

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.738320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:9b64b3122a5ea233aad2eacf3c7fff015a382b8eed8dd2ea25e5df3b2b3c6a83

Observation 2abeaca3-03fd-43da-b0ce-a732d94e8a36 · outbound

This paper cites Elrt: Efficient low-rank training for compact convolutional neural networks, 2024.

Efficient Pre-Training of LLMs through Truncated SVD Layers Elrt: Efficient low-rank training for compact convolutional neural networks, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:52ae0899277bd6105d094115dbb07f537c90fff8b8ae33344268426b8bf0c8ff

Observation 790e8e70-9ee6-484f-bc15-3ead2327b3be · outbound

This paper cites Llama: Open and efficient foundation language models, 2023.

Efficient Pre-Training of LLMs through Truncated SVD Layers Llama: Open and efficient foundation language models, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:15e7f8bcdd12c4bcbdf87de3a8d16e2334a17aa8512852b706c875efa6cd94f5

Observation 23a9d67c-b4c8-44ae-af57-8144981f721b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Efficient Pre-Training of LLMs through Truncated SVD Layers Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.729620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:44a0b83ce019e07ae93b3725acd17e0546dea7e326725b41407dd31b06903a54

Observation 1f155ceb-2721-4cd2-a698-aceedf5e7b84 · outbound

This paper cites BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models.

Efficient Pre-Training of LLMs through Truncated SVD Layers BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.734144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:2538cfe4d05477513c2b18a6d00b7e767c6744a705d1a92966cd9427249163b5

Observation c4922ffc-7840-41ad-bfa1-9a3c4f178240 · outbound

This paper cites Investigating low-rank training in transformer language models: Efficiency and scaling analysis, 2024.

Efficient Pre-Training of LLMs through Truncated SVD Layers Investigating low-rank training in transformer language models: Efficiency and scaling analysis, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:cfc1c6b214b892c9455e04ddca1bb656da4c27f176f12e0a039bcbee977f1f59

Observation f8ba27e5-d6b7-4ddb-8ddd-7ee8a7e031ca · outbound

This paper cites Transformers: State-of-the-art natural language processing.

Efficient Pre-Training of LLMs through Truncated SVD Layers Transformers: State-of-the-art natural language processing

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:ab17f2900674e40bbd67b8fbf28f476dcc947a45d091ff4efb53991eb23f59dd

Observation 5342ad7d-e7ae-4b99-be1f-0251853b6497 · outbound

This paper cites Learning low-rank deep neural networks via singular vector orthogonality regularization and singular value sparsification, 2020.

Efficient Pre-Training of LLMs through Truncated SVD Layers Learning low-rank deep neural networks via singular vector orthogonality regularization and singular value sparsification, 2020

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:9776d64bf44a9174d1f774eb6d58f4c0ee62e415245e85e0be6b1d2312e154f0

Observation ea414b5f-f6cc-46e0-9436-fbf34691dc6f · outbound

This paper cites InRank: Incremental Low-Rank Learning.

Efficient Pre-Training of LLMs through Truncated SVD Layers InRank: Incremental Low-Rank Learning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.741721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:c53bc30a99538ee3d13509f6389a06d7b1ca5210c58224d6b3ea24ec781d50af

Observation e1e5c323-c1c1-42e6-8c8d-901080d194ed · outbound

This paper cites Galore: Memory-efficient llm training by gradient low-rank projection, 2024.

Efficient Pre-Training of LLMs through Truncated SVD Layers Galore: Memory-efficient llm training by gradient low-rank projection, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:44d8f02b8da27cba0ecbc5724baa760b1f613dbf25382209116366868d9721c5

Observation 56fa5527-1306-4c1f-a86c-3bbd767502ba · outbound

This paper cites A survey of large language models, 2026.

Efficient Pre-Training of LLMs through Truncated SVD Layers A survey of large language models, 2026

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:320ecddc2df9f4796fb3cd0961c9e175a73bc8cd224cc763eeb11202aca38fa3

Observation 2649f54a-9539-4b9a-bdf8-7d6532208ee3 · outbound

This paper cites Hence the forward operator norm is controlled exactly byσ max.

Efficient Pre-Training of LLMs through Truncated SVD Layers Hence the forward operator norm is controlled exactly byσ max

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:6013bc0f9b616a2f7e24b0f2f5b26db3ea7dc3e52c530c0a2c2d991a3ab5230b

Observation fac4b32b-a6ed-4524-b5c2-c69b443f8280 · outbound

This paper cites Hence the backward operator norm is controlled by the same quantity.

Efficient Pre-Training of LLMs through Truncated SVD Layers Hence the backward operator norm is controlled by the same quantity

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:59ae9fd8751ec728549c8fcf92b7330f475853498c81986187c876815c32d4c2

Observation 80cd6486-57d4-4de3-aeb7-5e3cb76b4c21 · outbound

This paper cites Likewise, ifδlies in the represented output subspacespan(U), then σmin√r ∥δ∥2 ≤ ∥W ⊤δ∥2 ≤ σmax√r ∥δ∥2.

Efficient Pre-Training of LLMs through Truncated SVD Layers Likewise, ifδlies in the represented output subspacespan(U), then σmin√r ∥δ∥2 ≤ ∥W ⊤δ∥2 ≤ σmax√r ∥δ∥2

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:0f74acee4eb52b08ac406f4d8eb08909e88748450b6ea871b6151696ecc2316a

Observation 55bc3c9b-b5f1-4372-b960-4aad0885c27a · outbound

This paper cites If ui and vi denote the i-th columns of UandV, then W= 1√r rX i=1 σiuiv⊤ i , ∂L ∂σi = 1√r u⊤ i Gvi.

Efficient Pre-Training of LLMs through Truncated SVD Layers If ui and vi denote the i-th columns of UandV, then W= 1√r rX i=1 σiuiv⊤ i , ∂L ∂σi = 1√r u⊤ i Gvi

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:199eaa85ba784a6e854489d9ce40d722118d09b349c7c67ca33145e9f3437221

Observation eea9d66c-12e0-4a25-8437-458ef6f9f0e5 · outbound

This paper cites an unresolved cited work.

Efficient Pre-Training of LLMs through Truncated SVD Layers Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:0032a0e20c0f33694226980a02e78f14044223125bfc46d93800bcf8ec6481fa

Observation 6bd55ed0-fcae-4ba4-8119-e0d6361c1273 · outbound

This paper cites an unresolved cited work.

Efficient Pre-Training of LLMs through Truncated SVD Layers Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:53c86fab982388fd10861d800e6793f33d91e8eb235269aeaa8bc555d00a2238

Observation 205b66b8-5fff-461d-8e1e-577004d1e6f5 · outbound

This paper cites Orthonormality does not remove this fundamental low-rank bottleneck, but it does prevent additional instability caused by badly scaled basis factors.

Efficient Pre-Training of LLMs through Truncated SVD Layers Orthonormality does not remove this fundamental low-rank bottleneck, but it does prevent additional instability caused by badly scaled basis factors

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:1760116cba22bce371f902ad9724ba5be4e46fd6baad7015b0b624429c36037c

Observation e642d319-2108-4ce2-bbd2-e2cbca30cb20 · outbound

This paper cites an unresolved cited work.

Efficient Pre-Training of LLMs through Truncated SVD Layers Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:018b10d11209350659ae24047890350a3b871daf7d2cf821f8a058cc7f255c69

Observation dab857e3-bf4c-43e2-a22c-c7b94da8d6c4 · outbound

This paper cites an unresolved cited work.

Efficient Pre-Training of LLMs through Truncated SVD Layers Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:7dfb423fc9ffc2f6d8d40eeabba34b901bc40e861a837f6a3eb9187951243f6b

Observation bbc9a5c2-7d97-4b38-b1c8-f6dd55d208a8 · outbound

This paper cites an unresolved cited work.

Efficient Pre-Training of LLMs through Truncated SVD Layers Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:2a2b0182c2f616e99bbe68d4b17b07a85f20063fb18587a2e3b64220f1beddcf

Observation 2a62f5f2-1ac0-42e2-8e5b-16376d31b118 · outbound

This paper cites an unresolved cited work.

Efficient Pre-Training of LLMs through Truncated SVD Layers Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-29T13:39:53.206470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T13:39:53.206470Z digest=sha256:8e2d31f8083b7e1bcc25489fb5d971076052416390b97f86b3e6cf67efbb7383

Pith citing papers

No inbound Pith citation observations are available.