Pith. sign in

Paper Citation Record · LEDGER

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models

As of 21 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2508.01506.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.01506 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:40:49.339784Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02a2d03c-8b93-430a-8e3c-929950cd6abf · outbound

This paper cites Chang, W.-C.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Chang, W.-C

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:40:51.382691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T05:40:46.406532Z digest=sha256:bc370f8514c2e6b9a480924a067f4b95a788baf5fb0c70daaab0f5cb340004e7

Observation 12b7e2a0-0fa2-41d1-af57-7845b222685e · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:46.524655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:46.524655Z digest=sha256:2230d0a821cb75367366cc713ad3db09559cc1fa5b7d09abe3075278921cb2f2

Observation ac6995a2-594c-4b5c-8f74-29c1b6306a04 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:46.650824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:46.650824Z digest=sha256:4cdca681e045c7623d95efaee97e6254d5dbca20bb9228139d9554a22bf50261

Observation cb48aa86-83d4-4a31-ae1f-b0b52fbcd233 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models QLoRA: Efficient Finetuning of Quantized LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:46.791518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:46.791518Z digest=sha256:4329baf81d4a5f440ae4c40c3565aa08afe14c0453ea00348a44dccac96feb1b

Observation 3f2cea1c-7ece-4b4f-a12e-fbac5f37f672 · outbound

This paper cites Devlin, M.-W.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Devlin, M.-W

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:40:51.373721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T05:40:46.941855Z digest=sha256:ff15873b9b9b93d7e0671499ad7d01d1b695f9abeae10f7887c51492e16c0641

Observation 6a9ead9c-ae27-4054-b095-2b270f11c45b · outbound

This paper cites Theapproximationofonematrixbyanotheroflowerrank.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Theapproximationofonematrixbyanotheroflowerrank

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:40:51.363956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T05:40:47.083537Z digest=sha256:db2dd877a3c63259b05cb21e6631d24a469582dbcd5ed74feda506f9cf710e54

Observation 87868ae5-c470-4989-9cb1-0fa42813b9fb · outbound

This paper cites MorphNet: Fast & Simple Resource-Constrained Structure Learning of Deep Networks.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models MorphNet: Fast & Simple Resource-Constrained Structure Learning of Deep Networks

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:40:50.020452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T05:40:47.172004Z digest=sha256:463c840ac9ae64166c5d81e34057ac810541c0a90341784f88f23eb404b3c94d

Observation c2b9f613-7930-42f5-b2ce-5827759955f6 · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:47.324327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:47.324327Z digest=sha256:5ba5800c23d994c26a033f69e35c1796c22b8c210b366a76e513e745bebe40eb

Observation f36365dd-9df5-4447-86ac-798ca56512a3 · outbound

This paper cites an unresolved cited work.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:40:51.355346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T05:40:47.417020Z digest=sha256:26eaecdc26e63ce2ef095c2fcb7dec7abba89b13426631089ded261ecf9b076e

Observation bb1e7e88-02f2-4590-9ec6-ab79e78161e2 · outbound

This paper cites Language model compression with weighted low-rank factorization.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Language model compression with weighted low-rank factorization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:47.501518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:47.501518Z digest=sha256:6d7f8c4ba246c675a8c02b4ef601c27905d3f505507fdda70f88274cd03dc514

Observation 9e5c1f0a-2a34-431a-9e76-ca71f6863a0d · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:47.559961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:47.559961Z digest=sha256:3185fbebb058fe3ac3bee4c375bad8ec9854894a21fb94efccad08c3db0e2997

Observation 167e7114-f44c-444d-9923-7a8080a07728 · outbound

This paper cites Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:47.644420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:47.644420Z digest=sha256:a3b8eef0cc20a3896a3b3b8ac9be2ee6e860d37ce33bb1938bafd4a28f31743d

Observation 27a60a9a-c4d9-4de3-98ba-65c7d7dbc28e · outbound

This paper cites Pruning Filters for Efficient ConvNets.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Pruning Filters for Efficient ConvNets

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:47.741507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:47.741507Z digest=sha256:dfe876be934db45ff872579de24499bee26eb4483ae18a8e5dbf7a3d13b8d218

Observation c0536e03-c946-41fc-b1b8-42cd38adebcf · outbound

This paper cites Lin and Colleagues.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Lin and Colleagues

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:40:51.322437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T05:40:47.785776Z digest=sha256:386e5ef70de4f092384c88ab3bbe300df38051d53b2d1c1957583b283a9695b8

Observation a5e7bd7e-5369-4753-a220-fc18924e8d53 · outbound

This paper cites Towards Compact ConvNets via Structure-Sparsity Regularized Filter Pruning.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Towards Compact ConvNets via Structure-Sparsity Regularized Filter Pruning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:40:49.713172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T05:40:47.886142Z digest=sha256:89ac767e74a7443761575869c17c7041ec2948a7d31f2b748006eb8b93a6a4ed

Observation 1cfa2c0e-7705-4dd1-9ec3-16a67d9e115a · outbound

This paper cites Star Attention: Efficient LLM Inference over Long Sequences.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Star Attention: Efficient LLM Inference over Long Sequences

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:47.963422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:47.963422Z digest=sha256:c56d2fda89b40cd9a362b6eba05590164f72cd91fd9d9a17d493b38465935ea2

Observation e970d6f5-f3ab-49c5-bdf1-470b616bc3d0 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:48.111295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:48.111295Z digest=sha256:5d80dd48ce95cedf3abde09e087f377a001adc466332c36f757749bc1f01e3b8

Observation 2387da24-0e47-44d9-b467-581819e55b25 · outbound

This paper cites An Entropy-based Pruning Method for CNN Compression.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models An Entropy-based Pruning Method for CNN Compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:48.179747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:48.179747Z digest=sha256:c36ad8d6b1db7686c717fca9e960fb53f547a9057ecf49d900d76cf0ba294fd7

Observation abcc3029-9469-431f-99cb-34bace7a357c · outbound

This paper cites ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:48.270944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:48.270944Z digest=sha256:ce5a80e68a83259619e6dba6784870c520e227e79e1726be15142c66ba79e3f7

Observation f16036ae-f65c-4c4f-9de8-b912414712ff · outbound

This paper cites an unresolved cited work.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:40:51.175906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T05:40:48.341204Z digest=sha256:0049a23b5a68c48183b8d9d4dc545e31e22fa907476e33d34db3d279418e9663

Observation bda61b5f-d27c-4f1b-8437-137c8858ae69 · outbound

This paper cites an unresolved cited work.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:40:51.000642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T05:40:48.419107Z digest=sha256:91f0963528c486713e5717ec4a1f9824599b1c66855ce548b37f0fbba072b0b4

Observation b221776f-d390-4df0-a0c5-ed8e677a54b6 · outbound

This paper cites Shi and Team.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Shi and Team

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:40:50.787293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T05:40:48.456722Z digest=sha256:5b3478d0da2ab3823b948714ae99a0befa1895b4af0a166db0359ce87204618e

Observation a4fac312-3ee2-4ee4-bc9b-51d8b7f754c5 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:48.546549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:48.546549Z digest=sha256:ac9105d1451a3bf69d5cbfd148c240f392844bfa29b7863f9c4e70c1980763f7

Observation 65c5ff95-ec2e-4041-9c48-308848f3356f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models LLaMA: Open and Efficient Foundation Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:48.636130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:48.636130Z digest=sha256:18b6a82f6c1435eb97ebe63a0d7708dec4a9833513c28e2fc727636ae6518f1e

Observation eecf0104-23ac-4a8d-af97-6a19508b4fb2 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:48.697248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:48.697248Z digest=sha256:073d57436d76006728a544a465ae075271b3a9cc9537cb9a602f9049fd5253e9

Observation e14407b0-92ae-4892-ac3d-5e423cd16e21 · outbound

This paper cites an unresolved cited work.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T05:40:50.466047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T05:40:48.789902Z digest=sha256:fd7253ab3db5591dcac971c24debbed0eb26c9e36b5e488740629b4af601baec

Observation e9a7622b-e1e6-48bd-87af-b33efc4df0f4 · outbound

This paper cites SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:48.880739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:48.880739Z digest=sha256:b210de4d49d8e4ef7f470ad1d09e9207e52d472bfedf919bd6886995705532d4

Observation 5fc6813c-c647-47b9-908e-3d9cfb59dd32 · outbound

This paper cites SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:48.949595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:48.949595Z digest=sha256:232c53a2be4fbfcbe3a4718358aa435a68c690a77b2f5a1dbf255efa22038b24

Observation 45840cbb-72ee-4c4f-8e37-e9f5d9e50cfb · outbound

This paper cites Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:49.038808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:49.038808Z digest=sha256:7c0d3b5577a71356ec5861374600cae33eea6c44533b9db10f69b6d156e03f70

Observation 34e39958-89bc-4ddd-9793-a74cf1c8f4f7 · outbound

This paper cites Yuan and Others.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Yuan and Others

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:49.101917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:49.101917Z digest=sha256:23e827d65751183fb4ad7754b7d69b752c99d4bb0b1cab2a99d1e2743bdd66d3

Observation 0564b6f0-695a-4df4-a83a-9f47494ca087 · outbound

This paper cites Zhang, Y.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Zhang, Y

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:40:50.277503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T05:40:49.205494Z digest=sha256:8f2f940262b5956158e6fc1f0a5ad94d57cc7ce7b59db80b50b87c05f7704810

Observation 55795cd0-504f-4703-848e-50bc2f9c2627 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models OPT: Open Pre-trained Transformer Language Models

Reference 32

Resolution
malformed identifier
no resolver link, observed 2026-08-06T05:40:49.339784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:49.339784Z digest=sha256:ccde39dfd17e82308f087f116ea8967cfc715cf2c705834e105ba44c893c474e

Pith citing papers

No inbound Pith citation observations are available.