Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:40:49.339784Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2508.01506.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T05:40:49.339784Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 02a2d03c-8b93-430a-8e3c-929950cd6abf · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Chang, W.-C
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 12b7e2a0-0fa2-41d1-af57-7845b222685e · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac6995a2-594c-4b5c-8f74-29c1b6306a04 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb48aa86-83d4-4a31-ae1f-b0b52fbcd233 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models QLoRA: Efficient Finetuning of Quantized LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f2cea1c-7ece-4b4f-a12e-fbac5f37f672 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Devlin, M.-W
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6a9ead9c-ae27-4054-b095-2b270f11c45b · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Theapproximationofonematrixbyanotheroflowerrank
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 87868ae5-c470-4989-9cb1-0fa42813b9fb · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models MorphNet: Fast & Simple Resource-Constrained Structure Learning of Deep Networks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c2b9f613-7930-42f5-b2ce-5827759955f6 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f36365dd-9df5-4447-86ac-798ca56512a3 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bb1e7e88-02f2-4590-9ec6-ab79e78161e2 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Language model compression with weighted low-rank factorization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e5c1f0a-2a34-431a-9e76-ca71f6863a0d · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 167e7114-f44c-444d-9923-7a8080a07728 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27a60a9a-c4d9-4de3-98ba-65c7d7dbc28e · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Pruning Filters for Efficient ConvNets
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0536e03-c946-41fc-b1b8-42cd38adebcf · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Lin and Colleagues
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a5e7bd7e-5369-4753-a220-fc18924e8d53 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Towards Compact ConvNets via Structure-Sparsity Regularized Filter Pruning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1cfa2c0e-7705-4dd1-9ec3-16a67d9e115a · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Star Attention: Efficient LLM Inference over Long Sequences
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e970d6f5-f3ab-49c5-bdf1-470b616bc3d0 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2387da24-0e47-44d9-b467-581819e55b25 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models An Entropy-based Pruning Method for CNN Compression
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abcc3029-9469-431f-99cb-34bace7a357c · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f16036ae-f65c-4c4f-9de8-b912414712ff · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bda61b5f-d27c-4f1b-8437-137c8858ae69 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b221776f-d390-4df0-a0c5-ed8e677a54b6 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Shi and Team
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a4fac312-3ee2-4ee4-bc9b-51d8b7f754c5 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65c5ff95-ec2e-4041-9c48-308848f3356f · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models LLaMA: Open and Efficient Foundation Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eecf0104-23ac-4a8d-af97-6a19508b4fb2 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e14407b0-92ae-4892-ac3d-5e423cd16e21 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e9a7622b-e1e6-48bd-87af-b33efc4df0f4 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc6813c-c647-47b9-908e-3d9cfb59dd32 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45840cbb-72ee-4c4f-8e37-e9f5d9e50cfb · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34e39958-89bc-4ddd-9793-a74cf1c8f4f7 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Yuan and Others
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0564b6f0-695a-4df4-a83a-9f47494ca087 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Zhang, Y
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 55795cd0-504f-4703-848e-50bc2f9c2627 · outbound
FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models OPT: Open Pre-trained Transformer Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.