Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T23:31:27.783419Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 0 inbound Pith citation observations for arXiv:2605.30728.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T23:31:27.783419Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 100 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 58be9c57-58c9-4bc0-9f23-453c20e8f45a · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https: //docs.nvidia.com/cuda/cuda-c-programming-guide#compressible- memory
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 480b1c84-24c2-46a0-825c-395fcf9625ae · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://docs.nvidia.com/cuda/cuda-c-programming-guide/#device- memory-accesses
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd5618a2-0317-459a-ac84-e1571df69b9a · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://docs.nvidia.com/cuda/cuda-c-programming- guide/#hardware-implementation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ce86f4-3a34-4b03-85e4-44e725cd109d · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://docs.nvidia
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edd49fed-8a93-4b38-b361-70fe5cd6f868 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://catalog.ngc.nvidia.com/orgs/nvidia/teams/ dle/models/dlrm_base_tf2_ckpt_ds-criteo-fl15
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3f46a52-7d0d-48ce-9f6d-4b13a972d88f · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https:// developer.nvidia.com/nvcomp
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9421c787-6802-4ed6-91f7-9811c919da51 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb1ede28-9520-4da8-93f0-304e4d149a7a · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://resources.nvidia.com/en-us- tensor-core/nvidia-tensor-core-gpu-datasheet
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d811ba6d-25ec-41ca-8a15-f0e29468fd74 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://images.nvidia.com/content/ technologies/volta/pdf/volta-v100-datasheet-update-us-1165301- r5.pdf
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c244c94d-5fa2-49c1-923d-becefcb4df64 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://www.nvidia.com/en-in/data-center/ nvlink/
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 933516ca-e192-4959-9604-4ca54d7e8e01 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended https://labs.criteo.com/2013/12/download- terabyte-click-logs-2/
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc2e4bdc-ca83-411b-bb8d-6dfa9371be66 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Understanding training efficiency of deep learning recommendation models at scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c138396-e6c8-4561-8843-48da8f83ef8c · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Accelerating gpu data processing using fastlanes compression
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d74e3866-f4d7-4aa5-8b46-90a9e0cb6582 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Bagpipe: Accelerating deep recommendation model training
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f89cf1-d351-4faa-8a8e-0d06b1ffe084 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Graph neural network training systems: A performance comparison of full-graph and mini-batch
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eff29eb0-b998-453a-842d-b52dbb8b014e · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Aware: Workload-aware, redundancy-exploiting linear algebra
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 789f85ef-ebca-4da1-b721-e1a34d2b3bea · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Deep gaussian embedding of graphs: Unsupervised inductive learning via ranking
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3abfc8c3-2f8a-427a-9b8f-9b01c21cd119 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Molecular generative graph neural networks for drug discovery.Neurocomputing, 450:242–252, 2021
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d4e8955-bfff-4347-94b6-fcd9a52d7157 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Fcbench: Cross-domain benchmarking of lossless compression for floating-point data
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db428f5c-016d-426f-bbfc-03506bdcd199 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Learned image compression with discretized gaussian mixture likeli- hoods and attention modules
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0839d043-6779-4dd2-a5cd-91caea7330ea · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended The trade-offs of model size in large recommendation models : 100gb to 10mb criteo-tb dlrm model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3425b2b-04c2-49db-85a9-76a9451c2f02 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Gpt3.int8(): 8-bit matrix multiplication for transformers at scale
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ad0f06f-e73c-455e-a7e8-b6fd9441dfe8 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Qlora: Efficient finetuning of quantized llms
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a7fad3-6c85-4f36-99cf-b3fba532ce35 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Accuracy is not all you need
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f38b486-b791-47bf-95d4-8253d4e00fd9 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Haas, Frederick R
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 774f680c-1754-45d6-be16-5237fb861c2b · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Zipserv: Fast and memory-efficient llm inference with hardware-aware lossless compression
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d4e7e4-6dd8-4ec6-ba07-e17ee40e0cd7 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended A Frequency-aware Software Cache for Large Recommendation System Embeddings
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d956aa1e-e3d1-4c1d-81da-2739e64cccf0 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Sahu, Marco Canini, and Amedeo Sapio
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3460820b-b559-4f5c-b189-27e3f38ae811 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Accelerating Communication in Deep Learning Recommendation Model Training with Dual-Level Adaptive Lossy Compression
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45aa004e-7674-4d6a-920c-977a87340210 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Mahoney, and Kurt Keutzer
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e78ee4b7-bc97-4783-a6f0-1747a1a7a707 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Lee, David Brooks, and Carole- Jean Wu
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c6980d-5dec-4d42-aec9-338892be3f14 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Inductive representa- tion learning on large graphs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7905fbc9-36b0-4c2e-a9d4-335cc43ea1ab · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended How to optimize data transfers in cuda c/c++
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3958feea-a664-4238-be44-abe85260ca37 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Natural compression for distributed deep learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f32cc364-16fd-41b4-8f85-ca1632d963b2 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Open graph benchmark: Datasets for machine learning on graphs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4d4085d-2806-4c81-aade-4ba18d4596ca · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d3aa107-1505-4a63-8729-ea11f554b084 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Grow: A row-stationary sparse-dense gemm accelerator for memory-efficient graph convolutional neural networks
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee1a08a2-4a0f-446b-bceb-801e6129435d · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended An extended compression format for the optimization of sparse matrix-vector multiplication.IEEE Trans- actions on Parallel and Distributed Systems , 24(10):1930–1940, October 2013
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76a167d5-cc2c-45f5-bcfa-bb4e648c50f5 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Tensorfloat-32 in the a100 gpu accelerates ai training, hpc up to 20x
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d3ae069-df77-43e2-99ae-2a4a9b7fbd28 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Datasets for benchmarking floating-point compressors, 2020
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fa23bbb-2e43-410b-9024-6e9e46c9b815 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended ndzip-gpu: ef- ficient lossless compression of scientific floating-point data on gpus
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd6c6375-e65d-4a89-b5dc-47e929ce8943 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Webb, Xin Wang, Marcel Nassar, Arjun K
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06166886-e07d-4b9f-b8b1-41d2e548f4e2 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Splitrpc: A Control + Data path splitting rpc stack for ml inference serving
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 243aeb13-81db-4714-aa62-4c1739260fe1 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Efficient memory management for large language model serving with pagedattention
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90515bfd-b669-44db-8b7c-527daed8ff94 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended InfiniGen: Efficient generative inference of large language models with dynamic KV cache management
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eaa13f5-e8f6-4bf1-a4f2-e39cd06a4583 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Naughton, and Jignesh M
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6dc5995-db65-4a6b-a31b-5787a22441c2 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47b697ea-fe90-4f34-bb49-121875260143 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Colossal-ai: A unified deep learning system for large-scale parallel training
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1fcb082-3cb3-438b-bc1b-015b8b5c7786 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64a46fc8-ccb1-425b-9cdd-73fd58363cc9 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Recoil: Parallel rans decoding with decoder-adaptive scalability
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e858a7ae-6d0c-4d64-ae23-29f83b8b43ca · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Using cuda warp-level primitives
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6cb9f79-afc7-4fb5-a368-35b07155cdde · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccb87448-fd90-4747-a8ab-632fcd89b0f9 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Pa- graph: Scaling gnn training on large graphs via computation-aware caching
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c8a26e-a4e1-4b31-b86e-e6d0ba77c44b · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Indigo: Gnn-based inductive knowledge graph completion using pair-wise encoding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ba372c9-549c-4169-8b4e-ac1574d381ad · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended BGL: GPU-Efficient GNN training by optimizing graph data I/O and preprocessing
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7acec39a-9a2e-47b7-8801-4e1dc12cf63c · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Pick and choose: A gnn-based imbalanced learning approach for fraud detection
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 388b03f8-fa5f-4a2f-bf1b-13e7391b0bd3 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Cachegen: Kv cache compression and streaming for fast large language model serving
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30868c71-8ae6-494c-8488-2327e840b407 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Dvc: An end-to-end deep video compression framework
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f31eabe6-84e5-4a95-a311-e0ccba012da0 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Eliminating data processing bottlenecks in gnn training over large graphs via two-level feature compression
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff28a625-a2f4-4629-a8ad-a492ad724d7e · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended BiFeat: Supercharge GNN Training via Graph Feature Quantization
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 53c41487-f68c-4495-8354-554d37f3d024 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Emogi: efficient memory- access for out-of-memory graph-traversal in gpus
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 434671fd-1b3a-4c62-b9fb-e90ecab31030 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Unresolved cited work
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a40dea05-0faf-480b-855c-9f5d5ae8d850 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Unresolved cited work
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ddab374-b9a9-4f7e-b610-fe31c9c53caf · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Query-driven active surveying for collective classification
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a783d1e3-caea-43f3-ad0a-34df0e53354b · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Patel, Yao Zhang, Jason Mak, Andrew Davidson, and John D
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d23c0cfd-cc64-49aa-93f4-70d90dca7ff1 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Gpu-initiated on-demand high-throughput storage access in the BaM system architecture
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 875c4bd6-3e20-4484-908b-de79aaf456c4 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Real-time adaptive image com- pression
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c77b34-affa-4749-a7f1-afb9083063b4 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Faster across the pcie bus: a gpu library for lightweight decompression: including support for patched compression schemes
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7715261c-b98c-4ec9-ad5b-28f4e732f1c0 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Collective classification in network data
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40d51d3c-ae4f-411f-b3ba-80e68ee7f4af · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Scalable graph neural network training: The case for sampling
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31c0e8c1-c916-442d-bc98-09c0c76c1955 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Yogatama, Xiangyao Yu, and Samuel Madden
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc214dc-efb8-431b-8baf-4a2289901aea · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended FlexGen: High-throughput generative inference of large language mod- els with a single GPU
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dff30bca-721a-445c-bbfc-4963b008c41c · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Ugache: A unified gpu cache for embedding-based deep learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 649b7f85-b853-4efb-b489-f0fee07b7797 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Roformer: Enhanced transformer with rotary position embedding
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc495989-c00e-4e7a-8b4b-75abe4291eb5 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Legion: Automatically pushing the envelope of Multi-GPU system for Billion- Scale GNN training
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c82750b2-14ce-4140-8579-06f10d847d15 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Controlling data move- ment to boost performance on the nvidia ampere architec- ture
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8789fc65-98b0-4656-9cf7-eea646aaa210 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Unresolved cited work
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0c47773-57e0-413e-b0f5-0bd4c64c23c7 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Attention is all you need
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55836ff8-7788-4019-bbd9-e5dad7699ef4 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Mariusgnn: Resource-efficient out-of-core training of graph neural networks
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4931efc0-0154-4b32-ac72-10a779bb8be0 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended ZeRO++: Extremely Efficient Collective Communication for Large Model Training
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00eb0ab9-8e4e-4cf3-a011-439056655ce0 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation da586fe3-1ab0-4b7a-a99d-78cd59d14e5f · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Bfloat16: The secret to high per- formance on cloud tpus
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 161265b2-6c2f-4a08-9bf7-aab0bd0725e6 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended A gpu-specialized inference parameter server for large-scale deep recommendation models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edbd7f26-0ea9-4e15-9a5e-ff92b4e015c1 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Massively parallel inverse block-sorting transforms for bzip2 decompression on gpus
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bb53cea-deb6-4830-9a7e-aa3763d5216c · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended MLaaS in the wild: Workload analysis and scheduling in Large-Scale heterogeneous GPU clusters
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 964e1d95-68d8-493f-b387-4e513ce72c62 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Widrow, I
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b088b425-1ec4-4f1a-8d78-9d937a094051 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Abdelmoniem, Aritra Dutta, El Houcine Bergou, Konstantinos Karatsenidis, Marco Canini, and Panos Kalnis
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ec46941-92f6-4096-842c-4c0cab3a09d4 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Mpc: A massively parallel compression algorithm for scientific data
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1917737e-b805-4d3b-aadf-ea01c6e07444 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Gnnlab: a factored system for sample-based gnn training over gpus
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a6698af-23dc-441c-afae-4f51cc6da5d4 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Lossy image compression with conditional diffusion models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dca569c2-248b-4097-8a2f-653f38319afb · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Spargnn: Efficient joint feature-model sparsity exploitation in graph neural network acceleration
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eade4f5c-45c0-40b5-9621-65cb4722a893 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Hamilton, and Jure Leskovec
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66ca2141-80e0-472f-848b-0594027d9ac1 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended SGCN: Exploiting Compressed-Sparse Features in Deep Graph Convolutional Network Accelerators
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9c02e50-0c37-451d-8f89-9bd62460bc61 · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended OPT: Open Pre-trained Transformer Language Models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b2dfe8ad-33e0-41a5-a700-4c19e1ebb98f · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended 70% size, 100% accuracy: Lossless LLM compression for efficient GPU inference via dynamic- length float (DFloat11)
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 942f0507-fa9a-4a9a-9730-f7b5dc42558d · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Ducati: A dual- cache training system for graph neural networks on giant graphs with the gpu
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08442e00-035e-479c-a085-b7b29f8f283e · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended H2o: heavy-hitter oracle for efficient generative inference of large language models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54eaea14-c732-4b68-81c6-b33253925b5f · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Silod: A co-design of caching and scheduling for deep learning clusters
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e2f71d-3e78-405c-bfad-436a4651580c · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended Atom: Low-bit quantization for efficient and accurate llm serving
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4739a5d-6675-43b4-b467-68655052e63a · outbound
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended ${CONDA_PREFIX}
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
No inbound Pith citation observations are available.