Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T17:18:51.559897Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 7 inbound Pith citation observations for arXiv:2502.00922.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T17:18:51.559897Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T13:34:08.346536Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T16:25:49.759745Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0b6211d1-20cb-4f17-aac3-77fa40756c16 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 497d3a6e-d881-4038-a9ba-c831f6a46599 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference B., Muralimanohar, N., Shafiee, A., and Srinivas, V
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 39175179-70b4-42d0-a441-12f5585cd0d6 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., and Sze, V
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0df24379-621e-4e64-bbd0-786c816a7e30 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference E., Stoica, I., and Xing, E
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 186dc68c-d741-4e6d-9d13-4ae72cd31c86 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference B., O’Connor, M., Erez, M., Pool, J., Nellans, D., and Keckler, S
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5d581117-3a82-41e7-92a0-6d96e4d5d955 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d98a464e-7071-45a2-a9d3-a9a054ff5b11 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db0ee736-faeb-4c5d-bb65-6c7abe8b9028 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0da55fd-4ea4-4a84-965c-c9b1d98d74dd · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Accuracy is Not All You Need
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88038bfb-7d00-4cc3-a2d0-d1b65072ce7d · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39d89a7e-e826-4183-bfae-52ca8ac13d3f · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Does reduced precision hurt? Blog post, 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 52288120-1fcf-4d33-8116-68bb996df7f7 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5cb69f3-2a45-4c8c-a276-0dcbc753bf1a · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24224d4b-9e44-4677-86db-9cca85038bbb · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Measuring Massive Multitask Language Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13f2b916-e420-45f9-86c6-38e9454a6290 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference ZipNN: Lossless Compression for AI Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf66d4a-73f0-4672-bcef-a72f4ce154ff · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b83baf2-131f-4e76-bdd9-97215c84da27 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Choi, Y., Kim, C., Kim, Y., Yu, H., Abdel-Aziz, H., Park, J.-S., Lee, H., Lee, D., Kim, M
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c5d10bca-e45a-493e-856c-d855a22457bf · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference G., Zimmer, B., Dally, W
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 06fd14a6-0be5-47d4-8133-13c3a327a963 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Bit-plane compression: Transforming data for better compression in many-core architectures
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 544b551d-9f22-46a9-92d3-77aac1f6f33c · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Cerebras architecture deep dive: First look inside the hardware/software co-design for deep learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 89a7859f-4c67-47aa-8cbf-fcdf1e3a6091 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76ad177f-f306-4596-b79c-60b751e0cc29 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference How Does Quantization Affect Multilingual LLMs?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c3a6913-b377-46a5-acec-ed3e0f8d94b7 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference and Mutyam, M
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eccdb6a4-6572-4238-b29b-1f48900cdbc1 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Chen, Y.-H., Ying, V
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dbd0fc17-7f86-4d95-ba8a-5b423af1eec2 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Arrayflex: A systolic array architecture with configurable transparent pipelining
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5ca734ab-e616-4df4-83f6-da58ce0372d8 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference M., Zhu, Y., Whatmough, P., Mattina, M., and Krishna, T
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a4da60c4-3f85-4183-8837-66e36c6d6340 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Reagen, B., Wei, G.-Y., and Brooks, D
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 48ac90d6-2c7c-4844-b235-e3c4b65aed84 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Clemons, J., Venkatesan, R., Zimmer, B., Fojtik, M., Jiang, N., Keller, B., Klinefelter, A., Pinckney, N., Raina, P., Tell, S
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6ddde6e5-8229-4f70-9da0-ad8bb8de5360 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference The nvidia deep learning accelerator
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3e699c8-7458-4341-bb9c-e14c2b39b7e5 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Q., Gomez, J., Khwa, W.-S., Sarwar, S
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 84d6e910-77b1-4637-8770-53fafc06b513 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Google coral edge tpu board vs nvidia jetson nano dev board hardware comparison, 2020
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 28fef51e-1b3b-4d36-9759-1671562d4210 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Llama3.1 model quality evaluation: Cerebras, groq, sambanova, together, and fireworks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 13900705-0c5e-4824-b43d-9cf2bc913a95 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Wang, M., Clemons, J., Dai, S., Fojtik, M., Keller, B., Klinefelter, A., Pinckney, N., Raina, P., Zhang, Y., Zimmer, B., Dally, W
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f6d8a82c-08b5-4783-b062-57f30f0e6988 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Spatten: Efficient sparse attention architecture with cascade token and head pruning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4aa3a4ab-ec52-4e97-a1b2-a5744c3ace34 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2add9484-be78-4852-8edb-f4ea43b8cc38 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference The roofline model: A pedagogical tool for program analysis and optimization
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f05c421a-88d6-45f8-a4cf-923223872d26 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294af4e8-a8bf-4134-b83c-e97dec1d7de7 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Qwen2.5 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12ec46ef-eeb0-4e0b-9780-4187f5e6b99f · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a00658f7-e833-4828-ae7b-f6cbf34132b1 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference 15.1 a 0.795 fj/bit physically-unclonable function-protected tcam for a software-defined networking switch
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a3fa07f6-ec20-41a8-b086-fd63779593c6 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference OPT: Open Pre-trained Transformer Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a9fe5e5-78ad-43ad-a1e4-79127df8e422 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Catastrophic Failure of LLM Unlearning via Quantization
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a7a4bc4-b2e2-4205-a348-932c9ca58099 · outbound
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference write newline
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1dfcaf3-3b2c-476f-8f79-2ca47a43a852 · inbound
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fb484879-642c-4585-888e-3487ebe49ece · inbound
ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 403395a6-6800-428f-b105-edc52095be6c · inbound
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8f590228-6923-4cd6-9777-d30678fc3d88 · inbound
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 983e8689-bffa-4257-8f33-69b430b335a2 · inbound
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d2d39926-d8ed-4696-a386-0251e0c1bfae · inbound
Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 99a6d1d4-d55e-4205-8c80-fc89ebb29591 · inbound
Lossless Tensor Compression as Program Synthesis Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.