Pith. sign in

Paper Citation Record · LEDGER

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

As of 18 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 7 inbound Pith citation observations for arXiv:2502.00922.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00922 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:18:51.559897Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:34:08.346536Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T16:25:49.759745Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact9
  • verified fuzzy14
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b6211d1-20cb-4f17-aac3-77fa40756c16 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.401236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.401236Z digest=sha256:31521cf1d9cb60debd5c1738fe2314742e0554bc3faff7570ba01699ba99ee28

Observation 497d3a6e-d881-4038-a9ba-c831f6a46599 · outbound

This paper cites B., Muralimanohar, N., Shafiee, A., and Srinivas, V.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference B., Muralimanohar, N., Shafiee, A., and Srinivas, V

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.153310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.406018Z digest=sha256:0f1eb35913a0a19d7bce9f693bfd6fb803bc6425d2e6c5104c5000b83a9b2f71

Observation 39175179-70b4-42d0-a441-12f5585cd0d6 · outbound

This paper cites S., and Sze, V.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., and Sze, V

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.142123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.410365Z digest=sha256:fc249db8346faad75c1018b886a8452219b692dcb60c0914cc65218cd55967a9

Observation 0df24379-621e-4e64-bbd0-786c816a7e30 · outbound

This paper cites E., Stoica, I., and Xing, E.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference E., Stoica, I., and Xing, E

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.414132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.414132Z digest=sha256:b849e163c10a167a53990db48101b35a5d8155978be8a2663e841d6a6b25e77b

Observation 186dc68c-d741-4e6d-9d13-4ae72cd31c86 · outbound

This paper cites B., O’Connor, M., Erez, M., Pool, J., Nellans, D., and Keckler, S.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference B., O’Connor, M., Erez, M., Pool, J., Nellans, D., and Keckler, S

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.123443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.417709Z digest=sha256:782d3b0939addc8dbb259f7396f99b57c892d69bb8bc9e31f14a2f91ec9a4a93

Observation 5d581117-3a82-41e7-92a0-6d96e4d5d955 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.421395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.421395Z digest=sha256:04275c6c4394c564241d805718df5541c1299388dd20fc27429484091c1bb20e

Observation d98a464e-7071-45a2-a9d3-a9a054ff5b11 · outbound

This paper cites an unresolved cited work.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.425873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.425873Z digest=sha256:29a858eea7cc33976b4b0d843d88aa850505ef2b99906d5bce42e40e4a216aa8

Observation db0ee736-faeb-4c5d-bb65-6c7abe8b9028 · outbound

This paper cites The Llama 3 Herd of Models.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.429444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.429444Z digest=sha256:e0c51d943eefcaecffd8ec99906f1408ae56d39b295f4f96cc469a5f1cf479a2

Observation f0da55fd-4ea4-4a84-965c-c9b1d98d74dd · outbound

This paper cites Accuracy is Not All You Need.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Accuracy is Not All You Need

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.432553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.432553Z digest=sha256:49c5351e2b39ee4adc2270f548f2b5bea1e1fbb8449918e61ffa058c27d36dae

Observation 88038bfb-7d00-4cc3-a2d0-d1b65072ce7d · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.436104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.436104Z digest=sha256:4c03bd91ce7b06f005e7521332d19597900239282844bcbc636d308d2dcfea26

Observation 39d89a7e-e826-4183-bfae-52ca8ac13d3f · outbound

This paper cites Does reduced precision hurt? Blog post, 2024.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Does reduced precision hurt? Blog post, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.104745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.440045Z digest=sha256:e94a513aed0364fccf4017a75baa531fd39b17c5889c529a94fcfd4532eaa39e

Observation 52288120-1fcf-4d33-8116-68bb996df7f7 · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.443586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.443586Z digest=sha256:b25b8aaf077fdaea35f3e984b2c91607a14ddca5df39e92112a0a0feea8d58de

Observation c5cb69f3-2a45-4c8c-a276-0dcbc753bf1a · outbound

This paper cites NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.447148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.447148Z digest=sha256:3ee5debab511e2ed2ea12136eb7a82ea86239f6df1000b5fb491051e0a9aebcf

Observation 24224d4b-9e44-4677-86db-9cca85038bbb · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Measuring Massive Multitask Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.450902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.450902Z digest=sha256:de93408a9cd2e624a9f6b9d80ba610c525172f3e9a6ba29b92f95d2d4be2c96e

Observation 13f2b916-e420-45f9-86c6-38e9454a6290 · outbound

This paper cites ZipNN: Lossless Compression for AI Models.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference ZipNN: Lossless Compression for AI Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.455245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.455245Z digest=sha256:8197a569177e084c6067a33572a88f5d004f834fcaba72b7a7cb1f65e7bb4203

Observation faf66d4a-73f0-4672-bcef-a72f4ce154ff · outbound

This paper cites Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.459650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.459650Z digest=sha256:e8b652231fc7cf64d4ab051e35e21c786fd32fa40f4814e94a84cb24cacb6685

Observation 9b83baf2-131f-4e76-bdd9-97215c84da27 · outbound

This paper cites S., Choi, Y., Kim, C., Kim, Y., Yu, H., Abdel-Aziz, H., Park, J.-S., Lee, H., Lee, D., Kim, M.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Choi, Y., Kim, C., Kim, Y., Yu, H., Abdel-Aziz, H., Park, J.-S., Lee, H., Lee, D., Kim, M

Reference 17

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:53.851033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.464502Z digest=sha256:8cb4d8f91a250f5b0b6a8fe7ff15e028e5e5871abbdda88715e3a350f3100873

Observation c5d10bca-e45a-493e-856c-d855a22457bf · outbound

This paper cites G., Zimmer, B., Dally, W.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference G., Zimmer, B., Dally, W

Reference 18

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:53.677618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.468359Z digest=sha256:604d220031f4184a29851d838bb06369af782df1143358667b709b1678a85cfb

Observation 06fd14a6-0be5-47d4-8133-13c3a327a963 · outbound

This paper cites Bit-plane compression: Transforming data for better compression in many-core architectures.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Bit-plane compression: Transforming data for better compression in many-core architectures

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.093391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.472747Z digest=sha256:dec76afad6a8172854215983fde6b7f6815695637fa588c83667b4dc1b963508

Observation 544b551d-9f22-46a9-92d3-77aac1f6f33c · outbound

This paper cites Cerebras architecture deep dive: First look inside the hardware/software co-design for deep learning.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Cerebras architecture deep dive: First look inside the hardware/software co-design for deep learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.082633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.476815Z digest=sha256:9f5beb8a4bed2e9c33e7a363777b942522c777a6f92c6f6115bd9ac8fdde9ba1

Observation 89a7859f-4c67-47aa-8cbf-fcdf1e3a6091 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.480864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.480864Z digest=sha256:6dcb86a4894525a9b833c4bb00caa148239916748ed589aef9695f5d35509b6d

Observation 76ad177f-f306-4596-b79c-60b751e0cc29 · outbound

This paper cites How Does Quantization Affect Multilingual LLMs?.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference How Does Quantization Affect Multilingual LLMs?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.484484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.484484Z digest=sha256:ca27deaeb1b0d53db4dd17994a076bde5bfbd57d347ab6de0f936a2c1df46a12

Observation 2c3a6913-b377-46a5-acec-ed3e0f8d94b7 · outbound

This paper cites and Mutyam, M.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference and Mutyam, M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.065251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.488227Z digest=sha256:62c9f5d10831ffa3f2698140a778f6174cf52301f1b92649c809d757946f3425

Observation eccdb6a4-6572-4238-b29b-1f48900cdbc1 · outbound

This paper cites S., Chen, Y.-H., Ying, V.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Chen, Y.-H., Ying, V

Reference 24

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:53.514198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.491858Z digest=sha256:474bcb2b457ecfa14cfb6b2842e73975906f10fe57dfb1edca5176b4b1276e42

Observation dbd0fc17-7f86-4d95-ba8a-5b423af1eec2 · outbound

This paper cites Arrayflex: A systolic array architecture with configurable transparent pipelining.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Arrayflex: A systolic array architecture with configurable transparent pipelining

Reference 25

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:53.348934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.495866Z digest=sha256:07b567a84a5a9d4284ecac52e010d10034320424ca3619deb508133faf500d3b

Observation 5ca734ab-e616-4df4-83f6-da58ce0372d8 · outbound

This paper cites M., Zhu, Y., Whatmough, P., Mattina, M., and Krishna, T.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference M., Zhu, Y., Whatmough, P., Mattina, M., and Krishna, T

Reference 26

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:53.196253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.499709Z digest=sha256:c469b6cf64cbf4c11d0a4fca993f8378266cec99c1e080f9af9100c68de33dfd

Observation a4da60c4-3f85-4183-8837-66e36c6d6340 · outbound

This paper cites S., Reagen, B., Wei, G.-Y., and Brooks, D.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Reagen, B., Wei, G.-Y., and Brooks, D

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.054810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.503608Z digest=sha256:afbc9de5aa6c8a7b051d6b7419b0e4cc2953d6ac24d9b421fff5daac8a57c954

Observation 48ac90d6-2c7c-4844-b235-e3c4b65aed84 · outbound

This paper cites S., Clemons, J., Venkatesan, R., Zimmer, B., Fojtik, M., Jiang, N., Keller, B., Klinefelter, A., Pinckney, N., Raina, P., Tell, S.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Clemons, J., Venkatesan, R., Zimmer, B., Fojtik, M., Jiang, N., Keller, B., Klinefelter, A., Pinckney, N., Raina, P., Tell, S

Reference 28

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:53.032808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.507466Z digest=sha256:db81dbb83a943a29f63daff53f71658ee378ca05239a7eed441294116b2d8b56

Observation 6ddde6e5-8229-4f70-9da0-ad8bb8de5360 · outbound

This paper cites The nvidia deep learning accelerator.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference The nvidia deep learning accelerator

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.043556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.510907Z digest=sha256:894a812580deee3704ee65c3af6c4a61f077ed5e1a0f7310aa01578353a7db52

Observation f3e699c8-7458-4341-bb9c-e14c2b39b7e5 · outbound

This paper cites Q., Gomez, J., Khwa, W.-S., Sarwar, S.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Q., Gomez, J., Khwa, W.-S., Sarwar, S

Reference 30

Resolution
verified exact
doi, observed 2026-08-09T17:18:51.593843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.514086Z digest=sha256:c5b3c5f63cc51f187c4140820aec79c46342fe28d9b53c3cf0d855d3e2dfe75a

Observation 84d6e910-77b1-4637-8770-53fafc06b513 · outbound

This paper cites Google coral edge tpu board vs nvidia jetson nano dev board hardware comparison, 2020.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Google coral edge tpu board vs nvidia jetson nano dev board hardware comparison, 2020

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.032560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.518081Z digest=sha256:73a5fe06c6bea8d0264af559c5ca0cb85fd386440275e9316f985661b4524179

Observation 28fef51e-1b3b-4d36-9759-1671562d4210 · outbound

This paper cites Llama3.1 model quality evaluation: Cerebras, groq, sambanova, together, and fireworks.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Llama3.1 model quality evaluation: Cerebras, groq, sambanova, together, and fireworks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.020042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.521058Z digest=sha256:2824cf05e718c960d51cd0c8c079c7dc1f66df4b4f31a4f58e018e5278d4213c

Observation 13900705-0c5e-4824-b43d-9cf2bc913a95 · outbound

This paper cites S., Wang, M., Clemons, J., Dai, S., Fojtik, M., Keller, B., Klinefelter, A., Pinckney, N., Raina, P., Zhang, Y., Zimmer, B., Dally, W.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference S., Wang, M., Clemons, J., Dai, S., Fojtik, M., Keller, B., Klinefelter, A., Pinckney, N., Raina, P., Zhang, Y., Zimmer, B., Dally, W

Reference 33

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:52.752480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.524265Z digest=sha256:aa5b95df5d2d87fe9d96a415f3ec56a14917bd5972aea31352d764878bd88f41

Observation f6d8a82c-08b5-4783-b062-57f30f0e6988 · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Spatten: Efficient sparse attention architecture with cascade token and head pruning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:54.008849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.527838Z digest=sha256:806cc81f4366314c2f1a4aa982fdb8c9de06ba261170db8520817e8ae853f71d

Observation 4aa3a4ab-ec52-4e97-a1b2-a5744c3ace34 · outbound

This paper cites an unresolved cited work.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-09T17:18:53.997080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.531468Z digest=sha256:7aad07e078fad2f51b4642f634c9f5921626a343b0c3792eb16e2c6211c5d5f2

Observation 2add9484-be78-4852-8edb-f4ea43b8cc38 · outbound

This paper cites The roofline model: A pedagogical tool for program analysis and optimization.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference The roofline model: A pedagogical tool for program analysis and optimization

Reference 36

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T17:18:51.788060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.535139Z digest=sha256:32296e1b9c006074a2ff94d2ee368484f6773b94e4b070dcc8a763bd20639eb3

Observation f05c421a-88d6-45f8-a4cf-923223872d26 · outbound

This paper cites Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.538529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.538529Z digest=sha256:1bdccd06d3f680dd69e210589a9b2d5ffbe04e1aec8c80e18bf7ab5ab1e1034d

Observation 294af4e8-a8bf-4134-b83c-e97dec1d7de7 · outbound

This paper cites Qwen2.5 Technical Report.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.542155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.542155Z digest=sha256:ae31d8e51c6de97ec119a7d2afd79e8b371c524ef40b646b29fe9ffe2f42f25e

Observation 12ec46ef-eeb0-4e0b-9780-4187f5e6b99f · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:53.985520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.545813Z digest=sha256:024d4b4883bf70118352568e31e68662e4afa92e889f01bf8a9555d38842d4e0

Observation a00658f7-e833-4828-ae7b-f6cbf34132b1 · outbound

This paper cites 15.1 a 0.795 fj/bit physically-unclonable function-protected tcam for a software-defined networking switch.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference 15.1 a 0.795 fj/bit physically-unclonable function-protected tcam for a software-defined networking switch

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:18:53.974529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T17:18:51.549387Z digest=sha256:2654c8f307374cdeaf2297641cdae92f92b350791c38a5f298a07482f1d303d8

Observation a3fa07f6-ec20-41a8-b086-fd63779593c6 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference OPT: Open Pre-trained Transformer Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.552602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.552602Z digest=sha256:476c7cb964ae260eb1dd6f544334feef766bee621e62714697ca09efb7bfcd09

Observation 8a9fe5e5-78ad-43ad-a1e4-79127df8e422 · outbound

This paper cites Catastrophic Failure of LLM Unlearning via Quantization.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference Catastrophic Failure of LLM Unlearning via Quantization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.556165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.556165Z digest=sha256:340b9065ea74155c8a7490b30770ddf78b881702462950a6f02e51f378e68dde

Observation 3a7a4bc4-b2e2-4205-a348-932c9ca58099 · outbound

This paper cites write newline.

Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference write newline

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T17:18:51.559897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:18:51.559897Z digest=sha256:e3c10cf83e813059abf066bad6ef074bb1b8602d94017a5e369f5a03d5483de7

Pith citing papers

Observation e1dfcaf3-3b2c-476f-8f79-2ca47a43a852 · inbound

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling cites this paper.

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.195963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T07:43:24.937953Z digest=sha256:4b7a2e12a4879b83815b075eb98d3be92c3854f13a0e06a52af952a77bc1a0ac

Observation fb484879-642c-4585-888e-3487ebe49ece · inbound

ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs cites this paper.

ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:58:03.058558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T21:56:38.684104Z digest=sha256:cdbb8273589c09a7f2df65640ee7c804e8d1da4c66ee20dd5a46c6513775ecf4

Observation 403395a6-6800-428f-b105-edc52095be6c · inbound

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving cites this paper.

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:26:08.754295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T16:59:58.809897Z digest=sha256:bb3a95be1d9d1089b22e0993d34d9cf4fb29aa5293e419f2958301d8231e38fd

Observation 8f590228-6923-4cd6-9777-d30678fc3d88 · inbound

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving cites this paper.

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:21:23.650906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T01:15:06.859289Z digest=sha256:7a9742e6c51ce872bd93346d0b479c255890cde3d0cb048474501cf9596278e7

Observation 983e8689-bffa-4257-8f33-69b430b335a2 · inbound

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving cites this paper.

SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:15:45.375310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T00:58:28.798369Z digest=sha256:8d446d686b29ec44790e5bed76654d2c5993fcfab96f128751827f78c3778fb5

Observation d2d39926-d8ed-4696-a386-0251e0c1bfae · inbound

Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding cites this paper.

Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:25:49.761209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T16:19:54.910430Z digest=sha256:dd9e5da1a3c1f2c77524597ba656c075d25fef244e648d6556aa1519ce2aeb80

Observation 99a6d1d4-d55e-4205-8c80-fc89ebb29591 · inbound

Lossless Tensor Compression as Program Synthesis cites this paper.

Lossless Tensor Compression as Program Synthesis Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T13:34:08.346536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:34:08.346536Z digest=sha256:8948a498dd7d1024a4186b76f29ffd20152f2dabb3a6258709870298399263e1