Pith. sign in

Paper Citation Record · LEDGER

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding

As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2505.18758.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18758 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:32:20.014194Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy51
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40af015c-b43e-4cbc-81e1-a8cdcc552ddf · outbound

This paper cites Croci, Bo Li, Pashmina Cameron, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Croci, Bo Li, Pashmina Cameron, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:28.797709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:16.190894Z digest=sha256:3b7938ac99a8dfb5a6f0fcbd21bf92bc17064e75834fde30ee3ce0e939426737

Observation c51b9444-a97c-4766-a862-6c45b7db9d9c · outbound

This paper cites GPTVQ: The blessing of dimensionality for LLM quantization.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding GPTVQ: The blessing of dimensionality for LLM quantization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:28.654010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:16.241300Z digest=sha256:b5e56f87781d1f1cd58ca0a46f7870029f52edbade38e6dece31230cc543b41a

Observation e4eba4ad-b0b8-47bc-a21b-5627f02ccf0b · outbound

This paper cites ONNX: Open neural network exchange, 2019.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding ONNX: Open neural network exchange, 2019

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:28.545358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:16.358834Z digest=sha256:7440ac31fbbbe12f55900fb265946cbbc9a921b44f30b2d01dc78847175f56bc

Observation b84cb3ec-0010-460b-a007-edfffe9f8604 · outbound

This paper cites Understanding Entropy Coding With Asymmetric Numeral Systems (ANS): a Statistician's Perspective.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Understanding Entropy Coding With Asymmetric Numeral Systems (ANS): a Statistician's Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:16.426671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:16.426671Z digest=sha256:ace5500157294555a705021f2d1fbab1d8fe941348559cf157c4cad815e299c4

Observation c74d72d4-3de6-4958-b936-3e752748c3cb · outbound

This paper cites Bronstein, and Avi Mendelson.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Bronstein, and Avi Mendelson

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:28.368746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:16.511794Z digest=sha256:69480510b4bd522d21ff0cbb03cb0787b2292bea02dad0e8818044990acf5502

Observation dc197d84-ff2f-4ee5-8f6e-447a5cd3dde0 · outbound

This paper cites NNCodec: An Open Source Software Implementation of the Neural Network Coding ISO/IEC Standard.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding NNCodec: An Open Source Software Implementation of the Neural Network Coding ISO/IEC Standard

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:28.181108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:16.583499Z digest=sha256:d5ae31ca9f587033520b1d1b1635e00db0c38961ff8ce77e58d03d7d68ab5653

Observation fcc570e4-20a8-4a20-ac16-835174b597f0 · outbound

This paper cites EfficientQAT: Efficient Quantization-Aware Training for Large Language Models, October 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding EfficientQAT: Efficient Quantization-Aware Training for Large Language Models, October 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:27.999615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:16.640320Z digest=sha256:831ba6357c774520f9528c2b461f94bb1d76da2a0a866685945668a7fbc61140

Observation 98f4c2c9-2ba9-42fb-823e-77ad995f07db · outbound

This paper cites Bronstein, and Avi Mendelson.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Bronstein, and Avi Mendelson

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:27.831968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:16.729492Z digest=sha256:b949e5b5a4f0af474be43cb3b7377d3309bfed41e298a7747707f9352b1a7544

Observation ba0676f6-8c18-4488-aa96-6dc25e01e812 · outbound

This paper cites Universal Deep Neural Network Compres- sion.IEEE Journal of Selected Topics in Signal Processing, 14(4):715–726, May 2020.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Universal Deep Neural Network Compres- sion.IEEE Journal of Selected Topics in Signal Processing, 14(4):715–726, May 2020

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:27.628117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:16.835954Z digest=sha256:91dc032553d0a386a85dfce4456687f9d01a270d5ec047b754405d031eea9a89

Observation 2c2e90cb-97c4-455d-b258-6aac7cd8bb31 · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Imagenet: A large- scale hierarchical image database

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:16.916800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:16.916800Z digest=sha256:4dcadb4d3e364bbf97f217a30e514fb4627e6bb2d5df75c032aa3cee25f5e993

Observation 2c31a623-0c1b-40ab-8873-f52a6e545d96 · outbound

This paper cites GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale.Neural Information Processing Systems (NeurIPS), January 2022.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding GPT3.int8(): 8-bit Matrix Multiplication for Transformers at Scale.Neural Information Processing Systems (NeurIPS), January 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:27.442812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:16.971706Z digest=sha256:898df0b9e5563421450a09758381a5f09faca7297eaca53ff5cd83d3c914895c

Observation b61ccbcd-d1ed-45a4-96b9-10487a13ebdc · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs, May 2023.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding QLoRA: Efficient Finetuning of Quantized LLMs, May 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:27.226382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.038296Z digest=sha256:b83584f2c33c043e9ff4e0f7d5bf88cca8d9955c4939d68c5f049d70ad70f8e2

Observation 08b72a21-f583-4eba-b37f-99687fdc40e1 · outbound

This paper cites Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:27.050348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.101095Z digest=sha256:f15099d5d0f0de22b4adca420cf85bfc2d6a8ea13390d4ac1db3533d3012a47a

Observation 7cb383c7-385d-49cc-8c08-2913a1704ae8 · outbound

This paper cites The case for 4-bit precision: K-bit Inference Scaling Laws.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding The case for 4-bit precision: K-bit Inference Scaling Laws

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.869093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.167832Z digest=sha256:e5094684b953c01c0b6d0e697dd79b5e2b67383cddbbe017c0fc56dbe54a5bc9

Observation efdb50bb-466a-4a36-b625-99f1d7e75c22 · outbound

This paper cites STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs, August 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs, August 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.747986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.246176Z digest=sha256:7d7d3cb874101a6a7c1649531f995e980f0d94c77e940c6915da36aa80059cfa

Observation 50f7a082-196c-4b17-b9ab-92101700f1d1 · outbound

This paper cites The use of asymmetric numeral systems as an accurate replacement for huffman coding.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding The use of asymmetric numeral systems as an accurate replacement for huffman coding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.630506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.325605Z digest=sha256:6c6e731fbe9f602eb6bca59410de046d74ec5bb5027c132a24119be75afae130

Observation 966c4542-3a45-410b-93a2-27b59bc06d25 · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization, September 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Extreme Compression of Large Language Models via Additive Quantization, September 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.543235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.399995Z digest=sha256:28081d7a76a78ad0251cfe8e91491a5c31dd8707f0eb1422d5132f2ad0e44b36

Observation a7835023-4e6f-4a04-b6a4-30af7208a150 · outbound

This paper cites Optimal Brain Compression: A Framework for Accurate Post- Training Quantization and Pruning.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Optimal Brain Compression: A Framework for Accurate Post- Training Quantization and Pruning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.389323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.461915Z digest=sha256:3405201220489bacce7b808869be62d9887f1c45dcd511a7f7305115bab47862

Observation 97d2e906-e906-4fd3-a48c-fb5c6128f318 · outbound

This paper cites SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, March 2023.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, March 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.279903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.553273Z digest=sha256:eb181935341307aecabdbabb1c572b36e760581f69d7a52e82ebabc531fd7839

Observation 11561f6f-8600-4d3d-bc06-6a2290aac634 · outbound

This paper cites OPTQ: Accurate quan- tization for generative pre-trained transformers.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding OPTQ: Accurate quan- tization for generative pre-trained transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.200374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.633391Z digest=sha256:0ba943d14958ac25f6942d4e0988b0c498fc560baf8f827cb6ae2d72ca87a4ea

Observation a0f1d2a1-98dd-4ac6-8a23-f0827d426d36 · outbound

This paper cites Compression Scaling Laws:Unifying Sparsity and Quantization, February 2025.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Compression Scaling Laws:Unifying Sparsity and Quantization, February 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:26.118695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.690251Z digest=sha256:9128847ef304254c467ff981610c6f3800cdeee42726750e27a28072414fdb31

Observation 3805198f-2f45-4af3-a9e4-4bd38bd0eb37 · outbound

This paper cites MiniLLM: Knowledge Distillation of Large Language Models.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding MiniLLM: Knowledge Distillation of Large Language Models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.979526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.769125Z digest=sha256:3971294f2a3bd32d6c1f8add9deb562ee64bddc8107d4d6a550deeabef2e97a7

Observation e19cf533-d423-479f-9af8-f067ea7af6d8 · outbound

This paper cites an unresolved cited work.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:25.863479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.836993Z digest=sha256:24390bd437219b205ffdf6c4c303aff13a7e5c2302fca626268378fb39b4e4ab

Observation 29015da6-6314-4e3e-862d-e23f06486f50 · outbound

This paper cites NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks, October 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding NeuZip: Memory-Efficient Training and Inference with Dynamic Compression of Neural Networks, October 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.743356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.905015Z digest=sha256:dbf0b710522bfc37cd2d33000b8a5e563c26b806b4c70f1010706ea6cf890de3

Observation b56910ea-1476-4b58-98fd-403b3e979413 · outbound

This paper cites Hassibi, D.G.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Hassibi, D.G

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.630947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:17.992323Z digest=sha256:4bf76c2bf6ca03b9d0d716341e7dbb474aff942a84b39fe449a5d22ce1987ba1

Observation c06c72ef-c27c-4b0d-9d36-b769f6272758 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Deep Residual Learning for Image Recognition

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.469500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:18.053367Z digest=sha256:1d1e73f3dd382b78e7d89a329074c39b7cda4e2413dcfea9b32e364bbc741206

Observation f52202bb-ac7e-4232-94c8-eca58327caa6 · outbound

This paper cites Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experi- ences.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experi- ences

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.359504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:18.110251Z digest=sha256:d671b7b3bbd611899910526fe9f6fdc045ee88f17df78560eb24ff15ccaf276f

Observation 04780eec-bda4-4ae5-a8b2-4e05c0ba1871 · outbound

This paper cites Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.189720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:18.187867Z digest=sha256:9c19a36378f4c7572718ccb40b1126b9f54a1bf7e607c9eb47e1035eb0306068

Observation 07cebcb8-7ea8-4c33-a4ed-fc6ea06072be · outbound

This paper cites Le, and Hartwig Adam.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Le, and Hartwig Adam

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:25.077667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:18.289283Z digest=sha256:35d17e30e65d0f705ac9dc44263b6720f915e4f04306f34f532e69cbadbcd50e

Observation 9a27fcbb-35a0-41ce-8a89-2b75d6c075cc · outbound

This paper cites Accurate Post Training Quantization With Small Calibration Sets.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Accurate Post Training Quantization With Small Calibration Sets

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.916723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:18.361634Z digest=sha256:e3a355d630d0faa42df8a940c56be553e296cd0ba50185abef00f9c6973316da

Observation cb6914b4-e72e-44da-912b-88c05c479446 · outbound

This paper cites Mahoney, and Kurt Keutzer.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Mahoney, and Kurt Keutzer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.770296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:18.411007Z digest=sha256:0ee144e8530d9fc2cea4f9099956a53026ed549a9c7a997f4da6db940cfd4fec

Observation 7b2c153c-0e14-4d83-8602-23851cdba37d · outbound

This paper cites Aksu, Miska M.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Aksu, Miska M

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.589078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:18.463769Z digest=sha256:b9f7d97b2c513480523423434840285cc4bd66ba526b68b702288f7f7826f917

Observation 678cbc9e-cf74-4e3f-9650-dce56d55ccca · outbound

This paper cites Adaptive weight compression for memory-efficient neural networks.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Adaptive weight compression for memory-efficient neural networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.467567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:18.535752Z digest=sha256:160b616d3f53db7728bbef3e137ed8e53450bc3ebd97ad6d87376c6bd601086c

Observation 0c3747bb-9ab5-48b5-95a3-dcb1e6d96250 · outbound

This paper cites Cifar-10 (canadian institute for advanced research).

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Cifar-10 (canadian institute for advanced research)

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:18.611847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:18.611847Z digest=sha256:5469facc8217609cfca7439a9ed3419dd2e5882a04eb099c5880d9af05d08a55

Observation fedf0c03-7b2e-4efa-b589-abb654016c75 · outbound

This paper cites Energy-Efficient Model Compression and Splitting for Collaborative Inference Over Time-Varying Channels.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Energy-Efficient Model Compression and Splitting for Collaborative Inference Over Time-Varying Channels

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.309768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:18.661478Z digest=sha256:3fca73810eeecdebb7ef9f382087b8fd7f7e037f25f0aab19943d3a7a0c76a76

Observation f19237d7-95f6-4def-ad28-268ca16b2138 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Gonzalez, Hao Zhang, and Ion Stoica

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:18.725195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:18.725195Z digest=sha256:636f3c233af72cbad8e59c122c2ad3545c06749587164973cd5a3031c123b9b2

Observation 49b08133-2996-49c4-9b5a-a3a463600a6f · outbound

This paper cites Memory Efficient Optimizers with 4-bit States.Advances in Neural Information Processing Systems, 36:15136–15171, December 2023.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Memory Efficient Optimizers with 4-bit States.Advances in Neural Information Processing Systems, 36:15136–15171, December 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.193718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:18.777724Z digest=sha256:d58581d71ae12a9f167f3d479b44508a0a69ccee4c9ea48c40352f3415b8246e

Observation ea0ff6d2-256e-4297-bb5e-61e19044e52c · outbound

This paper cites PENNI: Pruned Kernel Sharing for Efficient CNN Inference.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding PENNI: Pruned Kernel Sharing for Efficient CNN Inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:24.038562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:18.879989Z digest=sha256:890b4e728d6161b37c5f9bad7321f91d7556350fd0c73dc04c7f0fe64b83576a

Observation acbd46d8-c6fd-4c0b-9d62-f3bfd2794fb1 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration, April 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration, April 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:23.928089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:18.942144Z digest=sha256:06ed72ef55c27962cd2257c464a49753ee6ab676dd6fca3a32f5d746f465e716

Observation 47fc9a82-5229-4550-931c-928db5094e91 · outbound

This paper cites Cambridge university press, 2003.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Cambridge university press, 2003

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:23.749967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.009752Z digest=sha256:d079c7bfe8e7d006aae57fdcce8e8a5371af4805e6b2f667c030f2c7602aaaa3

Observation 2812bc00-7f8b-45f5-a989-3b49903e3ae6 · outbound

This paper cites Range encoding: an algorithm for removing redundancy from a digitised message.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Range encoding: an algorithm for removing redundancy from a digitised message

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:23.570244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.078862Z digest=sha256:d687886e0e34e792fd149fae26b600dfe1be0f878813cb655858872a53e8dca2

Observation b6c3659e-7fcb-4b0e-baf2-a7b7ec128f50 · outbound

This paper cites Up or Down? Adaptive Rounding for Post-Training Quantization.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Up or Down? Adaptive Rounding for Post-Training Quantization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:23.422163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.142752Z digest=sha256:ce97a3c5778013b322eaebe2a425a6332a3ef766055228b723995ab38429c9b4

Observation 82cadcdb-dc92-4d6c-a9ed-4b73d51897fe · outbound

This paper cites A White Paper on Neural Network Quantization, June 2021.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding A White Paper on Neural Network Quantization, June 2021

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:23.259329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.217951Z digest=sha256:a8f3296fbd0ec07b060b9fb16587959742cc1081517c9da33ebdb423bdfee8ac

Observation 83cc05eb-0e6b-4970-9389-f3788a22d706 · outbound

This paper cites Kübler, Jiaji Huang, Matthäus Kleindessner, Jun Huan, V olkan Cevher, Yida Wang, and George Karypis.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Kübler, Jiaji Huang, Matthäus Kleindessner, Jun Huan, V olkan Cevher, Yida Wang, and George Karypis

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:23.089866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.307019Z digest=sha256:f42a78b16a3807963066972de003c021d6c491862dca504936e24e26f7f072be

Observation 0a6ff35c-4c2c-4566-8937-aa287d1d91a3 · outbound

This paper cites PhD thesis, Stanford University CA, 1976.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding PhD thesis, Stanford University CA, 1976

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:22.904461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.342907Z digest=sha256:13bfaeddf052e3783900f3ed778b16c7be7fd5c82315a7477c7c94a2a26700e3

Observation 1bd47c1e-0df3-47ac-8a80-3fd74b0e4fb3 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library, December 2019.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding PyTorch: An Imperative Style, High-Performance Deep Learning Library, December 2019

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:22.735567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.398920Z digest=sha256:cc1033909c80f1e94a6621a5906a943337ee7e0dc83e9af66dec9f60252424f9

Observation 8f6cb52b-a8e6-426b-84ce-e94e6f2d0d79 · outbound

This paper cites Accurate LoRA-Finetuning Quantization of LLMs via Information Retention, May 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Accurate LoRA-Finetuning Quantization of LLMs via Information Retention, May 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:22.532664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.468784Z digest=sha256:2291d96d304954844ed1f317446bf59959bdf3bf1f333b6bcfa2e7a7a63dfbf1

Observation d191c42e-5b74-4010-8c30-85b192be3aa7 · outbound

This paper cites Arithmetic coding.IBM Journal of research and development, 23(2):149–162, 1979.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Arithmetic coding.IBM Journal of research and development, 23(2):149–162, 1979

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:22.305396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.525797Z digest=sha256:24a93d58bb975946eb57cec40b60db4808752f39ee416c8d533aeca7a5ad94f2

Observation 8e4c1547-574d-463d-8438-b1084a7d111c · outbound

This paper cites an unresolved cited work.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:32:22.087711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.581936Z digest=sha256:1dd23b1a5d4f2ca80bd2bd3060b48de9abb4adba62e2fb50ed043a7347077ad5

Observation 6a844d03-f3cc-46d2-b71b-58b23dce5c2c · outbound

This paper cites an unresolved cited work.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:19.697752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:19.697752Z digest=sha256:332fd1ceb241d3ce425a99b211dd9c4e14dea37b7b4724e5bedc5e57d6300d68

Observation dfbe890b-d91c-4e24-ab30-706c230288ce · outbound

This paper cites FlexGen: High-Throughput Generative Infer- ence of Large Language Models with a Single GPU.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding FlexGen: High-Throughput Generative Infer- ence of Large Language Models with a Single GPU

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:21.889948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.716104Z digest=sha256:044041fd2950503a31a4c2e8472b18375cc04db049ea2f443c73364fa81f08ea

Observation a30e67ad-c023-49aa-b021-828605eae64c · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:21.699086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.719781Z digest=sha256:e80da649e7b6b57df8ddf4a89331b7b55c81730cd037a476ee1755b959db72b2

Observation 72310331-4608-482e-aaa2-4249c9bdf1dc · outbound

This paper cites Zico Kolter.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Zico Kolter

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:21.525600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.756556Z digest=sha256:b44635436be14d72f781715b810ceb0f8c43a901b5e0850725b1a9260b2f3cb8

Observation 7748fb8f-49b8-4b7e-9827-2715c1895a00 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, June 2024.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, June 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:21.316087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.797300Z digest=sha256:191d8736c62050bd8c8eb721ea3ecd04431f42ffc740574a763c2d4d11d4cd99

Observation caf8093d-eec4-4c4a-9e22-05a1ea09391a · outbound

This paper cites DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks.IEEE Journal of Selected Topics in Signal Processing, 14(4):700–714, May 2020.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks.IEEE Journal of Selected Topics in Signal Processing, 14(4):700–714, May 2020

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:21.114329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.843328Z digest=sha256:c0f8a99c57b9b286b5d6de634fdb298df420125983137566e0a4fac53d0d9e4d

Observation 11aceea7-f138-4598-bd53-90a71d68a4c1 · outbound

This paper cites Compact and computationally efficient representation of deep neural networks.IEEE Transactions on Neural Networks and Learning Systems, 31(3):772–785, 2020.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Compact and computationally efficient representation of deep neural networks.IEEE Transactions on Neural Networks and Learning Systems, 31(3):772–785, 2020

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:20.902614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.886008Z digest=sha256:ddbc95ce86092038f89d66eb3bfa4e5fa49473eb45430f016edc1107a3b1a3ed

Observation 085bb2b5-222b-4ca7-90dd-5f27afcb3923 · outbound

This paper cites Variational bayesian quantization.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding Variational bayesian quantization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:20.620742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:19.927150Z digest=sha256:55d4ad33dc3ddc9c6c62e4544897e11828a7fb5eb52abe9df7a6f48dbe625170

Observation cc116959-51fc-4bd1-8ffe-bdb6979cacb7 · outbound

This paper cites optimally compensate.

Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding optimally compensate

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:32:20.341654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:32:20.014194Z digest=sha256:45fba1ff99e148e4327e0101dff6b647f8417fc5427aafd14676fe82275d53d0

Pith citing papers

No inbound Pith citation observations are available.