Pith. sign in

Paper Citation Record · LEDGER

QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 74 inbound Pith citation observations for arXiv:2402.04396.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04396 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 74 of 74 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:43:49.765633Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 28701e4b-a19a-4ee3-8cbf-e3743c6ad444 · inbound

The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits cites this paper.

The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:11:43.597896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:11:43.559035Z digest=sha256:4f0d4c42b5b465db5e77ce115764f0fa4b96e182433d2a87dc5131e94d0625d8

Observation a9f02f46-12f3-4abb-b454-9281721a0e04 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 216

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.271712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:5a522c0aa11047ef83d861c9c9317f0d15e9a37736e078f876916730ce7d261e

Observation 6952bc00-5e45-44c9-bd43-e9a65c78c65e · inbound

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision cites this paper.

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:45:36.456215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T19:45:36.337956Z digest=sha256:b6f88757bb7c4fd42507a9818076485209829aa067108f8aebf7f4e1089a67c8

Observation f6311aaf-0b51-436d-9fa9-18924762d0fc · inbound

Diffusion Product Quantization cites this paper.

Diffusion Product Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:49:45.779731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:49:45.779731Z digest=sha256:daa6f47e0d49f921fd218afc41f3f87f514149008fbebaa3dbf0d1dba4400ec0

Observation 46eab912-d74a-47a8-b1b4-90ff15b5743c · inbound

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format cites this paper.

Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T13:46:01.246476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:46:01.246476Z digest=sha256:084d9d3a17eb00c180a7ffd5e772ef1691e472a3ebbb833f2575588d5e8355be

Observation 17d524fe-736b-4d72-921d-460b2db5530e · inbound

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem cites this paper.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.231141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.231141Z digest=sha256:c4dbff397e4b5e471dd31f94c2aa967aea3edb9975ef2381fcd29ba10851570f

Observation be99e24f-3f5e-45d4-a5c0-0b4fa7646747 · inbound

DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation cites this paper.

DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:23.916899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:20:23.916899Z digest=sha256:dca8f29a4963e53843cfcecdcd0a8ceed3c418f1a0a494698ad49df5903a3f03

Observation ece6ac9f-a26a-4754-8d2f-c35e67c3dbb9 · inbound

Low-Rank Correction for Quantized LLMs cites this paper.

Low-Rank Correction for Quantized LLMs QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T18:35:26.820572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:35:26.820572Z digest=sha256:81ab806d144e902fce5deedf1e31db2f252fb8477c8b6b349418f44d6e676ac7

Observation 5e9c0803-25cc-40ff-8915-b81ef49d95d1 · inbound

HadaCore: Tensor Core Accelerated Hadamard Transform Kernel cites this paper.

HadaCore: Tensor Core Accelerated Hadamard Transform Kernel QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:39.753354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:39.753354Z digest=sha256:52cfdea04e5a97fac962c86771e43bd753a011ed84d443bd40b7601d577169ef

Observation e6a3954c-345d-457a-a317-1c830cb69079 · inbound

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals cites this paper.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.946066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.946066Z digest=sha256:7422c49425fbe0f46c6fed2f03daed30c8a5be8a378ee53b9b5d82c746c797e7

Observation 8116dac2-c1ee-4493-a86f-e24760ad3bdb · inbound

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference cites this paper.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.175978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.175978Z digest=sha256:166bd468fd11c15a8fea6c71fc7a2d6a592292c33d4169ddf88c5a6636353b39

Observation e9811bac-a8e1-4076-997c-5feed5c2e8e6 · inbound

Dissecting Bit-Level Scaling Laws in Quantizing Vision Generative Models cites this paper.

Dissecting Bit-Level Scaling Laws in Quantizing Vision Generative Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:04:24.308679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:04:24.308679Z digest=sha256:f0a0e38721b660753c7b98017b4bc5edca2e7cbd6797b67ca307c9488dac5559

Observation 6537f63c-1669-4b9b-a292-4bb7a9402611 · inbound

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting cites this paper.

OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:19.981883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:19.981883Z digest=sha256:2cf08651e9c0dcac1dc79b942a81c792e8754fdd55e93ff229469992c627b914

Observation aff4718b-c0a3-4b33-bcaf-1f30955985b8 · inbound

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models cites this paper.

AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:02.642735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:02.642735Z digest=sha256:8b8ab0be6430f78238c96f839b4f3f118b2361dc313c9f10a3d619940f2c320c

Observation 9926d273-81d8-46b8-9d2d-605d88f75458 · inbound

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models cites this paper.

Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.320318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.320318Z digest=sha256:fd27e3d05973994c0529e5acdd00c7976732748b4c28f7e446f7b93416636f25

Observation b18d4993-1044-44d9-a6ac-83d11e28ebb1 · inbound

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models cites this paper.

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:03:44.677581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:03:44.677581Z digest=sha256:6fd44c663888b4ec2cd6229667f24f67c64c93a943c192be8db06e38b5406cb9

Observation 60e0babf-5767-441d-9c9d-7bd97b5ee134 · inbound

ICQuant: Index Coding enables Low-bit LLM Quantization cites this paper.

ICQuant: Index Coding enables Low-bit LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:49.765633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:43:49.765633Z digest=sha256:cd0cb38c16a27a36bec379c6d5baa88a9231becb8775452a932d29e67ccfeadf

Observation d6712259-3d05-4354-8759-50afda2488a3 · inbound

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models cites this paper.

One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:32.123193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:32.123193Z digest=sha256:3d18e0219110c774f33aa974c766e57e7f481e00fa4b08ddd45713c12c6506bc

Observation 960df4e4-6bd5-4ba9-a664-f4c54df000b8 · inbound

High-Rate Nested-Lattice Quantized Matrix Multiplication with Small Lookup Tables cites this paper.

High-Rate Nested-Lattice Quantized Matrix Multiplication with Small Lookup Tables QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:27:29.653760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:27:29.653760Z digest=sha256:f0551f02fbc89a1c7684da3fde44b83e77806964259be2fd4d529bf88f76f812

Observation 8f1435f5-aada-4861-9453-291e952d78cf · inbound

Rethinking the Outlier Distribution in Large Language Models: An In-depth Study cites this paper.

Rethinking the Outlier Distribution in Large Language Models: An In-depth Study QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:40.665688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:30:40.665688Z digest=sha256:3d2745149d4b7d2feea042fa0d294c91f39f28df63eb0f5e71bbb8f510691954

Observation abf25eb5-8f0d-4593-965a-6a9964b135d3 · inbound

FPTQuant: Function-Preserving Transforms for LLM Quantization cites this paper.

FPTQuant: Function-Preserving Transforms for LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.223987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.223987Z digest=sha256:28804843242c0fd747e4b1bbea1f5a43630925921b0b99f773629e944593b3fd

Observation f9803ddb-fe65-4a43-a073-ffdc263e8117 · inbound

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling cites this paper.

PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:41.017031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:42:41.017031Z digest=sha256:4567a51754c16f706e81c8e1cf59756b06bdc0389ca1374fb2d4c6d839aed4d2

Observation 653af550-633e-4c12-99bf-49198969a537 · inbound

LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation cites this paper.

LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:09.141264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:09.141264Z digest=sha256:b9a429825590376483aef9c1e50dfb6fdfe4d4a150c2b352bddbd8bec9f19073

Observation fa5f84dc-5fa2-48b3-b38d-220288c5e4af · inbound

BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook cites this paper.

BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:07:20.504672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T14:03:35.214840Z digest=sha256:bb2b1e758b272d521108c4d802cf419e14a5b025978bb8ba991b23c7b1a491cc

Observation 76718352-586f-4a20-830d-36091d474f31 · inbound

BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models cites this paper.

BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:24.880486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:24.880486Z digest=sha256:320d6f060498b44c3089de950adbd4ed2629c609c78d5a0d6d4415064c6d94db

Observation 844df51a-8dcc-4883-97e2-cff04b4e25a6 · inbound

CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs cites this paper.

CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:16.104514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:09:16.104514Z digest=sha256:90ae565b85ca87e7ce141e6ee6b4620e83177ccfb07a789e35cf05e14c48184a

Observation b468cf42-94a4-4060-b90c-7ff0b58c9d84 · inbound

Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos cites this paper.

Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:52:53.059020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T23:52:06.036879Z digest=sha256:1cce99d9adb4c823c1ca869f9521cc8d003a332694fdfd91ac4670eaf1c503cf

Observation bc2d4450-97ef-4fca-ad2c-fd26dad50cbd · inbound

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models cites this paper.

From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T18:28:19.977137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:28:19.977137Z digest=sha256:13ef4f8c5174faa57d9df15d012e29885acd746744fc721ceefc2f0bff34f207

Observation 334fd9c3-d10f-459a-86ea-b26c2e77af96 · inbound

High-Rate Quantized Matrix Multiplication I cites this paper.

High-Rate Quantized Matrix Multiplication I QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:20:53.235633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T11:18:41.456313Z digest=sha256:65e049632a22d293996ddae4c5e6401125711fe1656753ebfc26e756360fa67e

Observation e535b9b4-fc0d-45ed-b93e-4299509bb3ba · inbound

VersaQ-3D: Architecture Support for Visual Geometry Grounded Transformers via Versatile Quantization cites this paper.

VersaQ-3D: Architecture Support for Visual Geometry Grounded Transformers via Versatile Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T15:45:32.789339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:45:32.789339Z digest=sha256:5f9468ba325ee770ebfaa934773caa8d643002350e99a1c972dc2833e8de1fc3

Observation c4b4a7bf-04bf-43f8-9125-e543a17f50a8 · inbound

SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model Inference cites this paper.

SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:01:15.959552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:01:15.959552Z digest=sha256:5e64b7b30ece2377144aa54a399cc78b48676f34a3d7e522f595a609ce95fb29

Observation 29b21c0b-812c-4736-98e7-520c6166b220 · inbound

Price of metric universality in vector quantization is at most 0.11 bit cites this paper.

Price of metric universality in vector quantization is at most 0.11 bit QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:10:43.337679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:10:43.337679Z digest=sha256:62f388bfbdea16e29bb4576adac6789f75faf6358207dd407d879f6e3dd86bdc

Observation 5850f303-8b23-4b49-8c14-9523306ebce5 · inbound

Leech Lattice Vector Quantization for Efficient LLM Compression cites this paper.

Leech Lattice Vector Quantization for Efficient LLM Compression QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T23:10:46.775151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:10:46.775151Z digest=sha256:cc8f78c355b9e8000b8654e77ca179a7f6842b6d2dc6002985bddaf5a0d2baa7

Observation 979a4726-743f-460c-995d-f138547b3c09 · inbound

Rethinking Residual Errors in Compensation-based LLM Quantization cites this paper.

Rethinking Residual Errors in Compensation-based LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:01.232262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:10:54.439287Z digest=sha256:8606f389d4315ed2c0ab9049f2039d6597b44dbab03c65c95ddc4891409991ab

Observation a82fc49e-4cc3-4433-a214-14351a05f5a7 · inbound

DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization cites this paper.

DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:01:13.345365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T05:59:23.738476Z digest=sha256:e4a28f176dcd2aa629757d2c556788dfc985939e2982b4ecee50f8d7807925c2

Observation 1c84b740-b831-4017-9c08-7a1429a8d62b · inbound

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling cites this paper.

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.529038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T05:29:51.182114Z digest=sha256:64274a2e76a1f22cb1ead16a109e06a8edf641a07264c7c4b93f048c42ce740a

Observation ded0d6f2-9429-4b4d-9d74-40da99bc914d · inbound

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling cites this paper.

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:02:42.158526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T18:01:08.514022Z digest=sha256:5a2119f3a14677a1d66df6b3870052293499c92155af2d5bb9407b82bc370e3c

Observation 7faa9924-73d0-49a2-ab8e-b55ec9c3b411 · inbound

From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization cites this paper.

From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:38:17.179684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T02:32:50.182859Z digest=sha256:412fccfe6bbf1c1cb48baaaa69be9eb9b5a984dddcf8540a4a759ba74cba9133

Observation d90fc254-7403-4cc7-9e95-459904760672 · inbound

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales cites this paper.

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:04.520977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T01:44:42.989053Z digest=sha256:4978eee9302ee562701b52e9a50c0a86aeed5d95c2960d85754c30f6841aebb5

Observation 65084898-a928-49f3-8b7d-d021bd7234ed · inbound

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment cites this paper.

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:43.051795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T04:23:26.079298Z digest=sha256:f38a621b343cbe3b614e48cfe30672caa2ed6049088754575c71d181614f0495

Observation 1277e456-2dab-422e-8fe1-41b5f467ee08 · inbound

When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon cites this paper.

When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:14.373234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T03:16:00.868964Z digest=sha256:08ad584b5310339c48d382a98fd5b937dff6832770e304b9624c5741d9f316f4

Observation e74c1b17-4971-445f-9e69-2044a2204c1e · inbound

Search Your Block Floating Point Scales! cites this paper.

Search Your Block Floating Point Scales! QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:24.139246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T05:52:26.558984Z digest=sha256:c087bed07b74346c0283a4b8f4a458cec1b7c2b5ac80ca6147d7a240542149f2

Observation fa5e0692-584d-4ebb-a1c8-c1a8e06a9df5 · inbound

High-Rate Quantized Matrix Multiplication II cites this paper.

High-Rate Quantized Matrix Multiplication II QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:32:52.654039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T19:29:22.250363Z digest=sha256:4e17b83a713ee31dd6cf952aa0bfa9591c9a3a4605b84d2ef0440f7045b9170d

Observation 8e2d83e8-3f0b-4d30-a204-174f4a9231be · inbound

High-Rate Quantized Matrix Multiplication II cites this paper.

High-Rate Quantized Matrix Multiplication II QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.991846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T21:25:26.862900Z digest=sha256:a31d0af0a6452063d3ba709f3b9c50ed31889b28bd369efb44dd69bc3ae5e2c8

Observation 69e0ff21-96d9-40f6-a763-1e2993afd757 · inbound

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference cites this paper.

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.699300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T21:28:36.358474Z digest=sha256:db6a685f62edf6873f0f92a8d23459dcf1d8ba8c67e275fa7ebf4b6c82f3f6f4

Observation 7afb0a30-8a77-4ba0-ba68-76932eac5c88 · inbound

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets cites this paper.

GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:28:16.800354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T12:25:39.417436Z digest=sha256:9e2833293aa34ebe01719a69402767478a2e0f1d735d76a148023e91eb2cae52

Observation 5c711e35-6b5a-40c7-aeaa-e3055fc33baa · inbound

Theory-optimal Quantization Based on Flatness cites this paper.

Theory-optimal Quantization Based on Flatness QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:39:10.010380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T22:38:06.888665Z digest=sha256:53391781e84e89cf0c5073fc6cf40f724e98a249372697012f0dde1a30463338

Observation 66a62db8-dd20-40f8-84a5-1e764908bff2 · inbound

Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU cites this paper.

Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:53:55.288364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T02:52:40.923230Z digest=sha256:76f0bf36b8bb0f0ff01c4f9b4fa6ef1c6ab20d9d8b3885e6277d2e304334adbd

Observation 9a056b76-2bae-4243-beaf-d2045c7cc745 · inbound

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs cites this paper.

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.854470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T05:33:06.719954Z digest=sha256:07e85b93550bc69584e4d8b0a483e50eb25ef1d0b309b21a616af5a7d127bc54

Observation a922941e-80d6-4ce0-b902-4094c5e34ef0 · inbound

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization cites this paper.

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:05:48.338378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T17:36:45.807397Z digest=sha256:9f66e5f2aa6eee7e303683f7c3efe73e09efcd7ad7434558a33671ed89740afe

Observation a829676f-80bb-44d5-9680-467e235be0fc · inbound

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture cites this paper.

EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.610682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T14:35:02.484377Z digest=sha256:be71ae3591ecb3c110cf0a1287621a4f955d639afd59f10524cff481c1b94b68

Observation c625abbc-a448-4950-a656-559e37105655 · inbound

Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization cites this paper.

Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:14:39.052951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T12:13:18.805668Z digest=sha256:f8e914ebeeb67c8c5a54e1980e0a14087c05c41958df1400fd45cac8a7cf613b

Observation 66a62b39-59be-4d24-821d-bb8eb1562754 · inbound

Quantized Reasoning Models Think They Need to Think Longer, but They Do Not cites this paper.

Quantized Reasoning Models Think They Need to Think Longer, but They Do Not QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:00.108641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T23:05:00.401365Z digest=sha256:641d3c3b1e020377da9c70f44dced1e6d52cb50b10669d07c871da82283175d8

Observation c13b1c42-78a6-46b0-958b-7e8ed3a4b80c · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.620068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:3e70a659b5e1833aaad47bf7c5f919ba0051b0d0c5e2d04f65550d4d9d628e80

Observation c9ce07cf-3804-46ab-85a5-3a24dc1b8fa8 · inbound

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models cites this paper.

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.668492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T19:09:29.347284Z digest=sha256:f2d05f76ea23928d581da3ee2234608fcb79b220650f6aecbbd40295cc919931

Observation bde79f78-77dc-4e2e-b4b9-1ee44d2b52f9 · inbound

Qift: Shift-Friendly No-Zero W2 Post-Training Quantization for Rotated W2A4/KV4 LLM Inference cites this paper.

Qift: Shift-Friendly No-Zero W2 Post-Training Quantization for Rotated W2A4/KV4 LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:15.982040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T15:49:41.836888Z digest=sha256:21a6aab37fb2fc53094363093f39e1cb5800313051b784d785b011a1221b368f

Observation d3d7e18c-cfd3-49cb-8c9a-7aed78fa2441 · inbound

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection cites this paper.

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:26.980422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T11:14:03.535306Z digest=sha256:0aab0891b002e73583f3a7760db99a1cbaa0560bb14bc811f67b3569ff4af100

Observation d37ff3ed-95ea-4597-a431-61d0017580b9 · inbound

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection cites this paper.

LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:24:38.273775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T11:17:53.736872Z digest=sha256:4f5b32acb98f183b0f86912c450c02f850465672da8fcf70905e331459f527b4

Observation 24c04631-cb20-4438-892f-108233728ee6 · inbound

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models cites this paper.

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:16:48.288085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T06:09:42.838355Z digest=sha256:590941eda22c34acbed9c863d3b5ce204e3da9a31f92ee3084529624777e94a3

Observation 93ff78ee-bb36-434a-b4af-36eaf89af560 · inbound

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models cites this paper.

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.577768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T02:35:00.563647Z digest=sha256:5c15292bd1311768150fa948be11aab6b7a215ec1f87e6c9a4c0df3e1e6d81a5

Observation c0132601-3a29-4c84-96e1-d0977ae20abc · inbound

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models cites this paper.

FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T12:23:13.879878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:23:13.879878Z digest=sha256:8c7afe0ddfe3909d33d897ca4fe97621b74b3b16a25ddfc2fe31304a08f25b2e

Observation 70b88107-7b85-4d66-a397-62ab0c9862d1 · inbound

HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models cites this paper.

HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.532847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T08:46:01.500880Z digest=sha256:0251aa0a3e5700e195db760d724b456409c0ad6edff01e571d3e4cd0f720bf73

Observation 2f57c8e7-0ff3-4570-96ce-312bc4869769 · inbound

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation cites this paper.

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:45.441374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T08:38:32.577228Z digest=sha256:e36ab0425cfa7c206c3614d63bbf0ae0d4b1ce3ef550f00e8d667d05835b2c45

Observation 75cbd6cf-71ce-45fd-bc84-c577b9e28297 · inbound

Reliability Scaling Laws for Quantized Large Language Models cites this paper.

Reliability Scaling Laws for Quantized Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 160

Resolution
unresolved
no resolver link, observed 2026-07-14T08:45:52.855783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T08:45:52.855783Z digest=sha256:57e63c952a996611e57febb430eb69eba16562df11fa79fdc6dba313ea3bb6bd

Observation 6f7f2d25-95d6-4df6-a7d3-18d90b9d3ac6 · inbound

Cross-Layer Error Compensation and Finite-Sample Feature-Statistics Matching for Extreme Low-Bit Quantization of Large Language Models cites this paper.

Cross-Layer Error Compensation and Finite-Sample Feature-Statistics Matching for Extreme Low-Bit Quantization of Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:35:43.851034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:35:43.851034Z digest=sha256:cede2afad82870ff3b35236592641c21788e906452fb21f03e816f1597895a92

Observation 0d47cdbc-5d5d-4238-836f-77290468ea3e · inbound

Break Through the Compression Bottleneck: From Theory to Practice cites this paper.

Break Through the Compression Bottleneck: From Theory to Practice QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T14:29:28.509827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:29:28.509827Z digest=sha256:2afc564b489a9b1d7a556bc4a9dad4ea48a3d88fa1861342542470bb4d46c176

Observation 829af20d-2a27-471e-a618-badf62ddb760 · inbound

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference cites this paper.

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T22:36:29.341483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:36:29.341483Z digest=sha256:9d25e44b87a4735be84c5bb4ae1c716313f785e46d754253077f51eb83825697

Observation 34d4b659-a33f-4912-8ed3-172af34aa219 · inbound

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference cites this paper.

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T03:16:46.931878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:16:46.931878Z digest=sha256:25156ba6c6bfd2d2f83416bbbcd605f5cc99deb6440363c20d1fef48e9158c39

Observation c2ee7334-ce78-4440-80e0-3553412fa3c8 · inbound

TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning cites this paper.

TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T21:59:39.992784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:59:39.992784Z digest=sha256:0b06967a63de4e15466438c23eeaefd6b959c28f75095137d89f60da1631ba47

Observation cfb16197-e846-4ae6-9451-ab02bc3c76a6 · inbound

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights cites this paper.

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:51.921431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:51.921431Z digest=sha256:b0abdf4e39fd8563482bffd385000a5a39cbb85d8a748c9fe11ad07add944987

Observation 86d060aa-a84a-4244-8599-ba7357fc836f · inbound

Hidden Language Consistency Phenomena in Reasoning LLMs cites this paper.

Hidden Language Consistency Phenomena in Reasoning LLMs QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 246

Resolution
unresolved
no resolver link, observed 2026-08-14T04:40:10.278392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:40:10.278392Z digest=sha256:3121073dbaa0a30f2d7e34b9f6b12f94ffe49ae768259c3ab7b482ba82f9f315

Observation 24826879-a826-4232-ac42-41c68bffba9d · inbound

Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving cites this paper.

Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:25:48.989424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:25:48.989424Z digest=sha256:74a0f21c33cc87a4480264f5d9ecc19b943e956394d18794f1b97c4189f23bb5

Observation 3d65f970-4d23-4d0e-a014-63b7c6df67a4 · inbound

SoftWater: Class-Aware Rate Allocation for Softmax Quantization cites this paper.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.575124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.575124Z digest=sha256:1ea408e9bde361fff1741c266cdb84afcf8d34db156e299bd0c5bce778c7be05

Observation 772d053a-e0ab-4e07-b7bf-fd32b065b019 · inbound

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation cites this paper.

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-14T12:54:26.608490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T12:54:26.608490Z digest=sha256:36bde568886eb5075c43bb8ad7c436c07c9e0610ba1a216b4ab6c3b2aed7ecca