Pith. sign in

Paper Citation Record · LEDGER

SqueezeLLM: Dense-and-Sparse Quantization

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 78 inbound Pith citation observations for arXiv:2306.07629.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.07629 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 78 of 78 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:00:43.084578Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:18:37.340082Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fe5067cf-dab7-4166-aa91-a11026099e93 · inbound

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models cites this paper.

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T13:49:33.850130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T13:49:33.747672Z digest=sha256:8ad5415ba8742a2a215a41b16e9f9812dd266edfbc99e16a5180f0a04334a8b9

Observation 652a7ae9-34ed-4e07-bc1c-d8e964f9b301 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads SqueezeLLM: Dense-and-Sparse Quantization

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:36:18.341205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:b293bed2f2b385f9e733991a5db9ed306f2b770aa308463e06f1e7da24cb09f5

Observation 37c8d920-1703-41cf-9a2c-069a3b3898ab · inbound

KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache cites this paper.

KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache SqueezeLLM: Dense-and-Sparse Quantization

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:53:12.337562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T08:53:12.253243Z digest=sha256:e1451fdd8af57fa229d7dc4be11453e0b07d1eabf59e596bcc95707913fe9722

Observation 8c46ab81-0c82-4bf6-82d4-4f4f738c1f12 · inbound

RouterBench: A Benchmark for Multi-LLM Routing System cites this paper.

RouterBench: A Benchmark for Multi-LLM Routing System SqueezeLLM: Dense-and-Sparse Quantization

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:47:31.094875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T10:47:31.006944Z digest=sha256:31ad1c5801c16d7da01f0181caa546a66aed8553f83ae830955f937adaffe5d9

Observation dc0a389b-ad0f-4564-b0d7-fec01552fe30 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 197

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.214075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:2e428e63780bc245d689bfd9768f36ae3473a6026c482d45cb2a6d0e00873054

Observation c2b6fba9-5454-4565-b948-54fe9cad1fec · inbound

SpinQuant: LLM quantization with learned rotations cites this paper.

SpinQuant: LLM quantization with learned rotations SqueezeLLM: Dense-and-Sparse Quantization

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:52:34.700791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T15:52:34.606853Z digest=sha256:c14be9a145c5151c5bf40c6479df116062857b7b459bcc10188fc333ce2458f5

Observation 2944c500-c29d-442f-b280-76f842316fb5 · inbound

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem cites this paper.

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem SqueezeLLM: Dense-and-Sparse Quantization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:07:30.170714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:07:30.170714Z digest=sha256:d940009634937eb3540dce9c3e31ec2aea359be4c878bb9af52d4dd4a259455b

Observation 7a58fbbb-e45d-483c-9b3b-619bad0672ae · inbound

CPTQuant - A Novel Mixed Precision Post-Training Quantization Techniques for Large Language Models cites this paper.

CPTQuant - A Novel Mixed Precision Post-Training Quantization Techniques for Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:50:28.874894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:50:28.874894Z digest=sha256:6913ff32a3c7ad1315b20d83894a6a8ecab1186296ceaf78b57298bbc80f31c7

Observation 7621d1ce-c7a9-4454-835a-11f8a9a705e5 · inbound

SKIM: Any-bit Quantization Pushing The Limits of Post-Training Quantization cites this paper.

SKIM: Any-bit Quantization Pushing The Limits of Post-Training Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:16.624860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:16.624860Z digest=sha256:7b6d4b9d7f9d9dec1c7c79f02a7878e467c72143920c812f87dd205212829f62

Observation 04f5fb6c-f226-4659-bf54-4ba5affb2bb7 · inbound

Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization cites this paper.

Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:08:46.631080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:08:46.631080Z digest=sha256:b8c4c113258bc9f7126a625f649ffe6c81dfd14505581bfd88c53b2b246d2667

Observation 5194d093-8a59-44b6-8c34-a4f59fa04511 · inbound

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals cites this paper.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.853192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.853192Z digest=sha256:18586fe43e24b6537ba9ee9531ee4d32291199dd45383c94a69fbd2ea46117cc

Observation 2006d8ae-b07d-40b6-9803-43fa6828a3b0 · inbound

1.58-bit FLUX cites this paper.

1.58-bit FLUX SqueezeLLM: Dense-and-Sparse Quantization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:39:36.579296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:39:36.579296Z digest=sha256:4bb77a4a3fbae4c21eb1cc3812814623aa78975d02e2be1734b2f11b9df2ae0e

Observation 7d957495-3395-43b2-b49d-98a464d2b749 · inbound

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference cites this paper.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.128860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.128860Z digest=sha256:ed501c475bb486607d4ae4952918ec83ecf5831b3dc743a40afe170a1baf355b

Observation 546c6e91-2378-4d7f-b935-d5c1767416df · inbound

FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices cites this paper.

FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices SqueezeLLM: Dense-and-Sparse Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:20.735632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:56:20.735632Z digest=sha256:75634d2531cf7e16fa7a143d34a72e36c1b15d1f4409417ea06f935450c7ee85

Observation b321e75c-9bc7-4ccf-b55a-f6121396b34b · inbound

PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks cites this paper.

PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks SqueezeLLM: Dense-and-Sparse Quantization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:43.264913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:43.264913Z digest=sha256:3a68ca9cdd39c7a0f9de4b56a7b6ba5518972ec2c6376868016fcab73c8221b8

Observation e4589b22-bab1-4780-bff3-fd0e67da1f41 · inbound

Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring cites this paper.

Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring SqueezeLLM: Dense-and-Sparse Quantization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:57.446387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:57.446387Z digest=sha256:3de1ceaeb2136f2de75032ff843153beecee6a87fbc0a60ac68a58642c05f5d0

Observation 8a3ad5eb-fb11-437b-a842-f2b5f1e3381a · inbound

Elucidating Subspace Perturbation in Zeroth-Order Optimization: Theory and Practice at Scale cites this paper.

Elucidating Subspace Perturbation in Zeroth-Order Optimization: Theory and Practice at Scale SqueezeLLM: Dense-and-Sparse Quantization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T21:23:48.574835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:23:48.574835Z digest=sha256:9430c6df5ea2731ae2613e2549fea06386be3ed9d4feddadddabe4625f06d175

Observation 8b308851-e505-4df8-ad27-7ad194f53a52 · inbound

Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives cites this paper.

Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives SqueezeLLM: Dense-and-Sparse Quantization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:28:42.183775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:28:42.183775Z digest=sha256:c3f66db9db84a05635a7442263d67b3e553065a8e1ed96eeddeb84226a96addf

Observation 82638e70-4319-4cad-85ea-af7a65b3dd55 · inbound

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding cites this paper.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding SqueezeLLM: Dense-and-Sparse Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.010420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.010420Z digest=sha256:7629154cb17ee3406b9784b39d9315a83e89a72434a683f1fa2153165e3a9e9c

Observation 82fa24fe-d6f5-4e80-a3b1-bf4e25e3196a · inbound

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters cites this paper.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters SqueezeLLM: Dense-and-Sparse Quantization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.544467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.544467Z digest=sha256:bfb896d10beaeb867f661d5a359226aa7490f0b1ea4b88e6fb2a4ea381065963

Observation 4eee505b-e77f-4694-be24-ab55ae4e2279 · inbound

On multi-token prediction for efficient LLM inference cites this paper.

On multi-token prediction for efficient LLM inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.760608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.760608Z digest=sha256:a3cccbc8a853b452e130db63b40e607bb97b0e0854c3d10d46ffd1d2e68f47bd

Observation 9949b154-5c0b-4936-8e4c-06b655bb5173 · inbound

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache cites this paper.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache SqueezeLLM: Dense-and-Sparse Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.938479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.938479Z digest=sha256:40573ea8e02700e69b61f61626140598ac3a215840d33f6178342eff6eb774c3

Observation a609dd30-4a14-47d9-b2e7-a04c77294da5 · inbound

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices cites this paper.

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices SqueezeLLM: Dense-and-Sparse Quantization

Reference 174

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T01:05:16.310998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T01:03:26.037233Z digest=sha256:d3cf5295044d247f9b969b8d843a6e95ededcbaea42e674517efa9ac3b8a6437

Observation 52cc938e-9e87-487e-b765-9b181799675b · inbound

FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference cites this paper.

FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:00:43.084578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:00:43.084578Z digest=sha256:64e5991f930a7735a2af6ca30b8bf5ce0fcacbd28b5e3999f480d8c2cfa9f814

Observation 96011ff3-6b96-4f6e-86b1-e5309a87cf7c · inbound

TeleSparse: Practical Privacy-Preserving Verification of Deep Neural Networks cites this paper.

TeleSparse: Practical Privacy-Preserving Verification of Deep Neural Networks SqueezeLLM: Dense-and-Sparse Quantization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T06:06:57.425786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:06:57.425786Z digest=sha256:281690f290ca7450140be7c59d3ed5041a2823f1f9b11cf7cecbffb1836884a9

Observation 49ef2c4b-10db-4fd8-b047-447b5e02ca89 · inbound

R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference cites this paper.

R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:46.632513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:46.632513Z digest=sha256:d232d065239a1a3b00520ae29fbc2926d727f0fb422f894745e452224300d1fc

Observation 8c4c853e-fc2a-4b1b-b7df-37648d77476c · inbound

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs cites this paper.

FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T05:52:21.244939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:52:21.244939Z digest=sha256:adfb1f48fbdf4c32b1a1be2677dc2c7d5215047cf5c19f701e212d5963736d23

Observation e28853ba-9226-4caf-bf9b-35df3a4991a8 · inbound

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate cites this paper.

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate SqueezeLLM: Dense-and-Sparse Quantization

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:09:22.277661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T08:09:22.226608Z digest=sha256:5d597da9d4c557e12ac586e292f69b08c85d3127a562b1a74e3ecc1ef9ff81e3

Observation a965507b-7a49-44ba-a3a6-67f775b63726 · inbound

ICQuant: Index Coding enables Low-bit LLM Quantization cites this paper.

ICQuant: Index Coding enables Low-bit LLM Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:49.735389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:43:49.735389Z digest=sha256:916e9174a79f3eede5c75a2eff236050f98f1cfdf1671cee864202f57a77d40a

Observation 929dcb5e-aab5-4020-aca7-84de10b3c0bc · inbound

EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices cites this paper.

EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices SqueezeLLM: Dense-and-Sparse Quantization

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T16:34:59.261977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T16:34:37.083239Z digest=sha256:6abae3b00e6ca4be86a8bb13399fa380bf5a86a3bd5e1766220bd6c043c63bd3

Observation 47088b09-0a0a-4c35-a053-7c53473f74be · inbound

MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance cites this paper.

MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance SqueezeLLM: Dense-and-Sparse Quantization

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-16T04:31:56.870632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:31:56.870632Z digest=sha256:12eb3ea7341d3afd257942289d0662f069a97f5d51f231e5e7b457e4cf3b0854

Observation 41c61031-866a-4907-bf8c-50104bd6c46d · inbound

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design cites this paper.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design SqueezeLLM: Dense-and-Sparse Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.100645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.100645Z digest=sha256:17ee08778fc0333b5f3f0fb309be8c18ae651e441ef553cd2945020603366494

Observation bfaae37a-2dbb-4bac-9dff-6cb3a04119de · inbound

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression cites this paper.

Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression SqueezeLLM: Dense-and-Sparse Quantization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:36:42.294573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:36:42.294573Z digest=sha256:98cdf966925e34b753bfefbeb673e76cc6c36c15f058dec64c35d7a2e74746d7

Observation 776bbb39-0683-4ba8-a428-617b2fe31cc3 · inbound

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference cites this paper.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.656359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.656359Z digest=sha256:7d4624b2240a0e21ea48620fee7ba5b544d74ce46b4c622f81dec2e701ce71d5

Observation 61b27dfa-542c-4af4-b9dc-1325f222f8d7 · inbound

NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics cites this paper.

NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics SqueezeLLM: Dense-and-Sparse Quantization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:28.498086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:28.498086Z digest=sha256:d974f18473c5a5b408870ca1ba9ae3aa257979c93cf728d9b20e55500c7ec0ab

Observation c33b91b4-7076-4fed-839a-2d4ea3d92494 · inbound

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache cites this paper.

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache SqueezeLLM: Dense-and-Sparse Quantization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:21.069575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:21.069575Z digest=sha256:44536c618b866387b92dc2dba00162b99d5f3059e6419ba08ff4ebab3b6db4af

Observation 522070b8-6cf0-4e0d-a01a-26bae17ff451 · inbound

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression cites this paper.

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression SqueezeLLM: Dense-and-Sparse Quantization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.372336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.372336Z digest=sha256:1f7bbebae7213b0e8e23efcf3bf403f6891ea885bd141d568a3dd6693db2463e

Observation d2d1cc3b-7958-4a42-afa3-9fa22aa93fd5 · inbound

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression cites this paper.

Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:56.464514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:56.464514Z digest=sha256:3cf3c289173bd63454e33f686de874de7592066b4b10f9f21e53110b2a21cdf7

Observation 8667e534-2518-4bf2-954d-47bdcf62d9c7 · inbound

FPTQuant: Function-Preserving Transforms for LLM Quantization cites this paper.

FPTQuant: Function-Preserving Transforms for LLM Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.395454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.395454Z digest=sha256:8bf60bceb0a2be25a5a00ea4982d2464772fed253e10496c9a373dda9b4d239f

Observation 35aa94be-e2a4-4b81-aa83-0b824d01744a · inbound

BAQ: Efficient Bit Allocation Quantization for Large Language Models cites this paper.

BAQ: Efficient Bit Allocation Quantization for Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.624716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.624716Z digest=sha256:99b0db87459882f045c9f7ea2222c0fc382abbf451dae196c6f1d73e4ffd5656

Observation 0c4f7916-3238-492b-bb86-7a95c6562580 · inbound

Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs cites this paper.

Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:53:04.144810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T08:52:45.818050Z digest=sha256:2ac747377f294d0ed1484ca59d21df4694f90d9f84698e11c7bc38fded17889a

Observation 7b9aea9f-f074-4b9c-83ed-13b6c938936a · inbound

Efficient Serving of LLM Applications with Probabilistic Demand Modeling cites this paper.

Efficient Serving of LLM Applications with Probabilistic Demand Modeling SqueezeLLM: Dense-and-Sparse Quantization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:46.512823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:46.512823Z digest=sha256:a1c1a4f9211f9c67a578f228356c5fd6ab018e972acb53828466722fe7d4f479

Observation 53dfb406-e291-4a4b-badf-f4e1a106241a · inbound

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models cites this paper.

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:19.558832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:19.558832Z digest=sha256:f75e36b12655b953fad75454c62f8d863d4d3d646d6b19ab901fda09a3cd1055

Observation deb9f2da-2284-4ec8-bf66-743303f6bb47 · inbound

Information-Bottleneck Driven Binary Neural Network for Change Detection cites this paper.

Information-Bottleneck Driven Binary Neural Network for Change Detection SqueezeLLM: Dense-and-Sparse Quantization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:24.158945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:24.158945Z digest=sha256:aad5b5536bc68968873de8e58117974084ea3a551f1d64aa6d428dc95625a30d

Observation 14813428-cb77-415d-9713-4f4d6f3f670b · inbound

CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs cites this paper.

CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:14.550295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:09:14.550295Z digest=sha256:8d00d31459fe3dc2adfe30c18a670c938949f177e0eba3c3291d7124e2328929

Observation e84da4c0-1e3d-4d94-82c6-e8450d5c288a · inbound

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration cites this paper.

OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference Acceleration SqueezeLLM: Dense-and-Sparse Quantization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:15:33.956446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:15:33.956446Z digest=sha256:3b43fe0f27bfc1e2d44aa90261286c86b61a6940422ece3765a3fae3c5df4a47

Observation 7aa8ef5c-c2e7-459f-b440-804064a16c33 · inbound

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs cites this paper.

Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:44.950132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:52:44.950132Z digest=sha256:b0ef37066e707f6f16187400fefea66c45795c40ce26690afe29983919337102

Observation 5345532e-07ac-4710-acc3-c0145d49c1b6 · inbound

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference cites this paper.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.877669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.877669Z digest=sha256:2b895a2609aa0d9f70f4b7540013bdbcb9e0e2cb084a3dd375dab34b1b4e7584

Observation 0985625e-76f9-437f-ad97-54c08b11d733 · inbound

Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation cites this paper.

Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation SqueezeLLM: Dense-and-Sparse Quantization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T06:38:22.663076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:38:22.663076Z digest=sha256:e8a6974a899fa69dec0423db27f1ca1aeebe6e71331149f3f39805604674be8a

Observation bd74cafc-1369-4735-b1e1-8dc004cf81fc · inbound

CoreQ: Learning-Free Mismatch Correction and Successive Rounding for Quantization cites this paper.

CoreQ: Learning-Free Mismatch Correction and Successive Rounding for Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:52:28.393534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T06:51:18.629467Z digest=sha256:382e7e3860c11343b3268e797ed1c8a830f14114b287ab1a48f8d8a695355ade

Observation 4f1281b2-6e40-4ab6-8b75-236e1bbd952f · inbound

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks cites this paper.

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:24:47.003271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T00:21:30.101748Z digest=sha256:d72a1cf2cd75786ff0fd2e0089057e5129251fc6eadbcccaed33a0637036221c

Observation 4de3b8e2-c80f-463d-a5d6-a86773e97d0d · inbound

Coverage-Based Calibration for Post-Training Quantization via Weighted Set Cover over Outlier Channels cites this paper.

Coverage-Based Calibration for Post-Training Quantization via Weighted Set Cover over Outlier Channels SqueezeLLM: Dense-and-Sparse Quantization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:16.200257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T04:35:02.120009Z digest=sha256:3fd54b6f4a4b82285285c63252da8e18ac7f5bce43195b6b621c306731ae9279

Observation e07df099-eae8-4bee-9a7e-d68669a9ec3a · inbound

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment cites this paper.

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment SqueezeLLM: Dense-and-Sparse Quantization

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:42.443324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T04:23:26.079298Z digest=sha256:e5fb6b58c538536fe9ede2eb5f4b380bf4b64438e216da744862fe13f414addf

Observation 8f50f0c4-d94e-433f-ab19-6021234e0e44 · inbound

DurableUn: Quantization-Induced Recovery Attacks in Machine Unlearning cites this paper.

DurableUn: Quantization-Induced Recovery Attacks in Machine Unlearning SqueezeLLM: Dense-and-Sparse Quantization

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:10:42.160072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T18:49:53.432959Z digest=sha256:178b8743094976d5be3300f0e8246f94d05c2138c95b3699b826df924f77def6

Observation 39e3d22d-1c77-4575-9115-5a4f1c462bc9 · inbound

DurableUn: Quantization-Induced Recovery Attacks in Machine Unlearning cites this paper.

DurableUn: Quantization-Induced Recovery Attacks in Machine Unlearning SqueezeLLM: Dense-and-Sparse Quantization

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:50:51.165324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T01:49:51.021741Z digest=sha256:d32120d524e4a0b097387d58173e23197bf86b690c39efe1b048df616c9891dc

Observation 99e832e7-3ec7-4588-b1ac-415de378cc29 · inbound

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization cites this paper.

WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization SqueezeLLM: Dense-and-Sparse Quantization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:08.029916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T16:06:26.450483Z digest=sha256:3781393c26359dabd752ac50eba34df34bdccd79b18df9588e00329dce3ebaf7

Observation 0d0da963-e2f9-4554-b344-1f115737f3f5 · inbound

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning cites this paper.

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning SqueezeLLM: Dense-and-Sparse Quantization

Reference 121

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:01:10.321373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T17:19:59.247074Z digest=sha256:88019a7b4af06c2fd4fed574c5fbc8ca148a406a1268561abd59bfdab09caddf

Observation ee2ac829-f9e2-4fd6-ad98-c96f3e2bc34f · inbound

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization cites this paper.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:30:44.142303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T18:23:14.935801Z digest=sha256:577f686f3b2a8aa48622ca1930bd6b8739cfe3ea5fb03037951a1af83bdd7802

Observation 2740b635-abd5-496a-a944-9bdd0dd5b5ea · inbound

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization cites this paper.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:18.210120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T02:59:00.997742Z digest=sha256:48629d69f215cf776b9ed42ad98ce821c027552fdada0f90bb0f780631a46d6e

Observation d16da3e7-de78-4192-8548-66862fc2ff1b · inbound

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference cites this paper.

XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.690279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T21:28:36.358474Z digest=sha256:1fe3ae9032a167841753bba038683ed01a684814747f38c3790166c8ac17bbc0

Observation 7ec28d36-0207-4a96-98b0-9217129dccff · inbound

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models cites this paper.

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T05:23:03.625863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T05:20:45.264341Z digest=sha256:ac669aa000d76f40d987babbdfe8b754fd92a41fe6e5cecf19b011ac6907761e

Observation 9bd3ab0c-6349-4222-a3c0-7f1dfd3eba6d · inbound

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs cites this paper.

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.872360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T05:33:06.719954Z digest=sha256:69d713561c3e640819e9fa004fb9c9066b199e46ab7ebe5d0919b1bcb401b0c6

Observation bf0cc52c-1c31-4da3-9a8e-9cba87bf348f · inbound

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models cites this paper.

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:04:58.491515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T17:54:56.386488Z digest=sha256:a52a8c7c3ab26c06401794f32f53cf12ed560a7dd199d3403d4a906dd32999eb

Observation 9f262736-e8cc-4a2d-ada2-88c261a3f7b2 · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation SqueezeLLM: Dense-and-Sparse Quantization

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:34.617720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:6cb099a71f1c8b152ab013fb9fbd7ec0f301382000ae58bda2505e86c979f4da

Observation d4b4ba6f-6db8-4589-807f-f054a40c3feb · inbound

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation cites this paper.

GPTQ-intrinsic LoRA: A Near-optimal Algorithm for Low-precision Quantization with Low-rank Adaptation SqueezeLLM: Dense-and-Sparse Quantization

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:14.475000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T17:28:14.160341Z digest=sha256:25b3cb6933bafc5e450a994e5e927ffb59f1be758beb57bb288d1f41995c1a2f

Observation cb5672a2-5a67-43de-ac03-95457cefd64a · inbound

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models cites this paper.

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:16:48.293704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T06:09:42.838355Z digest=sha256:ab05b169f6bc558163a5a8a817627dddc624c242ca9896659f1681168f45b2d8

Observation cc0d0cdb-359f-4fbf-a642-c00e4281e965 · inbound

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs cites this paper.

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 141

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:25:48.408417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T01:29:42.919461Z digest=sha256:168fd12b313c3f218158ff1fbbc4672d117be45e31d634e42fc0154b03094b4a

Observation 1bac6c64-6ee8-4ee4-a0bd-310a5e10279c · inbound

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache cites this paper.

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache SqueezeLLM: Dense-and-Sparse Quantization

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:57:06.403971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-02T15:55:40.177742Z digest=sha256:bf0110357b97078491240abcfaef5e13431bb0581a65e1ebd33e53b303c94da1

Observation 807090fd-1b0e-484d-b714-c291eb299548 · inbound

SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models cites this paper.

SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:18:37.341488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T16:14:03.717787Z digest=sha256:4a3a53f16bf611ff40c33f9073d11f0ec8c53bd2dd53e05beeb7e61eb2ad28af

Observation 0e56bdc7-8bc2-4799-9fed-944eb3b39dbb · inbound

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers cites this paper.

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers SqueezeLLM: Dense-and-Sparse Quantization

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:58:32.417750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T14:56:10.553212Z digest=sha256:34ce67b3726945f4ea65ef68cdce97b1e395cf4a102feb02e2e532235cb7b3d9

Observation 91ab3514-5ab9-417b-bd6b-a58cc4aea220 · inbound

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration cites this paper.

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration SqueezeLLM: Dense-and-Sparse Quantization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T01:10:03.032181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:10:03.032181Z digest=sha256:61bfd6e425a67ad3f6220f3d062a37f32673b3a68223315b096e84bad2640013

Observation 25b40a43-1701-4822-b146-c3fd2fdc511c · inbound

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference cites this paper.

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T17:12:27.941291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:12:27.941291Z digest=sha256:1012f256660c4e905a724160bacb71106409df81a33c8d61117455ce1206ed39

Observation a32228b5-c235-4152-960b-a3ebecdb9bf7 · inbound

Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications cites this paper.

Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications SqueezeLLM: Dense-and-Sparse Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T12:04:14.348367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:04:14.348367Z digest=sha256:6b3300882cae2ff378f4b62509dafd21bba4b576c5ff8a697ee8c57059469f53

Observation 97eec831-accd-4d0d-81d9-b292476d3de8 · inbound

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction cites this paper.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction SqueezeLLM: Dense-and-Sparse Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.561194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.561194Z digest=sha256:00c75bdfaad79d0b5f6cbd644d1430c8482833ae8afc53d11dec52376734eb46

Observation ae555642-7422-4b52-bcb6-322e78b476d7 · inbound

NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory cites this paper.

NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:07:27.121364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:07:27.121364Z digest=sha256:014a729d586ea576361b7235ae501de3adfe566c3988e0b3ce4171ceef250cf3

Observation c46894f7-5422-494f-8ba4-abcadcbcafc3 · inbound

Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs cites this paper.

Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs SqueezeLLM: Dense-and-Sparse Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:50:44.530504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:50:44.530504Z digest=sha256:bf463c094eb642da55d23c2dea8caa8d4d29def9a46a53f5e3edd7735bd06271

Observation a7416f9b-af4f-4111-8886-f14a698dd9e2 · inbound

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights cites this paper.

CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights SqueezeLLM: Dense-and-Sparse Quantization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:37:51.911058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:37:51.911058Z digest=sha256:0aa5aeb3b6c6933fa1941da2197d12c1a8415aae0c84cb8e42f5fb77adff4250

Observation 97c6bcb4-8125-4bfa-aff4-e3ba4efc8a6c · inbound

Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry cites this paper.

Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry SqueezeLLM: Dense-and-Sparse Quantization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T14:29:49.272723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:29:49.272723Z digest=sha256:6ff31f5a306e3f59e3120aa3c814e12b9805ce586be0a2130dc61a807854bf67