Pith. sign in

Paper Citation Record · LEDGER

QQQ: Quality Quattuor-Bit Quantization for Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2406.09904.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.09904 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:20:23.960857Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:37:36.355740Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 27167c75-c01d-4d21-8d73-2964c3c4e91a · inbound

DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation cites this paper.

DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T05:20:23.960857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:20:23.960857Z digest=sha256:141ce625fa00472cc1628b72ed6f9fac981144b186ac81e95ab8e736fa538532

Observation 1b4fc94e-1dd0-4338-b2a7-526e6d2d1e66 · inbound

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization cites this paper.

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T19:12:56.330350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:12:56.330350Z digest=sha256:82384c8b0feb6cbdec4c4d6741c860e056cfec13adc6d8039d508f676e364157

Observation f168e218-e86b-4e7b-932e-e33431ef6247 · inbound

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference cites this paper.

Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:05.720980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:05.720980Z digest=sha256:7b00dcfb1273157994fa920b2af1ccf4cc13e58ddb0bc4eece63d82ac5e75718

Observation 428558c0-df21-44d1-b1de-79ebf8a71218 · inbound

Is (Selective) Round-To-Nearest Quantization All You Need? cites this paper.

Is (Selective) Round-To-Nearest Quantization All You Need? QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:02.080691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:02.080691Z digest=sha256:8423b4b946f47285b09a11c77ccf5470367ad2b46327dbd7ac06bd93e5650f8c

Observation 290d779b-6d71-411c-bd9d-2c5226e91831 · inbound

Rethinking the Outlier Distribution in Large Language Models: An In-depth Study cites this paper.

Rethinking the Outlier Distribution in Large Language Models: An In-depth Study QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:41.859346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:30:41.859346Z digest=sha256:b1ebbc0d7533cf41245fc1ff929be5fbfd378277170647e1255804d2ee16a884

Observation 714b484a-6a34-4218-8b9a-9bb2ae8603df · inbound

Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design cites this paper.

Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:18:31.964329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:18:31.964329Z digest=sha256:9dbf7f2dbaff85c42eea0251801ebf4511e97ab2ebcc10f550d36d6683c38951

Observation f2029150-7fba-44c7-8a93-aaaa1bfb8ad2 · inbound

CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking cites this paper.

CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:18:30.219200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:18:30.219200Z digest=sha256:3c6d67b441fd5c09cd2e6892ab7bc1aa2a96f16965c2f49bc57b0236b7311d91

Observation 49c6fe9f-3712-4dab-becc-206455f21cbc · inbound

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving cites this paper.

LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T12:51:01.231907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:51:01.231907Z digest=sha256:5568c42b2e0bb6df54a2ad14b63a7e0701cd81c763e35ac4e24e92f7f98f1d52

Observation c1ec7b89-61ca-468c-aee0-8311e69a110e · inbound

Baichuan-M2: Scaling Medical Capability with Large Verifier System cites this paper.

Baichuan-M2: Scaling Medical Capability with Large Verifier System QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T11:50:17.356679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:50:17.356679Z digest=sha256:c8a61aed1197c7ada651eb8a78da6023d2525e8d73c924e77b789497a1b55dc0

Observation c8da9949-a1d9-4937-9305-6d5d391d2c20 · inbound

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models cites this paper.

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:23:03.636300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T05:20:45.264341Z digest=sha256:1f91f9385d1d5adca59057704b57eca381786afce21f77f150898eff84193f33

Observation 9d03236a-51af-4588-8506-902322be98a6 · inbound

APEX4: Efficient Pure W4A4 LLM Inference via Intra-SM Compute Rebalancing cites this paper.

APEX4: Efficient Pure W4A4 LLM Inference via Intra-SM Compute Rebalancing QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:47:28.467228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T17:48:12.167724Z digest=sha256:5d43fe4f398dbc0397a7d8bddd6082591c504d4f15c11473233b4024041e186e

Observation 7b79d209-e441-4b82-869a-0488b8de8e9c · inbound

APEX4: Efficient Pure W4A4 LLM Inference via Intra-SM Compute Rebalancing cites this paper.

APEX4: Efficient Pure W4A4 LLM Inference via Intra-SM Compute Rebalancing QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:24:38.664691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-30T11:08:53.878181Z digest=sha256:c2ead41127603d103eb3f5b905da6b5757d8f02964f0ea64120f4461a332666e

Observation aae66742-337e-4318-91a8-594888654095 · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:04:45.295952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T05:37:13.211613Z digest=sha256:84e7479554cf8e3f7e1a7aa17d29b215ee36d607b391943ce0315917578081d6

Observation febd8adb-ab86-45b8-b719-385085d3dbec · inbound

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving cites this paper.

Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and Serving QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.112157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T07:06:53.318182Z digest=sha256:8a9e1c41911589d0894fe057933a7366a87f20141484edf9f293445be79640d7

Observation e0591a8a-cea4-434f-9ced-41a4f45daba0 · inbound

MxGLUT: A Reconfigurable LUT-Centric Broadcast Dataflow Accelerator for Mixed-Precision GEMM cites this paper.

MxGLUT: A Reconfigurable LUT-Centric Broadcast Dataflow Accelerator for Mixed-Precision GEMM QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:36.357500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-03T04:33:39.771377Z digest=sha256:53ffbb9d29b3136c9e3880f3d2e21c3545a61ab45b3c9575b02a46b77845fc01

Observation 534bfa3a-3d78-41fb-bcf1-69c21c86017b · inbound

Studying quantization trade-offs for efficient inference deployment in machine translation cites this paper.

Studying quantization trade-offs for efficient inference deployment in machine translation QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T07:51:22.191497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T07:51:22.191497Z digest=sha256:5b8f712846bde3d989c909dca58d5bfcc53bf13cdfac386ea247131d39b484cd

Observation 96a2a206-7936-4f44-87df-514c8351bf6a · inbound

Studying quantization trade-offs for efficient inference deployment in machine translation cites this paper.

Studying quantization trade-offs for efficient inference deployment in machine translation QQQ: Quality Quattuor-Bit Quantization for Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T04:25:22.049357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:25:22.049357Z digest=sha256:dd37ab4108b163248326741c55c0880e19d20d33da3fb3436a458af670b9937b