Pith. sign in

Paper Citation Record · LEDGER

QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2310.09259.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.09259 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:43:23.877823Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:59:52.313183Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 561a2594-52b9-46dd-8a6f-22701e488092 · inbound

Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration cites this paper.

Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:43:23.877823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:43:23.877823Z digest=sha256:a4e265681cd36faedd9176fd3bec6eafc185c503673b746aaf881b3cb27ee244

Observation 569864b3-9418-4831-83eb-7a5c6e432dd5 · inbound

FlashDP: Private Training Large Language Models with Efficient DP-SGD cites this paper.

FlashDP: Private Training Large Language Models with Efficient DP-SGD QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:00.968022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:06:00.968022Z digest=sha256:a32d6df0e473ef2285bc6da8d71799e91ac10d8c6a0d0c623af9fcc2de1d4960

Observation ad1243e9-dcf7-459e-8899-8dde1d289ec0 · inbound

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models cites this paper.

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:16.723597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:16.723597Z digest=sha256:1fb3a4bfdb485019bd46a2a0784b03ef2418e4fe5dd5f408552da626777a550a

Observation b8ade73f-d59f-4d64-a380-6d64298dae77 · inbound

ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs cites this paper.

ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T11:10:52.639298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T11:10:52.639298Z digest=sha256:441694f21d9be24263e62e899d7cefa621a7cff12c11aad166d46860d0b27351

Observation 6fad93e2-9670-448b-b5a2-8768c2d7be71 · inbound

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation cites this paper.

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:45.466022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T08:38:32.577228Z digest=sha256:29831607c78347d73776a5459594495237b2f4aa0cec09fda3700b4912b361d4

Observation 168084e3-149b-418c-887b-450aed8e7270 · inbound

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference cites this paper.

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:52.314559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T05:41:39.052865Z digest=sha256:87c73d7919d052c07bc2456f0a69eb23adb232c9a915e1347566d85dc4ae3e1a