Pith. sign in

Paper Citation Record · LEDGER

Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2310.19102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.19102 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:10.289122Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

23
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fbdf8ea6-4f0a-4542-8838-cace19618111 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 212

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.259219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:a7ea47f119e8ca047d6c379244952d97d7dbab4e824f70a69cf06aae80cb2f5f

Observation 14b49881-a09e-490f-b4a6-f518a0782182 · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:26.468374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:0878ee0e17f3fb6b549af88f5dfd534bc59bb87b4b3934d8dfa7728d3d9a40a6

Observation 084a6871-9427-4e92-a23b-92f0da99b653 · inbound

SpinQuant: LLM quantization with learned rotations cites this paper.

SpinQuant: LLM quantization with learned rotations Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:52:34.799916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T15:52:34.606853Z digest=sha256:c294a778eecfea653204b2f0ca17795c0b802b1058accce3edeb3aa941e09a22

Observation a1b4a0cb-dae7-4f47-9d34-20ef2809089e · inbound

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching cites this paper.

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:44.165708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:44.165708Z digest=sha256:9ddd67ff7531aef2d30c885ea33c05ce9cd3b3d39539e01142e2127cf501f337

Observation fa460b24-0a05-46b9-a05d-70d9a349967a · inbound

TurboAttention: Efficient Attention Approximation For High Throughputs LLMs cites this paper.

TurboAttention: Efficient Attention Approximation For High Throughputs LLMs Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T17:50:17.710071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:50:17.710071Z digest=sha256:a898097160402299dd1bffeae060253afe9ac4f74749712c171a42c852461509

Observation 701fc2dd-82e0-4125-8523-63f72ed3fca3 · inbound

Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring cites this paper.

Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:57.666122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:19:57.666122Z digest=sha256:3ca232f2ee15fe79d27cca49d0c37032176744e8d08a8c08f378d3c595dec4e8

Observation 7e3f4934-d3b9-49bb-a126-fac795bbde37 · inbound

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations cites this paper.

QuEST: Stable Training of LLMs with 1-Bit Weights and Activations Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T20:44:48.486215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:44:48.486215Z digest=sha256:ec519ff11172bb854ea4a86ab7282ab5ceadd441e77cda63a6d79d1605fb1f52

Observation d9d73e46-be51-4b75-9116-763a59c93c58 · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 138

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:10.289122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:10.289122Z digest=sha256:ff930e41d2f8ce06bc6e54db1ee3898f01d43c324fedc0ae84d4d69073d405f4

Observation da10163d-d38c-4102-ba5e-e01fcc73c35f · inbound

Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs cites this paper.

Win Fast or Lose Slow: Balancing Speed and Accuracy in Latency-Sensitive Decisions of LLMs Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:47.771772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:47.771772Z digest=sha256:07e94ad269ab17c0cec7421795a81958786c8f13b87f87f4173431b46ad5c61a

Observation 2ed7aee6-40e6-4c0c-9e9a-517ce38ae763 · inbound

Turning LLM Activations Quantization-Friendly cites this paper.

Turning LLM Activations Quantization-Friendly Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:32:56.650567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:32:56.650567Z digest=sha256:88991fc47de7589c8b30530fb28118b05cba72d4bfc4ed54637e95f8dc905b09

Observation c3432691-8733-4ffd-aab5-4797d67bc395 · inbound

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models cites this paper.

SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:47:08.829325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:47:08.829325Z digest=sha256:d279e760004b5b7d03251a0017720a33fde7fc69592f0825e109cbea1a47d57b

Observation 69192696-a043-4bd7-8fd5-10dd46887f6e · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:47:37.208524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T08:47:29.236561Z digest=sha256:d7c3a0f1eb52a5fc37983fbf2b3eab87982a7a444eae8aae259db2225b2d0ecf

Observation 9b258fc4-ad6b-41c4-a2b2-984d63fe9bdb · inbound

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse cites this paper.

Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T05:50:25.663791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:50:25.663791Z digest=sha256:daef61caa4649f662f6029acbe307839ba3f3ec26c84a16cf7ca76ba380eeb54

Observation c8184ae7-d659-4ff8-9a4a-a3a514c14322 · inbound

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning cites this paper.

The Quantization Trap: Breaking Linear Scaling Laws in Multi-Hop Reasoning Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:30:22.043454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T22:27:49.943714Z digest=sha256:cb8f78ecf549a234957921271be1f10de898bed123d043a6e792829d0c33031c

Observation d27ba4df-1089-46e6-9ff2-471ec59aa2f9 · inbound

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs cites this paper.

Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:49.914962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T07:49:07.128814Z digest=sha256:95094fae881210adf9c22a943c3910b196d2fb9be6ff249774a3868e1cbbded5

Observation 82316228-15c4-40da-99e9-3211db6570f9 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 154

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.154251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:877e08f8445c659371bd4d6adf491269c5e463db2755ce6f1ce3599495612634

Observation 2f586a10-8e87-4812-83ad-77e986131c3c · inbound

PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference cites this paper.

PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:09.483062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:09.483062Z digest=sha256:6b90728df3c0f5196460e62ce4872c3323aaecde3e148acdcd5365737a640e7a