Pith. sign in

Paper Citation Record · LEDGER

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs

As of 15 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2506.09104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09104 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:04:03.961067Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T23:28:12.790404Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:33:28.258102Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 21764a78-b4eb-47e1-b8d1-ce16dd3bed6c · outbound

This paper cites Paretoq: Scaling laws in extremely low-bit llm quantization, 2025a.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Paretoq: Scaling laws in extremely low-bit llm quantization, 2025a

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:02.768971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:02.768971Z digest=sha256:e770fd67d0f9962f2b7d64c71d11945469cb097585d734fbf1f663e1a8222a04

Observation 8c4cb3a6-46e1-4183-bac2-86cdaabd1031 · outbound

This paper cites Evaluating Quantized Large Language Models.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Evaluating Quantized Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:02.889275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:02.889275Z digest=sha256:904c9a43c1e379bbb179e6384cb416acad01afeb07eb06df2d576b9857cb8be0

Observation 87080d35-f0c3-421d-aa94-a84a33b5ce14 · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.029383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.029383Z digest=sha256:686ab8dcb68be19ffe2d2212fa66dd186a72adcf14674f4bc88151b1795cfc2e

Observation 8e479b64-4769-4745-8936-23e0b4357318 · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.157167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.157167Z digest=sha256:6f156ac70e3a941e0e991fe85473a24d70d2d8fc15a1c4ff4e324dfbdaea06c9

Observation fd29aa23-4d4d-4923-ac68-3e36efdde6b6 · outbound

This paper cites BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.290534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.290534Z digest=sha256:3d861b04a4e6e25fb59c65c36e25dbd70d29732ee1be6d6e11f10194b9ec4d3e

Observation ea7a7258-9f9b-4821-b58d-3b2a190050b8 · outbound

This paper cites Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.353964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.353964Z digest=sha256:86863ed86bf565f642b653c859ba02f820c30903770e1ce09737bbb4fae7c32f

Observation 3a60f12e-39ee-4ef5-8bf1-2a97c447cfb6 · outbound

This paper cites The Llama 3 Herd of Models.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.408635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.408635Z digest=sha256:649851a967d234d4f100d08c41793626012442a2f34e179a32de6f3bc5738829

Observation d16d2e05-dc17-4e7e-afc3-1719d2005b6e · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs DataComp-LM: In search of the next generation of training sets for language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.563789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.563789Z digest=sha256:d704be5377ad628d765862e51f6561d59003bd394bbf37647b03f03af944f79a

Observation 07f79d10-4c5c-4fd6-9026-00fa5c75fbbf · outbound

This paper cites Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.737130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.737130Z digest=sha256:8bf374b18a4d2754561cf0420f989eead6e2a2d5aeae226b7ccc0563bbef35eb

Observation 4c4e4d04-5021-47de-b862-c01e66c963de · outbound

This paper cites A Practical Mixed Precision Algorithm for Post-Training Quantization.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs A Practical Mixed Precision Algorithm for Post-Training Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.801268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.801268Z digest=sha256:40861cf7e90f1102ece785b901f1c465907e1af0d75bd3f5aa0e1534da5287d8

Observation 35e5ee0c-d156-4529-be51-ea568a4697ae · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.858593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.858593Z digest=sha256:ce0ce5f0b67b5776538dec2894d2b8b7fb94dd9140960cdca307446aeda34216

Observation fea0784d-eb2d-47e7-b56c-067165aa9c26 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.905843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.905843Z digest=sha256:0f8140f29080e7f9c694090f837189e244391414d5d9259a5eb21546a7c79dca

Observation fb904247-fbb2-49d0-99c1-9eaac1453974 · outbound

This paper cites an unresolved cited work.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:04:04.513967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T05:04:03.961067Z digest=sha256:1119f21dc19381a26548f64a2c4b4051b23f97a90c1554b41d80b20611ffc2b8

Observation e30bb9b0-cc2f-4e27-b647-e26ed7790c2b · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.637756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.637756Z digest=sha256:426232c5cd0f6e76cd7312f11c36227ccb26ddd6fa014ea99907eac0752891c5

Observation 9587349f-bf95-48ff-98dc-0462929cb415 · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.689898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.689898Z digest=sha256:56dae28098eb9b18dd98a59f65d615c54eab9504828e2caa5034e631e1ecd49b

Observation e385ac98-30e5-4afb-b474-7c28198c8c49 · outbound

This paper cites doi: 10.18653/v1/2020.emnlp-main.494.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs doi: 10.18653/v1/2020.emnlp-main.494

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.239711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.239711Z digest=sha256:fc08cd36f8fdbd406e3068c3a8f582e0e340e42c0310fb8c5b5c247cb0fbc81c

Observation c180ca09-9f15-48a2-b7ec-15c3a2e27e4e · outbound

This paper cites Sharpness-aware Quantization for Deep Neural Networks.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Sharpness-aware Quantization for Deep Neural Networks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:02.956190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:02.956190Z digest=sha256:743579dadd64ea26d930fa0230738e5223077df8c43acd8c20d2209107a87102

Observation 05946ffe-d495-459e-abf9-48c568fa523b · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:02.811644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:02.811644Z digest=sha256:9d721175ca98ff0665389c973216f8ab5d714630a4d43cf81bd45bede574190a

Observation ab2c83ec-7cc4-4bd6-b8aa-0fcee475aed7 · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.098057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.098057Z digest=sha256:137d7a1f7bace1a7f5760e91717a231a32d5771c7f1951c7cb3bc98ce75f9d5d

Observation 6b5bce84-5d51-42b6-9aa0-c786512bb809 · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:03.467515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:03.467515Z digest=sha256:0a3b67a01dc5e7adf75de3a0ea7b7e6146fa2bc062aab8b0ee27ebc6c0610c1c

Pith citing papers

Observation a37eb065-df34-407a-ad1f-1c5b1d403cc9 · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs

Reference 148

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:e78bf82044e3cbd224720632250496e199ad589b5d242726b9bc6727bbf60c5f

Observation f89623b7-ee66-4c66-9230-e0a86ebf87bd · inbound

Apertus LLM Family Expansion via Distillation and Quantization cites this paper.

Apertus LLM Family Expansion via Distillation and Quantization Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.259669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T13:27:47.476242Z digest=sha256:e50552bace1679a6bdb339b4c44ea2d19540d9166c138e35a8a97f8358aaa473