Pith. sign in

Paper Citation Record · LEDGER

BAQ: Efficient Bit Allocation Quantization for Large Language Models

As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2506.05664.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05664 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:55.423637Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-05T05:41:13.869451Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-05T05:50:43.590536Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d542fac4-a47a-4dfa-93b8-fb721f8ba1f6 · outbound

This paper cites Introducing ChatGPT.OpenAI Blog.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Introducing ChatGPT.OpenAI Blog

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:58.981600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:18:51.125122Z digest=sha256:50f050e443a6c942312e0db28970f37831cf31802e393ceae11a4c62832ab42c

Observation 01ae5068-8aa3-42a2-bb3f-e8dc07648b74 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.Advances in neural information processing systems, 36:10088– 10115, 2023.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Qlora: Efficient finetuning of quantized llms.Advances in neural information processing systems, 36:10088– 10115, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:51.215053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:51.215053Z digest=sha256:0ccd72497dc19389731db04b171b849227f54f5f3d01697f877faabebb2e0790

Observation d37587dc-eb2b-445b-8843-73e5bf796cd0 · outbound

This paper cites OPTQ: Accurate quantization for generative pre-trained transformers.

BAQ: Efficient Bit Allocation Quantization for Large Language Models OPTQ: Accurate quantization for generative pre-trained transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:51.356767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:51.356767Z digest=sha256:b03046351330445111e226828dc748f365ca7995d7eb455f38cb9a02a865c486

Observation c4638107-943c-4efa-842d-ece7443e4e08 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

BAQ: Efficient Bit Allocation Quantization for Large Language Models BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:51.557589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:51.557589Z digest=sha256:ac990fe2e0bc71584e76b2fe29206fe7cae2c34a8817f38788733219d94dabf6

Observation adbbb647-40a9-4e9f-b99c-bc4e58e2d344 · outbound

This paper cites QuIP: 2-bit quantization of large language models with guarantees.

BAQ: Efficient Bit Allocation Quantization for Large Language Models QuIP: 2-bit quantization of large language models with guarantees

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:58.643512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:18:51.709490Z digest=sha256:93255ce1563205587eefc94e4f512d1adcfb8b17dd93de67d80a891496bc6e9c

Observation 56f249c7-4c10-407a-bcfc-4a6227bbbff0 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:51.833564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:51.833564Z digest=sha256:95be978c93da96246e8d068a25207987f2cee0cba11b5234a4fa6225093f9ddf

Observation 825ef0da-5999-4b78-9e0e-5247cf497052 · outbound

This paper cites Omniquant: Omnidirectionally calibrated quan- tization for large language models.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Omniquant: Omnidirectionally calibrated quan- tization for large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.035060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.035060Z digest=sha256:4bf6f4fd4b8e43daa342f5a80630bc13c3e79ca979d12e883cdb4c736cc90d22

Observation 991ba312-d4ad-4764-8e7e-2ebcca51e8a8 · outbound

This paper cites Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.178887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.178887Z digest=sha256:a03ad822cb6c3d205d912faaeefe0efeb97a1c4518dbaf53cc91c73a0bb56017

Observation f28f6709-fb4c-4d74-b87c-bdbdfe264093 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.315397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.315397Z digest=sha256:045a19111ad56e27ff7f172a1f24d225b7858fffd4e9b2be9193a2c654e1dd8e

Observation c08e9e6a-75e3-4440-a082-d8b03450aba9 · outbound

This paper cites Owq: Outlier- aware weight quantization for efficient fine-tuning and inference of large language models.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Owq: Outlier- aware weight quantization for efficient fine-tuning and inference of large language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.490197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.490197Z digest=sha256:9c73ef52f9af2414bbdd2341558ddd00374101269307093328bbd6b25a297fc6

Observation 35aa94be-e2a4-4b81-aa83-0b824d01744a · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

BAQ: Efficient Bit Allocation Quantization for Large Language Models SqueezeLLM: Dense-and-Sparse Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.624716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.624716Z digest=sha256:a7935635aa85071f7a5dc9c1559a38453d75a183bfc21fab7f54549c62104cd1

Observation 6b7e28c6-9be5-4042-b20e-ccb2a8fb9961 · outbound

This paper cites Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:52.749496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:52.749496Z digest=sha256:cb0cccd20d08ac77f833a4f7cc917e2a8630e42158a825abccc7ecc136b44298

Observation 52fbcbee-8dc5-48c3-b17d-615644050114 · outbound

This paper cites Optimal brain surgeon and general network pruning.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Optimal brain surgeon and general network pruning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:58.376113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:18:52.920639Z digest=sha256:6f43cf3c321eca9c2578386d4f72784442600cd6db0808d638e1aab402cf28b1

Observation c011a7f5-878c-40ab-836e-7053d7aa0c89 · outbound

This paper cites Optimal brain damage.Advances in neural information processing systems, 2, 1989.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Optimal brain damage.Advances in neural information processing systems, 2, 1989

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:53.069213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:53.069213Z digest=sha256:7b194f6dd240c682973955a9b842fbbcde711186d9fa729e15c035fb89ef254f

Observation e00c5e90-9f8a-4caf-8480-2c0e09f4f89c · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:58.061585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:18:53.189419Z digest=sha256:dd29e0a74b84e2cfd1c6de9414adaccbfc88b23477faa1f6f4d6210eba6df17e

Observation 20cf8c7a-20cf-4c73-a47d-3e33a7b449c2 · outbound

This paper cites Optimal brain compression: A framework for accurate post- training quantization and pruning.Advances in Neural Information Processing Systems, 35:4475– 4488, 2022.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Optimal brain compression: A framework for accurate post- training quantization and pruning.Advances in Neural Information Processing Systems, 35:4475– 4488, 2022

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:53.356799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:53.356799Z digest=sha256:abe461cf939d4a7a3d04eec90dec86734b146689651a7b0b2c7eff7e49274a21

Observation 1e388151-715e-4228-b3cf-11399f60e860 · outbound

This paper cites BiLLM: Pushing the Limit of Post-Training Quantization for LLMs.

BAQ: Efficient Bit Allocation Quantization for Large Language Models BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:53.534133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:53.534133Z digest=sha256:6b768a2bfedb618e9c205725b36b9d9258b0a15ac824409510faafa8990d77bd

Observation 767b367f-c3e5-41fd-9517-9335576483a1 · outbound

This paper cites Gray and David L.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Gray and David L

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:57.715317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:18:53.742757Z digest=sha256:c3d810dcd65a96a2d11474977173f647d79b5555b5f7afe4ef9d7e2973b7eaa3

Observation a549fd08-e043-4463-a65f-5907fc04c75e · outbound

This paper cites John Wiley & Sons, 1999.

BAQ: Efficient Bit Allocation Quantization for Large Language Models John Wiley & Sons, 1999

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:53.878901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:53.878901Z digest=sha256:f1bdb87e0fb47a63da3e8062b6164eca324408ea7ca04172512808b85229e877

Observation c8d1f918-2e0c-40cb-aefc-0f4f22b8207d · outbound

This paper cites Jpeg2000: Image compression fundamentals, standards and practice.Journal of Electronic Imaging, 11(2):286–287, 2002.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Jpeg2000: Image compression fundamentals, standards and practice.Journal of Electronic Imaging, 11(2):286–287, 2002

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:57.352361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:18:54.023364Z digest=sha256:ebc36da1113ec328329d34d0fc4966e540ec05b6bca4c2d0767eb54677d4beca

Observation f486d404-4f12-4790-b780-fea2a955d48d · outbound

This paper cites Springer Science & Business Media, 2012.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Springer Science & Business Media, 2012

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:54.104520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:54.104520Z digest=sha256:4c8e467fb2e1711d01605d682edb98d4daeba64089be9bddc01f2a6b4271aa42

Observation dde2d55b-4767-4ea0-bb7d-46e6631a94c3 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

BAQ: Efficient Bit Allocation Quantization for Large Language Models OPT: Open Pre-trained Transformer Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:54.262302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:54.262302Z digest=sha256:bc769d7584c4802affb5024dbe0f8bc00fdb70d1a8d0eb693b56f5256ab24556

Observation cc1f0635-fd71-45ba-9a48-a4c83394244d · outbound

This paper cites an unresolved cited work.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:54.443826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:54.443826Z digest=sha256:ddf854cbebdf5c333303fb5d4c9518c93efbc96eba08ccab44d5523a33df7b92

Observation 403ab597-f344-4cff-9ee5-8511a7f80c6f · outbound

This paper cites Pointer Sentinel Mixture Models.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Pointer Sentinel Mixture Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:54.621632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:54.621632Z digest=sha256:c7e2783031994d265bf768179dcf671f48acd15201aa0bb931b4a7206d95539f

Observation 34702f25-e960-4259-8b8a-e4a69673bc55 · outbound

This paper cites The Penn Treebank: Annotating predicate argument structure.

BAQ: Efficient Bit Allocation Quantization for Large Language Models The Penn Treebank: Annotating predicate argument structure

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:57.082435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:18:54.807264Z digest=sha256:78286be7dd875b7b6c6649cb1b68cc0e10043d3f02f64ab7745d924b561fa8b3

Observation 8e0adfe8-8572-40f5-a731-c30a19440f50 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

BAQ: Efficient Bit Allocation Quantization for Large Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:56.608156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:18:55.087121Z digest=sha256:47acfc5074e022d3bf1f78c8edb512b5929d8e01efe27816f7857ae200d495eb

Observation 3d4a6eec-b3dc-41ef-84eb-9abcfb0959b8 · outbound

This paper cites Piqa: An algebra for querying protein data sets.

BAQ: Efficient Bit Allocation Quantization for Large Language Models Piqa: An algebra for querying protein data sets

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:56.247548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:18:55.236048Z digest=sha256:ebf201b784d7174511aa6f88603b33e21f9e19eb208dbe8de87d3fa5921c20e7

Observation e378cbbd-1a98-4b0f-9598-387bd629d825 · outbound

This paper cites A systematic classification of knowl- edge, reasoning, and context within the ARC dataset.

BAQ: Efficient Bit Allocation Quantization for Large Language Models A systematic classification of knowl- edge, reasoning, and context within the ARC dataset

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:55.902652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:18:55.423637Z digest=sha256:8b3247c4881a5f6f37bc0ca289a287a63e4f6cbd2ec4606921a80911cbad469e

Pith citing papers

Observation 196ce1c7-1a55-4e9f-98e8-714d0a7fa3c2 · inbound

SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs cites this paper.

SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs BAQ: Efficient Bit Allocation Quantization for Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:10:24.851205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T17:08:05.945757Z digest=sha256:9abf53418a0cb8e227f22e80522013c09a4c30752a534d483019632c2ab605cc

Observation db559af4-13dd-4bac-9c2a-f077de7e7051 · inbound

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory cites this paper.

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory BAQ: Efficient Bit Allocation Quantization for Large Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:40:59.669846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:09:21.811344Z digest=sha256:c230af0e4395f2ab05d21298397ec65298f778fbc0cd922eb5d43ffe2e7bb423

Observation 6267c1b7-0ebd-46ce-94bf-35493f06b28f · inbound

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory cites this paper.

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory BAQ: Efficient Bit Allocation Quantization for Large Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-05T05:50:43.592486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-05T05:41:13.869451Z digest=sha256:aaf0c954ae712f5a09178200cd9a63bb409c2e450f553db6cc673b32c3a7d6e3