Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:55.423637Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2506.05664.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:55.423637Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-05T05:41:13.869451Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-05T05:50:43.590536Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d542fac4-a47a-4dfa-93b8-fb721f8ba1f6 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Introducing ChatGPT.OpenAI Blog
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 01ae5068-8aa3-42a2-bb3f-e8dc07648b74 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Qlora: Efficient finetuning of quantized llms.Advances in neural information processing systems, 36:10088– 10115, 2023
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d37587dc-eb2b-445b-8843-73e5bf796cd0 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models OPTQ: Accurate quantization for generative pre-trained transformers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4638107-943c-4efa-842d-ece7443e4e08 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models BitNet: Scaling 1-bit Transformers for Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adbbb647-40a9-4e9f-b99c-bc4e58e2d344 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models QuIP: 2-bit quantization of large language models with guarantees
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56f249c7-4c10-407a-bcfc-4a6227bbbff0 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Smoothquant: Accurate and efficient post-training quantization for large language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 825ef0da-5999-4b78-9e0e-5247cf497052 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Omniquant: Omnidirectionally calibrated quan- tization for large language models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 991ba312-d4ad-4764-8e7e-2ebcca51e8a8 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Kvquant: Towards 10 million context length llm inference with kv cache quantization.Advances in Neural Information Processing Systems, 37:1270–1303, 2024
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f28f6709-fb4c-4d74-b87c-bdbdfe264093 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of Machine Learning and Systems, 6:87–100, 2024
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c08e9e6a-75e3-4440-a082-d8b03450aba9 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Owq: Outlier- aware weight quantization for efficient fine-tuning and inference of large language models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35aa94be-e2a4-4b81-aa83-0b824d01744a · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models SqueezeLLM: Dense-and-Sparse Quantization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b7e28c6-9be5-4042-b20e-ccb2a8fb9961 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52fbcbee-8dc5-48c3-b17d-615644050114 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Optimal brain surgeon and general network pruning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c011a7f5-878c-40ab-836e-7053d7aa0c89 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Optimal brain damage.Advances in neural information processing systems, 2, 1989
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e00c5e90-9f8a-4caf-8480-2c0e09f4f89c · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Sparsegpt: Massive language models can be accurately pruned in one-shot
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20cf8c7a-20cf-4c73-a47d-3e33a7b449c2 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Optimal brain compression: A framework for accurate post- training quantization and pruning.Advances in Neural Information Processing Systems, 35:4475– 4488, 2022
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e388151-715e-4228-b3cf-11399f60e860 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 767b367f-c3e5-41fd-9517-9335576483a1 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Gray and David L
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a549fd08-e043-4463-a65f-5907fc04c75e · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models John Wiley & Sons, 1999
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8d1f918-2e0c-40cb-aefc-0f4f22b8207d · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Jpeg2000: Image compression fundamentals, standards and practice.Journal of Electronic Imaging, 11(2):286–287, 2002
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f486d404-4f12-4790-b780-fea2a955d48d · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Springer Science & Business Media, 2012
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dde2d55b-4767-4ea0-bb7d-46e6631a94c3 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models OPT: Open Pre-trained Transformer Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc1f0635-fd71-45ba-9a48-a4c83394244d · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 403ab597-f344-4cff-9ee5-8511a7f80c6f · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Pointer Sentinel Mixture Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34702f25-e960-4259-8b8a-e4a69673bc55 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models The Penn Treebank: Annotating predicate argument structure
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8e0adfe8-8572-40f5-a731-c30a19440f50 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3d4a6eec-b3dc-41ef-84eb-9abcfb0959b8 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models Piqa: An algebra for querying protein data sets
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e378cbbd-1a98-4b0f-9598-387bd629d825 · outbound
BAQ: Efficient Bit Allocation Quantization for Large Language Models A systematic classification of knowl- edge, reasoning, and context within the ARC dataset
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 196ce1c7-1a55-4e9f-98e8-714d0a7fa3c2 · inbound
SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs BAQ: Efficient Bit Allocation Quantization for Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation db559af4-13dd-4bac-9c2a-f077de7e7051 · inbound
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory BAQ: Efficient Bit Allocation Quantization for Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6267c1b7-0ebd-46ce-94bf-35493f06b28f · inbound
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory BAQ: Efficient Bit Allocation Quantization for Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.