Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:05.850300Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2505.14638.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:05.850300Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 236c17b8-9918-47c9-aa7d-3e09aa1d7f80 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Learned Step Size Quantization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b1badb-5fc8-4880-ac05-496904854c23 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Ef- fective training of convolutional neural networks with low- bitwidth weights and activations,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c1403f1-ef5d-45e0-9004-6e41e8b4630e · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Cluster- promoting quantization with bit-drop for minimizing net- work quantization loss,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c5fd896-d33d-4a13-8799-4efc3d93369b · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89a56030-cdbd-4dc1-94f1-e2a0167bf1b3 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Data-free quantization through weight equalization and bias correction,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0c27305-2a9d-404e-a070-7be9039062a8 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13814e29-fb9e-4965-9a5a-ca9090c6622c · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Accurate post training quantization with small calibra- tion sets,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c581bb25-a830-4af4-a641-a331f5e4e420 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Post-training quantization for vision transformer,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0475f67b-30fc-4443-97e9-829742ba1e62 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Loss aware post-training quantization,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 080915ef-394e-4e00-845d-b555b13a803e · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e6e918e-5112-4acb-8998-657589850c16 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Up or down? adaptive rounding for post- training quantization,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26777c0e-02f8-4ef8-8aca-08f2bc01fa26 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb585fe7-2705-479a-84eb-424dbe2362c1 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Awq: Activation- aware weight quantization for on-device llm compression and acceleration,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d959deeb-80d8-4ca0-9b45-3fab8ec4f7a9 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Q-vlm: Post-training quantization for large vision-language models,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d0470db3-9213-4550-b031-f30ff1dead43 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Advancing multimodal large language models with quantization-aware scale learning for efficient adaptation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71c0371c-cbd3-4427-9d06-8d0ae17c2a4b · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Reg-ptq: Regression-specialized post-training quantization for fully quantized object detector,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 246f09ef-fe89-4379-8f6c-7cdabb71c9fb · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Ptq4sam: Post- training quantization for segment anything,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16077d71-deda-48f9-8cb1-53adac73f52d · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference FP8 Formats for Deep Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bd83e60-759d-4ada-924b-038855fe6362 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Faster Inference of LLMs using FP8 on the Intel Gaudi
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 838360a2-baa4-4155-99d9-7c9a44345f88 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Fp8 quantization: The power of the expo- nent,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c798dbe1-7d59-4c39-bb16-ae3069122cd4 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference An Inquiry into Datacenter TCO for LLM Inference with FP8
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb5b9e2-2d15-4d7e-a109-12e46ba393d6 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference FP8 versus INT8 for efficient deep learning inference
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 737d520a-c7aa-464c-97e5-ec41adcffdce · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2041b6a0-c748-49ed-bdc0-c59c99c7781a · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference The Super Weight in Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 256a2434-d6c1-47b5-af54-542cc5f7bff8 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Smoothquant: Accurate and efficient post-training quan- tization for large language models,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ff6de07-f271-49c8-89c7-15423dd1ddcc · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 776bbb39-0683-4ba8-a428-617b2fe31cc3 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference SqueezeLLM: Dense-and-Sparse Quantization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3930d6bd-754d-4e32-b8df-b762c1043d1b · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Towards accurate post-training quantization for vi- sion transformer,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95cf1fa8-c207-4e47-bc0a-7d556ba4f0a5 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Noisyquant: Noisy bias-enhanced post-training activation quantization for vision transformers,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6c4e8709-e7a7-4e5d-9709-47d47966f795 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f4a9bee-0eb8-4885-ac3b-a90dbc2e4367 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 288a4b30-0a7e-44a2-9498-3fe66a8649ad · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Half-quadratic quantization of large machine learning models,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0acc76a7-1570-4324-9422-4a8600415e8c · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Pd-quant: Post-training quantization based on prediction difference metric,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93cd2ddc-5142-4fe7-b545-2e32eeb6e68b · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Pd-quant: Post-training quantization based on prediction difference metric,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1bc2262d-b2c5-4f0c-ba8e-28df625beb99 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Optimal brain sur- geon and general network pruning,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa35d4cf-af9e-4a6c-b26c-467877ef1d27 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Optimal brain compression: A framework for accurate post-training quantization and prun- ing,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24f09a99-4a9d-4cae-bad6-bd6e87e8ce7c · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Ptq4vit: Post-training quantization for vision transformers with twin uniform quantization,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f168e218-e86b-4e7b-932e-e33431ef6247 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference QQQ: Quality Quattuor-Bit Quantization for Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd3c5138-7ee3-4af2-8310-3fa74fe23239 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Qlora: Efficient finetuning of quantized llms,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45deaeb0-6ad2-445a-a43a-de4b52f0b1e6 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference SpinQuant: LLM quantization with learned rotations
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f2c8c5c-774a-4a46-9200-b00f149dacea · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Massive Activations in Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8afbdc7-d703-48ec-8fb4-9b0bbc7b8aec · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Mmmu: A massive multi-discipline multimodal understanding and rea- soning benchmark for expert agi,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4574bcbf-a4c9-4416-8eea-b5ec034b1866 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Mmbench: Is your multi-modal model an all-around player?
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3745ee05-ec80-4ae7-a844-0a136a981a5c · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00b31de8-d9bc-4270-b39f-07fc5ad2812a · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42cebc4f-9923-4baa-be4b-2fc40dc95cfb · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Vlmevalkit: An open-source toolkit for evaluating large multi-modality models,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bfbd2ce-a4ad-4242-abdb-26744edb88a7 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e06e537-d460-4808-8db1-67c7902ccab0 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference The Llama 3 Herd of Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6344a6fe-09f4-452f-8914-209142fd5ef1 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Pointer Sentinel Mixture Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c1037a0-f90d-49be-a677-3cf71978fae8 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Measuring Massive Multitask Language Understanding
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d12ea00-6193-4ffc-83d4-73c7f518e560 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bff160fa-db2b-4996-9a0b-bc3ed1d1035b · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Language models are unsupervised mul- titask learners,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 200b87c5-d284-4a8d-9322-28264c0cb5e2 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d74b1fd4-21b7-48a3-be62-12be5560fa0b · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 773700e1-cc7f-484f-9068-9c7681b95810 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Reason- ing about physical commonsense in natural language,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 307dd09e-be6c-4d85-8b7f-2c63168013cf · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference WinoGrande: An Adversarial Winograd Schema Challenge at Scale
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2267353-2366-4298-b1d5-bf7637c81f46 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14025751-8455-4e36-a77e-90ac1c071fa5 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference A framework for few-shot language model evaluation,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3973de9a-8f95-417c-9b88-ddac976c38c2 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference Semantic parsing on Freebase from question-answer pairs,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82a53ce8-7103-4aba-baf1-5aa3ba9503b1 · outbound
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.