Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:22:39.003914Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 9 inbound Pith citation observations for arXiv:2412.14363.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:22:39.003914Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:48.606027Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T12:59:52.356456Z
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fda058dc-91e5-4449-8546-40e713c32df8 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SliceGPT: Compress Large Language Models by Deleting Rows and Columns
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98d6d846-a6bf-4649-8968-14bffd0b998c · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QUIK : Towards end-to-end 4-bit inference on generative large language models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29ebbb2a-601c-4599-8f5c-79cef0d22026 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf70a50-b769-449d-a644-eacf1bbe934c · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals L ong B ench: A bilingual, multitask benchmark for long context understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4906daad-b726-4035-baa8-4ed936460af0 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals PIQA : Reasoning about physical commonsense in natural language
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 323c0ab9-4f1b-4b74-ad4e-9bb33b0d543a · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Palu: Compressing KV-Cache with Low-Rank Projection
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73f4de28-cf69-4dbd-9b85-e5de4284169e · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ea7bf735-f1ff-4498-b177-db544f4e548b · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals PACT: Parameterized Clipping Activation for Quantized Neural Networks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5a31317-a4a4-42c1-a1c2-c64253763a13 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 888b947e-c520-44d0-b28f-42183e4f8769 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca9a8a28-4d9b-4d4c-9079-ec3c659ea1e0 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Training Verifiers to Solve Math Word Problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9735236a-e432-4f9c-a318-640e13cfae42 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa5ce59d-3650-4b85-9f2b-1675639fee9f · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd55bf82-363b-46c6-ba82-1cd0ece397af · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QAQ: Quality Adaptive Quantization for LLM KV Cache
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 012ecf26-b6d9-431e-96fb-17a1f35b1b75 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Extreme Compression of Large Language Models via Additive Quantization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0300c697-a7cb-4e93-808d-7971eda845a7 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9320ed35-7913-479d-8256-8bc779be1619 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals A framework for few-shot language model evaluation, 07 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98a1aaa1-5cee-4e4c-b7d6-b28d76ef90f9 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals W., and Keutzer, K
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a187354c-1988-4510-bcea-69d363d219c8 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b491a92e-362b-4c19-8fb1-d7eee7902bd1 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals APTQ : Attention-aware post-training mixed-precision quantization for large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eb1cb692-bfda-4467-89e1-2f8988c49a39 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85af2621-dfc3-414b-8eb8-47ee1a297c0f · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Measuring massive multitask language understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aab0856-dddc-441b-bce0-76553e2988e1 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e41327-fd94-452c-9613-0033d004c734 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 180cdfcd-5a4f-449d-9476-ff9ccfa7946b · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Accurate post training quantization with small calibration sets
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8a50b5f7-d468-4e8b-bfbc-e2ed5fb7edd7 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5194d093-8a59-44b6-8c34-a4f59fa04511 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SqueezeLLM: Dense-and-Sparse Quantization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01e71233-e6d8-40b0-8ba8-b8bf66f08fbf · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals OWQ : Outlier-aware weight quantization for efficient fine-tuning and inference of large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3670b04e-1dd8-4b6e-adfc-e19f7341568e · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SVDQuant : Absorbing outliers by low-rank components for 4-bit diffusion models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b295b582-d3bf-47e4-b5c9-8ef67398a795 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8518c039-3763-4a23-8a5d-9a0328f83243 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Duquant: Distributing outliers via dual transformation makes stronger quantized llms
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f5882700-fdd5-418b-bdd4-ddb5ba719e1e · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals AWQ : Activation-aware weight quantization for on-device llm compression and acceleration
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 672889be-de83-492a-8334-d73da03b3c3f · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02159c26-b30b-4751-8b7a-edb233335d2d · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a409a826-96ac-4f73-a143-d94ac889b68e · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a384ac31-e490-4143-920b-a8521ea8b37b · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ada04707-734b-4730-a56a-be7f57fb5790 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SpinQuant: LLM quantization with learned rotations
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 657afab0-bffc-4382-9b2d-d373b37bc0e7 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Pointer Sentinel Mixture Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e525663-48d4-4e72-bfa3-8bf9528542f0 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Llama 3.2: Revolutionizing edge AI and vision with open, customizable models , 2024 a
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 71408979-c82a-4e96-84d7-4e1d34ae3197 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Introducing Meta Llama 3: The most capable openly available LLM to date
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bbfbb140-674e-4a7a-bc1f-b9094136d1c0 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Can a suit of armor conduct electricity? a new dataset for open book question answering
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2abd8d63-e3fd-4201-bb72-1a8e6559d684 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96cef960-686f-4279-9847-37ee47d223ca · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Pytorch: An imperative style, high-performance deep learning library
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51a6c593-6f2c-4792-9f37-ba3063208c1f · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals L., Bhagavatula, C., and Choi, Y
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2f89d0e0-1866-4a9c-be81-8e66a142ad62 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ESPACE: Dimensionality Reduction of Activations for Model Compression
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cd559363-2a2b-4f94-977b-17763261db36 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Social iqa: Commonsense reasoning about social interactions
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1ad94519-701c-4131-85b9-67b86ff644e7 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Eigen attention: Attention in low-rank space for KV cache compression
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c7b3e8d-46a8-40e0-ac33-5946b21554a6 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbdb286d-76dd-496f-968e-2e0daafb0ce1 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Post training quantization of large language models with microscaling formats
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c44c37e3-0c88-4b40-a9de-4b43cd10ed19 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals FlexGen : High-throughput generative inference of large language models with a single gpu
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0e18ae65-d51c-4214-8b61-b6d113c7b9a9 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals CUTLASS , January 2023
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ae6c5a15-afcb-4760-8200-43a06e826daf · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a3954c-345d-457a-a317-1c830cb69079 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a24d496d-dfa2-4b16-9fb1-8f5f6267e687 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f376dde5-0c87-4d47-9ce2-ec305c96d1ec · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9d1323a-5e62-499c-b10c-edb95c4170a4 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Training transformers with 4-bit integers
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d1cd3c1c-d13b-4952-9c5e-cf547ed10e8d · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Smoothquant: Accurate and efficient post-training quantization for large language models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cefbaddf-3a47-45ac-a991-1b22181e5166 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Qwen2.5 Technical Report
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ba31ebb-25e1-45fb-92bd-33a7d3d2f5d6 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1b72ddd-8ca1-4837-8e6d-45c1e7ce89af · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ZeroQuant : Efficient and affordable post-training quantization for large-scale transformers
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bfdc818f-8b94-49fe-811a-552d050a3e62 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals RPTQ: Reorder-based Post-training Quantization for Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 640b47c3-9b34-4981-8d9f-570908d32951 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62b9c7e8-dff2-47be-8ce9-34a270e56684 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6768f08d-6eb8-4a9f-b10d-ea272f8c4cba · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6ec41da-6b92-4b3d-97ed-afbeef829b46 · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba34eb41-461d-413b-98c4-71b9beb7e17e · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Atom: Low-bit quantization for efficient and accurate llm serving
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 82f4e82c-5f7c-4893-adf2-93f52d8046fc · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QMSum : A new benchmark for query-based multi-domain meeting summarization
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0b525f8a-5d78-4922-8207-1dc469f4a69c · outbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals write newline
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0989df39-3651-4f1f-b73d-8a0fd835d447 · inbound
A Survey on Large Language Model Acceleration based on KV Cache Management ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
Reference 182
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9354d518-b538-432f-b321-777feba70eff · inbound
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db04ed1b-9a1f-47cd-889a-6f9c31fb8190 · inbound
ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 203c50dc-9625-4f66-a193-588f56647cd9 · inbound
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2fa022f9-b499-44ef-8a7c-376882dc0c21 · inbound
dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 74e42f0f-80a9-4853-899d-dc176c6a364d · inbound
dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76203019-c512-4b06-928c-48c9af433a7d · inbound
Multi-Bitwidth Quantization for LLMs Using Additive Codebooks ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0c189f19-d0d7-4bf7-96ca-0de477a6d0e3 · inbound
SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 040d99fe-e452-4a6e-8205-a5dc4e315599 · inbound
GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.