Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 74 inbound Pith citation observations for arXiv:2402.04396.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:43:49.765633Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 28701e4b-a19a-4ee3-8cbf-e3743c6ad444 · inbound
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a9f02f46-12f3-4abb-b454-9281721a0e04 · inbound
A Survey on Efficient Inference for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 216
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6952bc00-5e45-44c9-bd43-e9a65c78c65e · inbound
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f6311aaf-0b51-436d-9fa9-18924762d0fc · inbound
Diffusion Product Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46eab912-d74a-47a8-b1b4-90ff15b5743c · inbound
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17d524fe-736b-4d72-921d-460b2db5530e · inbound
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be99e24f-3f5e-45d4-a5c0-0b4fa7646747 · inbound
DFRot: Achieving Outlier-Free and Massive Activation-Free for Rotated LLMs with Refined Rotation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ece6ac9f-a26a-4754-8d2f-c35e67c3dbb9 · inbound
Low-Rank Correction for Quantized LLMs QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e9c0803-25cc-40ff-8915-b81ef49d95d1 · inbound
HadaCore: Tensor Core Accelerated Hadamard Transform Kernel QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a3954c-345d-457a-a317-1c830cb69079 · inbound
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8116dac2-c1ee-4493-a86f-e24760ad3bdb · inbound
BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9811bac-a8e1-4076-997c-5feed5c2e8e6 · inbound
Dissecting Bit-Level Scaling Laws in Quantizing Vision Generative Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6537f63c-1669-4b9b-a292-4bb7a9402611 · inbound
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aff4718b-c0a3-4b33-bcaf-1f30955985b8 · inbound
AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9926d273-81d8-46b8-9d2d-605d88f75458 · inbound
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b18d4993-1044-44d9-a6ac-83d11e28ebb1 · inbound
RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60e0babf-5767-441d-9c9d-7bd97b5ee134 · inbound
ICQuant: Index Coding enables Low-bit LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6712259-3d05-4354-8759-50afda2488a3 · inbound
One-for-All Pruning: A Universal Model for Customized Compression of Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 960df4e4-6bd5-4ba9-a664-f4c54df000b8 · inbound
High-Rate Nested-Lattice Quantized Matrix Multiplication with Small Lookup Tables QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f1435f5-aada-4861-9453-291e952d78cf · inbound
Rethinking the Outlier Distribution in Large Language Models: An In-depth Study QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abf25eb5-8f0d-4593-965a-6a9964b135d3 · inbound
FPTQuant: Function-Preserving Transforms for LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9803ddb-fe65-4a43-a073-ffdc263e8117 · inbound
PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 653af550-633e-4c12-99bf-49198969a537 · inbound
LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa5f84dc-5fa2-48b3-b38d-220288c5e4af · inbound
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 76718352-586f-4a20-830d-36091d474f31 · inbound
BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 844df51a-8dcc-4883-97e2-cff04b4e25a6 · inbound
CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b468cf42-94a4-4060-b90c-7ff0b58c9d84 · inbound
Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bc2d4450-97ef-4fca-ad2c-fd26dad50cbd · inbound
From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 334fd9c3-d10f-459a-86ea-b26c2e77af96 · inbound
High-Rate Quantized Matrix Multiplication I QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e535b9b4-fc0d-45ed-b93e-4299509bb3ba · inbound
VersaQ-3D: Architecture Support for Visual Geometry Grounded Transformers via Versatile Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4b4a7bf-04bf-43f8-9125-e543a17f50a8 · inbound
SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b21c0b-812c-4736-98e7-520c6166b220 · inbound
Price of metric universality in vector quantization is at most 0.11 bit QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5850f303-8b23-4b49-8c14-9523306ebce5 · inbound
Leech Lattice Vector Quantization for Efficient LLM Compression QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 979a4726-743f-460c-995d-f138547b3c09 · inbound
Rethinking Residual Errors in Compensation-based LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a82fc49e-4cc3-4433-a214-14351a05f5a7 · inbound
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1c84b740-b831-4017-9c08-7a1429a8d62b · inbound
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ded0d6f2-9429-4b4d-9d74-40da99bc914d · inbound
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7faa9924-73d0-49a2-ab8e-b55ec9c3b411 · inbound
From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d90fc254-7403-4cc7-9e95-459904760672 · inbound
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 65084898-a928-49f3-8b7d-d021bd7234ed · inbound
BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1277e456-2dab-422e-8fe1-41b5f467ee08 · inbound
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e74c1b17-4971-445f-9e69-2044a2204c1e · inbound
Search Your Block Floating Point Scales! QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 157
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fa5e0692-584d-4ebb-a1c8-c1a8e06a9df5 · inbound
High-Rate Quantized Matrix Multiplication II QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8e2d83e8-3f0b-4d30-a204-174f4a9231be · inbound
High-Rate Quantized Matrix Multiplication II QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 69e0ff21-96d9-40f6-a763-1e2993afd757 · inbound
XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7afb0a30-8a77-4ba0-ba68-76932eac5c88 · inbound
GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5c711e35-6b5a-40c7-aeaa-e3055fc33baa · inbound
Theory-optimal Quantization Based on Flatness QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 66a62db8-dd20-40f8-84a5-1e764908bff2 · inbound
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9a056b76-2bae-4243-beaf-d2045c7cc745 · inbound
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a922941e-80d6-4ce0-b902-4094c5e34ef0 · inbound
MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a829676f-80bb-44d5-9680-467e235be0fc · inbound
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c625abbc-a448-4950-a656-559e37105655 · inbound
Influence-Inspired Spectral Rotations for Extreme Low-Bit LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 66a62b39-59be-4d24-821d-bb8eb1562754 · inbound
Quantized Reasoning Models Think They Need to Think Longer, but They Do Not QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c13b1c42-78a6-46b0-958b-7e8ed3a4b80c · inbound
DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c9ce07cf-3804-46ab-85a5-3a24dc1b8fa8 · inbound
LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bde79f78-77dc-4e2e-b4b9-1ee44d2b52f9 · inbound
Qift: Shift-Friendly No-Zero W2 Post-Training Quantization for Rotated W2A4/KV4 LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d3d7e18c-cfd3-49cb-8c9a-7aed78fa2441 · inbound
LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d37ff3ed-95ea-4597-a431-61d0017580b9 · inbound
LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 24c04631-cb20-4438-892f-108233728ee6 · inbound
Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 93ff78ee-bb36-434a-b4af-36eaf89af560 · inbound
FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c0132601-3a29-4c84-96e1-d0977ae20abc · inbound
FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70b88107-7b85-4d66-a397-62ab0c9862d1 · inbound
HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2f57c8e7-0ff3-4570-96ce-312bc4869769 · inbound
GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 75cbd6cf-71ce-45fd-bc84-c577b9e28297 · inbound
Reliability Scaling Laws for Quantized Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 160
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f7f2d25-95d6-4df6-a7d3-18d90b9d3ac6 · inbound
Cross-Layer Error Compensation and Finite-Sample Feature-Statistics Matching for Extreme Low-Bit Quantization of Large Language Models QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d47cdbc-5d5d-4238-836f-77290468ea3e · inbound
Break Through the Compression Bottleneck: From Theory to Practice QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 829af20d-2a27-471e-a618-badf62ddb760 · inbound
A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34d4b659-a33f-4912-8ed3-172af34aa219 · inbound
GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2ee7334-ce78-4440-80e0-3553412fa3c8 · inbound
TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfb16197-e846-4ae6-9451-ab02bc3c76a6 · inbound
CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86d060aa-a84a-4244-8599-ba7357fc836f · inbound
Hidden Language Consistency Phenomena in Reasoning LLMs QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 246
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24826879-a826-4232-ac42-41c68bffba9d · inbound
Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d65f970-4d23-4d0e-a014-63b7c6df67a4 · inbound
SoftWater: Class-Aware Rate Allocation for Softmax Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 772d053a-e0ab-4e07-b7bf-fd32b065b019 · inbound
When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.