Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T06:56:51.829741Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2412.14590.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T06:56:51.829741Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T06:20:07.112455Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T00:49:18.139563Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 37e62c2a-8bcb-451e-ae7e-285763dcdce8 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 80233a95-9765-431e-8da9-c8aec65df4f4 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Gulavani, Alexey Tumanov, and Ramachandran Ramjee
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3d05d206-a179-4663-894f-efe5f138528e · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a79ce376-f19e-4d98-b145-ba71dbc37eb5 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e2c2e61-c1ba-4ead-b6a2-fa5ad332ef80 · outbound
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b435e830-b05f-4aed-a269-f03613b4b5c5 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design A systematic classification of knowledge, reasoning, and context within the ARC dataset
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 328dbda5-aab3-4f19-9012-b69289a0ed10 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf5f2a50-3df0-4687-92a0-c22a5f281592 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Quip: 2-bit quantization of large language models with guarantees
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 44d0ce27-b02e-4ede-9552-4d15ae1f48d9 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b567af91-f51d-4045-b786-727f64249083 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 36e87cfe-e0b8-4882-80f6-23765d942231 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Spqr: A sparse-quantized representation for near-lossless LLM weight compression
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5dd118ca-c11b-4669-9c0a-5f7cfc8f4d10 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Mahoney, and Kurt Keutzer
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 625e4c1c-2724-4088-9700-81c4b2dfb807 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Optimal brain compression: A framework for accurate post-training quantization and pruning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d4685eb-7449-43f7-8289-f7c56f092f09 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c7e285a5-b683-40a4-abc5-450fa0e3c48c · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design A framework for few-shot language model evaluation, 07 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15924654-a6e9-4bb0-a154-96407715c746 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design The Unreasonable Ineffectiveness of the Deeper Layers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c2d872c-6dc6-4506-bb6f-d8a2b6a34909 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Qwen2.5: A party of foundation models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2b793956-7161-48b8-8df3-de9b75fb91c2 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Stork, and Gregory J
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c8721d8-0d08-4ff2-ad35-51587c21c02f · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dd189893-7f68-4131-ae81-4c7c4203e583 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Mistral 7B
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fccbe28f-94f4-442d-b8c0-fcc1435bc337 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Mahoney, and Kurt Keutzer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7ebcf684-9784-4b84-b16c-b8a1dc19ec57 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Scaling Laws for Precision
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5d490239-93d0-46b6-87d2-8ccca3958243 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 77a1fb6c-d6b1-455a-9e77-e80c3881cf1a · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Denker, and Sara A
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 46af8c01-ddc2-46c9-8737-68ebd542122a · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design OWQ: outlier-aware weight quantization for efficient fine-tuning and inference of large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96ff1376-e8d0-45d4-bd22-8ccefc5e0f74 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design AWQ: activation-aware weight quantization for on-device LLM compression and acceleration
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 25cf0132-d301-4547-b25e-42fbae3358b2 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5cc3f7dc-0408-44ef-8b6a-a465057f4807 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design SpinQuant: LLM quantization with learned rotations
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 082821a9-6969-4256-a07e-71703f945bdf · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Affinequant: Affine transformation quantization for large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5fa5e41-febb-4a3d-acb1-bc99b492e866 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 27706596-3357-4875-9fb2-efb2ad40f6f8 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Pointer sentinel mixture models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5925bd4b-8abd-4459-8de0-e2626acedd48 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7ee67852-e4db-4291-b213-f86024dfce4a · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 24699ae9-ba64-4717-8032-0503701a09f8 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 70aa6d0d-76cc-418b-9d3e-b091e345afaf · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df2d04e5-ba55-4ded-9db7-d4e2e52ad99f · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 150ed859-5ba6-42ea-b95d-902aab744314 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Omniquant: Omnidirectionally calibrated quantization for large language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a916006-04de-4242-a580-cce8adaaf96f · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Musr: Testing the limits of chain-of-thought with multistep soft reasoning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9ab5c030-5907-41f6-9520-ca5bbb367aad · outbound
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1af3be82-8754-4986-9e91-b642720089cf · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Tensorrt-llm
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14bf5fc2-f624-4d6f-8bad-89df8cbd4141 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eed79741-d8c4-4379-8170-d62df4d12977 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 76c4fb00-95ce-4142-b113-fc6415d154b3 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Flash-llm: Enabling low-cost and highly-efficient large generative model inference with unstructured sparsity
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a272b44b-2ee9-4610-80eb-f1e7cca57fee · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 852e474e-3a66-475d-af0f-0d40db242d5b · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Smoothquant: Accurate and efficient post-training quantization for large language models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b98ecc84-49b8-424c-806e-381e79472077 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 06da1fc8-e47f-4710-8f4f-da7a2c3af0d9 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Orca: A distributed serving system for transformer-based generative models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f4de0dba-4110-4ca8-9826-474856e95ec3 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design RPTQ: Reorder-based Post-training Quantization for Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1c153e90-38a1-464b-9575-64af6d836d91 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 288a7325-3656-4722-bda6-b4afa6948d63 · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Atom: Low-bit quantization for efficient and accurate LLM serving
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c7fae06-d6b4-4e1a-8f53-58f68161af7f · outbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f7421b5f-a3f4-4fec-90df-40557947d13d · inbound
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 357fa1c1-a011-493d-b621-bca2b712eb82 · inbound
Displacement Is Not Direction: Evaluating Fidelity Metrics for Quantized LLM Deployment MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 58e27d8e-bafb-43ad-9522-5e1bd0616350 · inbound
Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.