Pith. sign in

Paper Citation Record · LEDGER

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

As of 5 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2412.14590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14590 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T06:56:51.829741Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T06:20:07.112455Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T00:49:18.139563Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact14
  • verified fuzzy28
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 37e62c2a-8bcb-451e-ae7e-285763dcdce8 · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.068343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:bcec22e11430ba654a14262bb15a93b5eaa275ecc3837cb4250d632b783c1f57

Observation 80233a95-9765-431e-8da9-c8aec65df4f4 · outbound

This paper cites Gulavani, Alexey Tumanov, and Ramachandran Ramjee.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Gulavani, Alexey Tumanov, and Ramachandran Ramjee

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.060729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:e37b5f96b7e3c4a9b2959ed5acfa26b58f3a45d9f6be7c2878b4080b16dfc3a2

Observation 3d05d206-a179-4663-894f-efe5f138528e · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.303236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:c7001a740356577498f3951f958f773d92d4d9ac4357132b251160c606b4bb22

Observation a79ce376-f19e-4d98-b145-ba71dbc37eb5 · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.066006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:a3b496ee7db9b1e5307b1c003c8a66aa95b10aa0a077479831aa3a66aac49fe8

Observation 9e2c2e61-c1ba-4ead-b6a2-fa5ad332ef80 · outbound

This paper cites Autogptq.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Autogptq

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.063348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:663d4a985cf9546fda4af741cd50ce205fb003ba2f94bbde1d46b9c60a435569

Observation b435e830-b05f-4aed-a269-f03613b4b5c5 · outbound

This paper cites A systematic classification of knowledge, reasoning, and context within the ARC dataset.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design A systematic classification of knowledge, reasoning, and context within the ARC dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.076269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:c3bc3dc0148e504aa6f8a396f311de5d90c85e973c031dd4d55175f2e89613da

Observation 328dbda5-aab3-4f19-9012-b69289a0ed10 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T06:57:40.294413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:c0d3f834b7c1cefbf22281edccf277290f1c4c9e836d352e4342d921ff228357

Observation cf5f2a50-3df0-4687-92a0-c22a5f281592 · outbound

This paper cites Quip: 2-bit quantization of large language models with guarantees.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Quip: 2-bit quantization of large language models with guarantees

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.073634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:5f61e92859def719f7d74e95e132faa7934e14482ed54e8eafb2ced535569a58

Observation 44d0ce27-b02e-4ede-9552-4d15ae1f48d9 · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.071072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:1aef838984c167aefdaeb03af08dca4787d15cf0928986c8ff91c14d4541fa3c

Observation b567af91-f51d-4045-b786-727f64249083 · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.298346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:a06a27cc2ce021335b747561a2e9d6a1b4be2a1219a194a9c5521de65ff7e871

Observation 36e87cfe-e0b8-4882-80f6-23765d942231 · outbound

This paper cites Spqr: A sparse-quantized representation for near-lossless LLM weight compression.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Spqr: A sparse-quantized representation for near-lossless LLM weight compression

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.133770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:9a4f957371998df9d6fbf8ee5851b111a037135e395930735eeabafac68176f7

Observation 5dd118ca-c11b-4669-9c0a-5f7cfc8f4d10 · outbound

This paper cites Mahoney, and Kurt Keutzer.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Mahoney, and Kurt Keutzer

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.136046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:e355a9ddc5e740ef3ed848acda3df39750bab149c9d211f85b63e1082408a889

Observation 625e4c1c-2724-4088-9700-81c4b2dfb807 · outbound

This paper cites Optimal brain compression: A framework for accurate post-training quantization and pruning.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Optimal brain compression: A framework for accurate post-training quantization and pruning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.131468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:186e38742e2e204d591c5b41aa1e28f1211b6062003999e5445ce8c85c99cd96

Observation 7d4685eb-7449-43f7-8289-f7c56f092f09 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.248123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:689509499498f2e634b9b3a60945ccdc6aba961e80f7ec06b9b39bf54be16ecd

Observation c7e285a5-b683-40a4-abc5-450fa0e3c48c · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design A framework for few-shot language model evaluation, 07 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.145836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:9b5d6e48f4ade1f6a79a67042dd456533b4e056a63aee4bcd3fbb25a4d9876c4

Observation 15924654-a6e9-4bb0-a154-96407715c746 · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design The Unreasonable Ineffectiveness of the Deeper Layers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.285477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:3217933b848e3aad8adbe62b2eaee70eae4d5eeffb88cb28cd6efc6467b971da

Observation 3c2d872c-6dc6-4506-bb6f-d8a2b6a34909 · outbound

This paper cites Qwen2.5: A party of foundation models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Qwen2.5: A party of foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.127122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:14a3ad65d291a98bed13d97232147391623c33a3f71054524390fabe29bb9def

Observation 2b793956-7161-48b8-8df3-de9b75fb91c2 · outbound

This paper cites Stork, and Gregory J.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Stork, and Gregory J

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.129322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:7fa6134866d006e151bd6c2570074e9484bddb75b7a0be3415dbf8be89eed358

Observation 0c8721d8-0d08-4ff2-ad35-51587c21c02f · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.262382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:254d210240e4a4f1a15d03af68c9b4b1cf13d2ae6edfcd632cbc2147cc447887

Observation dd189893-7f68-4131-ae81-4c7c4203e583 · outbound

This paper cites Mistral 7B.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Mistral 7B

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.289405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:a595ea04b93bb70d3373565e001f78bc553e0661489fd9d8b9392bce4d5b8f5d

Observation fccbe28f-94f4-442d-b8c0-fcc1435bc337 · outbound

This paper cites Mahoney, and Kurt Keutzer.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Mahoney, and Kurt Keutzer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.140885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:a7cbcbb8b62e066f96d0fd0c846c2050121bfd16f1257ec39f3a25bc555976a2

Observation 7ebcf684-9784-4b84-b16c-b8a1dc19ec57 · outbound

This paper cites Scaling Laws for Precision.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Scaling Laws for Precision

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T06:57:40.270961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:180a09d26c679d9a91d84b18e2b731ddd8ecb0dae448b89b380047272f15f208

Observation 5d490239-93d0-46b6-87d2-8ccca3958243 · outbound

This paper cites Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.107546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:0bf20512fba6f805b3ff4d6ae23abc7f1a42e5043f624a112a66bb9174029721

Observation 77a1fb6c-d6b1-455a-9e77-e80c3881cf1a · outbound

This paper cites Denker, and Sara A.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Denker, and Sara A

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.124803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:2f30720e0b29d86d86f56904b011e404d2477e6ae880ea38d3e71a6dbe1c8120

Observation 46af8c01-ddc2-46c9-8737-68ebd542122a · outbound

This paper cites OWQ: outlier-aware weight quantization for efficient fine-tuning and inference of large language models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design OWQ: outlier-aware weight quantization for efficient fine-tuning and inference of large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.117685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:b2e9aa411c40fca679b5a9c1e20f2840ca8bca4f8b518f571c5b44e2f841f55a

Observation 96ff1376-e8d0-45d4-bd22-8ccefc5e0f74 · outbound

This paper cites AWQ: activation-aware weight quantization for on-device LLM compression and acceleration.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design AWQ: activation-aware weight quantization for on-device LLM compression and acceleration

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.119902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:2b5c5830d6ef03db9b4558e6028603331ca7cfcf9817b46bdb365137824ab382

Observation 25cf0132-d301-4547-b25e-42fbae3358b2 · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.275531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:ab171913abad54fc4c6cecdcd5b818ee6d7cb99afe55698abff55718e0c9c7e3

Observation 5cc3f7dc-0408-44ef-8b6a-a465057f4807 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design SpinQuant: LLM quantization with learned rotations

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.266681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:0229ced5cec057300b4c6a92268ea5dd2109b44ea95df993d01a0288ecadf161

Observation 082821a9-6969-4256-a07e-71703f945bdf · outbound

This paper cites Affinequant: Affine transformation quantization for large language models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Affinequant: Affine transformation quantization for large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.122100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:cb1831c26d1d57a1a5f28aa288f2572834462d273f9b6b3a9512ce8af71dbccb

Observation c5fa5e41-febb-4a3d-acb1-bc99b492e866 · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.238529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:45d5635f68869968b302115da783202a8feb315bcc21d71b8a390bb2bbf67396

Observation 27706596-3357-4875-9fb2-efb2ad40f6f8 · outbound

This paper cites Pointer sentinel mixture models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Pointer sentinel mixture models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.109864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:e3ace03855ca44befc80f133bb7a95ef26f6595fbadaf06d52e65bec80e46f9f

Observation 5925bd4b-8abd-4459-8de0-e2626acedd48 · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.138711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:7863ff6cb9bcdb349ede6c196578fd380d456978de8ce57692afae55db4b296b

Observation 7ee67852-e4db-4291-b213-f86024dfce4a · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.112699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:15d71477b3f023915b061002fbd3e4218a5e2198aaec54c5d774eb51a1c22fdf

Observation 24699ae9-ba64-4717-8032-0503701a09f8 · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.115047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:406e63e0251b4cc5b61cf123441247710aa2b7a44d3d14d77c47244aa1752847

Observation 70aa6d0d-76cc-418b-9d3e-b091e345afaf · outbound

This paper cites an unresolved cited work.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-23T06:57:41.105230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:3fdf9b53b416fab97bce7b376967dc59c34bcd64fdd78f6433007b40e80f610b

Observation df2d04e5-ba55-4ded-9db7-d4e2e52ad99f · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.233060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:a2e6ac4ab0613ace32f11d870cb06c588e67d19b9241c1047c8f983babd35e75

Observation 150ed859-5ba6-42ea-b95d-902aab744314 · outbound

This paper cites Omniquant: Omnidirectionally calibrated quantization for large language models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Omniquant: Omnidirectionally calibrated quantization for large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.100063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:522cee66918c5253d9115f6da5b0fa836d85e4102c6102de99ad429c06827e95

Observation 9a916006-04de-4242-a580-cce8adaaf96f · outbound

This paper cites Musr: Testing the limits of chain-of-thought with multistep soft reasoning.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Musr: Testing the limits of chain-of-thought with multistep soft reasoning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.097254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:9bc2ac44f696940db97621fc034d8f5d7cf4bcbcfa6ff0819d19d8affc269b72

Observation 9ab5c030-5907-41f6-9520-ca5bbb367aad · outbound

This paper cites Le, Ed H.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Le, Ed H

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.102699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:9e72889febc559a69f9694dabf9fee932e234498c8fadf7a1e3b01e76ff0caa7

Observation 1af3be82-8754-4986-9e91-b642720089cf · outbound

This paper cites Tensorrt-llm.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Tensorrt-llm

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.091846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:4c2f8424863dba89f2268c3e8199a7e4708f42d99135439885a841bcb2c7bdf0

Observation 14bf5fc2-f624-4d6f-8bad-89df8cbd4141 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.280362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:609b7795da0aaefe42a885551032370d76fb24299b8386a01a30c1a36418e622

Observation eed79741-d8c4-4379-8170-d62df4d12977 · outbound

This paper cites ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.258320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:606d4d26f64084614711bb30db45dfd27aa91cc356e3ad94698f163438bbd4cc

Observation 76c4fb00-95ce-4142-b113-fc6415d154b3 · outbound

This paper cites Flash-llm: Enabling low-cost and highly-efficient large generative model inference with unstructured sparsity.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Flash-llm: Enabling low-cost and highly-efficient large generative model inference with unstructured sparsity

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.094814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:e1741ea5ea3ab551e0060deffd3cbe7d75f5e19072c70c70303783fa30a51c1b

Observation a272b44b-2ee9-4610-80eb-f1e7cca57fee · outbound

This paper cites Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Quant-llm: Accelerating the serving of large language models via fp6-centric algorithm-system co-design on modern gpus

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.143164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:ead9bda18e80ddcdc01b408560a33c3866c8beb3c020f2f0efb6b1a34ceae808

Observation 852e474e-3a66-475d-af0f-0d40db242d5b · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.087377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:d37ff64fc6ab7e522fba3bd8e49119f0b69dc014beff4779536fc23f871a76d3

Observation b98ecc84-49b8-424c-806e-381e79472077 · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.084338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:8c393296ccc0e772f82c37967b489ab0fdc02da3ab0dce192593d41d26503116

Observation 06da1fc8-e47f-4710-8f4f-da7a2c3af0d9 · outbound

This paper cites Orca: A distributed serving system for transformer-based generative models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Orca: A distributed serving system for transformer-based generative models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.089673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:92770d4fe9f8e2bda875b7968a3ecc8f99bac510e048e9e502d0be24cc18c33d

Observation f4de0dba-4110-4ca8-9826-474856e95ec3 · outbound

This paper cites RPTQ: Reorder-based Post-training Quantization for Large Language Models.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design RPTQ: Reorder-based Post-training Quantization for Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.243433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:bb72faac5ac0b54492223c4fc4b1ca9dce4f1508d227b697faeba02c920e6350

Observation 1c153e90-38a1-464b-9575-64af6d836d91 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.079285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:a61d4b8b953438ba209597e5f7f4a6e3417803633af2084be274325e6ce8306c

Observation 288a7325-3656-4722-bda6-b4afa6948d63 · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate LLM serving.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design Atom: Low-bit quantization for efficient and accurate LLM serving

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T06:57:41.081983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:a407e32064ed59f0663eaa95d7476aff312e92d9e6453c442916b98f67eeefaf

Observation 3c7fae06-d6b4-4e1a-8f53-58f68161af7f · outbound

This paper cites BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.253161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:8e46a3ae84ff3c527d18a56ccca3b21b761f024b4837897dac572fa935f79eb8

Pith citing papers

Observation f7421b5f-a3f4-4fec-90df-40557947d13d · inbound

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models cites this paper.

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:23:03.651022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T05:20:45.264341Z digest=sha256:1b750f7b8d9af41215859c77aa1e9271b153c9c8b28ee30237f164fdffe05817

Observation 357fa1c1-a011-493d-b621-bca2b712eb82 · inbound

Displacement Is Not Direction: Evaluating Fidelity Metrics for Quantized LLM Deployment cites this paper.

Displacement Is Not Direction: Evaluating Fidelity Metrics for Quantized LLM Deployment MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T00:49:18.141527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T20:58:33.720040Z digest=sha256:3f3fa07c99e6d23d3329712670f1964749d120df3772bc8612a7be22b9042770

Observation 58e27d8e-bafb-43ad-9522-5e1bd0616350 · inbound

Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models cites this paper.

Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T06:20:07.112455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:20:07.112455Z digest=sha256:c2823d7ff03aa9a2897a1aed5308ba1168c34a031fd5e81b2a7ab6b7d3c5aedb