Pith. sign in

Paper Citation Record · LEDGER

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

As of 14 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 9 inbound Pith citation observations for arXiv:2412.14363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14363 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:22:39.003914Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:48.606027Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:59:52.356456Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fda058dc-91e5-4449-8546-40e713c32df8 · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.755084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.755084Z digest=sha256:b8192a20a6930d33467af7237ba945830690e2720aac8760726dc2dd449232c2

Observation 98d6d846-a6bf-4649-8968-14bffd0b998c · outbound

This paper cites QUIK : Towards end-to-end 4-bit inference on generative large language models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QUIK : Towards end-to-end 4-bit inference on generative large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.760201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.760201Z digest=sha256:30f65eea85fb70100fad432c2fd375c82b5c14672b0d9ee3f7a184ec04c5e4c6

Observation 29ebbb2a-601c-4599-8f5c-79cef0d22026 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.763953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.763953Z digest=sha256:0a0e6c873461391b76b9454839a80f474943ec2501743ebfb9e87dee336ae007

Observation 0cf70a50-b769-449d-a644-eacf1bbe934c · outbound

This paper cites L ong B ench: A bilingual, multitask benchmark for long context understanding.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals L ong B ench: A bilingual, multitask benchmark for long context understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.768700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.768700Z digest=sha256:a3b04dcf0522ab727ec575808d80f15fe23da8f8af3328dddb5d493b5cc6d1c1

Observation 4906daad-b726-4035-baa8-4ed936460af0 · outbound

This paper cites PIQA : Reasoning about physical commonsense in natural language.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals PIQA : Reasoning about physical commonsense in natural language

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.910642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.772381Z digest=sha256:bc07d774a157cb573ce95f992c03416f76a23676415b51b669197520732841d1

Observation 323c0ab9-4f1b-4b74-ad4e-9bb33b0d543a · outbound

This paper cites Palu: Compressing KV-Cache with Low-Rank Projection.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Palu: Compressing KV-Cache with Low-Rank Projection

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.776013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.776013Z digest=sha256:69d39acb5fe85fe1c7255d41290e3e60add86027a2b5070ef4bcdcfe42aa4908

Observation 73f4de28-cf69-4dbd-9b85-e5de4284169e · outbound

This paper cites an unresolved cited work.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:22:39.899036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.780084Z digest=sha256:70d7cbcdf48a470f040d24716609eb7983cebda8716bb2579375f59b07a94e55

Observation ea7bf735-f1ff-4498-b177-db544f4e548b · outbound

This paper cites PACT: Parameterized Clipping Activation for Quantized Neural Networks.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals PACT: Parameterized Clipping Activation for Quantized Neural Networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.783535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.783535Z digest=sha256:b1e00912cd6b2f42890d9c2dc4f22d5076ccf6f032146f40e0efa2d899dd9858

Observation b5a31317-a4a4-42c1-a1c2-c64253763a13 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.787784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.787784Z digest=sha256:ee3f55e4c4a8493f07a553721ef8c38e015dc8ebbed49ea4255444c12ac03ebb

Observation 888b947e-c520-44d0-b28f-42183e4f8769 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.792466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.792466Z digest=sha256:31e356a846840f2cf079a1f9944a8440315413c92832a9df216f9d650267263e

Observation ca9a8a28-4d9b-4d4c-9079-ec3c659ea1e0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.796468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.796468Z digest=sha256:0e11b9996f5854748734992dec422838bac089076d67e2fc1488967534536cca

Observation 9735236a-e432-4f9c-a318-640e13cfae42 · outbound

This paper cites an unresolved cited work.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.800077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.800077Z digest=sha256:e49c33da289ea6a07bcc656d15db14d865dc323fbcaf85f256133d7fa1bdcd0b

Observation fa5ce59d-3650-4b85-9f2b-1675639fee9f · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.803446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.803446Z digest=sha256:62237a11c5399e735e8d083c8f7f355653708b57353c39f86d6d6eed2d039202

Observation bd55bf82-363b-46c6-ba82-1cd0ece397af · outbound

This paper cites QAQ: Quality Adaptive Quantization for LLM KV Cache.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QAQ: Quality Adaptive Quantization for LLM KV Cache

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.806973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.806973Z digest=sha256:4799db9db27358ebb2cebb2d6bb4a91498eff29d6dc15da02dc410cf855ee73c

Observation 012ecf26-b6d9-431e-96fb-17a1f35b1b75 · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Extreme Compression of Large Language Models via Additive Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.810656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.810656Z digest=sha256:7e43357ee272b84fd20bb334c9b77444cf0ac96e999a66d9055ee8e0789f4c70

Observation 0300c697-a7cb-4e93-808d-7971eda845a7 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.814251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.814251Z digest=sha256:cb75e43aa187d65fb520973b576d6805ef2576c85eb51a24853021edd73716c8

Observation 9320ed35-7913-479d-8256-8bc779be1619 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals A framework for few-shot language model evaluation, 07 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.817747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.817747Z digest=sha256:47cc3888c99254129ef80d71ce153d40709cba33729e4a846d87e6ee1240441a

Observation 98a1aaa1-5cee-4e4c-b7d6-b28d76ef90f9 · outbound

This paper cites W., and Keutzer, K.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals W., and Keutzer, K

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.821257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.821257Z digest=sha256:e1b911d16e51f1ae4deb2b8a76d24c9cbc1300a03cdb1b29b339a9d26bb44487

Observation a187354c-1988-4510-bcea-69d363d219c8 · outbound

This paper cites SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.824457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.824457Z digest=sha256:da53928bedbd3817b2ab2b34c7705596aef0b9c9d8421e4541d1cbe0d19d65d5

Observation b491a92e-362b-4c19-8fb1-d7eee7902bd1 · outbound

This paper cites APTQ : Attention-aware post-training mixed-precision quantization for large language models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals APTQ : Attention-aware post-training mixed-precision quantization for large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.874845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.828156Z digest=sha256:6524ca8c3068d3b0e95eade906ce7d9e8f82c91175b4482c4d176a249e2d6130

Observation eb1cb692-bfda-4467-89e1-2f8988c49a39 · outbound

This paper cites ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.831351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.831351Z digest=sha256:2ab4f7cac75f0372d38f8ef1651d731a0d637d47dee74a991b6a47ea0b4fa495

Observation 85af2621-dfc3-414b-8eb8-47ee1a297c0f · outbound

This paper cites Measuring massive multitask language understanding.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Measuring massive multitask language understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.835297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.835297Z digest=sha256:f9fc9029727f6575ed7dd07f6c53aceb1f8623f6989d2df40568fc095cfd2bc7

Observation 6aab0856-dddc-441b-bce0-76553e2988e1 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.838579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.838579Z digest=sha256:f5ac05deb52ff427cb37e9a5dd5556cdd34ba3ae28e7f1cdbde7917defc456ba

Observation 94e41327-fd94-452c-9613-0033d004c734 · outbound

This paper cites SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.842294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.842294Z digest=sha256:7ff454393f7485a79bf75b8e1a121becbfa72beec8a3f179a23d13ad2e88fd67

Observation 180cdfcd-5a4f-449d-9476-ff9ccfa7946b · outbound

This paper cites Accurate post training quantization with small calibration sets.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Accurate post training quantization with small calibration sets

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.856640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.845886Z digest=sha256:ff01a3036025bdbe8e9b76134fa3fea6bb40034c964888df50e008c263637abd

Observation 8a50b5f7-d468-4e8b-bfbc-e2ed5fb7edd7 · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.849655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.849655Z digest=sha256:79c92f7dbfbccaaf1aa70922a64a01a28336e6214d02190deae8aaec88dc8ae1

Observation 5194d093-8a59-44b6-8c34-a4f59fa04511 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SqueezeLLM: Dense-and-Sparse Quantization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.853192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.853192Z digest=sha256:76069ae9979526b31f7ea761830af3ea4682286c687f0e1422956359e4643e75

Observation 01e71233-e6d8-40b0-8ba8-b8bf66f08fbf · outbound

This paper cites OWQ : Outlier-aware weight quantization for efficient fine-tuning and inference of large language models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals OWQ : Outlier-aware weight quantization for efficient fine-tuning and inference of large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.845115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.856813Z digest=sha256:ac25fce1684024c5d4754572bc6adacee7c073c75f065ed155f4cf3a1a6db0b6

Observation 3670b04e-1dd8-4b6e-adfc-e19f7341568e · outbound

This paper cites SVDQuant : Absorbing outliers by low-rank components for 4-bit diffusion models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SVDQuant : Absorbing outliers by low-rank components for 4-bit diffusion models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.860219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.860219Z digest=sha256:eecda3633fa8e2893b2a8f2f8a20c908626ff15beeba59044e0c07e1769587f7

Observation b295b582-d3bf-47e4-b5c9-8ef67398a795 · outbound

This paper cites MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.863760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.863760Z digest=sha256:358bb9f127d6fa76bad901fb4266121479017c0be862f7da2cdb5563a3856b7a

Observation 8518c039-3763-4a23-8a5d-9a0328f83243 · outbound

This paper cites Duquant: Distributing outliers via dual transformation makes stronger quantized llms.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Duquant: Distributing outliers via dual transformation makes stronger quantized llms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.834633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.867491Z digest=sha256:b69ec84a74593ad180e17078728986c3fccf977158b37926d226cf414d483974

Observation f5882700-fdd5-418b-bdd4-ddb5ba719e1e · outbound

This paper cites AWQ : Activation-aware weight quantization for on-device llm compression and acceleration.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals AWQ : Activation-aware weight quantization for on-device llm compression and acceleration

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.823746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.870930Z digest=sha256:3ba6a04214764e342d453e737f66c7d71a516a19f3ca1bf1eef4df59f7b27e00

Observation 672889be-de83-492a-8334-d73da03b3c3f · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.874322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.874322Z digest=sha256:c28668cf776418409951987c841876550c7d7fd5016ec0cf6048933817cd0c5d

Observation 02159c26-b30b-4751-8b7a-edb233335d2d · outbound

This paper cites QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.878253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.878253Z digest=sha256:0c4ea3171262ad9999f029a2c6cabfe70d05a5b400b45779bfffa87405466b53

Observation a409a826-96ac-4f73-a143-d94ac889b68e · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.882488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.882488Z digest=sha256:3b69f5669b0eb1c4f74e580a58957d8742ade96d0a65e611f015bfa48444c85e

Observation a384ac31-e490-4143-920b-a8521ea8b37b · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.885979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.885979Z digest=sha256:462a328b70995fee0f9a578ec7a8ca9cf403539bbe302357830069c9416eb015

Observation ada04707-734b-4730-a56a-be7f57fb5790 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals SpinQuant: LLM quantization with learned rotations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.889642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.889642Z digest=sha256:5e21c84ec703a8d2e99fda5df35689423e4a05a946a8e71cf5ec82b166492ab5

Observation 657afab0-bffc-4382-9b2d-d373b37bc0e7 · outbound

This paper cites Pointer Sentinel Mixture Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Pointer Sentinel Mixture Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.893145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.893145Z digest=sha256:a4edc6e4fbc914ee749a58dc38f218529c763856ba99d73802b8c7fce8e4fed6

Observation 7e525663-48d4-4e72-bfa3-8bf9528542f0 · outbound

This paper cites Llama 3.2: Revolutionizing edge AI and vision with open, customizable models , 2024 a.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Llama 3.2: Revolutionizing edge AI and vision with open, customizable models , 2024 a

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.813238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.896642Z digest=sha256:4596c0c6909c1bf778922b0ca970af59819f42384f059d134e2d48c48104742e

Observation 71408979-c82a-4e96-84d7-4e1d34ae3197 · outbound

This paper cites Introducing Meta Llama 3: The most capable openly available LLM to date.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Introducing Meta Llama 3: The most capable openly available LLM to date

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.801562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.900006Z digest=sha256:48934a94fcc1595ad0acd32d912c39d2ead77a29db92e33758311b2c8595cc7a

Observation bbfbb140-674e-4a7a-bc1f-b9094136d1c0 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.903241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.903241Z digest=sha256:ecdf968ac6de3b816175d9ed52e89719eceb59a57f16ee0742e68e5612f29542

Observation 2abd8d63-e3fd-4201-bb72-1a8e6559d684 · outbound

This paper cites LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.906855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.906855Z digest=sha256:739f012ba51289c86479145046155aa4a31f294d75f925698f14a58cda805e28

Observation 96cef960-686f-4279-9847-37ee47d223ca · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Pytorch: An imperative style, high-performance deep learning library

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.910471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.910471Z digest=sha256:8306f9d3788168a7e4ef4513b3a922b9abfb16989f62e9bb89a144d3d95c49c4

Observation 51a6c593-6f2c-4792-9f37-ba3063208c1f · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals L., Bhagavatula, C., and Choi, Y

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.777147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.913876Z digest=sha256:74a429fb68a94436a13409e6c0e2698b3e319946b9830a783f5180eaadf21472

Observation 2f89d0e0-1866-4a9c-be81-8e66a142ad62 · outbound

This paper cites ESPACE: Dimensionality Reduction of Activations for Model Compression.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ESPACE: Dimensionality Reduction of Activations for Model Compression

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:22:39.284687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.917322Z digest=sha256:f18d2a22843e3c3be5c8d7e4c3b292faeb27e66936e8f64ed7bce3ee793f0ac6

Observation cd559363-2a2b-4f94-977b-17763261db36 · outbound

This paper cites Social iqa: Commonsense reasoning about social interactions.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Social iqa: Commonsense reasoning about social interactions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.765854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.921067Z digest=sha256:d9935a2b76a4a277c8e6a59f7ec05a7ac5a2b5d49b1f285cdf4593f56534d4c0

Observation 1ad94519-701c-4131-85b9-67b86ff644e7 · outbound

This paper cites Eigen attention: Attention in low-rank space for KV cache compression.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Eigen attention: Attention in low-rank space for KV cache compression

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.924318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.924318Z digest=sha256:dbbccef14bc3cd36efd07b0349b27426ae6d172048ae1d81c4ce199cae514feb

Observation 1c7b3e8d-46a8-40e0-ac33-5946b21554a6 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.927992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.927992Z digest=sha256:e8707cd81f8cc44c6bdf9670e65ff773f0ad43142d37361c21e7abada3296628

Observation bbdb286d-76dd-496f-968e-2e0daafb0ce1 · outbound

This paper cites Post training quantization of large language models with microscaling formats.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Post training quantization of large language models with microscaling formats

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.755113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.931852Z digest=sha256:19d346403cbb45bfbb3633caa8007289d619edbc18824637fcd387cc13c9d677

Observation c44c37e3-0c88-4b40-a9de-4b43cd10ed19 · outbound

This paper cites FlexGen : High-throughput generative inference of large language models with a single gpu.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals FlexGen : High-throughput generative inference of large language models with a single gpu

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.743866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.935507Z digest=sha256:3534f81a0e03686a24537ea88b5b5e3611955afb98901993b24734852d91c514

Observation 0e18ae65-d51c-4214-8b61-b6d113c7b9a9 · outbound

This paper cites CUTLASS , January 2023.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals CUTLASS , January 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.732878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.938871Z digest=sha256:419f4486f72898874ac7494355e55eaaa4ddf54700e6e184d653886a1e633d9c

Observation ae6c5a15-afcb-4760-8200-43a06e826daf · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.942577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.942577Z digest=sha256:52a930a898e96b08b4442768c1797b270d366864603da7b60dec1f0894b8bc83

Observation e6a3954c-345d-457a-a317-1c830cb69079 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.946066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.946066Z digest=sha256:450c19486671ea2097038f857f6ccfbde5d7bc58975963cfc43b51ef62891743

Observation a24d496d-dfa2-4b16-9fb1-8f5f6267e687 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.949786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.949786Z digest=sha256:054f7dc0c389cd2920686cc00e195493e54d5d0c732a6d4fc9bebe48743b98c2

Observation f376dde5-0c87-4d47-9ce2-ec305c96d1ec · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.953613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.953613Z digest=sha256:89451bae0617c00fe1be4b323509caa81ea4e83ae09846cf5286811685b7f9d2

Observation b9d1323a-5e62-499c-b10c-edb95c4170a4 · outbound

This paper cites Training transformers with 4-bit integers.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Training transformers with 4-bit integers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.721622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.957266Z digest=sha256:1d7ff40ef5ac040e2198579f864b2a3a2c5f705b33142f815dedf51182cd0add

Observation d1cd3c1c-d13b-4952-9c5e-cf547ed10e8d · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.960869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.960869Z digest=sha256:097322a0210a6665a4c1d1a19d263a21b03b8767af1d297454c4963a605f239d

Observation cefbaddf-3a47-45ac-a991-1b22181e5166 · outbound

This paper cites Qwen2.5 Technical Report.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Qwen2.5 Technical Report

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.964652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.964652Z digest=sha256:196ef56b40bc9e5db9fb08e2e189c0bbf69281f180f4f487df17ca41e12a540b

Observation 3ba31ebb-25e1-45fb-92bd-33a7d3d2f5d6 · outbound

This paper cites No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.968304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.968304Z digest=sha256:6c4b29cd58966aa6a6abd3ceb518dde5c6792559f6feefe944b4cf7798542a76

Observation e1b72ddd-8ca1-4837-8e6d-45c1e7ce89af · outbound

This paper cites ZeroQuant : Efficient and affordable post-training quantization for large-scale transformers.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ZeroQuant : Efficient and affordable post-training quantization for large-scale transformers

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.704547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.972203Z digest=sha256:42faef637766258692eeef8a9f42df44ee0e5f1b2454884d6fb0d44bde65b5ed

Observation bfdc818f-8b94-49fe-811a-552d050a3e62 · outbound

This paper cites RPTQ: Reorder-based Post-training Quantization for Large Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals RPTQ: Reorder-based Post-training Quantization for Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.975957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.975957Z digest=sha256:0ebfcebb4f8a15323758cd5804906da26678506560aefbefbb7da8c4f484b9e9

Observation 640b47c3-9b34-4981-8d9f-570908d32951 · outbound

This paper cites ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.979639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.979639Z digest=sha256:196da467b38490fd1cfcce5d6f34039436f5aaa670eade92c235bd5ac48ee9bf

Observation 62b9c7e8-dff2-47be-8ce9-34a270e56684 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.983535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.983535Z digest=sha256:4c622396085eb6698f00d2d8103dc486b1612cb02fe58ec8f8a75b0d27066ed1

Observation 6768f08d-6eb8-4a9f-b10d-ea272f8c4cba · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.987961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.987961Z digest=sha256:f28ae71d29794be47d3a70b498fe098dfbf73ad761e55ef1b79235af5faea83e

Observation e6ec41da-6b92-4b3d-97ed-afbeef829b46 · outbound

This paper cites ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:38.991673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:38.991673Z digest=sha256:061cb6cb16c3c25e8bf95bbbdc08543ee48c36fb247f83c5a2b812b00b472c47

Observation ba34eb41-461d-413b-98c4-71b9beb7e17e · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals Atom: Low-bit quantization for efficient and accurate llm serving

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.693096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.995854Z digest=sha256:ccb8d99beeebc973194ee9b231e8456042cd8489e68393a029589de59cbd6e37

Observation 82f4e82c-5f7c-4893-adf2-93f52d8046fc · outbound

This paper cites QMSum : A new benchmark for query-based multi-domain meeting summarization.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals QMSum : A new benchmark for query-based multi-domain meeting summarization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:22:39.681667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T12:22:38.999640Z digest=sha256:f41c69b277cd6af3480dfecae7282578759c988f7843e9a57cae08a8e095ec7e

Observation 0b525f8a-5d78-4922-8207-1dc469f4a69c · outbound

This paper cites write newline.

ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals write newline

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T12:22:39.003914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:22:39.003914Z digest=sha256:be284696ba94bcf6078bae41db048917651c9480d287b9d5978679ee64628f0a

Pith citing papers

Observation 0989df39-3651-4f1f-b73d-8a0fd835d447 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:48.606027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:48.606027Z digest=sha256:f8cf315dbe95a5e24739f6e99fb36a002dad95eaa24cca842b1ab1e3ad1157ba

Observation 9354d518-b538-432f-b321-777feba70eff · inbound

RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations cites this paper.

RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T14:54:22.698198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:54:22.698198Z digest=sha256:14c45b0b9db2a3f07017a00dea21c0bb95994c1ab36d2801859ab86ace78f2f8

Observation db04ed1b-9a1f-47cd-889a-6f9c31fb8190 · inbound

ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs cites this paper.

ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T11:10:52.707779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T11:10:52.707779Z digest=sha256:bc069e7d5cd50a282df533a29587cdf999ed1abcaa140c10ddcf71225597dfe3

Observation 203c50dc-9625-4f66-a193-588f56647cd9 · inbound

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales cites this paper.

Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:04.544270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T01:44:42.989053Z digest=sha256:4355b8548c570308c7efb13ac1a18b08ac5241c16b643ca9c9a232daad162262

Observation 2fa022f9-b499-44ef-8a7c-376882dc0c21 · inbound

dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats cites this paper.

dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:26.105639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T11:20:06.292977Z digest=sha256:5cae53f47439ebd56a8a369efb8d386512b9e3d80f3d4bd8293403670d987b39

Observation 74e42f0f-80a9-4853-899d-dc176c6a364d · inbound

dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats cites this paper.

dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-15T10:56:54.155632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:56:54.155632Z digest=sha256:ea078d6d435009f8c2900a52017474b472aff0ba21a68612179e4e53d23740fc

Observation 76203019-c512-4b06-928c-48c9af433a7d · inbound

Multi-Bitwidth Quantization for LLMs Using Additive Codebooks cites this paper.

Multi-Bitwidth Quantization for LLMs Using Additive Codebooks ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:48:21.307051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T07:29:14.923431Z digest=sha256:280e4cc3d176952817d58aaee3c61bb7be110d49de13d49cdbc403ad6fd513b6

Observation 0c189f19-d0d7-4bf7-96ca-0de477a6d0e3 · inbound

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference cites this paper.

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:52.357889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T05:41:39.052865Z digest=sha256:17170066b8d88655474478fce889047e489a2730152a9467e6668e87aad36c75

Observation 040d99fe-e452-4a6e-8205-a5dc4e315599 · inbound

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference cites this paper.

GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T03:16:46.892829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:16:46.892829Z digest=sha256:197fefa85b181c10236f113f7ccb1de699b31c10ebf39315b3498ab95c4df68b