Pith. sign in

Paper Citation Record · LEDGER

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference

As of 11 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2501.01144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01144 v5

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:44:30.222315Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T19:31:28.053090Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 59c233a4-120d-470f-b98c-b007632f82a6 · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.059448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.059448Z digest=sha256:f3f46ac8b1bbe1829a62696a30c6f9eb73839e9eab4ac050465bfe6f3d13e526

Observation 5730738b-99fe-49ed-a07a-4295aaab48a5 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.075869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.075869Z digest=sha256:a29abf1aaef9af340e9db31836e229be3bc1ff370855b63caa3e359a8a99917d

Observation de82ca2d-16c6-454f-93cd-0ebcd8d9434d · outbound

This paper cites Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Learning from Students: Applying t-Distributions to Explore Accurate and Efficient Formats for LLMs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:44:30.653784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.091674Z digest=sha256:bfac1347a01641a3d63b2d2005f08a75ba746d7ed9b684ad2dc31c33b5603940

Observation 6845930e-22d5-46b2-b580-465d4fbe2a77 · outbound

This paper cites The Llama 3 Herd of Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.097173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.097173Z digest=sha256:5ef3df25b1dc57e70a839bcdd9969d3bc42d6445507e53b8ce50150ae0678e7b

Observation 66c5d3b4-f15c-4449-8cb2-7a52421e01cb · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Extreme Compression of Large Language Models via Additive Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.102307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.102307Z digest=sha256:da448a66b74e634d65caa61619cad9bfe46a693fcf8d03cefab313da1bc37c4c

Observation 44839155-e3cc-4ca0-9edd-e8f9df445a1a · outbound

This paper cites BCQ: Block Clustered Quantization for 4-bit (W4A4) LLM Inference.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference BCQ: Block Clustered Quantization for 4-bit (W4A4) LLM Inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.107005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.107005Z digest=sha256:579c6896c0fda5698f9dcdf4c868a7df9cd92f5c26871757aa24b0da241cab7c

Observation abd3db42-e086-469b-afd7-86e140747b2e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Measuring Massive Multitask Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.115643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.115643Z digest=sha256:d62b9e118375f35fb59592463807b08117015d327f2394d53991dac7d8109bd3

Observation 12508ec8-b29a-4e3d-b6eb-f5b092297b26 · outbound

This paper cites Mistral 7B.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.124358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.124358Z digest=sha256:996a386ded5c1f7fa9f08caad2073a839ae1cef9a99208486fdcc12096ba38e2

Observation 7d957495-3395-43b2-b49d-98a464d2b749 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.128860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.128860Z digest=sha256:351591188e286228bf6f5005e1e7164c38a2df3c6fd5ce044db5d3da57a4341d

Observation 845afe81-1327-4b44-ae0e-bd8772d76f91 · outbound

This paper cites FPTQ: Fine-grained Post-Training Quantization for Large Language Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference FPTQ: Fine-grained Post-Training Quantization for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.133415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.133415Z digest=sha256:c3e53af75ef95878e91372f60e73048ad9a22d72b395067d09530f95c4908c93

Observation ab32f4d4-f661-4a35-966a-79d4cfc4a8fe · outbound

This paper cites LLM-FP4: 4-Bit Floating-Point Quantized Transformers.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference LLM-FP4: 4-Bit Floating-Point Quantized Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.137644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.137644Z digest=sha256:2427570c2bc485bd0f44ef13fe30e1e70f3a0242220705cadcc2e8dc6902ea3d

Observation 6fd60780-c798-4c51-a749-3164d5e1d12e · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.142155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.142155Z digest=sha256:007369b8bcfcab8af75b2b8d8ac7ae155078e0e0c62c7fd104311487cb483589

Observation 620618b6-9e81-41a0-acc5-3d44f4f5853e · outbound

This paper cites Pointer Sentinel Mixture Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Pointer Sentinel Mixture Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.147521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.147521Z digest=sha256:3307b79a0996e79dedf6ecf4bae81104c71ad4e2125e36bb079ba63108be7b67

Observation 30a922f6-ae9a-4391-9fc2-12d018e30109 · outbound

This paper cites Microscaling Data Formats for Deep Learning.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Microscaling Data Formats for Deep Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.157115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.157115Z digest=sha256:0509df821213c933dd22535b525280dc3d15a2944bc327d05c3b9bdde0b268ea

Observation 45e330fa-b746-4e55-8786-03d9f1dd90ed · outbound

This paper cites Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.161918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.161918Z digest=sha256:5ebee28db4cca7b060c9f07ed5bea81e3a6f951e0f1ce8fdd0e5dfad8d103eab

Observation fa2e2a66-609a-43b6-91b9-f657944f403b · outbound

This paper cites Release Strategies and the Social Impacts of Language Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Release Strategies and the Social Impacts of Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.167061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.167061Z digest=sha256:e0b169a9e04ef0b55c03151c3640d8f252e3e4b5d7f21b146dd82374a5f1d3b0

Observation 6b41623d-4886-43d5-a8e0-207e4be94fc1 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.171823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.171823Z digest=sha256:df02b92ed552af1b8d8e947f4e283590f3bd19b22ce532f4a9feff42dba6e094

Observation 8116dac2-c1ee-4493-a86f-e24760ad3bdb · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.175978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.175978Z digest=sha256:423674bc9bbedfa507ed414ef22d7c67a995cc22edef16f09aaa5214805fe85e

Observation c9ee65de-9815-4904-a74b-f6315e105749 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.180277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.180277Z digest=sha256:d0d0a19e79154cab935f90049e8c861f661a29223c153be8e461fad7376f512c

Observation 8b644334-71ea-409b-bab9-a42091015223 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.189163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.189163Z digest=sha256:83a16a8fee815a4e20ed43ce513825e30441de11e91e4ae1be46dad407a3f467

Observation bae98df9-0dae-4ba3-826d-c91b58bb71fe · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.193758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.193758Z digest=sha256:b61ada0ac1421d63005536d55fa08ff3c288c968419a38a7e5fea260b5853fba

Observation 443379e9-08ad-40c8-a500-52f77b1244a7 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference OPT: Open Pre-trained Transformer Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.198608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.198608Z digest=sha256:79127700d51d44a4b9aa56fa57f884b384e5c793e32848d4bd5b79fe6a68bb45

Observation 08012266-ae3c-495e-b919-96016dba2b7f · outbound

This paper cites Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:30.792866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.203582Z digest=sha256:2b30e672e6ffa5d0485ba43ab5ea85f496c3f1bf5a0c62fc158921e9341f8f8d

Observation 1a341f18-8111-494f-8508-5b9ca07d02e4 · outbound

This paper cites BlockDialect quantizes matrices and vectors along their respective multiplication dimensions.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference BlockDialect quantizes matrices and vectors along their respective multiplication dimensions

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:30.779589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.207612Z digest=sha256:a0d1777422d388f1c994a7b4249765fe5964c98978f9e0bf53dd1e1b75993ee0

Observation bc91eacd-7966-4cd4-9cd6-af15ba681bde · outbound

This paper cites These approaches often dequantize data to FP16 before performing multiplications, which limits computational efficiency.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference These approaches often dequantize data to FP16 before performing multiplications, which limits computational efficiency

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:30.766708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.211562Z digest=sha256:12b72db3ad469477c3843e203ebb93d386ad9c8985df4363c529237b7018c645

Observation 291a412a-29bd-46d2-bdf7-84eb21efe9df · outbound

This paper cites an unresolved cited work.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Unresolved cited work

Reference 34

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:44:30.752293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.215113Z digest=sha256:185d2a4d908769b50a5ff0c2f5ba7ab3115f34bac3ab8174643017dccedafc60

Observation fe82128d-1d2a-4c50-8bde-56c642954bb2 · outbound

This paper cites Infinitesimal generators for a class of polynomial processes.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Infinitesimal generators for a class of polynomial processes

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T22:44:30.261174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.218614Z digest=sha256:8351df15291e5d5ee17a46421519e789f47daa884107b59bed955e1e11355895

Observation 0666a55e-aba5-4c73-a525-14dc7910b484 · outbound

This paper cites However, 2D block quantization generally results in higher perplexity.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference However, 2D block quantization generally results in higher perplexity

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:44:30.739617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:44:30.222315Z digest=sha256:c83e3d70b166b281f410803657ae6408e519f72234699d12ca002ce885111d86

Observation 56011b9a-eab4-4f73-8e41-f7b303c754cc · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.152331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.152331Z digest=sha256:788e4ed6f420a0380652834c17395885aa394b616835eda1e805aa4b17f8e84f

Observation 36f44717-6eb3-46df-80a9-6fa3c8c7ed8b · outbound

This paper cites ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference ZeroQuant-FP: A Leap Forward in LLMs Post-Training W4A8 Quantization Using Floating-Point Formats

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.184926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.184926Z digest=sha256:d5e329884b6dceb48bab3a173b4b79e468211f3a604a56cb5a50ef8c3bf8d126

Observation b4009a59-93ff-46a6-a113-6d43629015c9 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.081186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.081186Z digest=sha256:23c9f1c2adc3ffd4aeb892901d9a1ba00ad39be1cd13a029b9bd5603977c53c1

Observation e256fefb-690f-4ec1-9063-af310016def5 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.120356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.120356Z digest=sha256:698fa364ac520434af3597652be1d8844af0306a179eff1843305c53ec7c1711

Observation ba4d08ae-e315-4153-882f-88fc32a473a4 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.086096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.086096Z digest=sha256:1fc843ea5a510feb77d375c6e68e6cb1831a6b17f8d2cc5e2c8ca3f8cad9115e

Observation c86a6e86-becd-4432-97fa-457a4d37d62f · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.071130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.071130Z digest=sha256:bad458319712e3b4387ba43b41a1e5732f811fab35354bdac06bfed68f858e4a

Observation 46e991ba-1024-466c-94aa-34eb68889dd3 · outbound

This paper cites QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.065512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.065512Z digest=sha256:418f783ff41ecdf1c8dc34b037460f139e5d5a3020040973ad8b056c348e16bf

Observation cdbfc1e9-7671-4337-a652-aaff7238b9d6 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.111333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.111333Z digest=sha256:730a5a03df2a00b2de2367492b27c249c61c2af809a9032a12dd5458d00f2d00

Pith citing papers

Observation fe7fbe63-008d-49dd-99cb-af5d1d277a18 · inbound

SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference cites this paper.

SPARQLe: Sub-Precision Activation Representation for Quantized LLM Inference BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:32:34.309836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T19:31:28.053090Z digest=sha256:b9fe226dd58cea6f43ea828cad2307dce1924e52471318331e9afa5b7a2a21f0