Pith. sign in

Paper Citation Record · LEDGER

FPTQuant: Function-Preserving Transforms for LLM Quantization

As of 8 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 5 inbound Pith citation observations for arXiv:2506.04985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04985 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:43:50.212638Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T23:10:46.775151Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T22:11:14.387607Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved39
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ba8cc1a-738b-4380-916b-0bd6202341e8 · outbound

This paper cites Understanding and overcoming the challenges of efficient transformer quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Understanding and overcoming the challenges of efficient transformer quantization

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:57.626006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:43.392506Z digest=sha256:05a9466002bf3aba393ffab68e95130305b475d733e426c24c7fc19c19fa8890

Observation d8256077-1503-4508-a07d-c06095ff042e · outbound

This paper cites Bert busters: Outlier dimensions that disrupt transformers.

FPTQuant: Function-Preserving Transforms for LLM Quantization Bert busters: Outlier dimensions that disrupt transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:57.355244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:43.456213Z digest=sha256:40f098b1add6e10c2c5099ebb1d245d74fb72b064193be0be408d1d1681d8278

Observation 58f5443b-3af2-4d9a-a195-b157c627bcf6 · outbound

This paper cites an unresolved cited work.

FPTQuant: Function-Preserving Transforms for LLM Quantization Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.539598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.539598Z digest=sha256:2b8b709af135417452eb7a4d04a0b584e6c25f119c33c5302c141260f7b80a3e

Observation c7a78b1e-bcb8-4804-99bd-ffe86490b644 · outbound

This paper cites Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.611046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.611046Z digest=sha256:14d5355f8f72e196a4acbd3d572f9cc44b174d95088c1bb8a955df3e95520240

Observation dbc63c28-99f4-47d7-b4b6-d19a1ba1c380 · outbound

This paper cites Massive Activations in Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization Massive Activations in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.694378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.694378Z digest=sha256:6b23ad928800351ad0fbd77863267beafeea4429fd8e02e18c3efd0af52dea06

Observation b8559f51-d5d2-40ec-adc0-7fce677ae39b · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.819233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.819233Z digest=sha256:17ed3993f7f914b6e3ebca3b6806173880f9d92b82df1e1fef6a3337ed833b99

Observation 4d5c4acd-5ea6-4cb4-94c9-3db336267736 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

FPTQuant: Function-Preserving Transforms for LLM Quantization QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.915764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.915764Z digest=sha256:39394e31079e48af51f30c89d041c14bea6fbd7abe0fbd23d4ad39afd676b062

Observation dd047719-78d1-4eb2-be08-62d303645640 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

FPTQuant: Function-Preserving Transforms for LLM Quantization SpinQuant: LLM quantization with learned rotations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.996956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.996956Z digest=sha256:62daa1501c2c56c3fb8636864a3cdf894237b5d7950266ba7a6096a317a48b54

Observation 509b68ee-495b-49c2-b78c-d8b23520fd09 · outbound

This paper cites Quantizing deep convolutional networks for efficient inference: A whitepaper.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quantizing deep convolutional networks for efficient inference: A whitepaper

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.073059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.073059Z digest=sha256:26981a4cc0a56c635d09d329a99fa16156c049660a6a98cc88d4ebd799450f75

Observation cd97d06f-d500-4c0b-bfb6-2e87bb99de26 · outbound

This paper cites A White Paper on Neural Network Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization A White Paper on Neural Network Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.171780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.171780Z digest=sha256:b42b73e6afb6af44d6c2ee7d2adf8769fb27e2832d286fd60a17ed6c848a65cf

Observation 3c83e253-9f44-4951-9bd2-dd5921bec1df · outbound

This paper cites Post-training 4-bit quantization of convolution networks for rapid-deployment.

FPTQuant: Function-Preserving Transforms for LLM Quantization Post-training 4-bit quantization of convolution networks for rapid-deployment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.249922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.249922Z digest=sha256:4399c25d14e7132adbdc209a286f68d3f9f510c81e577b85121b36e1004db10e

Observation 50e04b74-1f22-4005-b250-0dacbfd07d84 · outbound

This paper cites Zeroq: A novel zero shot quantization framework.

FPTQuant: Function-Preserving Transforms for LLM Quantization Zeroq: A novel zero shot quantization framework

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:57.104781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:44.317595Z digest=sha256:aa1408a1893e71333e8138f77de768364ab09b1dd7368708d2ce0c74e267c94f

Observation 885d442f-996e-402c-bedb-82def12ce75b · outbound

This paper cites Low-bit quantization of neural networks for efficient inference.

FPTQuant: Function-Preserving Transforms for LLM Quantization Low-bit quantization of neural networks for efficient inference

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:56.881070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:44.407583Z digest=sha256:9b510c2fd9cbd79d3c206e9c72af9dafb6c22a5b812d1262c8aaffa91d4855aa

Observation 438057e3-01be-4626-a686-f32d716e6d88 · outbound

This paper cites Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming.

FPTQuant: Function-Preserving Transforms for LLM Quantization Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.490409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.490409Z digest=sha256:3a137fd797b0f3cbc816bf3a1fa6a99db0774329d142396663ad6a30704698d1

Observation 1f216aed-f293-45eb-a214-bb2a4a85655a · outbound

This paper cites Same, same but different: Recovering neural network quantization error through weight factorization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Same, same but different: Recovering neural network quantization error through weight factorization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:56.608405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:44.585531Z digest=sha256:b815d8dd044b08d5cf9ac8926b9f527fbc8a6b76ab826bf29f9127aebb0e237a

Observation 6b191482-579b-4f32-8e53-101cfa8e6b23 · outbound

This paper cites Improving neural net- work quantization without retraining using outlier channel splitting.

FPTQuant: Function-Preserving Transforms for LLM Quantization Improving neural net- work quantization without retraining using outlier channel splitting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:56.392159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:44.660979Z digest=sha256:1b59b3488b81b107e78ae91449cff409b70c01b28aab8193fe0ec27382dabbce

Observation 2aed43fc-d301-4c5a-917d-c633adb95ae8 · outbound

This paper cites Data-free quantization through weight equalization and bias correction.

FPTQuant: Function-Preserving Transforms for LLM Quantization Data-free quantization through weight equalization and bias correction

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:56.110221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:44.761658Z digest=sha256:055663122262422e06fb443676c5b785a25087612ea840515422ba7268b2c2f3

Observation b210edd9-11bb-41f6-af73-0a1ab55e1632 · outbound

This paper cites Up or Down? Adaptive Rounding for Post-Training Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Up or Down? Adaptive Rounding for Post-Training Quantization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.852414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.852414Z digest=sha256:3c86adf2bda1e381e865d89d5d7cc3dbaa11aa5d2b7b9d9f5b72dae6522556d3

Observation e388f3a5-d0be-46bb-93b9-73a199e95015 · outbound

This paper cites BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction.

FPTQuant: Function-Preserving Transforms for LLM Quantization BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.932932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.932932Z digest=sha256:26a40750273ac13733a5ec0139523995a1e5694e8a2b3ac3d3eb1c7c3d8e39e7

Observation c024df58-8b6e-4c68-be7e-3940bd5d7b7c · outbound

This paper cites Deep learning with limited numerical precision.

FPTQuant: Function-Preserving Transforms for LLM Quantization Deep learning with limited numerical precision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.882906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:45.012514Z digest=sha256:e7e256a236545eaa2ea90449d615067e6964efd63325d573bbeeb8a2a3239976

Observation 5cc967d0-211a-47cf-9b83-fad6bcc52476 · outbound

This paper cites Quantization and training of neural networks for efficient integer-arithmetic-only inference.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quantization and training of neural networks for efficient integer-arithmetic-only inference

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.735484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:45.059061Z digest=sha256:09b7da8dcb60f9346321cf615d70f4e37023e5caac4952dbebc655ab0ea7338d

Observation 969ea878-426c-48d8-bc60-02170c14f88e · outbound

This paper cites Esser, Jeffrey L.

FPTQuant: Function-Preserving Transforms for LLM Quantization Esser, Jeffrey L

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.535555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:45.133526Z digest=sha256:4c92a9b0621248178f35cd7742f45c681e2ecc14c65bbef00c0305ef36fdae9e

Observation 485f1515-d3bf-4677-986b-9f33a9ddafdc · outbound

This paper cites Lsq+: Improving low-bit quantization through learnable offsets and better initialization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Lsq+: Improving low-bit quantization through learnable offsets and better initialization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.246016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:45.211084Z digest=sha256:954ba964ebaf603efeb2706295dfff12cd357e4d32726e9be4c02d25f032d0b4

Observation d507e2b7-65cc-4069-b782-597c119b6fd5 · outbound

This paper cites Overcoming oscillations in quantization-aware training.

FPTQuant: Function-Preserving Transforms for LLM Quantization Overcoming oscillations in quantization-aware training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.037723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:45.287584Z digest=sha256:31701c38c2157c57e410a503a26f9c7c7b0b5fe1b90f0e44ae884c5cda796489

Observation 7da8cef2-4587-4574-81db-7b5309ac1165 · outbound

This paper cites LLM-QAT: Data-Free Quan- tization Aware Training for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization LLM-QAT: Data-Free Quan- tization Aware Training for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.412441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.412441Z digest=sha256:fb5fa5a682273cc4de82b1fa2436f6ad8b1f300d403395ab21e36d1f3bd8365b

Observation 3cb6aee2-6a07-482b-b315-3ac18a450963 · outbound

This paper cites BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation.

FPTQuant: Function-Preserving Transforms for LLM Quantization BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.503765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.503765Z digest=sha256:bbbfe61030bc929850cd69f723c6a6671645066468b40910d58fe741099ed3cb

Observation f030abb1-84cb-4181-94c3-7817bbfdee0e · outbound

This paper cites EfficientQAT: Efficient Quantization-Aware Training for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization EfficientQAT: Efficient Quantization-Aware Training for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.644520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.644520Z digest=sha256:4d8d81cc27b9b3ab539e0ba3a995c87fd3616d9376dbf0c1bfedb8f4944ac06b

Observation 657cfb93-0532-4c48-87ab-aec4df2bec6d · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

FPTQuant: Function-Preserving Transforms for LLM Quantization Qlora: Efficient finetuning of quantized llms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.723371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.723371Z digest=sha256:d61a4170399705bf734618bde250d3b32b077c92060c2aae4902515c8e4af189

Observation 289ff14b-bc35-4733-ae33-54f1fb574b65 · outbound

This paper cites QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.821751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.821751Z digest=sha256:93d473f0be63a8944696cb0ad18c6448cab8e2ce38b7b4aeefd8ea0f708baad1

Observation 1768de1f-a55d-48c9-be05-903aee41c92e · outbound

This paper cites Low-Rank Quantization-Aware Training for LLMs.

FPTQuant: Function-Preserving Transforms for LLM Quantization Low-Rank Quantization-Aware Training for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.907542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.907542Z digest=sha256:899424817873b02570c69cc2bf216fc65bec682d9846b0c5aae8522e8aa55acf

Observation e4a6e479-b504-4dfe-b7b0-2a07dfdb3d53 · outbound

This paper cites Paretoq: Scaling laws in extremely low-bit llm quantization, 2025.

FPTQuant: Function-Preserving Transforms for LLM Quantization Paretoq: Scaling laws in extremely low-bit llm quantization, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.985375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.985375Z digest=sha256:687088ae3a8c847136666ff9e2fb6852af56e34370ceb20b3a7eb49ae68f58ca

Observation 15fd83d1-ec0c-4244-85a5-e58b9eb4ac3b · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

FPTQuant: Function-Preserving Transforms for LLM Quantization GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.065076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.065076Z digest=sha256:5c46078abc98e50faa9389752f67ffd99b6d6a0d3d5ea1dcdd5cd91e6b2811c9

Observation 9ecc4d88-0f0c-4804-8227-bdda6b9b0470 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

FPTQuant: Function-Preserving Transforms for LLM Quantization SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.139211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.139211Z digest=sha256:54e97c2cd27aa31d5a0f1c5574cf149627033a86913f179b84926c52b1efa4e8

Observation 8fad7b8c-9a67-4e89-8b00-bbee994ae2b9 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

FPTQuant: Function-Preserving Transforms for LLM Quantization AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.226301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.226301Z digest=sha256:84d8684210c8877cecd7d1e260957d94e1ca22bb0d59177506488d2b0dde81c1

Observation c502772d-b90a-412c-8437-8958ce0fba68 · outbound

This paper cites Owq: Outlier- aware weight quantization for efficient fine-tuning and inference of large language models.

FPTQuant: Function-Preserving Transforms for LLM Quantization Owq: Outlier- aware weight quantization for efficient fine-tuning and inference of large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.309944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.309944Z digest=sha256:94537bbe0cd418da595769c47040f824abbce4f9804c674f7f70080810a0cf41

Observation 8667e534-2518-4bf2-954d-47bdcf62d9c7 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.395454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.395454Z digest=sha256:e2ae484a1fd1d248324954a7a6b362e079159a2743c0c62be3df68c0e13708a5

Observation f696d429-04b3-4cc4-a2a8-feec06fcb35b · outbound

This paper cites SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.491664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.491664Z digest=sha256:dc251c6e1feab89360fba76bf1b1d38e328869eb98012d7a8ded3ec5257acc6a

Observation 9f794de9-c001-4ba4-8925-f3957faa6e3a · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Extreme Compression of Large Language Models via Additive Quantization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.566561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.566561Z digest=sha256:16d54ecb3f0e1f31db1dc964650911e5f35bd3cc5403884c4cbc62768bc1deda

Observation ec96620d-493a-4815-8591-81342b0588be · outbound

This paper cites A frustratingly easy post-training quantization scheme for llms.

FPTQuant: Function-Preserving Transforms for LLM Quantization A frustratingly easy post-training quantization scheme for llms

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:54.828721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:46.675630Z digest=sha256:e1b544ad2bcc8e64e45de8c41357e2f593778b9b19025931933f7aa3101d1857

Observation da98e275-fbe4-41a6-ad7a-77f7f3f47aa0 · outbound

This paper cites Flexround: Learnable rounding based on element-wise division for post-training quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Flexround: Learnable rounding based on element-wise division for post-training quantization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:54.667231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:46.752521Z digest=sha256:27a04410a3f612b7187c3f0f184a0e980aece208cc45c9569cb289af6c8c1abd

Observation bf7d682b-5699-4ec9-bb14-7d3682149f37 · outbound

This paper cites Long- range zero-shot generative deep network quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Long- range zero-shot generative deep network quantization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:54.454601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:46.841044Z digest=sha256:41a25414c6d32ed69c8a68f88bf8495c5dafabc82fe2e23e541719fa1fb677c9

Observation 0a53fc79-93eb-4c1d-989b-0d38e74703a7 · outbound

This paper cites Quip: 2-bit quantiza- tion of large language models with guarantees.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quip: 2-bit quantiza- tion of large language models with guarantees

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:54.223241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:46.918388Z digest=sha256:2ee3f97ddbae5de268645d10b00a157002912e7247c77ca356216a176a6db6df

Observation a2cac9bb-05ea-4ff6-a91f-d0c1ecd524df · outbound

This paper cites Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling.

FPTQuant: Function-Preserving Transforms for LLM Quantization Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.980855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.980855Z digest=sha256:526b8a3caadc5a29f898f5c0829f9be67c5e983756bc7a02e330f2b933105d1d

Observation 69b59f67-87e3-4053-98ab-73f54a6393f9 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.032509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.032509Z digest=sha256:392de31118660c33da3b015bfe9705126df52495cc8730037d6b20920592b84a

Observation e26850cc-f1fa-42f3-8bd4-b7d17fd82973 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, February.

FPTQuant: Function-Preserving Transforms for LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, February

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:53.975213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:47.124919Z digest=sha256:8aa407589a01a94ba6093969da80801c84dfbd34a50ab75b9902405c979e4bb0

Observation 870a263e-334d-4d85-bcaa-7fa4cc1a2161 · outbound

This paper cites DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs.

FPTQuant: Function-Preserving Transforms for LLM Quantization DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.305676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.305676Z digest=sha256:0b44b9fd1d84fff80c579c993768ab0ea5aed58f2139868cec25b6af5eff5097

Observation 252b32eb-f7f4-4bcd-809a-0b2ae9fc5dcc · outbound

This paper cites FlatQuant: Flatness Matters for LLM Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization FlatQuant: Flatness Matters for LLM Quantization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.377076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.377076Z digest=sha256:d265cd1c8a912ccf92ef78aacd040de2092e1dd29c68e60bc6f8532b97484425

Observation 64dc9d2b-c2fd-430f-a0f7-a6a009f3ce69 · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

FPTQuant: Function-Preserving Transforms for LLM Quantization SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.464816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.464816Z digest=sha256:14fb71635486e82f5b1f9cc09bab3fc32bf7dba2417e1f5e6bf7e8b80adc2e7f

Observation 7287845d-6e7e-4085-9c68-9ffa36c6253a · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

FPTQuant: Function-Preserving Transforms for LLM Quantization Roformer: Enhanced transformer with rotary position embedding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.553216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.553216Z digest=sha256:109305af8431126850dfbb8d6b959e8189497af6042e2223964cea93d6d4bc45

Observation 43470d1e-cbcb-47c7-af8c-9615542eca02 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.634359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.634359Z digest=sha256:3c5ffc3f7a755d9600f9fe9a4e216c0720378acbfba2d68ef8827b8b7e17a397

Observation f00ec514-0be5-4c54-a336-416364dbeb24 · outbound

This paper cites The Llama 3 Herd of Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization The Llama 3 Herd of Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.714871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.714871Z digest=sha256:83c5b7e5163ffaf178f3b8c102fe832eb26942aa21cc23f8c928cf63cde52f9d

Observation 16655abc-63aa-47d1-bdef-c8b6b42e0071 · outbound

This paper cites Pointer sentinel mixture models.

FPTQuant: Function-Preserving Transforms for LLM Quantization Pointer sentinel mixture models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.806276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.806276Z digest=sha256:616b7ae9b46d0ba2bb9dd5c180b751e099349ef64ad7bf24409d667bf816e04e

Observation 19c78207-9ae1-4dc9-996b-c6d7aeeb92ed · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

FPTQuant: Function-Preserving Transforms for LLM Quantization PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:53.813234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:47.981754Z digest=sha256:9f3877ba4a4f94eed3a3449aeca69ab3c5bc6058564c34dc394eb2433dc214a1

Observation a34ff82a-8ca1-475f-b07b-e0983c96af3a · outbound

This paper cites WinoGrande: an adversarial winograd schema challenge at scale.

FPTQuant: Function-Preserving Transforms for LLM Quantization WinoGrande: an adversarial winograd schema challenge at scale

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:53.636507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:48.097905Z digest=sha256:325210a5fee2093607bd725840f648412bfcf484d94ebbb13affe1bd22c46a3d

Observation 1d04d0d8-2b3b-488a-99d5-7ca5f7ae42f5 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

FPTQuant: Function-Preserving Transforms for LLM Quantization HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.336766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.336766Z digest=sha256:7b3dc22322f3ea6c9fc979f54dd5a08207ccfb7f7b4a88abdab4326f459196ea

Observation e7ad51cb-3060-462f-87fc-3d24e2b28753 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

FPTQuant: Function-Preserving Transforms for LLM Quantization Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.475734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.475734Z digest=sha256:2c4a5bd8de4c572a65638473e47d1920ff38c16fa76837b0e7640a2dff356cc6

Observation f46dc064-b0cf-4035-a0c0-11d2169c7aa3 · outbound

This paper cites The lambada dataset: Word prediction requiring a broad discourse context.

FPTQuant: Function-Preserving Transforms for LLM Quantization The lambada dataset: Word prediction requiring a broad discourse context

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:53.430817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:48.616558Z digest=sha256:3c557cdd62869ca87d664381990f20461f8048cd3afab4c1fcd04bf6492d6a59

Observation f3c35bc1-b422-4358-baea-1ae5efa5ca4d · outbound

This paper cites Working with Quantized Types — NVIDIA TensorRT Doc- umentation.

FPTQuant: Function-Preserving Transforms for LLM Quantization Working with Quantized Types — NVIDIA TensorRT Doc- umentation

Reference 58

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:53.245750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:48.792727Z digest=sha256:9128ae1395346339fe4f3bfceb8c013f85f0c50121988832c714ecddd67d22eb

Observation e737f547-f15a-4453-a5aa-4dad17c7128f · outbound

This paper cites Quantization — PyTorch AO documentation.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quantization — PyTorch AO documentation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:53.033620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:48.899498Z digest=sha256:f3ce0b188cdac81890145df3b3e1e8bd2f839f84e0548ec66c9d500be40dc7af

Observation 30abc064-9048-4414-a3c4-459a113e0e72 · outbound

This paper cites AI Engine Direct SDK documentation.

FPTQuant: Function-Preserving Transforms for LLM Quantization AI Engine Direct SDK documentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:52.915494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:49.036004Z digest=sha256:5fafa69bdf4d24d8836e620fcea1d1e40ac6cc5a5deda1ab6c4abd4b2a9c08f3

Observation 9d768a9a-4f8f-4431-ba16-962eb3325d66 · outbound

This paper cites TensorRT operators documentation: DynamicQuantize not supported on DLA.

FPTQuant: Function-Preserving Transforms for LLM Quantization TensorRT operators documentation: DynamicQuantize not supported on DLA

Reference 61

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:52.766879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:49.215303Z digest=sha256:c5803d3567d0637173706acdb441a5596fd0031e1eefee70a66f87dbfb1bab45

Observation e72a981b-b49f-4efb-80c4-7bc224be7ec9 · outbound

This paper cites fast-hadamard-transform.

FPTQuant: Function-Preserving Transforms for LLM Quantization fast-hadamard-transform

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:52.513457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:49.347868Z digest=sha256:322550d3bda65c7c3f1995b769e9cb4aaa971a00b2043a8704d1dd58ab0caa83

Observation a265f935-0e7d-45f6-9090-3ef70eb62590 · outbound

This paper cites double-packed.

FPTQuant: Function-Preserving Transforms for LLM Quantization double-packed

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:52.383814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:49.457392Z digest=sha256:13b7eb52aec48ead3793f2c76e3ab1e4833984d61a7e18569aafee9cc0920ed0

Observation 389e9d3d-8215-4636-88c7-3c78d8a8139a · outbound

This paper cites Evaluate quantization error per quantizer placement (e.g.

FPTQuant: Function-Preserving Transforms for LLM Quantization Evaluate quantization error per quantizer placement (e.g

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:52.260528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:49.606480Z digest=sha256:438992d075dd210883e98ff47b509bc121083efbaf540d4ae269719aef2a18e0

Observation 0d901eaf-5cd8-4c22-9a12-499f44daeb7a · outbound

This paper cites Based on step 1, choose which FPTs to add: (a) Attention and FFN input.

FPTQuant: Function-Preserving Transforms for LLM Quantization Based on step 1, choose which FPTs to add: (a) Attention and FFN input

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.990152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:49.704829Z digest=sha256:4ee9c747a5f76b70c44ab272357aaa928da7570b9dd9c9e6d1c9a2e69a9771a2

Observation b57a2e1a-8f17-4c26-90bc-d603c2affe6a · outbound

This paper cites Initialize transforms, e.g.

FPTQuant: Function-Preserving Transforms for LLM Quantization Initialize transforms, e.g

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.744777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:49.859176Z digest=sha256:25ca9c249b72cb50d78643ea59355c8f85c86cb6d94775939ecc66b676f869d4

Observation fa9295aa-18e8-4de6-a6a1-cb6f538260c3 · outbound

This paper cites Locally optimizing transforms improves performance and reduces training time, whilst incurring very little cost (Appendix F.2.1).

FPTQuant: Function-Preserving Transforms for LLM Quantization Locally optimizing transforms improves performance and reduces training time, whilst incurring very little cost (Appendix F.2.1)

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.540897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:49.988735Z digest=sha256:a15e6909b11d6bf9c4b4e7e8cd47566ce4e1d591094c9d22f526d81229fe0c0e

Observation 33ec7427-9afd-4f78-a2fc-1c9b3d06c178 · outbound

This paper cites Set the initial quantization grid, e.g.

FPTQuant: Function-Preserving Transforms for LLM Quantization Set the initial quantization grid, e.g

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.372855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:50.099603Z digest=sha256:2a5a549090f556d4a9638050b2ffe123c38b1218fe484d96460f0bd596bd60f1

Observation b4549857-6b67-4f37-9003-499b98980c9d · outbound

This paper cites Train the FPTs and quantization grid end-to-end, with the unquantized outputs as target.

FPTQuant: Function-Preserving Transforms for LLM Quantization Train the FPTs and quantization grid end-to-end, with the unquantized outputs as target

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.087441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:50.212638Z digest=sha256:40da037e4eb6e0d0c8a377fedbae759be78ac27a0661483965e768511751dc34

Observation 9a0cd5c9-7041-469c-94cf-c7e320333fb7 · outbound

This paper cites doi: 10.1145/3474381.

FPTQuant: Function-Preserving Transforms for LLM Quantization doi: 10.1145/3474381

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.191111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.191111Z digest=sha256:9cdfe62492c9f3d6a0ee2364c273010a09cbcb646fd40f372515655ee2bd9bc2

Observation abf25eb5-8f0d-4593-965a-6a9964b135d3 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

FPTQuant: Function-Preserving Transforms for LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.223987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.223987Z digest=sha256:54e72a8512a4f144730f2a853110290a920b4ce0ea87a6bfcdd33864cfdf041b

Pith citing papers

Observation fbe7599b-8d5c-4a5e-b8b4-f20481b37886 · inbound

Leech Lattice Vector Quantization for Efficient LLM Compression cites this paper.

Leech Lattice Vector Quantization for Efficient LLM Compression FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T23:10:46.775151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:10:46.775151Z digest=sha256:c159e9827ca0d1fe9d96af77d56e149ee40d7670222cb66fcde9efd4a9d9498f

Observation 7bd8316e-ec7e-42e8-9b9d-7b05baa4b7ba · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 147

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:6d7924f417d0f7a611142598a3949ed616e8aecfa566863434defdc122910ffc

Observation 781beda3-9b01-4397-bf23-090a6a93836f · inbound

When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon cites this paper.

When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:36.313617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:16:00.868964Z digest=sha256:609e52b6bf3ea2539973fc6ddd6906ee1628961d354a391744e6dc818acf300b

Observation 70b001ba-0d1e-4e7a-8c1e-dd1c65847f4e · inbound

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation cites this paper.

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T09:35:12.908124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:35:12.908124Z digest=sha256:6bcb7e0df2664bd170434d483cce4b8281b8132e1f02c3063e247c2538ef3a8c

Observation f81c0d0d-901b-418c-bf2b-bff90bf5bbf0 · inbound

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation cites this paper.

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T16:53:34.398685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:53:34.398685Z digest=sha256:e146c86b65fee7e2b3b8b1e0b9010e5f1e95a907866b874be0682561e807311a