Pith. sign in

Paper Citation Record · LEDGER

FPTQuant: Function-Preserving Transforms for LLM Quantization

As of 22 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 6 inbound Pith citation observations for arXiv:2506.04985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04985 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:43:50.212638Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:54:26.612471Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T22:11:14.387607Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved39
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ba8cc1a-738b-4380-916b-0bd6202341e8 · outbound

This paper cites Understanding and overcoming the challenges of efficient transformer quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Understanding and overcoming the challenges of efficient transformer quantization

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:57.626006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:43.392506Z digest=sha256:c442a4be96cd7b013d521ac86dc2bce84adc1278977f9f30205e3c3c739f09d1

Observation d8256077-1503-4508-a07d-c06095ff042e · outbound

This paper cites Bert busters: Outlier dimensions that disrupt transformers.

FPTQuant: Function-Preserving Transforms for LLM Quantization Bert busters: Outlier dimensions that disrupt transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:57.355244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:43.456213Z digest=sha256:e8ceb7151e0965f7d18b4c99ab19c689864f22c5ba41a6ec496176b91ed9bd09

Observation 58f5443b-3af2-4d9a-a195-b157c627bcf6 · outbound

This paper cites an unresolved cited work.

FPTQuant: Function-Preserving Transforms for LLM Quantization Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.539598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.539598Z digest=sha256:b5c6702bc908fa5d1be22f72756025e50ce4c72dc5d57c7e24c39d0261970de7

Observation c7a78b1e-bcb8-4804-99bd-ffe86490b644 · outbound

This paper cites Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.611046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.611046Z digest=sha256:eaee0ee185d80afc8fb55047997e3647d7c44468bac0fb331a5a23a5305813d3

Observation dbc63c28-99f4-47d7-b4b6-d19a1ba1c380 · outbound

This paper cites Massive Activations in Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization Massive Activations in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.694378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.694378Z digest=sha256:6da9f511cdf9b02ec9065a3e5cb4d63136d20a87e4eb507eb284cae6cdf7c1ba

Observation b8559f51-d5d2-40ec-adc0-7fce677ae39b · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.819233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.819233Z digest=sha256:66ae9d8bbf3a8df125ba4b1c1a940620f23e57681542bf5453db396a460b7152

Observation 4d5c4acd-5ea6-4cb4-94c9-3db336267736 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

FPTQuant: Function-Preserving Transforms for LLM Quantization QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.915764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.915764Z digest=sha256:8aae3a520bb30b87cc45eaa4bb1e47bd32c7151f30d6a87370f0aa568cce268e

Observation dd047719-78d1-4eb2-be08-62d303645640 · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

FPTQuant: Function-Preserving Transforms for LLM Quantization SpinQuant: LLM quantization with learned rotations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:43.996956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:43.996956Z digest=sha256:e4aea2e8553e911bc9293437e85a366ccf996b0fdcbb4938a239ad632e3e0a6a

Observation 509b68ee-495b-49c2-b78c-d8b23520fd09 · outbound

This paper cites Quantizing deep convolutional networks for efficient inference: A whitepaper.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quantizing deep convolutional networks for efficient inference: A whitepaper

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.073059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.073059Z digest=sha256:4748ac5fb7f8efa210efed08c64bbbcaf742f4dce44c6a3d34c184d0f9180139

Observation cd97d06f-d500-4c0b-bfb6-2e87bb99de26 · outbound

This paper cites A White Paper on Neural Network Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization A White Paper on Neural Network Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.171780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.171780Z digest=sha256:93205021886e7c478c849957eaca28a83b5040636ca1d77b95482e51a48ad71f

Observation 3c83e253-9f44-4951-9bd2-dd5921bec1df · outbound

This paper cites Post-training 4-bit quantization of convolution networks for rapid-deployment.

FPTQuant: Function-Preserving Transforms for LLM Quantization Post-training 4-bit quantization of convolution networks for rapid-deployment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.249922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.249922Z digest=sha256:9faae1728a83157218977088c5d6f1ac2bad0066da6d0d4743c8eeb289a6a437

Observation 50e04b74-1f22-4005-b250-0dacbfd07d84 · outbound

This paper cites Zeroq: A novel zero shot quantization framework.

FPTQuant: Function-Preserving Transforms for LLM Quantization Zeroq: A novel zero shot quantization framework

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:57.104781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:44.317595Z digest=sha256:6ea7eb76945eb8e30afd5cd913ad8792613bc28f131e4e5d34a49f1611db881a

Observation 885d442f-996e-402c-bedb-82def12ce75b · outbound

This paper cites Low-bit quantization of neural networks for efficient inference.

FPTQuant: Function-Preserving Transforms for LLM Quantization Low-bit quantization of neural networks for efficient inference

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:56.881070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:44.407583Z digest=sha256:4e85bbf7c9cbe5e405d8847efa199309f1982392885c996d21b797fbbf3f21d8

Observation 438057e3-01be-4626-a686-f32d716e6d88 · outbound

This paper cites Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming.

FPTQuant: Function-Preserving Transforms for LLM Quantization Improving Post Training Neural Quantization: Layer-wise Calibration and Integer Programming

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.490409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.490409Z digest=sha256:23be671fd7a247efb94152bd6a914514ddf61932dde54875367e694caa9f04c1

Observation 1f216aed-f293-45eb-a214-bb2a4a85655a · outbound

This paper cites Same, same but different: Recovering neural network quantization error through weight factorization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Same, same but different: Recovering neural network quantization error through weight factorization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:56.608405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:44.585531Z digest=sha256:a4fa7d64f5d1b2ca4a212ba980dfe43a7310206fc88084cf18794f6224a55bcb

Observation 6b191482-579b-4f32-8e53-101cfa8e6b23 · outbound

This paper cites Improving neural net- work quantization without retraining using outlier channel splitting.

FPTQuant: Function-Preserving Transforms for LLM Quantization Improving neural net- work quantization without retraining using outlier channel splitting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:56.392159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:44.660979Z digest=sha256:14cea9255b56ef55d00394f106b12262c651c8fedb8cacf431f8eef25f76096a

Observation 2aed43fc-d301-4c5a-917d-c633adb95ae8 · outbound

This paper cites Data-free quantization through weight equalization and bias correction.

FPTQuant: Function-Preserving Transforms for LLM Quantization Data-free quantization through weight equalization and bias correction

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:56.110221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:44.761658Z digest=sha256:cef8fe14f408055c76e69e97ebfd05ec3b32d99bf55ce26c3d47396e00e7aa80

Observation b210edd9-11bb-41f6-af73-0a1ab55e1632 · outbound

This paper cites Up or Down? Adaptive Rounding for Post-Training Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Up or Down? Adaptive Rounding for Post-Training Quantization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.852414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.852414Z digest=sha256:c8308315b1b3964db27ddd691873f74d9c26e08c3eca47f5f575082776272ae0

Observation e388f3a5-d0be-46bb-93b9-73a199e95015 · outbound

This paper cites BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction.

FPTQuant: Function-Preserving Transforms for LLM Quantization BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:44.932932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:44.932932Z digest=sha256:619139094fbe45953e1420dd5f1476370dbaf8c1a2d8b351929f6ce172233bb0

Observation c024df58-8b6e-4c68-be7e-3940bd5d7b7c · outbound

This paper cites Deep learning with limited numerical precision.

FPTQuant: Function-Preserving Transforms for LLM Quantization Deep learning with limited numerical precision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.882906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:45.012514Z digest=sha256:7e3974e45571d7b23649a9b3e4e7126170cf068b82003b758f235c65247f6659

Observation 5cc967d0-211a-47cf-9b83-fad6bcc52476 · outbound

This paper cites Quantization and training of neural networks for efficient integer-arithmetic-only inference.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quantization and training of neural networks for efficient integer-arithmetic-only inference

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.735484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:45.059061Z digest=sha256:9506645d3c5a48d3104cbb71df935efc18c3cd1bd56a035d816df8e98aabdbf8

Observation 969ea878-426c-48d8-bc60-02170c14f88e · outbound

This paper cites Esser, Jeffrey L.

FPTQuant: Function-Preserving Transforms for LLM Quantization Esser, Jeffrey L

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.535555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:45.133526Z digest=sha256:2d408020e61ab1da20e92135928d688712066c29dfc3a9f1b768234ae7f21241

Observation 485f1515-d3bf-4677-986b-9f33a9ddafdc · outbound

This paper cites Lsq+: Improving low-bit quantization through learnable offsets and better initialization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Lsq+: Improving low-bit quantization through learnable offsets and better initialization

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.246016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:45.211084Z digest=sha256:362c512fb558321a7624e3419dfdd8949235b6241a6e8b509e33dd30100303ca

Observation d507e2b7-65cc-4069-b782-597c119b6fd5 · outbound

This paper cites Overcoming oscillations in quantization-aware training.

FPTQuant: Function-Preserving Transforms for LLM Quantization Overcoming oscillations in quantization-aware training

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:55.037723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:45.287584Z digest=sha256:89a3f7b2388f674efd90f1578af970b07f7997ac04c9bb871f9b680a768d258c

Observation 7da8cef2-4587-4574-81db-7b5309ac1165 · outbound

This paper cites LLM-QAT: Data-Free Quan- tization Aware Training for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization LLM-QAT: Data-Free Quan- tization Aware Training for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.412441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.412441Z digest=sha256:c449e021c3e1cdaa7368b9ed13c439c469b738e711b5613d2e9a37aa2d59ba9e

Observation 3cb6aee2-6a07-482b-b315-3ac18a450963 · outbound

This paper cites BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation.

FPTQuant: Function-Preserving Transforms for LLM Quantization BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.503765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.503765Z digest=sha256:8379edf4752c688be37849269cfec661d401d7477e1d416147696e0ee8901a9d

Observation f030abb1-84cb-4181-94c3-7817bbfdee0e · outbound

This paper cites EfficientQAT: Efficient Quantization-Aware Training for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization EfficientQAT: Efficient Quantization-Aware Training for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.644520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.644520Z digest=sha256:2368f1bfeb5eaf69ea960cfacdccd0552301c44928deaf1876f976a758378df7

Observation 657cfb93-0532-4c48-87ab-aec4df2bec6d · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

FPTQuant: Function-Preserving Transforms for LLM Quantization Qlora: Efficient finetuning of quantized llms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.723371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.723371Z digest=sha256:7f04bddd6786fceac2f036f4578c1c8bd554e8c39c6b80aa9e014538cc95beba

Observation 289ff14b-bc35-4733-ae33-54f1fb574b65 · outbound

This paper cites QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.821751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.821751Z digest=sha256:3488416c021a082abb05e2571964301e3c7573ddc84eaed0a3ea5cb86b21c5f3

Observation 1768de1f-a55d-48c9-be05-903aee41c92e · outbound

This paper cites Low-Rank Quantization-Aware Training for LLMs.

FPTQuant: Function-Preserving Transforms for LLM Quantization Low-Rank Quantization-Aware Training for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.907542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.907542Z digest=sha256:699500a3109935d98dd31a517e47fc33095b42a9f898196c46836eea91d240f5

Observation e4a6e479-b504-4dfe-b7b0-2a07dfdb3d53 · outbound

This paper cites Paretoq: Scaling laws in extremely low-bit llm quantization, 2025.

FPTQuant: Function-Preserving Transforms for LLM Quantization Paretoq: Scaling laws in extremely low-bit llm quantization, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:45.985375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:45.985375Z digest=sha256:c5338f3aa87e3b825d15687d928e38c8bb73a5fb25d206b4dc4549ff27bf42ac

Observation 15fd83d1-ec0c-4244-85a5-e58b9eb4ac3b · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

FPTQuant: Function-Preserving Transforms for LLM Quantization GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.065076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.065076Z digest=sha256:04a2c34b2a303a86d9799a7ee5f8ae138fc404e3aa78b14043b487ae6d187820

Observation 9ecc4d88-0f0c-4804-8227-bdda6b9b0470 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

FPTQuant: Function-Preserving Transforms for LLM Quantization SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.139211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.139211Z digest=sha256:a7d9ffeb697105e79c38c1dfdec9c91ed98c95e4772c59a6c20b1491c783349a

Observation 8fad7b8c-9a67-4e89-8b00-bbee994ae2b9 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

FPTQuant: Function-Preserving Transforms for LLM Quantization AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.226301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.226301Z digest=sha256:4ef26ac284a6a89036df91b81b5f9c701f51306b43e13e7dbc7f2210f904f7ac

Observation c502772d-b90a-412c-8437-8958ce0fba68 · outbound

This paper cites Owq: Outlier- aware weight quantization for efficient fine-tuning and inference of large language models.

FPTQuant: Function-Preserving Transforms for LLM Quantization Owq: Outlier- aware weight quantization for efficient fine-tuning and inference of large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.309944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.309944Z digest=sha256:c5d1d53f9a77aeae8fe75340f55d6750c8df983e793e5a1001f42a515edfb740

Observation 8667e534-2518-4bf2-954d-47bdcf62d9c7 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization SqueezeLLM: Dense-and-Sparse Quantization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.395454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.395454Z digest=sha256:8bf60bceb0a2be25a5a00ea4982d2464772fed253e10496c9a373dda9b4d239f

Observation f696d429-04b3-4cc4-a2a8-feec06fcb35b · outbound

This paper cites SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.491664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.491664Z digest=sha256:663e4175cf64675d5bc48894411ed4c6398175946c89b2e59068799ccf748e3f

Observation 9f794de9-c001-4ba4-8925-f3957faa6e3a · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Extreme Compression of Large Language Models via Additive Quantization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.566561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.566561Z digest=sha256:badac802ef5c2d872bfea5bda7fbcb8b9e36c2fc2cbddd22a8aa6f79a8b02795

Observation ec96620d-493a-4815-8591-81342b0588be · outbound

This paper cites A frustratingly easy post-training quantization scheme for llms.

FPTQuant: Function-Preserving Transforms for LLM Quantization A frustratingly easy post-training quantization scheme for llms

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:54.828721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:46.675630Z digest=sha256:c4c62ca6fd36e1b6bb76d9fdad5616110ec73aa349be76768b0e01e644783191

Observation da98e275-fbe4-41a6-ad7a-77f7f3f47aa0 · outbound

This paper cites Flexround: Learnable rounding based on element-wise division for post-training quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Flexround: Learnable rounding based on element-wise division for post-training quantization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:54.667231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:46.752521Z digest=sha256:66b057ddd00f461f7559e964f362a720dff5dbcf21d2adeb696cd77beb60d93a

Observation bf7d682b-5699-4ec9-bb14-7d3682149f37 · outbound

This paper cites Long- range zero-shot generative deep network quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization Long- range zero-shot generative deep network quantization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:54.454601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:46.841044Z digest=sha256:546470857f3e9fbdf68803444db4d5c0a4522300c9503421315490239c9fe496

Observation 0a53fc79-93eb-4c1d-989b-0d38e74703a7 · outbound

This paper cites Quip: 2-bit quantiza- tion of large language models with guarantees.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quip: 2-bit quantiza- tion of large language models with guarantees

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:54.223241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:46.918388Z digest=sha256:9a5068f19ea9566efab96ff7c30e5792b23da3b96b0d04806becc8ef20068cb3

Observation a2cac9bb-05ea-4ff6-a91f-d0c1ecd524df · outbound

This paper cites Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling.

FPTQuant: Function-Preserving Transforms for LLM Quantization Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:46.980855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:46.980855Z digest=sha256:ff561621514985cdc4cf44732a3f91578a30fd1f767d6e4ca0b37a5aa3d818e0

Observation 69b59f67-87e3-4053-98ab-73f54a6393f9 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.032509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.032509Z digest=sha256:cd542cf918a268903bc5a1607291dcd643508bd795c8052ce50137c645400820

Observation e26850cc-f1fa-42f3-8bd4-b7d17fd82973 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, February.

FPTQuant: Function-Preserving Transforms for LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks, February

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:53.975213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:47.124919Z digest=sha256:492f25bb22fed3846af280b4d32e536b22d8e88ecea75b6229e17d45260b61e0

Observation 870a263e-334d-4d85-bcaa-7fa4cc1a2161 · outbound

This paper cites DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs.

FPTQuant: Function-Preserving Transforms for LLM Quantization DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.305676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.305676Z digest=sha256:6960966f9de3aa7416d0f717280fdddd8d5fb316b896df043ceac98360ebdc7f

Observation 252b32eb-f7f4-4bcd-809a-0b2ae9fc5dcc · outbound

This paper cites FlatQuant: Flatness Matters for LLM Quantization.

FPTQuant: Function-Preserving Transforms for LLM Quantization FlatQuant: Flatness Matters for LLM Quantization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.377076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.377076Z digest=sha256:d1e1bec27dc07735239e89aa9703f766fd9b01138baffbbe32caafc401e9885b

Observation 64dc9d2b-c2fd-430f-a0f7-a6a009f3ce69 · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

FPTQuant: Function-Preserving Transforms for LLM Quantization SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.464816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.464816Z digest=sha256:7c61645c554fe720bd5938f433a714560efdf3b226a3d80fc250b28efed1f419

Observation 7287845d-6e7e-4085-9c68-9ffa36c6253a · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

FPTQuant: Function-Preserving Transforms for LLM Quantization Roformer: Enhanced transformer with rotary position embedding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.553216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.553216Z digest=sha256:514637ad798ec0b320b31a55952645710991a8bb69314aac188759a9872455e2

Observation 43470d1e-cbcb-47c7-af8c-9615542eca02 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.634359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.634359Z digest=sha256:bc84e59f9d2d51ed839d775e6180c90d301324c3526e8c6d8be9a7a89cc972e3

Observation f00ec514-0be5-4c54-a336-416364dbeb24 · outbound

This paper cites The Llama 3 Herd of Models.

FPTQuant: Function-Preserving Transforms for LLM Quantization The Llama 3 Herd of Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.714871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.714871Z digest=sha256:4eb0f7fde0ade3e6d2cb03f2b88bfe23a40cd79851f1a5dd09e697a2b65d709d

Observation 16655abc-63aa-47d1-bdef-c8b6b42e0071 · outbound

This paper cites Pointer sentinel mixture models.

FPTQuant: Function-Preserving Transforms for LLM Quantization Pointer sentinel mixture models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.806276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.806276Z digest=sha256:48549f0d7d27af27da546b8c21b2a1ca1cc275d6f75a9f45429dfeab7f3806a8

Observation 19c78207-9ae1-4dc9-996b-c6d7aeeb92ed · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

FPTQuant: Function-Preserving Transforms for LLM Quantization PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:53.813234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:47.981754Z digest=sha256:8a04242b273ea5536279070d52b78e0e70ac77a85654f4252509857cef978758

Observation a34ff82a-8ca1-475f-b07b-e0983c96af3a · outbound

This paper cites WinoGrande: an adversarial winograd schema challenge at scale.

FPTQuant: Function-Preserving Transforms for LLM Quantization WinoGrande: an adversarial winograd schema challenge at scale

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:53.636507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:48.097905Z digest=sha256:c1dbcbf2625ed26ea70b6a45547eceddd60026fe9a9e29e0e21445ef60125833

Observation 1d04d0d8-2b3b-488a-99d5-7ca5f7ae42f5 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

FPTQuant: Function-Preserving Transforms for LLM Quantization HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.336766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.336766Z digest=sha256:3162387d5a10bbada4cd2bd5e0b01292a519d2930f4c159ba2759e825f0d76ea

Observation e7ad51cb-3060-462f-87fc-3d24e2b28753 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

FPTQuant: Function-Preserving Transforms for LLM Quantization Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.475734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.475734Z digest=sha256:98c9247c3cbad0da5eebdd60be81f694c5ed0d217dcf33dc865e95defad7a6bf

Observation f46dc064-b0cf-4035-a0c0-11d2169c7aa3 · outbound

This paper cites The lambada dataset: Word prediction requiring a broad discourse context.

FPTQuant: Function-Preserving Transforms for LLM Quantization The lambada dataset: Word prediction requiring a broad discourse context

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:53.430817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:48.616558Z digest=sha256:039f91497e6b05d7b6e92b77b368a98e2f59fd0301d51804c00b92dead6b6e3e

Observation f3c35bc1-b422-4358-baea-1ae5efa5ca4d · outbound

This paper cites Working with Quantized Types — NVIDIA TensorRT Doc- umentation.

FPTQuant: Function-Preserving Transforms for LLM Quantization Working with Quantized Types — NVIDIA TensorRT Doc- umentation

Reference 58

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:53.245750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:48.792727Z digest=sha256:2fd34f0aa41cfbe632b3846aac98bb3019aa5fe38271b21b9bc0cce7dcdfd54c

Observation e737f547-f15a-4453-a5aa-4dad17c7128f · outbound

This paper cites Quantization — PyTorch AO documentation.

FPTQuant: Function-Preserving Transforms for LLM Quantization Quantization — PyTorch AO documentation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:53.033620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:48.899498Z digest=sha256:4e6c68df085dc2d4a73a41810513283797d8562f71a2d97c6f06af5fa135768f

Observation 30abc064-9048-4414-a3c4-459a113e0e72 · outbound

This paper cites AI Engine Direct SDK documentation.

FPTQuant: Function-Preserving Transforms for LLM Quantization AI Engine Direct SDK documentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:52.915494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:49.036004Z digest=sha256:cb4dd3e07f09020d6ca35a7214a037add51d869cc91f9cbcde7bc357e05899d5

Observation 9d768a9a-4f8f-4431-ba16-962eb3325d66 · outbound

This paper cites TensorRT operators documentation: DynamicQuantize not supported on DLA.

FPTQuant: Function-Preserving Transforms for LLM Quantization TensorRT operators documentation: DynamicQuantize not supported on DLA

Reference 61

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:52.766879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:49.215303Z digest=sha256:9c5dc77b0adb3b524edc24015484e73e3ddf7a6417f1379dfd62e8a8f28bf674

Observation e72a981b-b49f-4efb-80c4-7bc224be7ec9 · outbound

This paper cites fast-hadamard-transform.

FPTQuant: Function-Preserving Transforms for LLM Quantization fast-hadamard-transform

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:52.513457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:49.347868Z digest=sha256:c945b47ece5af36ac307b50dedb0e1375ce2ea8bf696c7409f1c5a7af63fd552

Observation a265f935-0e7d-45f6-9090-3ef70eb62590 · outbound

This paper cites double-packed.

FPTQuant: Function-Preserving Transforms for LLM Quantization double-packed

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:43:52.383814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:49.457392Z digest=sha256:30851deb139efed6090ea7dc0363a8596b230633a6a3fd62c9106f0fd4c4f7c7

Observation 389e9d3d-8215-4636-88c7-3c78d8a8139a · outbound

This paper cites Evaluate quantization error per quantizer placement (e.g.

FPTQuant: Function-Preserving Transforms for LLM Quantization Evaluate quantization error per quantizer placement (e.g

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:52.260528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:49.606480Z digest=sha256:457cabd35bc08f07aeceb5121f1ec4681583313aa928b4eab59e82451235cadd

Observation 0d901eaf-5cd8-4c22-9a12-499f44daeb7a · outbound

This paper cites Based on step 1, choose which FPTs to add: (a) Attention and FFN input.

FPTQuant: Function-Preserving Transforms for LLM Quantization Based on step 1, choose which FPTs to add: (a) Attention and FFN input

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.990152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:49.704829Z digest=sha256:42c55a0006b98102ae1c21eb88152ca9cb2e8451acb0d4e74ae8c8c2137303ce

Observation b57a2e1a-8f17-4c26-90bc-d603c2affe6a · outbound

This paper cites Initialize transforms, e.g.

FPTQuant: Function-Preserving Transforms for LLM Quantization Initialize transforms, e.g

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.744777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:49.859176Z digest=sha256:0c420cdd2bfb5bcc43a65b25122c95bce612af6846f2eb9a30f5f5d5d22f1ef9

Observation fa9295aa-18e8-4de6-a6a1-cb6f538260c3 · outbound

This paper cites Locally optimizing transforms improves performance and reduces training time, whilst incurring very little cost (Appendix F.2.1).

FPTQuant: Function-Preserving Transforms for LLM Quantization Locally optimizing transforms improves performance and reduces training time, whilst incurring very little cost (Appendix F.2.1)

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.540897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:49.988735Z digest=sha256:cdc442ecbb7a4ab85422ee120c7258c4d3024331f246dbb5a8264e13fcaa16aa

Observation 33ec7427-9afd-4f78-a2fc-1c9b3d06c178 · outbound

This paper cites Set the initial quantization grid, e.g.

FPTQuant: Function-Preserving Transforms for LLM Quantization Set the initial quantization grid, e.g

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.372855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:50.099603Z digest=sha256:af5116f3600358eeb5618836dd913076edf50e53b885f79c768f9b23e0c281b2

Observation b4549857-6b67-4f37-9003-499b98980c9d · outbound

This paper cites Train the FPTs and quantization grid end-to-end, with the unquantized outputs as target.

FPTQuant: Function-Preserving Transforms for LLM Quantization Train the FPTs and quantization grid end-to-end, with the unquantized outputs as target

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:43:51.087441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T10:43:50.212638Z digest=sha256:69462b66a38f777febbe7426a78f081ee87231c31026648c150b99e4c62b93f4

Observation 9a0cd5c9-7041-469c-94cf-c7e320333fb7 · outbound

This paper cites doi: 10.1145/3474381.

FPTQuant: Function-Preserving Transforms for LLM Quantization doi: 10.1145/3474381

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:48.191111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:48.191111Z digest=sha256:511a1718fa72b8705079bbbae276f9fc39a2eb383c834ce5c333c07b94241c93

Observation abf25eb5-8f0d-4593-965a-6a9964b135d3 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

FPTQuant: Function-Preserving Transforms for LLM Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:47.223987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:47.223987Z digest=sha256:28804843242c0fd747e4b1bbea1f5a43630925921b0b99f773629e944593b3fd

Pith citing papers

Observation fbe7599b-8d5c-4a5e-b8b4-f20481b37886 · inbound

Leech Lattice Vector Quantization for Efficient LLM Compression cites this paper.

Leech Lattice Vector Quantization for Efficient LLM Compression FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T23:10:46.775151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:10:46.775151Z digest=sha256:9b2e1ed2d92979a62a9cefaa1cf2060ef252ef6d8d1f5e96951cb7c9d53450f7

Observation 7bd8316e-ec7e-42e8-9b9d-7b05baa4b7ba · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 147

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:93a5adbe5ec267bc56ac3dc453d1db804bcf58bfda1045f3fad37f3858d733c4

Observation 781beda3-9b01-4397-bf23-090a6a93836f · inbound

When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon cites this paper.

When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:36.313617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T03:16:00.868964Z digest=sha256:3623d1c74ccde84e1e2875efcf799156be568d78c46d804a5d0ffd5258fb71b7

Observation 70b001ba-0d1e-4e7a-8c1e-dd1c65847f4e · inbound

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation cites this paper.

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T09:35:12.908124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:35:12.908124Z digest=sha256:33d61d2732799a13b7fa44c784b593c9935b0f2e186f2bda4db5acc804b7a67b

Observation f81c0d0d-901b-418c-bf2b-bff90bf5bbf0 · inbound

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation cites this paper.

RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T16:53:34.398685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:53:34.398685Z digest=sha256:3e112dc87fcb16f9374550ce786cfc8e0b4cc2e659a0a3bf71145f1efcfb4e1b

Observation e06b035f-5455-42b9-a447-a8d42814148c · inbound

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation cites this paper.

When Local Variance Optimality Is Not Enough: RoPE-Aligned Q/K Rotations for Dynamic 4-Bit Quantisation FPTQuant: Function-Preserving Transforms for LLM Quantization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-14T12:54:26.612471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T12:54:26.612471Z digest=sha256:d6091d253918f934df2b29509f23df24c9ac6d9acd4ccb990ec608aa77dfffb2