Pith. sign in

Paper Citation Record · LEDGER

SoftWater: Class-Aware Rate Allocation for Softmax Quantization

As of 19 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.12026.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12026 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:26:51.778411Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact10
  • verified fuzzy1
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d23ad11c-ac73-4c34-8ebe-2034a1314552 · outbound

This paper cites AXELRAM: Quantize Once, Never Dequantize.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization AXELRAM: Quantize Once, Never Dequantize

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.500088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.500088Z digest=sha256:f0fcd39ede517a0a3614f300182f773c9dba9dc7d3ab6e4e1b45df09e2d04b61

Observation 2b86eb8a-201c-4f20-a999-e14b9a084ba2 · outbound

This paper cites TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.504805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.504805Z digest=sha256:6f061ff1a198dc5427aded54c4016e4bafeb0d7b9d17fde2917b401cd48c8c47

Observation a4ee110a-22d5-442b-8012-e351fb0510af · outbound

This paper cites QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.509442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.509442Z digest=sha256:1850151cc413435126f3077d96920052835f72826a29fc186286d18297b0169e

Observation 1b6a5747-1954-4ae4-bf7c-4c3d28b66a8a · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.513703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.513703Z digest=sha256:28946fd1983a02a5e6a1d5296a2c3cf07d116a9ac63e8f789f65d82544f0d8ed

Observation 235f8011-2932-4530-a385-a1158e7055cd · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.518037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.518037Z digest=sha256:e6f40dff5deb76fbf9b6847353fdbf4fa4ffcb864b09763ea7443024ab29aeb3

Observation b2766b5d-917f-4429-9267-a68cece3f823 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.522547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.522547Z digest=sha256:b117e5895e63987d0d79b07c06c219f42f7c0090b5cb814ffb5859bb88ff18ab

Observation 5663b2f8-f976-49f4-b2b9-c8df948b83c0 · outbound

This paper cites an unresolved cited work.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Unresolved cited work

Reference 7

Resolution
verified exact
doi, observed 2026-08-16T00:26:53.041368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.526747Z digest=sha256:bf6a85f6af549e0e4ecc357c2bcaaa97e26755ff346abbba77e6b1afe55874aa

Observation a7cc9b04-0212-42f1-861a-cc6d3e373acd · outbound

This paper cites Dissecting.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Dissecting

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.531203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.531203Z digest=sha256:dc3df08c75e6b02a5f79b110f87766b252a2130d0f26312f372638b52c352bca

Observation a87882a0-d631-46a5-aa02-dfc168befcf4 · outbound

This paper cites doi:10.48550/arXiv.2511.04063 , abstract =.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization doi:10.48550/arXiv.2511.04063 , abstract =

Reference 9

Resolution
verified exact
doi, observed 2026-08-16T00:26:52.914324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.535210Z digest=sha256:efdf02cd9373cef1b2443745fa97ba069ada52ec0a9fae3442236e7f75d10e9d

Observation d2b58a0a-0968-473b-a683-e1feb7c4afe9 · outbound

This paper cites doi:10.48550/arXiv.2511.10645 , abstract =.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization doi:10.48550/arXiv.2511.10645 , abstract =

Reference 10

Resolution
verified exact
doi, observed 2026-08-16T00:26:52.844119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.538852Z digest=sha256:120a7a0b96791045b2902b68123e81739dc2d206a995e0cd7de11a990e9991c9

Observation 97d65aba-40d7-4841-b922-d8b2986223cf · outbound

This paper cites SpinQuant: LLM quantization with learned rotations.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization SpinQuant: LLM quantization with learned rotations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.543169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.543169Z digest=sha256:71288f6402645cf70b1a3f6e479fecfc2339e832ef452063ed576872e6413d2c

Observation caf38226-9d49-482e-952b-e44c219443e1 · outbound

This paper cites doi:10.13140/RG.2.2.28167.37282 , abstract =.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization doi:10.13140/RG.2.2.28167.37282 , abstract =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.547272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.547272Z digest=sha256:0f9b36d790978aa9146ce978e366b9e93828f14df5861bcae0ec98154758443c

Observation 87d1159e-4129-4b3d-921a-5cd31ba1d80a · outbound

This paper cites State-Free Inference of State-Space Models: The Transfer Function Approach.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization State-Free Inference of State-Space Models: The Transfer Function Approach

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.550784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.550784Z digest=sha256:f704be63a5bb7acdbe625513099a07a17584957e1d8de5563adc89e1075030f7

Observation 087e4c9f-a18b-4f6b-b0b0-cc6c0789185e · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.554740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.554740Z digest=sha256:9f0e2a518c0ed20b6d4d85eb56d29733884756b90a3a1d794a831eabdfada1dd

Observation 831da614-fa9b-4ee5-9cad-f063f484c55e · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.558596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.558596Z digest=sha256:a146b260e3aa6d272d42c1a3a58f39aa9e263316d8bc535b128feb7f899f4abc

Observation e1c4e79b-6d5f-4fa9-bc71-0e09d4070e1a · outbound

This paper cites Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and Pruning.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and Pruning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.562222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.562222Z digest=sha256:f958e79e5d355fb0d5104e62ad49db4b317b00bf4036de8f0e28baa3f08b17c8

Observation 24b5c727-5a02-4075-92e2-fabc432917ce · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.566355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.566355Z digest=sha256:86ce199df9df882b3f7e69b36d076f3811868b59ca2ab1dc1fcb3f44eb0b31e3

Observation 3539eb1d-cc1c-4224-804f-874405855c10 · outbound

This paper cites QuIP: 2-Bit Quantization of Large Language Models With Guarantees.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization QuIP: 2-Bit Quantization of Large Language Models With Guarantees

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.570859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.570859Z digest=sha256:9b09264bc8283a8156cfaac80f2e5d6b9e98f4d591c71f6951cb2436c32dacfa

Observation 3d65f970-4d23-4d0e-a014-63b7c6df67a4 · outbound

This paper cites QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.575124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.575124Z digest=sha256:1ea408e9bde361fff1741c266cdb84afcf8d34db156e299bd0c5bce778c7be05

Observation 21c60d54-ce2e-48c9-8cd4-b389263751a6 · outbound

This paper cites Alignment-.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Alignment-

Reference 20

Resolution
verified exact
doi, observed 2026-08-16T00:26:52.693420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.579166Z digest=sha256:2c9227ce00070dd82aa8dcec5b2ed9afcd4c6ae9f3739c7ceb929cdf50d34a89

Observation ed628f4e-df81-4a71-a841-828e89b04848 · outbound

This paper cites Catastrophic Failure of LLM Unlearning via Quantization.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Catastrophic Failure of LLM Unlearning via Quantization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.583018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.583018Z digest=sha256:28bb59b0f1a62d511e83c999c79cf204880daa6fb22615f8e16bb1c35cc13501

Observation 5c93bedd-ff1c-4c83-b3a4-13aa4f6ab65b · outbound

This paper cites and Colombo, Maurizio and Damiani, Ernesto and Asal, Rasool and Almemari, Al Anoud and Alhammadi, Yousof , month = dec, year =.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization and Colombo, Maurizio and Damiani, Ernesto and Asal, Rasool and Almemari, Al Anoud and Alhammadi, Yousof , month = dec, year =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.586838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.586838Z digest=sha256:7031759921e062ce6a07c6f3edf909e1beec8fc644ef2e24836a13568cf2f9b0

Observation 36add2ab-5b41-4b9e-935b-7dca4107d409 · outbound

This paper cites QTIP: Quantization with Trellises and Incoherence Processing.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization QTIP: Quantization with Trellises and Incoherence Processing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.590478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.590478Z digest=sha256:07919e289f44685a9a767f21167b484373439744ca398ca6e10bd472097e98dd

Observation 28cce917-4049-48f3-9968-73e21502bab9 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization QLoRA: Efficient Finetuning of Quantized LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.594380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.594380Z digest=sha256:04c61656074d2638057a350800a663fe5a354293882dc71ff516bcbef1d7dba2

Observation ccb1293b-471a-47cc-b055-fd365cefbbab · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.598410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.598410Z digest=sha256:27bc66d58a1243be3dd9d7d14628c3b4c51905d78d3e2f30c7a7e5fec4ba99e8

Observation 19507ccb-c5c4-4a5b-b5aa-0353fc2272f0 · outbound

This paper cites Extreme Compression of Large Language Models via Additive Quantization.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Extreme Compression of Large Language Models via Additive Quantization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.602354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.602354Z digest=sha256:d8c557f76723a459f57b1145dc40e26207d4528ef1e8384364f253ab442051ac

Observation af8f2fe6-fa14-4a05-b213-4163f423ebaf · outbound

This paper cites Killing Two Birds with One Stone: Quantization Achieves Privacy in Distributed Learning.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Killing Two Birds with One Stone: Quantization Achieves Privacy in Distributed Learning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:26:52.563153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.606515Z digest=sha256:79bb078d4ba6b52d496230e871cc104748f0f99b553c0528711cc86b923fe212

Observation 74a6596e-74ea-4225-8938-7d94d2bdf475 · outbound

This paper cites Recipes for Pre-training LLMs with MXFP8.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Recipes for Pre-training LLMs with MXFP8

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.610461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.610461Z digest=sha256:3db74e4d2f307e22401ed8971f9c6ff04bbcf436fe5c624682ea510a38ec069a

Observation db3b9bd9-9be9-461c-8789-c16c952a9579 · outbound

This paper cites Pretraining.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Pretraining

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.614400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.614400Z digest=sha256:759d2fabbceb584ea45170faec8a899e22cb4fc3b914c42fa64ff027b5ddf567

Observation 3636cfac-6dc4-4c69-8671-613198abb1dc · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.618005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.618005Z digest=sha256:35f8f9a8b3453760e41e6ffaa7d8c83c7c15cc9c458afe2305512f7343fc8f10

Observation d9e3e1fb-59bf-4c8b-a7aa-d605dfa6987d · outbound

This paper cites an unresolved cited work.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Unresolved cited work

Reference 31

Resolution
verified exact
doi, observed 2026-08-16T00:26:52.500001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.622002Z digest=sha256:ebf244db14c1ce28bd396224aaebd49dee50f22e725a1e0a7087e521580a0004

Observation 2e73f742-8acc-4034-aea7-93e1458efc6e · outbound

This paper cites Up or Down? Adaptive Rounding for Post-Training Quantization.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Up or Down? Adaptive Rounding for Post-Training Quantization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.626399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.626399Z digest=sha256:1552472a7ba77bd8cb51af61b45a36f6530ff78fe2fb847c4060ef97a1015b8b

Observation 4e2d3b31-60ab-4525-ba3a-e70eaa4e5e4d · outbound

This paper cites MLoRQ: Bridging Low-Rank and Quantization for Transformer Compression.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization MLoRQ: Bridging Low-Rank and Quantization for Transformer Compression

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:26:52.425831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.630883Z digest=sha256:aa9fbcf41ab3159450ad7fc9e90b787db81a8552192a5a2111cced5fbc2f0871

Observation b0169f28-116a-4918-bb90-a3924ce3dbd9 · outbound

This paper cites SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.634894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.634894Z digest=sha256:fba6d54db35dde5054792ad1700d9ea69c6c5795b00b28f91dbd3e150545844d

Observation 4013b1a7-1956-40f5-abfc-d00e97b20717 · outbound

This paper cites NestQuant: Nested Lattice Quantization for Matrix Products and LLMs.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization NestQuant: Nested Lattice Quantization for Matrix Products and LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.638644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.638644Z digest=sha256:1e2c03bd5833eb791c9b979205088367fbee2b9604fa706244ec5c4d51046481

Observation ade6ce60-90a3-4f87-939f-0291b703176d · outbound

This paper cites GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.642528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.642528Z digest=sha256:a97e16e38924249b92b309c654b406485f55e364ebc9d8576155518e9a07e671

Observation 26bc2b3a-4cd1-406e-b065-8706d7b8fb36 · outbound

This paper cites Rethinking Residual Errors in Compensation-based LLM Quantization.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Rethinking Residual Errors in Compensation-based LLM Quantization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.646615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.646615Z digest=sha256:828b2de1344afb0474ec158057c1c84324b82d874f5ab85ce94571a27267d2ff

Observation 2d4e0277-9f39-4269-bd06-bd8485746398 · outbound

This paper cites WaterSIC: Information-Theoretically (Near) Optimal Linear Layer Quantization.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization WaterSIC: Information-Theoretically (Near) Optimal Linear Layer Quantization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.650581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.650581Z digest=sha256:8106d13fa2d3c501f0ac8f180904558532888eb1896f3667a058b5b4d0d39d4f

Observation 3b6534d7-e62d-4c74-bf4e-84377acf9c71 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.654544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.654544Z digest=sha256:b20a66ff149aac2977a0fb29d8a41b6ee0e087b1d06b2633b3f9ce1c93de434a

Observation 1be444a1-20c3-480d-afd0-44bccfaf8617 · outbound

This paper cites LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.658533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.658533Z digest=sha256:d5399eb7d9be621b698dfaaa94c11b06e39c1054bc40fca7bc93c2f09744b885

Observation 2d9b3d8e-5cc4-473a-a2e8-45b1615d9559 · outbound

This paper cites and Bich, Philippe and Zhuang, Jiawei and Çelik, Ahmet and Benfenati, Luca and Cavigelli, Lukas , month = jan, year =.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization and Bich, Philippe and Zhuang, Jiawei and Çelik, Ahmet and Benfenati, Luca and Cavigelli, Lukas , month = jan, year =

Reference 41

Resolution
verified exact
doi, observed 2026-08-16T00:26:52.344767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.663885Z digest=sha256:ece8e116e305e0d6b84c24ccee873d281865bc27d30da85ce17085b338e1395f

Observation 993250b8-cdda-4420-a0bc-670ead4e7926 · outbound

This paper cites The Super Weight in Large Language Models.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization The Super Weight in Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.667494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.667494Z digest=sha256:308137c3d25e0fbeb7669c9a10d697a1a5427b8a1b2d76ecb0aceb686a79892d

Observation 3756099b-bceb-44ae-b4ba-7ae27cd418c1 · outbound

This paper cites Massive Activations in Large Language Models.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Massive Activations in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.671329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.671329Z digest=sha256:aa2bbad74e25ad11421b616b9c490d5234ea3d515d821e98829a28402c137718

Observation e1585feb-4a47-4952-98c2-ddb46aae4078 · outbound

This paper cites OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.675323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.675323Z digest=sha256:8880d76fbe8d454957debddf3b68fab74e089b87888d655004576a67e3f1ce17

Observation f9c5dd66-575a-4b60-b814-0b88228afa18 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Efficient Streaming Language Models with Attention Sinks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.679402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.679402Z digest=sha256:5158324b90fd36812319388c46efc4c84528d8879a49ce71b10fc0e0818e2c93

Observation aa933d45-7c55-40f8-acb6-a0d17c2b0c51 · outbound

This paper cites Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.684121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.684121Z digest=sha256:ac7b9e0b97c9d0a86764baad54461e9dad40fe796e30a63800e62fa73bc5aff1

Observation f91ee2ba-c361-4795-b166-3c2d3b8d50a2 · outbound

This paper cites Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.687990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.687990Z digest=sha256:82e6cdd0147c748497d23e3c43aa78576e81c7aaee70bab1ba51fcc55de90937

Observation 49f88765-0c58-4e08-906e-aae74434e448 · outbound

This paper cites Round and Round We Go! What makes Rotary Positional Encodings useful?.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Round and Round We Go! What makes Rotary Positional Encodings useful?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.691945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.691945Z digest=sha256:ab5e55d7d1d015d02999da43e7c73afaba7f5baeb0499922477e4c0829fb6cef

Observation 8e031988-e728-467b-afc3-bde3ec8892f7 · outbound

This paper cites Selective Rotary Position Embedding.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Selective Rotary Position Embedding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.696105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.696105Z digest=sha256:2a5f8d71a69973d0b86b9cf5e39386ad34324c9d551dba789ae0bfb73132c139

Observation 9ca675b6-9056-4e84-b843-023de9cb0037 · outbound

This paper cites Base of RoPE Bounds Context Length.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Base of RoPE Bounds Context Length

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.699966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.699966Z digest=sha256:ae9ddfa5335b9112f47c8212e4f51f5fea58f3df45e75bac18fb51b283355507

Observation 507c1f6e-ec95-4dfa-b3dd-21e28d38e172 · outbound

This paper cites LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:26:52.179385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.703987Z digest=sha256:e699d4cb60530dea2491d07e090072c2656fe3a097ff1f7f6cbf3cbf0f0997a6

Observation 985d0610-4d2d-4a68-92e6-acb10f7dc5cf · outbound

This paper cites VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:26:52.160396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.707959Z digest=sha256:3c25c3f5aa7ff136eb087097b0e7cf9ec03e70b88c3ef2a71ad1a894ff3336c4

Observation 49d27bf2-290a-4511-8082-8ad6c7caf822 · outbound

This paper cites doi:10.48550/arXiv.2603.14591 , abstract =.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization doi:10.48550/arXiv.2603.14591 , abstract =

Reference 53

Resolution
verified exact
doi, observed 2026-08-16T00:26:52.145900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.712012Z digest=sha256:f480e805d9eaf59b26159d1d29456a424adee1942ef488772b035756eb2c5dce

Observation ca14a1bc-e680-4ae9-bc85-9b6c6aff31e2 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.715866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.715866Z digest=sha256:2363f3dc98204d14c83758f0686e2870e408253f29e01a9b592c3f302a35464e

Observation 8a5b88ee-d84b-44d9-8265-e6561511cb08 · outbound

This paper cites Efficient softmax approximation for GPUs.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Efficient softmax approximation for GPUs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.719917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.719917Z digest=sha256:180beb7ed256c1a983e18b0bc4cff7a788ea37c83609fbc6a7aed5d5dec2e225

Observation 254ee6af-47ac-4045-91d5-99b157a0eb0d · outbound

This paper cites The Curious Case of Neural Text Degeneration.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization The Curious Case of Neural Text Degeneration

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.723834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.723834Z digest=sha256:669d3acbdbdc80ae55030ada004a9c1b7efadad449be79710272cec6ee236e16

Observation ea009249-871f-4384-805e-5d62b270eeba · outbound

This paper cites CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:26:52.046228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.728639Z digest=sha256:90d31ccde346c158ed58145fcbec597dec50f887a9ff66f62ebb57965d28cec3

Observation ffb199e8-2cbe-4df4-ab01-62132dd81eca · outbound

This paper cites an unresolved cited work.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.732590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.732590Z digest=sha256:a7d4cab0538dc4b914aa8ac388ea70edd4807b0fe860da826dc335e88b963dbb

Observation 80c75c3f-d275-4cc9-955d-c78ea2415877 · outbound

This paper cites High-Rate Quantized Matrix Multiplication I.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization High-Rate Quantized Matrix Multiplication I

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.736302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.736302Z digest=sha256:6bc7949afa50f25ba62843a4957640a633cfc1966a74445d784214fcedb33a3a

Observation f53b61fb-25a5-4b64-8e05-b2918e630807 · outbound

This paper cites Model-Preserving Adaptive Rounding.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Model-Preserving Adaptive Rounding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.740209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.740209Z digest=sha256:53a898815d16caace4609b6be66fda0d3a4714ba27b3b8cea9c4912e2a9cd374

Observation da77bd52-72ff-49b5-8878-2eebeaa89c42 · outbound

This paper cites and Song, Hyun Oh , month = sep, year =.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization and Song, Hyun Oh , month = sep, year =

Reference 61

Resolution
verified exact
doi, observed 2026-08-16T00:26:51.994116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.744611Z digest=sha256:5b5ca6a491ab258b851aa8b4ee729c85de9a4b4250502e21ffbde5001afa7cb7

Observation d4ce992d-b695-48fe-88a0-391f7cc1ca3a · outbound

This paper cites Human behavior and the principle of least effort.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Human behavior and the principle of least effort

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:26:53.251815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.748379Z digest=sha256:85371ea9e4ebc4f4779050d16f897127c1d2fb1e67318389c6b79ece7f175049

Observation 6ad7c461-1b02-43b8-9624-efe60819ca52 · outbound

This paper cites Psychonomic Bulletin & Review , author =.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Psychonomic Bulletin & Review , author =

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.751706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.751706Z digest=sha256:ffac64e83ff6841f777b65c00c4d6586f5c03094792b9059319a1228319ad094

Observation b240474b-dd8a-4af9-b7eb-a28a2a3334d2 · outbound

This paper cites Optimizing Neural Networks with Kronecker-factored Approximate Curvature.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Optimizing Neural Networks with Kronecker-factored Approximate Curvature

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.755983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.755983Z digest=sha256:3f5995ed68ece43bca67aecd2229687eff9b8b7302f08645bf56dbd4231bec84

Observation 3ed13b82-8516-466a-a84b-f993095d7624 · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.759776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.759776Z digest=sha256:dfa19647b617c127bcc291887702f93f09d5e43338e3ec4b9a717140a0e84341

Observation adbdb7b3-ce6f-4f39-880d-f56aaf6f7f9b · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.764631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.764631Z digest=sha256:826efcd4bb247e147450ceb834d32cc4dbd1005cee5b15f1175d3ccf4b16d9a6

Observation 50cc1f59-62da-453f-82f3-8a06c8113563 · outbound

This paper cites The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization The Geometry of LLM Quantization: GPTQ as Babai's Nearest Plane Algorithm

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.769211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.769211Z digest=sha256:9acf75037488c282f25cfb8b2fd336ae8eed8d1e3337bbd6b8acda244fe71380

Observation 8b5471b6-96d2-4dea-8d8d-f04fa5c628a1 · outbound

This paper cites an unresolved cited work.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Unresolved cited work

Reference 68

Resolution
verified exact
doi, observed 2026-08-16T00:26:51.872148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T00:26:51.773395Z digest=sha256:7eeb4b66407a562c1304bc88aa3a9136f70d675dc268ad38ca7e7b5913d3ec94

Observation f3fde395-b9e3-4a47-9e28-ab6125f5474c · outbound

This paper cites Journal of Computational and Applied Mathematics , author =.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Journal of Computational and Applied Mathematics , author =

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.778411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.778411Z digest=sha256:a51d7816a6e96a406f1d0662093918bfae38c4943545bcb71e40428ebdc4f0c3

Pith citing papers

No inbound Pith citation observations are available.