Pith. sign in

Paper Citation Record · LEDGER

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models

As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2504.21553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.21553 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:05:12.928814Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:09:57.976654Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T12:09:58.199009Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45c07fd1-c0f4-478a-bab6-089220b261ec · outbound

This paper cites Advances in Neural Information Processing Systems36, 34278–34294 (2023).

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Advances in Neural Information Processing Systems36, 34278–34294 (2023)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:05:13.234203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:05:12.838529Z digest=sha256:8df981fc2201a2a9325a137370d14f6d5f121d772e91af035365d9dd59f1396f

Observation 8f4fde32-7388-4520-90fe-067769996c82 · outbound

This paper cites Understanding and Overcoming the Challenges of Efficient Transformer Quantization.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Understanding and Overcoming the Challenges of Efficient Transformer Quantization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.843611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.843611Z digest=sha256:6a56f2ae74c8f473fc00397e0f930916190d9af910089ae2e605538ffc7926cb

Observation eda4d2bc-0b03-4630-9a0f-72709eda8d96 · outbound

This paper cites int8 (): 8-bit matrix multiplicationfortransformersatscale.AdvancesinNeuralInformationProcessing Systems 35, 30318–30332 (2022).

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models int8 (): 8-bit matrix multiplicationfortransformersatscale.AdvancesinNeuralInformationProcessing Systems 35, 30318–30332 (2022)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.848749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.848749Z digest=sha256:a1f5e1d939bd2d622d5ebba3a12eb45eae90e37b2db4f8b815c5097c6a0d28e5

Observation fe0c0ad9-39a8-4277-802b-8f5a05ec22ae · outbound

This paper cites Advances in Neural Information Processing Systems36 (2024).

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Advances in Neural Information Processing Systems36 (2024)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.853226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.853226Z digest=sha256:9828c0b49b34646cec09095ba758b6f35f2ecea6285711a1f36df98c9990288e

Observation 337645d5-2733-468e-8c36-8c2b374579cc · outbound

This paper cites In: International Conference on Machine Learning.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models In: International Conference on Machine Learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:05:13.201046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:05:12.857680Z digest=sha256:1a81926fa9a21a19c1996f49edb6ba0d78d5e8d6dea73a41faaf025f570b4361

Observation 1ae82f16-f726-40cd-be1f-3b2282f3fd38 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.861936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.861936Z digest=sha256:5a9cafedbaa7a9cecb398fef46bb33d085088a083e804995bf10ddaf88023ca1

Observation 520e269d-5eea-4679-ab18-969282ee0dcf · outbound

This paper cites Understanding and Minimising Outlier Features in Neural Network Training.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Understanding and Minimising Outlier Features in Neural Network Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.867019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.867019Z digest=sha256:a2d28d7589b87524ad4b8534c73e67ea8b779e8cca7daff0e5d6ec9ac8c14fc6

Observation ddfe4a3b-c127-478a-a440-5e7577d41207 · outbound

This paper cites an unresolved cited work.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:05:13.187308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:05:12.872026Z digest=sha256:a01b2a562c037a74b588184a883a37a5b94b09eb49dd110a81c057513fcadb15

Observation c35f00c6-773c-4eac-a5d3-9b6bfc21fe71 · outbound

This paper cites Mistral 7B.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Mistral 7B

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.876317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.876317Z digest=sha256:e39b290877f219745be03e6e31352259a9a93e6b790f720ec0b1bb1e4a1800f7

Observation 7c4dff9b-d604-4e10-b11d-1e28b491087b · outbound

This paper cites an unresolved cited work.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:05:13.173293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:05:12.880599Z digest=sha256:a6435b9284d4e829c2b8f8264bebfb6c99fa1cdb4e2f5cd2e457fb763950e7dc

Observation 6f0aa712-e0ad-4e49-a513-41b0ae1e7d55 · outbound

This paper cites The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.885301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.885301Z digest=sha256:46a16b680318180c96a4d0bfc95e955a6b8434140e14f5b87e4bca32165a1732

Observation 33c4ffb8-f0d8-4383-ae7e-e6cd0eb63ca0 · outbound

This paper cites FP8 Formats for Deep Learning.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models FP8 Formats for Deep Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.890019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.890019Z digest=sha256:e82bdb63fa3d93242ee535ae2701ae4b657774608a6e2534bb79cc8905fe1485

Observation 6b44d73b-7bf7-42d4-9b94-3d55ac77e13c · outbound

This paper cites Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.894637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.894637Z digest=sha256:4cc5a8181ee6199fd458a86c7eeb568f6595c0efb038d4f607609c97aa5c9428

Observation 3cb75199-7fba-4c16-a011-927ca194a9da · outbound

This paper cites Proceedings of Machine Learning and Sys- tems 6, 483–498 (2024).

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Proceedings of Machine Learning and Sys- tems 6, 483–498 (2024)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:05:13.158237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:05:12.898893Z digest=sha256:a693da99e3fa8344988cf83721948be9e21a16224450843b0ca439e71f2a2a0c

Observation 12ebaf82-965f-4019-8fe3-47c6875bdf78 · outbound

This paper cites Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.903254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.903254Z digest=sha256:1e63ebebed1bef813bbcbb758f8e5e309560fcc46165f50d9ba454eb22df9d4a

Observation 45cac19a-40ed-437a-98ad-ca2e1c43b986 · outbound

This paper cites Massive Activations in Large Language Models.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Massive Activations in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.907873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.907873Z digest=sha256:b6ff0b2f18ceacdbc72c555538fcde23df6090efb9d455838cd15bfa647cd511

Observation c33236ec-3ad5-4f97-8a55-7ddcb28b3c77 · outbound

This paper cites https://doi.org/10.5281/zenodo.10256836, https://doi.org/10.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models https://doi.org/10.5281/zenodo.10256836, https://doi.org/10

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.912149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.912149Z digest=sha256:765ed2f872b412ce62ee48cfc692e336a976a9828b3d467ae361ed4e6c2c6575

Observation 1ee496d5-c5ee-46a8-9f10-5f9e9b15a31f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.916285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.916285Z digest=sha256:da58b278eb17fadebd8f5195deca5658d8f30ee28426bacb28f025955092355b

Observation a59e6372-493e-4fba-9906-368a752eeb44 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.920433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.920433Z digest=sha256:63cddeefe8f8f05f94af318f62300a2ee03ab8284d88abd77e2ce63cbaa92b6c

Observation 2c5c0b76-c9f0-4067-93de-6e1e8c64449c · outbound

This paper cites In: International Conference on Machine Learning.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models In: International Conference on Machine Learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:05:13.144216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T05:05:12.924694Z digest=sha256:3e5e6b2c9d337c1893a3114bceed7f317d140233c5cdde8381240406b037892f

Observation f28bf1b0-3d02-4e9d-a8cb-8af9175b062f · outbound

This paper cites Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs.

Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:05:12.928814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:05:12.928814Z digest=sha256:9907898b951f3f2ce42c84e3ffc634e375ee6e0c67ab0ecbe81e9d80c0951841

Pith citing papers

Observation 98677ebc-dbe7-41d4-a815-9166739c8b99 · inbound

Gradual Binary Search and Dimension Expansion : A general method for activation quantization in LLMs cites this paper.

Gradual Binary Search and Dimension Expansion : A general method for activation quantization in LLMs Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:09:58.204110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:09:57.976654Z digest=sha256:610ae5ac3768c12ef64982d433bc3655809722a372558788e69688879cd40681