Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:05:12.928814Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2504.21553.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:05:12.928814Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:09:57.976654Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-16T12:09:58.199009Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 45c07fd1-c0f4-478a-bab6-089220b261ec · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Advances in Neural Information Processing Systems36, 34278–34294 (2023)
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8f4fde32-7388-4520-90fe-067769996c82 · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Understanding and Overcoming the Challenges of Efficient Transformer Quantization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eda4d2bc-0b03-4630-9a0f-72709eda8d96 · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models int8 (): 8-bit matrix multiplicationfortransformersatscale.AdvancesinNeuralInformationProcessing Systems 35, 30318–30332 (2022)
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe0c0ad9-39a8-4277-802b-8f5a05ec22ae · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Advances in Neural Information Processing Systems36 (2024)
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 337645d5-2733-468e-8c36-8c2b374579cc · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models In: International Conference on Machine Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1ae82f16-f726-40cd-be1f-3b2282f3fd38 · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 520e269d-5eea-4679-ab18-969282ee0dcf · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Understanding and Minimising Outlier Features in Neural Network Training
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddfe4a3b-c127-478a-a440-5e7577d41207 · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c35f00c6-773c-4eac-a5d3-9b6bfc21fe71 · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Mistral 7B
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c4dff9b-d604-4e10-b11d-1e28b491087b · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6f0aa712-e0ad-4e49-a513-41b0ae1e7d55 · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c4ffb8-f0d8-4383-ae7e-e6cd0eb63ca0 · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models FP8 Formats for Deep Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b44d73b-7bf7-42d4-9b94-3d55ac77e13c · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cb75199-7fba-4c16-a011-927ca194a9da · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Proceedings of Machine Learning and Sys- tems 6, 483–498 (2024)
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 12ebaf82-965f-4019-8fe3-47c6875bdf78 · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45cac19a-40ed-437a-98ad-ca2e1c43b986 · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Massive Activations in Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c33236ec-3ad5-4f97-8a55-7ddcb28b3c77 · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models https://doi.org/10.5281/zenodo.10256836, https://doi.org/10
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee496d5-c5ee-46a8-9f10-5f9e9b15a31f · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a59e6372-493e-4fba-9906-368a752eeb44 · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models BitNet: Scaling 1-bit Transformers for Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c5c0b76-c9f0-4067-93de-6e1e8c64449c · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models In: International Conference on Machine Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f28bf1b0-3d02-4e9d-a8cb-8af9175b062f · outbound
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98677ebc-dbe7-41d4-a815-9166739c8b99 · inbound
Gradual Binary Search and Dimension Expansion : A general method for activation quantization in LLMs Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.