Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T10:01:56.131253Z
Paper Citation Record · LEDGER
As of 3 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 4 inbound Pith citation observations for arXiv:2510.04212.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T10:01:56.131253Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:54:47.068638Z
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 84ef2b0d-6d00-4360-8707-47ea35c43375 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Scalify: scale propagation for efficient low-precision LLM training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c8ce8386-332b-4824-a827-d4101c1410d5 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention u-$\mu$P: The Unit-Scaled Maximal Update Parametrization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation acad3c74-c1c2-4dc6-bcbf-cf9f912f814c · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2747e26a-663e-407e-9239-46432204b867 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Scaling FP8 training to trillion-token LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a45bdfba-62a7-40cd-8e79-c5196ab742e5 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0a266c88-bd65-4dac-bdbe-67731ecbcadc · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Is Flash Attention Stable?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 51ef9cb5-1de1-4065-b39d-0102ff8bb07f · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation cc065a55-4627-413b-9d94-158782c9e583 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Query-Key Normalization for Transformers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8f6b5e54-03fc-4290-a857-7034d2da5616 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Training Compute-Optimal Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 88aff18f-3fb0-4d42-b513-6972e031ff0f · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 52e00fa9-ec25-4b75-8a7a-7be7ccc302f3 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention A Study of BFLOAT16 for Deep Learning Training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 53f66d0e-6c7b-41f6-96be-f52c7003805a · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Kimi K2: Open Agentic Intelligence
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0d9e66a6-2abb-4607-b5bf-4c19b90c21b2 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention DeepSeek-V3 Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 3e77396a-2006-4532-94df-1611f448be03 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Mixed Precision Training With 8-bit Floating Point
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7b607e40-1368-4393-880c-4fb464af087f · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Mixed Precision Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 267a6714-9fda-4947-8fd6-a76fb131a427 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention FP8 Formats for Deep Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 89f76f7b-7d55-4d47-bbd3-2d7e6cf46b53 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention A Theory on Adam Instability in Large-Scale Machine Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1e1dbe79-4ae7-41dc-939e-c2a6900e0c92 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention nanoGPT Issue
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 93c4ca0f-0f0f-492e-9f4b-9839f0659c4e · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention 8-bit Numerical Formats for Deep Neural Networks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 65c71bf6-d9cd-4002-9fb2-6a0f447ec470 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention FP8-LM: Training FP8 Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 56d5338f-e883-4b21-a256-22388e7fa288 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Training and inference of large language models using 8-bit floating point
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bf83668c-6f1c-4e1e-b0be-bd887d7e2282 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 35f406cc-da44-44be-9695-590dda5c8203 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Qwen3 Technical Report
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f223dff7-69e1-434c-8aa5-d4a7d67c0139 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 22a56c48-f98b-43e8-84e5-c06205cf2fcc · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Methods of improving LLM training stability
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ab409d59-1954-493a-bdfc-a17037221042 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention LLaMA: Open and Efficient Foundation Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 0ee60707-4b4f-478a-ab91-f55567e38442 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Training LLMs with MXFP4
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7eee6e6f-7efe-42c7-a26d-c059d5f17b33 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Optimizing Large Language Model Training Using FP4 Quantization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 8f603f96-661d-4b81-a8c9-86a00ab49069 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Mitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari S
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7b776de5-23a8-4a73-bd64-03fed27455ee · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Efficient Streaming Language Models with Attention Sinks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation ab5518ee-fc94-4220-a46e-72af56d69016 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 94f1e29d-3a57-407c-8f9e-e46d13956016 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention A Spectral Condition for Feature Learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 2afeead0-f78f-4c56-b46a-2fd1adf104bb · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 7d863fef-6e0c-4126-91d7-be66321cb39a · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention gradient spikes
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a5f9a6f2-e73d-488c-b77f-1eed422f3313 · outbound
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Seg en à st 're ich s ho hem S oh ne / Un ser m Kaiser Ferdinand !
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d8e3f90e-889b-4d2a-94c3-39e22dfc17bc · inbound
Reference Traces for Auditing Invisible Weight Updates and Guiding Exact-Budget Protection Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab61e964-02b0-4497-874c-767d0fae2e8b · inbound
Reference Traces for Auditing Invisible Weight Updates and Guiding Exact-Budget Protection Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55e5c0eb-4757-49d1-9eb8-3d238c96c5ab · inbound
An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f11b71b5-ce65-43a5-9ff8-5b4734344cb7 · inbound
Automated Numerical Stability Analysis of Deep Learning Operators Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.