Pith. sign in

Paper Citation Record · LEDGER

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

As of 3 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 4 inbound Pith citation observations for arXiv:2510.04212.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.04212 v3

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T10:01:56.131253Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:54:47.068638Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact23
  • verified fuzzy4
  • unresolved0
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84ef2b0d-6d00-4360-8707-47ea35c43375 · outbound

This paper cites Scalify: scale propagation for efficient low-precision LLM training.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Scalify: scale propagation for efficient low-precision LLM training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.677451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:b433e2d57cd9f8b8c4a6144a7dc568e34ae598fa53408bc119b05eaa34398aa0

Observation c8ce8386-332b-4824-a827-d4101c1410d5 · outbound

This paper cites u-$\mu$P: The Unit-Scaled Maximal Update Parametrization.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention u-$\mu$P: The Unit-Scaled Maximal Update Parametrization

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.733884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:625fe9d70727de67d4266ab0eeea0be77a43ec01b576a902825198718055d959

Observation acad3c74-c1c2-4dc6-bcbf-cf9f912f814c · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T10:02:31.810103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:bce2a07199c686059b69d135fd3b389005437c61d0a76acdcbab59678147de59

Observation 2747e26a-663e-407e-9239-46432204b867 · outbound

This paper cites Scaling FP8 training to trillion-token LLMs.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Scaling FP8 training to trillion-token LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.710568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:79bddcd718a0eaded3f3abc00e3c7eccb933d27ca172280c6a5b6efa55351618

Observation a45bdfba-62a7-40cd-8e79-c5196ab742e5 · outbound

This paper cites Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T10:02:31.806824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:0bfa48c5403a62123914a6e00d0fe11b79b144fc7c624ea1a693a2c82a663cd8

Observation 0a266c88-bd65-4dac-bdbe-67731ecbcadc · outbound

This paper cites Is Flash Attention Stable?.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Is Flash Attention Stable?

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:02:31.719792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:59b857fe4e28d76b34859330afb6ee9531898036ce75780ed3d5bc492309ac17

Observation 51ef9cb5-1de1-4065-b39d-0102ff8bb07f · outbound

This paper cites Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-30T01:18:51.459998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:d4ab36e2cc8b0b34436c039b5a84306e727486038ff5e5d4ffb112532244a056

Observation cc065a55-4627-413b-9d94-158782c9e583 · outbound

This paper cites Query-Key Normalization for Transformers.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Query-Key Normalization for Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.673927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:02bf01b1b34b9568fc97ce204da12748914bf88ea41ae3ced02745fd017d6225

Observation 8f6b5e54-03fc-4290-a857-7034d2da5616 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Training Compute-Optimal Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.752653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:983a085c521f75f78cecae12cb6f859b45fbf26c70e4d0b06f3390afa4c96bd8

Observation 88aff18f-3fb0-4d42-b513-6972e031ff0f · outbound

This paper cites SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.706582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:001df523ffc1f238fb6510700b33dd57e8bec9e74802b8c58c8af833c5471e1b

Observation 52e00fa9-ec25-4b75-8a7a-7be7ccc302f3 · outbound

This paper cites A Study of BFLOAT16 for Deep Learning Training.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention A Study of BFLOAT16 for Deep Learning Training

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.738309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:ad38cbe3f9e1588fde96512d2eaf95efe893bbeb81d81b0a2f855d7c78e89e35

Observation 53f66d0e-6c7b-41f6-96be-f52c7003805a · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Kimi K2: Open Agentic Intelligence

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.756129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:70bb9aeb560181224f9e62750323a9d7b9a85f4ceeff05ec044d55dd5ca90742

Observation 0d9e66a6-2abb-4607-b5bf-4c19b90c21b2 · outbound

This paper cites DeepSeek-V3 Technical Report.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention DeepSeek-V3 Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.702310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:9db0bd87dde91dd92da373a3a1193465c786bf8293b5e33b8bd0264643207911

Observation 3e77396a-2006-4532-94df-1611f448be03 · outbound

This paper cites Mixed Precision Training With 8-bit Floating Point.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Mixed Precision Training With 8-bit Floating Point

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.684327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:6f698c727f1ed090cda4d363aaf9c1c168601c5ca3e5817b61fbf3e36dce4fb3

Observation 7b607e40-1368-4393-880c-4fb464af087f · outbound

This paper cites Mixed Precision Training.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Mixed Precision Training

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.654634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:2c6a01aecab70589ca17320df0068a830207a00f30ca757591ec1ce4703d61de

Observation 267a6714-9fda-4947-8fd6-a76fb131a427 · outbound

This paper cites FP8 Formats for Deep Learning.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention FP8 Formats for Deep Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.723095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:29deccc9e098351e8fc424b6d5ebd12d77230b03d380ad3d7b5dfb02e66b136c

Observation 89f76f7b-7d55-4d47-bbd3-2d7e6cf46b53 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention A Theory on Adam Instability in Large-Scale Machine Learning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:02:31.691756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:3810e45ce1d79ea89c67442c9873ab85ab9464ad996a1f517c38c8e0b0ee1509

Observation 1e1dbe79-4ae7-41dc-939e-c2a6900e0c92 · outbound

This paper cites nanoGPT Issue.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention nanoGPT Issue

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T10:02:31.803305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:c90e08222c8a112d5c2f7112f360c1b384ed260af5a04dc87b03d1e465751bdc

Observation 93c4ca0f-0f0f-492e-9f4b-9839f0659c4e · outbound

This paper cites 8-bit Numerical Formats for Deep Neural Networks.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention 8-bit Numerical Formats for Deep Neural Networks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.639481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:d1ee40305e79a2c227f8e708fae0860a7c76fd4eba9ad96e0f8d5ac1aca7bbd1

Observation 65c71bf6-d9cd-4002-9fb2-6a0f447ec470 · outbound

This paper cites FP8-LM: Training FP8 Large Language Models.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention FP8-LM: Training FP8 Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:02:31.727622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:81bca6fc54f51ca457bbc702d5596b7ceb4cfc8a3b10c69a176a44464329d784

Observation 56d5338f-e883-4b21-a256-22388e7fa288 · outbound

This paper cites Training and inference of large language models using 8-bit floating point.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Training and inference of large language models using 8-bit floating point

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.696207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:945a614dc02c9d80124e2fd5f0724c5fdebb5adae03d8d9a6caf82c2e21032b8

Observation bf83668c-6f1c-4e1e-b0be-bd887d7e2282 · outbound

This paper cites Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.643403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:a1250db78a1952451b745f6e1cabdadfee3f127a08f490ae9e7cf33450fd8973

Observation 35f406cc-da44-44be-9695-590dda5c8203 · outbound

This paper cites Qwen3 Technical Report.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Qwen3 Technical Report

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T10:02:31.647731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:20c3b564755d7a0efab3e9fbe1858c0b03222d1202317b739dbc789525feaa3e

Observation f223dff7-69e1-434c-8aa5-d4a7d67c0139 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.746740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:f573004e88cac9bc409a3d333500681e8fc3c52b9e663be099eddb899abb3302

Observation 22a56c48-f98b-43e8-84e5-c06205cf2fcc · outbound

This paper cites Methods of improving LLM training stability.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Methods of improving LLM training stability

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.742547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:871b716f995cb649a5f002fd43bccd4dc3f8a696a41e41d4326aff6139e13618

Observation ab409d59-1954-493a-bdfc-a17037221042 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention LLaMA: Open and Efficient Foundation Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:02:31.651207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:7281f6aeb25edbf9f6a99cff13973b13935e16657e8377b454e837e5c79bd95e

Observation 0ee60707-4b4f-478a-ab91-f55567e38442 · outbound

This paper cites Training LLMs with MXFP4.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Training LLMs with MXFP4

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.760017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:7b0548ee916c66d5035639eb9772b5c362740a2a83620fffadb18577262dac17

Observation 7eee6e6f-7efe-42c7-a26d-c059d5f17b33 · outbound

This paper cites Optimizing Large Language Model Training Using FP4 Quantization.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Optimizing Large Language Model Training Using FP4 Quantization

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:17.330307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:83781c0ae8abb3bc102480e56e7d00281670ea774e230c58ea6fa3818d1eea2e

Observation 8f603f96-661d-4b81-a8c9-86a00ab49069 · outbound

This paper cites Mitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari S.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Mitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari S

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T10:02:31.799841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:71f140b8e50e4fe124516554424c7e321903f2093a193c6c9068b3d260d5b38d

Observation 7b776de5-23a8-4a73-bd64-03fed27455ee · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Efficient Streaming Language Models with Attention Sinks

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T10:02:31.687763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:b2c8984e806d01fe17f732a3028cd86e3044f21064f271b821720c0ecfafd2c0

Observation ab5518ee-fc94-4220-a46e-72af56d69016 · outbound

This paper cites Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.664367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:f30e1f77f785bf63052f8d735a53da00d2dca889a97a5d7e478933c327d2733d

Observation 94f1e29d-3a57-407c-8f9e-e46d13956016 · outbound

This paper cites A Spectral Condition for Feature Learning.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention A Spectral Condition for Feature Learning

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:02:31.669691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:ddcb24b30998fdcebef7d505050dbd00c0037893c4b7c8865d6daa13192d0dc9

Observation 2afeead0-f78f-4c56-b46a-2fd1adf104bb · outbound

This paper cites Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.635398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:a24a0d36dee52f50e24c517e32c2bf3eb8553218f85b4a69d1f41190e0f0d8ec

Observation 7d863fef-6e0c-4126-91d7-be66321cb39a · outbound

This paper cites gradient spikes.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention gradient spikes

Reference 35

Resolution
malformed identifier
raw_fallback, observed 2026-05-18T10:02:31.796249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:4ab150e8f9eb5423ca71d6ea4311124dcc356dd088caa97c01a877746c3eb69f

Observation a5f9a6f2-e73d-488c-b77f-1eed422f3313 · outbound

This paper cites Seg en à st 're ich s ho hem S oh ne / Un ser m Kaiser Ferdinand !.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Seg en à st 're ich s ho hem S oh ne / Un ser m Kaiser Ferdinand !

Reference 36

Resolution
malformed identifier
raw_fallback, observed 2026-05-18T10:02:31.793196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:c80e06605518b0cc57f527874785d8d2257603b397062f6cd96685ffa8b09bd5

Pith citing papers

Observation d8e3f90e-889b-4d2a-94c3-39e22dfc17bc · inbound

Reference Traces for Auditing Invisible Weight Updates and Guiding Exact-Budget Protection cites this paper.

Reference Traces for Auditing Invisible Weight Updates and Guiding Exact-Budget Protection Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T15:32:26.691504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:32:26.691504Z digest=sha256:c854fc300eefe640ef0d8163a2ac8cfa5831479a6de14eb474ab92bcb0105412

Observation ab61e964-02b0-4497-874c-767d0fae2e8b · inbound

Reference Traces for Auditing Invisible Weight Updates and Guiding Exact-Budget Protection cites this paper.

Reference Traces for Auditing Invisible Weight Updates and Guiding Exact-Budget Protection Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T07:54:47.068638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:54:47.068638Z digest=sha256:375a404f953dc47cf1352653432b559941c379454bca18c3488f1b79dc7356e7

Observation 55e5c0eb-4757-49d1-9eb8-3d238c96c5ab · inbound

An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies cites this paper.

An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:35.779079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:35.779079Z digest=sha256:31e4bd0bd2fcbfc0b93b7ecdf7d6d0094473d8c9de0b859c897babffa47c0ea3

Observation f11b71b5-ce65-43a5-9ff8-5b4734344cb7 · inbound

Automated Numerical Stability Analysis of Deep Learning Operators cites this paper.

Automated Numerical Stability Analysis of Deep Learning Operators Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T02:20:39.957037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:20:39.957037Z digest=sha256:ff73d5f54b42e274466c042c5c4b4edb7ca41d254345331d893a8f8f18520bee