Pith. sign in

Paper Citation Record · LEDGER

Scaling FP8 training to trillion-token LLMs

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2409.12517.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.12517 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:30:56.446105Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T07:27:44.431731Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7b2e7694-2335-4d9d-b5b8-edca112359ab · inbound

AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference cites this paper.

AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference Scaling FP8 training to trillion-token LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T20:16:46.594244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:16:46.594244Z digest=sha256:5ed9e08075a58e74d80e210f505fc0088199088b03896698bd6a2b066c25f502

Observation c4b9be6e-0317-4925-8dc5-709c656c726e · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Scaling FP8 training to trillion-token LLMs

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.040130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:a8299312f1f04023987dca5ca3879753f20b186721e3375cd3758c1e0a588a22

Observation f804b551-985d-46b4-9dc7-9f50336ff2f0 · inbound

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture cites this paper.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Scaling FP8 training to trillion-token LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.929720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.929720Z digest=sha256:298a8a6896008609e3a73ea28428395d793dc2c8310ca042ff55d89524d96faf

Observation 67aba63f-2798-4df1-9f36-1e7f316667f5 · inbound

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache cites this paper.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Scaling FP8 training to trillion-token LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.922252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.922252Z digest=sha256:58593eeec527730f8c378c1107826af57f1476583983e4e699ce8613c2fe5bd1

Observation f56f30da-9d65-4ae5-8287-53efb66152f9 · inbound

Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities cites this paper.

Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities Scaling FP8 training to trillion-token LLMs

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-16T04:30:56.446105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:30:56.446105Z digest=sha256:19671f4b8d4bd03fedb842e424f39953bd3742fbf44c50846527b13fa91ef10e

Observation b4e070fd-686e-400f-b9d1-a0231fd220cb · inbound

Gaussian Weight Sampling for Scalable, Efficient and Stable Pseudo-Quantization Training cites this paper.

Gaussian Weight Sampling for Scalable, Efficient and Stable Pseudo-Quantization Training Scaling FP8 training to trillion-token LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:32.759252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:32.759252Z digest=sha256:5cc3b3775b633707d3efb0a00ba8479dbc9bb4a489f3c7bd8aa25807335015e3

Observation 4cac8f33-ef05-41ca-aeeb-e6fc175a5a72 · inbound

Scaling Law for Quantization-Aware Training cites this paper.

Scaling Law for Quantization-Aware Training Scaling FP8 training to trillion-token LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:06.104682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:06.104682Z digest=sha256:9f23e986be20c62a90a0a2b9f2cd9e886358492e4694cdce2e59e02329ffaf68

Observation f59b69c3-0296-407c-b2e7-12e45dd8d740 · inbound

FP4 All the Way: Fully Quantized Training of LLMs cites this paper.

FP4 All the Way: Fully Quantized Training of LLMs Scaling FP8 training to trillion-token LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:38.069250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:38.069250Z digest=sha256:f8ff123b7912e4407ebc3d42b6090f9e5f1ec498d348fad4aba01132ed528975

Observation 6873f7de-f463-4129-8d2a-ad77c575243b · inbound

Recipes for Pre-training LLMs with MXFP8 cites this paper.

Recipes for Pre-training LLMs with MXFP8 Scaling FP8 training to trillion-token LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.282884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.282884Z digest=sha256:74c78d727521a239445d196954fb27549a1b9d1bb0f30864ff83c30961dae7a1

Observation 87602fef-a9ad-4628-a28a-e57900c9ba43 · inbound

Characterization and Mitigation of Training Instabilities in Microscaling Formats cites this paper.

Characterization and Mitigation of Training Instabilities in Microscaling Formats Scaling FP8 training to trillion-token LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:47:52.213010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:47:52.213010Z digest=sha256:f39fce580eb7e5b72d4fe3c5dbd3d14e94eab4b349296d3a7bd9fa922e11e823

Observation d5b58414-1fa1-4c1b-b30d-a0aa36c878c4 · inbound

Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources cites this paper.

Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources Scaling FP8 training to trillion-token LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:58:09.687331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:58:09.687331Z digest=sha256:59bbee5e357f6b69335551c45ad675003741541de8b382de6b116139bb1df6f1

Observation 0ef52389-2791-4e33-9958-2e8aac0f444e · inbound

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models cites this paper.

DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models Scaling FP8 training to trillion-token LLMs

Reference 128

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:40.337409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:40.337409Z digest=sha256:616577752aa092b0a3075a400b013df8adeca9efd415dbd0a6388db4f72b777e

Observation a17c0e55-6fb2-471e-9476-971700957f1f · inbound

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models cites this paper.

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models Scaling FP8 training to trillion-token LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T14:51:29.281759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:51:29.281759Z digest=sha256:cc1fba7d4f387fd841fa28cfc62bcd0cff0e184d963e2f3e4a5c37a4c5c7ca67

Observation 2747e26a-663e-407e-9239-46432204b867 · inbound

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention cites this paper.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Scaling FP8 training to trillion-token LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.710568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:4bc7d584b0f7c0feb16e40729bc24d3b9ee15e731c66218a7b3a5c098b70ed26

Observation 16d8a70c-f9d1-4dae-b920-fb03d4795f40 · inbound

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs cites this paper.

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs Scaling FP8 training to trillion-token LLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:15.382428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T02:10:43.006238Z digest=sha256:7bbbd28aa343b8e3dc79f3b9248bfc633f3b222b2c21d40a19d452956a3ed51f

Observation ceea8870-8af1-47ab-941f-aaac89e01fb9 · inbound

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs cites this paper.

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs Scaling FP8 training to trillion-token LLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.637821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T22:57:31.085889Z digest=sha256:841a88b433294a9f34ebdd81461858d14807cae2eaba94f8ed78c381d9ec2695

Observation 8732e7ae-e736-45f3-b579-645a99dd585c · inbound

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale cites this paper.

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale Scaling FP8 training to trillion-token LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:28.075687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:33:41.411292Z digest=sha256:aab9c97f5e59d964c3668321f0e846e86b583b3f91348e1d39599bfd657a82af

Observation aee8c9c7-516b-4276-869f-b4e7a917cc26 · inbound

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale cites this paper.

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale Scaling FP8 training to trillion-token LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:59:46.257916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T04:55:01.973832Z digest=sha256:c083621fdd134b0872f3ab8f9905b3221fc1077c15851fd004c0dc86006576e0

Observation ec427814-b049-427b-889c-79e1561f93b9 · inbound

Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in Recommendation cites this paper.

Expand More, Shrink Less: Shaping Effective-Rank Dynamics for Dense Scaling in Recommendation Scaling FP8 training to trillion-token LLMs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:30:22.650040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T05:29:49.474591Z digest=sha256:0f90d80f76ee5e567aa9f9ae7bf18cae58c92a498f768ad654b6e1b285218ab2

Observation 04e1bcac-80e3-4b73-b29a-26833937f021 · inbound

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design cites this paper.

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design Scaling FP8 training to trillion-token LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:27:44.433192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T12:06:46.806138Z digest=sha256:34784207172b312183ed525f7ff2e9cb4868e098403127e34cc1f54edf4e82c2

Observation 7ae5b301-63ae-49e1-8d4e-0c59f242f88d · inbound

Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention cites this paper.

Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention Scaling FP8 training to trillion-token LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T19:17:59.044982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:17:59.044982Z digest=sha256:3b6d6f4f045ada8ae177029955759efe85b15c56bf5e03825a240fef52ed2f32

Observation fc699b5a-1423-467d-977a-07cd651a73b9 · inbound

One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse cites this paper.

One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse Scaling FP8 training to trillion-token LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T15:20:15.926212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:20:15.926212Z digest=sha256:fb43ffca3e6856b0c3082f42610816823bbc23bec0eb5300ebe092754b71239b

Observation e4741f26-8a1d-4a78-a568-913c2c8b3ca9 · inbound

K-EXAONE 2.0 Technical Report cites this paper.

K-EXAONE 2.0 Technical Report Scaling FP8 training to trillion-token LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:43:07.425705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:43:07.425705Z digest=sha256:39d30ece6f130f9297011e17878b035f6835361619272947cbc0b4352dd9bdac

Observation 2a71be8d-0421-4ef3-897e-dd6dcc881746 · inbound

Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU cites this paper.

Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU Scaling FP8 training to trillion-token LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T10:23:35.884910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:23:35.884910Z digest=sha256:bafa5f652ad0c3fd200df0c0c2c03477bfd128a51120deef07235fef1b888af6