Pith. sign in

Paper Citation Record · LEDGER

Recipes for Pre-training LLMs with MXFP8

As of 22 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 11 inbound Pith citation observations for arXiv:2506.08027.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08027 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:47.881130Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:26:51.610461Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:26:13.675967Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e77c830a-1820-49fd-ae46-2e13d579e519 · outbound

This paper cites Ocp microscaling (mx) specification.

Recipes for Pre-training LLMs with MXFP8 Ocp microscaling (mx) specification

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.996595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:12:45.489420Z digest=sha256:28fa94538f6ce7ca583f3236cd7928af88f05995575b2900e557de151395ce10

Observation d2778806-4b3d-4b72-8898-02fb0cef4683 · outbound

This paper cites URL https://resources.nvidia.com/ en-us-blackwell-architecture.

Recipes for Pre-training LLMs with MXFP8 URL https://resources.nvidia.com/ en-us-blackwell-architecture

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.848468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:12:45.576876Z digest=sha256:061fe036632d1c38571b01d8772ddfcb301c4d6a1a5d1c3215009e7438bf6dfd

Observation ea954705-fd04-42aa-92d2-feba37489914 · outbound

This paper cites Microscaling Data Formats for Deep Learning.

Recipes for Pre-training LLMs with MXFP8 Microscaling Data Formats for Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:45.790981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:45.790981Z digest=sha256:2a5b8a80004fc48df49d4130199ed0fc1fdacc7a05c9d596b0aa80d1523b38d0

Observation c91acf53-9320-43cb-890b-4de84268977c · outbound

This paper cites With Shared Microexponents, A Little Shifting Goes a Long Way.

Recipes for Pre-training LLMs with MXFP8 With Shared Microexponents, A Little Shifting Goes a Long Way

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.032229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.032229Z digest=sha256:1c5b81379bc57a72613e5cfd3e2de9e427d3cd51fe81f7ec2fb57c084de570db

Observation 978acedf-b124-4ddd-9747-37dc484cf5b0 · outbound

This paper cites VS-Quant: Per-vector Scaled Quantization for Accurate Low-Precision Neural Network Inference.

Recipes for Pre-training LLMs with MXFP8 VS-Quant: Per-vector Scaled Quantization for Accurate Low-Precision Neural Network Inference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.101980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.101980Z digest=sha256:0829944635203b94afc9d96f0fc55904bef4d4f47ce7403a8bb886e8b854b364

Observation 5004dd90-74fe-4df4-9bd4-241d6565ebf4 · outbound

This paper cites IEEE Std 754-2008 , pages 1–70, 2008.

Recipes for Pre-training LLMs with MXFP8 IEEE Std 754-2008 , pages 1–70, 2008

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T12:12:48.576575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:12:46.111933Z digest=sha256:4ea9081553905806d1247139ab04f9eb76d3c251816036f37f64685ed92443c3

Observation 97beda48-3844-43fc-b362-5c5cab22791b · outbound

This paper cites FP8 Formats for Deep Learning.

Recipes for Pre-training LLMs with MXFP8 FP8 Formats for Deep Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.118074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.118074Z digest=sha256:5e357ffa581545bd24c2e6718716ae74f4271b1fb115ba5c535d0251b276ed95

Observation f0aed45f-126b-44e2-9d42-acb6f5b6ea54 · outbound

This paper cites Nemotron-4 15B Technical Report.

Recipes for Pre-training LLMs with MXFP8 Nemotron-4 15B Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.131912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.131912Z digest=sha256:d40eb106a6ab9304bf7ef07e10ea0a93b903b8847d08a5e3105a76642bf8f691

Observation b5193df3-7e9b-400a-9b83-9666c5fe7341 · outbound

This paper cites Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models.

Recipes for Pre-training LLMs with MXFP8 Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.195871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.195871Z digest=sha256:e8beae7dd9ae02b331d425688e71630744729c94eb554858e878b2186b3ac794

Observation c911db37-3acf-4e46-ad04-4fe11bdbcc4d · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Recipes for Pre-training LLMs with MXFP8 Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.269897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.269897Z digest=sha256:baf57e180fb2a428a82bb8d8cc22dab44f807e2df866a680da61a6bc34c8b14f

Observation 13320e44-8583-430e-80bf-3fb3072fbb25 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Recipes for Pre-training LLMs with MXFP8 Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.342345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.342345Z digest=sha256:f9a06e2b32d4a8b47c8c8cc0b159b4b76eb3a7ab303a7b46732b0469bede5845

Observation d51b2a97-8a5a-4e00-acf8-e89e75597ef9 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Recipes for Pre-training LLMs with MXFP8 Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.423646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.423646Z digest=sha256:89225577bfa46d0fb8408d8d23b08a3a7142606def70aeda5650b7b9eeea4ff6

Observation aa12abb8-26b2-4213-9838-7d379e44d8e6 · outbound

This paper cites RACE: Large-scale ReAding Comprehension Dataset From Examinations.

Recipes for Pre-training LLMs with MXFP8 RACE: Large-scale ReAding Comprehension Dataset From Examinations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.493873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.493873Z digest=sha256:70b803e67a1d3466c93d1d31ea1be6ced7c3b4fc6753d5f2e01dbbce1b2d7db4

Observation eb1c1a33-5e07-4ef4-b502-57f205bd6a5f · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

Recipes for Pre-training LLMs with MXFP8 PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.582966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.582966Z digest=sha256:9781661fcd36bd1dccd8b430dab9b09d269f2a8851c6a15dada31e2a6a108c45

Observation 576341e0-8c74-421d-84d1-797c1d9cb0b9 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale, 2019.

Recipes for Pre-training LLMs with MXFP8 Winogrande: An adversarial winograd schema challenge at scale, 2019

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.617897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.617897Z digest=sha256:aaa38e9de1dfb1c68bf97629703e3161b726a5d9b42f9c93f98487c90c478efc

Observation ac0d46a1-cc6d-418c-a915-b00240ae8c9c · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Recipes for Pre-training LLMs with MXFP8 HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.646932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.646932Z digest=sha256:1bbeff373123e10858cd0064e6bc1f2bed460f6354e7ff760e41373d4f43f4c6

Observation 8158c597-b84e-42aa-b16f-721c584852f3 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Recipes for Pre-training LLMs with MXFP8 Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.664591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.664591Z digest=sha256:e8bd1e0bf9ec81c6f033c66ac3c165b7171276a483a5a60f9ecfaa7ab3c3fa27

Observation 47481e3d-4ead-439e-b2a9-1a932ff3223a · outbound

This paper cites Socialiqa: Com- monsense reasoning about social interactions, 2019.

Recipes for Pre-training LLMs with MXFP8 Socialiqa: Com- monsense reasoning about social interactions, 2019

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.697169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:12:46.684428Z digest=sha256:788b5dcb927510ede7b249e4dc0bed4700c3080cf8b8087aee75bb98a3a440c8

Observation 27417057-aa1c-4011-b874-ef31b5cef6b2 · outbound

This paper cites CommonsenseQA: A question answering challenge targeting commonsense knowledge.

Recipes for Pre-training LLMs with MXFP8 CommonsenseQA: A question answering challenge targeting commonsense knowledge

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.703553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.703553Z digest=sha256:2b4630f66eaee08c4539d96beddf482428282f03c8ee45840ed79e3119ba5624

Observation 4324b597-0cae-48ec-97f6-5ad225c24c3c · outbound

This paper cites The Llama 3 Herd of Models.

Recipes for Pre-training LLMs with MXFP8 The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.730511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.730511Z digest=sha256:947946f09780243bb6d28f3c2f275b6e03f3f16146f5b9cc35b40ce731e6beda

Observation 8c86409a-67f0-46ea-a62c-353d2b508dcc · outbound

This paper cites 8-bit numerical formats for deep neural networks, 2022.

Recipes for Pre-training LLMs with MXFP8 8-bit numerical formats for deep neural networks, 2022

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.576743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:12:46.736164Z digest=sha256:eb75b73483c3bc91e53b93c633a6a7433eacbbff75d0b527a9f6d17b1e5bb89d

Observation c7d49f70-6b50-4e37-9bf3-15da04787a62 · outbound

This paper cites Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

Recipes for Pre-training LLMs with MXFP8 Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.754388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.754388Z digest=sha256:4ab614fe9600cb618999acd9eea2f74c723bcbb741eced88a4a2b919b55e8040

Observation 8abe41fd-7c90-4bf4-af27-f82a37759078 · outbound

This paper cites Ocp 8-bit floating point specification (ofp8).

Recipes for Pre-training LLMs with MXFP8 Ocp 8-bit floating point specification (ofp8)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.395984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:12:46.890232Z digest=sha256:198ff589be4aaf434d4b25e9abe1bd7fa830e8b65c2d1c54e16f5e7c2d6f0b95

Observation 073bdb9f-4131-4710-9eb0-861788e9ff79 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models,.

Recipes for Pre-training LLMs with MXFP8 Smoothquant: Accurate and efficient post-training quantization for large language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.253153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:12:46.946348Z digest=sha256:4f990d731036316e7e08b396393d9a29a5d55a072051483ccff3f73de0b97949

Observation 97cacf03-4797-4ebe-bce0-b18e0910712d · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

Recipes for Pre-training LLMs with MXFP8 QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.074402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.074402Z digest=sha256:49ed3037d89c18bf4484ce1cf4be1006ef7c517221bfc09fd7f3bdb9a0091493

Observation 28f760c5-22f8-4eb8-bbd2-b581dedb1e4c · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Recipes for Pre-training LLMs with MXFP8 GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.159556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.159556Z digest=sha256:3b2584b2cf3cb60f9fac5401e420cd105f9b1a49444c6671abb270af10023db6

Observation 5a0ae3d2-6aab-4a30-8b62-7852f67537a9 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

Recipes for Pre-training LLMs with MXFP8 AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.221979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.221979Z digest=sha256:ec8ed1291c9f76368f0fab469ea1fb68dda3645483b7057d7089b144812ebd05

Observation 6873f7de-f463-4129-8d2a-ad77c575243b · outbound

This paper cites Scaling FP8 training to trillion-token LLMs.

Recipes for Pre-training LLMs with MXFP8 Scaling FP8 training to trillion-token LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.282884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.282884Z digest=sha256:74c78d727521a239445d196954fb27549a1b9d1bb0f30864ff83c30961dae7a1

Observation b808aec9-2b83-4fe6-b3cd-43066f20400f · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innova- tion.

Recipes for Pre-training LLMs with MXFP8 The llama 4 herd: The beginning of a new era of natively multimodal ai innova- tion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:49.124436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:12:47.338504Z digest=sha256:3aee734e408070d7e5d39f2a03913bfca023b32df5951f4112b5cb214ad17748

Observation 2ad6dd0c-2c4d-450a-8a9c-19ee36aabc99 · outbound

This paper cites Training LLMs with MXFP4.

Recipes for Pre-training LLMs with MXFP8 Training LLMs with MXFP4

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.436064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.436064Z digest=sha256:bf458dccbb0e1286079dba6b48f6ae6e9a35818afacbcd12a948e6b891c538ad

Observation faf32e08-be52-45c7-a1cb-f11c842a2909 · outbound

This paper cites Optimizing Large Language Model Training Using FP4 Quantization.

Recipes for Pre-training LLMs with MXFP8 Optimizing Large Language Model Training Using FP4 Quantization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.457958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.457958Z digest=sha256:07a017608d597755a52d04b330f208b9468075e3d0b1ebde819f35876372ee9f

Observation 40625447-23d6-4b9f-b3ca-81594a08b54c · outbound

This paper cites Transformer engine.

Recipes for Pre-training LLMs with MXFP8 Transformer engine

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:48.996727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:12:47.463802Z digest=sha256:1f56a8786a015fdb77a36322374de91ca8becd9ce0ef70beb109d0685ea7f3e7

Observation 0f338323-0bc3-41c1-b8b8-2a77fa8253e5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Recipes for Pre-training LLMs with MXFP8 DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.507189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.507189Z digest=sha256:2b033bf1de32e3d5de93a1a62f6a25148c054d651536f937895736ddd81cf7e5

Observation cefd3c8d-4ce7-45c9-99ff-acb5919ea260 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

Recipes for Pre-training LLMs with MXFP8 Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.569722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.569722Z digest=sha256:74c82c02253aae47423c71e731235e0c4fb6b172693fd8be414b964b88f47d62

Observation caef3419-ee7f-44e3-ac4c-c9d34222be05 · outbound

This paper cites Nemotron-4 340B Technical Report.

Recipes for Pre-training LLMs with MXFP8 Nemotron-4 340B Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.627819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.627819Z digest=sha256:885c5cb297018831b1cfe7dea216fa01b3656ac51bf4b9d727a7bd8a518dfe25

Observation be684594-0e13-4231-98fc-0ab4833e81eb · outbound

This paper cites an unresolved cited work.

Recipes for Pre-training LLMs with MXFP8 Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:12:48.896269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:12:47.734191Z digest=sha256:d427f3293a985da4f768f88228cbba3397aa8d951d09a8b11850a032d4a18a71

Observation 802c0213-cdba-4429-bcf3-8cbef9468f7f · outbound

This paper cites an unresolved cited work.

Recipes for Pre-training LLMs with MXFP8 Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:12:48.785664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:12:47.807734Z digest=sha256:344e71e030f2b5fb2cb0d061ba5eacc875b07b8f5c0940a1f73ed0a018c9dadc

Observation 91e9d15e-83dd-480e-89f7-a4fca49d9d1b · outbound

This paper cites By construction amax/destmax never exceeds 2127 (which is the largest value representable in UE8M0) with FP8, FP6 or FP4 formats.

Recipes for Pre-training LLMs with MXFP8 By construction amax/destmax never exceeds 2127 (which is the largest value representable in UE8M0) with FP8, FP6 or FP4 formats

Reference 40

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:12:48.698298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:12:47.881130Z digest=sha256:e5de49259f88e8ede0e83119f8e0f5d6967db914538fd12a12b43b4c2dbd1666

Observation ff2704d2-c5dc-4343-a25d-6be97f90f558 · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

Recipes for Pre-training LLMs with MXFP8 SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:47.016932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:47.016932Z digest=sha256:7d5a38afdf367602360471c5615175c086697bc6a7c8dbf7882bb28e95095aba

Observation 80d5696e-8679-4c52-a166-52d569aefcda · outbound

This paper cites DeepSeek-V3 Technical Report.

Recipes for Pre-training LLMs with MXFP8 DeepSeek-V3 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:46.816128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:46.816128Z digest=sha256:f8bb7f656f3dd38eeae9dc99f035d95cda9348ce60774a936e87f1dbfc9f9fd9

Pith citing papers

Observation d238393d-c84c-4717-8e97-4a6fc1071923 · inbound

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models cites this paper.

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models Recipes for Pre-training LLMs with MXFP8

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T14:51:29.915567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:51:29.915567Z digest=sha256:1f3d7677b929579bfba7c847190e1b8501732b475b9b2ff526a214972f7e43a6

Observation 886ff8dc-5d2d-408b-bfc2-f9ebd0f55fcf · inbound

Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling cites this paper.

Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling Recipes for Pre-training LLMs with MXFP8

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:23:52.769839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T02:23:01.845123Z digest=sha256:0cec4e4bf1434bb4870669484296a45228247d54b9a856b1d9456ba32e8d7a55

Observation 5bc29d3f-438b-4a33-b922-6dfd36cb0907 · inbound

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models cites this paper.

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models Recipes for Pre-training LLMs with MXFP8

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:15:22.229993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T12:10:44.802059Z digest=sha256:77adee637aaa353bd0124ceeaee4ed46bf502bb6c349f9abc64c9141f4d91cf1

Observation a50e61fd-5ca4-4c24-8e7c-b3e0301cdfbb · inbound

OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning cites this paper.

OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning Recipes for Pre-training LLMs with MXFP8

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:13:27.553824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T13:06:02.201607Z digest=sha256:3337b952b4af7975f5ec854b4f2d351e6b3426ff1eea0cd442834adb1f628070

Observation 81579add-036b-46f1-b340-2027e20a9d29 · inbound

Stochastic Rounding Increases Small Singular Values cites this paper.

Stochastic Rounding Increases Small Singular Values Recipes for Pre-training LLMs with MXFP8

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.679444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T21:02:34.955394Z digest=sha256:70c9ef305f70764e54e29fe3a6958321a52bb48bcd7e10c151bedbbe595fa3b0

Observation cc3f4a9d-f5f0-4d3c-9c6d-a1906fda7d6a · inbound

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales cites this paper.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Recipes for Pre-training LLMs with MXFP8

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:18.277482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:18.277482Z digest=sha256:13a3dda6185d2a566b1a3fc5c852eca4a8b1b31ff4c684aa5c799f30b0beca3b

Observation 35ce8b5d-d972-4296-9c31-5c3f40a2dbc8 · inbound

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning cites this paper.

ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning Recipes for Pre-training LLMs with MXFP8

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T23:13:13.763909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:13:13.763909Z digest=sha256:058e33f376873f93c36bcf6f68ba76e46010914fed4f4d03d3582467ee3decb2

Observation 6bc22111-1ccd-450d-9be1-e93ba5b19779 · inbound

Stable FP4 Training via Transposition-Invariant Block Quantization cites this paper.

Stable FP4 Training via Transposition-Invariant Block Quantization Recipes for Pre-training LLMs with MXFP8

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T05:04:00.365982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:04:00.365982Z digest=sha256:2e30dbde91297e2632d06805928b00bef1eec00c5310167cdbb06a76efb850a7

Observation 4524645d-eafa-4379-bd0b-c0b96d1facf0 · inbound

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference cites this paper.

Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference Recipes for Pre-training LLMs with MXFP8

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:30.832174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:54:30.832174Z digest=sha256:57ede5bb63a93463426c66cf146f1c20db64606ab39fbd89507bff64b71e4f62

Observation b5f5423a-0c9f-4548-9faa-4ef49be4751d · inbound

Motif 3: Technical Report cites this paper.

Motif 3: Technical Report Recipes for Pre-training LLMs with MXFP8

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:12:10.029200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:12:10.029200Z digest=sha256:f56236a4c2807b6392bf1299a5698ff21551d120af0ac993011d7571ba18f27b

Observation 74a6596e-74ea-4225-8938-7d94d2bdf475 · inbound

SoftWater: Class-Aware Rate Allocation for Softmax Quantization cites this paper.

SoftWater: Class-Aware Rate Allocation for Softmax Quantization Recipes for Pre-training LLMs with MXFP8

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:26:51.610461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:26:51.610461Z digest=sha256:3db74e4d2f307e22401ed8971f9c6ff04bbcf436fe5c624682ea510a38ec069a