Pith. sign in

Paper Citation Record · LEDGER

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 2 inbound Pith citation observations for arXiv:2601.18150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.18150 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T11:07:18.839899Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T03:13:14.384567Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-15T03:14:52.571593Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact7
  • verified fuzzy14
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1812e0d-54ce-4e72-aed1-0004af5a4f7f · outbound

This paper cites Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:07:47.112397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:4376cf13ee5a183824d8308dbd774cab4316e7e6f516a97be3a0081812eb94d4

Observation 7819ca9a-cdc5-45ef-8686-25b09cba5c31 · outbound

This paper cites TensorRT LLM.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning TensorRT LLM

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.906622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:5e8f6048275493ab5482ef1b215599e42ba577816f7675b1c34789f6b64ea9ee

Observation 5e2a1471-9d5d-4d12-8ef0-63f0d0af4507 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAtten - tion.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Efficient Memory Management for Large Language Model Serving with PagedAtten - tion

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.901339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:ddac99d9510c9e6121acac5b7333d2c9babf2ac627672bda4413dccaf85da593

Observation a14b7fc8-1490-4e63-8d09-8bbc71786235 · outbound

This paper cites Available: https://github.com/sgl-project/sglang.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Available: https://github.com/sgl-project/sglang

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.928087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:a0514337b376817c8f577f785551d1a1b6dd47e8f3bb8e2827b09fc54b50ffb3

Observation 04bf0a96-7197-468b-a166-31a5da24b390 · outbound

This paper cites When Speed Kills Stability: Demystifying RL Collapse from the Training-Inference Mismatch.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning When Speed Kills Stability: Demystifying RL Collapse from the Training-Inference Mismatch

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.904224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:2409791c67f3e87cb578b81679678ffe416ea9ab8da6c2b1637850c879a0ccc3

Observation 84c9e339-d854-43ef-b0e5-388f622760f3 · outbound

This paper cites FlashRL: 8Bit Rollouts, Full Power RL.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning FlashRL: 8Bit Rollouts, Full Power RL

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.923197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:9561d653cb35a9bfdaafb498e0f0cbb5d2ad6e6bd32db299af1f5192d17cd730

Observation 29a1a3f3-4afa-45bd-87b2-12bcd3962041 · outbound

This paper cites Your Efficient RL Framework Secretly Brings You Off-Policy RL Training.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Your Efficient RL Framework Secretly Brings You Off-Policy RL Training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.896140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:a2065e75f460b13bfeb9051debd9a10e03dd9ef071ceb13290202a15fa99c850

Observation 2a4821ba-b753-4866-b0d0-e3bdcd52ba2d · outbound

This paper cites DeepSeek-V3 Technical Report.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning DeepSeek-V3 Technical Report

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.925889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:d3456ee6a7b8d41a0026a82df5207c0d9e0a0d06bd71259e59a3f5b487d7bf8d

Observation e8f2a045-5ed5-4ea7-afdc-2ea26f0da161 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:07:47.088515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:59aa5c5ed7df60c3d44706ade7c6068ceaac3e6c283d4b6e6e505d91ab573754

Observation b3f1567e-d1b4-4c92-802d-cfeda2296eae · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.918601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:0ef49e013cbee54c96b41cfc575efe586dc066e89f4d706b6dfb0fcf0ca7ceec

Observation 11acfa0e-579d-4b2c-b2ba-1ade549b10e7 · outbound

This paper cites NeMo RL: A Scalable and Efficient Post-Training Library.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning NeMo RL: A Scalable and Efficient Post-Training Library

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.920681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:a7275ab15747cea72695eb677524b7028bc9bdbc8f64853078c5c979238c1bcc

Observation 9072d075-1b32-41f7-b6ef-a6f3ce81af98 · outbound

This paper cites FP8 Formats for Deep Learning.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning FP8 Formats for Deep Learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.916167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:0fffa11069131c7a3b587437f8c96fa886f27821c0c4407c30e66a8fdcbb1707

Observation b03e7fe6-fdf2-4310-93f7-039ae18c397e · outbound

This paper cites FP8-LM: Training FP8 Large Language Models.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning FP8-LM: Training FP8 Large Language Models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.898697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:fd99a875a07e4ce629894ef253b21016f71388ea65235a1461c541fc028424fd

Observation 6a8f4821-a09c-4daa-a10a-b230ba7d5b2e · outbound

This paper cites DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:07:47.104402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:f1c10c84d8a70e136a2bf9c7ab268c921d553e5a304dc4a490654520d0226c92

Observation 156a7ee1-7416-45ab-b851-7fef121056b8 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:07:47.108354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:9c6ed715bee02a45eeca595971a6901128a6e9f6884502d5438a6e029c81dffb

Observation a890e8d4-eb8d-409d-8cd3-39bcc6009c17 · outbound

This paper cites Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:07:47.084927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:c4f0d97bdc5e2fbe62205ef9301805037d7df97bcbf546467a212e960ff7f302

Observation cb54d84e-2a83-486e-b70a-fe5bda2cccfd · outbound

This paper cites slime: An LLM post-training framework for RL Scaling.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning slime: An LLM post-training framework for RL Scaling

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.911466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:f6488a14d2e7c3a730aeaa788ccbb0d04e90cebd3e78e22fd79e01c5934a78de

Observation 26e97cb3-bbe2-4ffc-aa2d-49d925b892c4 · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:07:47.092261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:07616edf025fc80150ff93fc0bcda8dd2c5e3f94a0e2f8845ab8f8d9560d2f2b

Observation ab328949-9340-4e9b-a113-4afae7ab48cf · outbound

This paper cites Small Leak Can Sink a Great Ship–Boost RL Training on MoE with IcePop!.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Small Leak Can Sink a Great Ship–Boost RL Training on MoE with IcePop!

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.914046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:84c9a5b57bc480011184bb390759bfde5807468ffbcb40bc832472dd44f49046

Observation 90735c17-72d1-4819-ba3c-517e0749ae07 · outbound

This paper cites Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:07:47.100142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:f0ee4a553d3e85d7c881f1166e15448f9cac612c1f57595effff7deed8a41568

Observation 4d0bf14a-325d-4424-b52d-b4508ac02842 · outbound

This paper cites No More Train-Inference Mismatch: Bitwise Consistent On-Policy Reinforcement Learning with vLLM and TorchTitan.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning No More Train-Inference Mismatch: Bitwise Consistent On-Policy Reinforcement Learning with vLLM and TorchTitan

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T11:10:52.908965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:6c50a7c9a4eb09d35b1ecacf522496c6b66aa73de08054c73133f008034b56ce

Observation bfc898b0-4975-4f2c-9624-557041151ca3 · outbound

This paper cites Defeating the training-inference mismatch via fp16.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Defeating the training-inference mismatch via fp16

Reference 22

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T11:07:47.096305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:d108e96b53ab68b824ea395681985e00e1fbe996ec51981c3786a0e3bed63245

Pith citing papers

Observation c5ce8129-5817-47e9-9c82-6ab934b73057 · inbound

FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling cites this paper.

FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:21:08.293476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:10:06.994557Z digest=sha256:9bf6044c7bc7bd3f1d592307e4c1a0806adb7c0a35d4ae7856467871b628aa1b

Observation 5128f4a1-330b-493a-b192-b16bd07de1ce · inbound

AIS: Adaptive Importance Sampling for Quantized RL cites this paper.

AIS: Adaptive Importance Sampling for Quantized RL FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:14:52.575296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T03:13:14.384567Z digest=sha256:d9481a4436d533d9c844a6822a82a46054aa0a34896a1a2904757c0297d7852d