Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T11:07:18.839899Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 2 inbound Pith citation observations for arXiv:2601.18150.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T11:07:18.839899Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-15T03:13:14.384567Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-15T03:14:52.571593Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e1812e0d-54ce-4e72-aed1-0004af5a4f7f · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7819ca9a-cdc5-45ef-8686-25b09cba5c31 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning TensorRT LLM
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5e2a1471-9d5d-4d12-8ef0-63f0d0af4507 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Efficient Memory Management for Large Language Model Serving with PagedAtten - tion
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a14b7fc8-1490-4e63-8d09-8bbc71786235 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Available: https://github.com/sgl-project/sglang
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04bf0a96-7197-468b-a166-31a5da24b390 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning When Speed Kills Stability: Demystifying RL Collapse from the Training-Inference Mismatch
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 84c9e339-d854-43ef-b0e5-388f622760f3 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning FlashRL: 8Bit Rollouts, Full Power RL
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29a1a3f3-4afa-45bd-87b2-12bcd3962041 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Your Efficient RL Framework Secretly Brings You Off-Policy RL Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2a4821ba-b753-4866-b0d0-e3bdcd52ba2d · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning DeepSeek-V3 Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e8f2a045-5ed5-4ea7-afdc-2ea26f0da161 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3f1567e-d1b4-4c92-802d-cfeda2296eae · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 11acfa0e-579d-4b2c-b2ba-1ade549b10e7 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning NeMo RL: A Scalable and Efficient Post-Training Library
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9072d075-1b32-41f7-b6ef-a6f3ce81af98 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning FP8 Formats for Deep Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b03e7fe6-fdf2-4310-93f7-039ae18c397e · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning FP8-LM: Training FP8 Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6a8f4821-a09c-4daa-a10a-b230ba7d5b2e · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 156a7ee1-7416-45ab-b851-7fef121056b8 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a890e8d4-eb8d-409d-8cd3-39bcc6009c17 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb54d84e-2a83-486e-b70a-fe5bda2cccfd · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning slime: An LLM post-training framework for RL Scaling
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 26e97cb3-bbe2-4ffc-aa2d-49d925b892c4 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ab328949-9340-4e9b-a113-4afae7ab48cf · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Small Leak Can Sink a Great Ship–Boost RL Training on MoE with IcePop!
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 90735c17-72d1-4819-ba3c-517e0749ae07 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4d0bf14a-325d-4424-b52d-b4508ac02842 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning No More Train-Inference Mismatch: Bitwise Consistent On-Policy Reinforcement Learning with vLLM and TorchTitan
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bfc898b0-4975-4f2c-9624-557041151ca3 · outbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning Defeating the training-inference mismatch via fp16
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5ce8129-5817-47e9-9c82-6ab934b73057 · inbound
FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5128f4a1-330b-493a-b192-b16bd07de1ce · inbound
AIS: Adaptive Importance Sampling for Quantized RL FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.