Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:07:52.017037Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 5 inbound Pith citation observations for arXiv:2604.11554.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:07:52.017037Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T21:32:02.473074Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T03:05:55.340981Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1ae3668c-a270-4107-8c10-7735919be773 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 714e39aa-db5f-4aef-aeb5-1c9337d642b9 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b274a53d-5f5f-4918-a1b5-a213f762e5d9 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Proximal Policy Optimization Algorithms
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation db06b6f2-ac58-48de-bc7c-201a510860e9 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 389b0b5c-4d69-4e40-ab41-6bacfd400e4b · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 69cea288-5a44-4e2d-9c81-121c0dbd02d6 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale HybridFlow: A Flexible and Efficient RLHF Framework
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b83aa4c6-ea2d-4b2d-ab97-d57879b95d3a · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a80d0625-6c83-4771-ae17-5c4bf08115af · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa820f69-54c1-4543-b167-7b144dec4266 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3a2c3f59-b823-4fc0-9825-6df5dccc7e83 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3b3b2430-9bd8-4b90-b944-7c5de00d4679 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale type": "function
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 591aac5d-eb66-468d-bf3e-096d54507a30 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale slime: An llm post-training framework for rl scaling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c5e76a10-61fa-4300-8191-63716353c914 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Ray: A Distributed Framework for Emerging AI Applications
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 323d3a8a-39db-4ade-95b4-8f65f6b6faac · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Deep reinforcement learning from human preferences
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b387c5ab-57a5-4f31-8afe-d48cc3ad668d · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Training language models to follow instructions with human feedback
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa881d94-6290-4ec3-84c3-a3051ebe5043 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8ec5f994-e0ca-49e6-95ef-2cd0bf0003bb · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e29febe5-7e34-4270-a0c3-f8dbc3898095 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale SGLang: Efficient Execution of Structured Language Model Programs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b2f940ef-fad4-45de-a12e-be7f400cce3d · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 609eb919-2099-427a-962b-6cb59eedc62b · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Inference-time scaling for generalist reward modeling
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 87eebf59-fac5-45cc-b15c-4399c9236ab6 · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71c61e6f-ee4f-4fe5-a518-e46df2326e7a · outbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f287ec78-98c1-4c7c-a4b5-d6bdf99027d9 · inbound
MinT: Managed Infrastructure for Training and Serving Millions of LLMs Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 884a07e3-0ce2-459d-b434-52d7660279fd · inbound
MinT: Managed Infrastructure for Training and Serving Millions of LLMs Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6e2c39ff-7e1f-479a-b9c1-8b983d4ab74c · inbound
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0e0cbcdd-cab6-4b7a-adf8-2a79aacbd96a · inbound
JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b849eeb-c909-49a0-82db-ab62884f9c63 · inbound
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.