Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 73 inbound Pith citation observations for arXiv:2505.24298.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:17:54.142833Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
50 of 50 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation f1b5f65d-5401-46eb-a59d-e09f50458e72 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Dota 2 with Large Scale Deep Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 04c98360-1a8c-4d4b-9851-275de73d2520 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Evaluating Large Language Models Trained on Code
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c68c561a-5cbb-46c8-b5b1-78b1f6320f9f · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f6726aac-2bc3-4aa1-867d-592e16976e21 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 907e8104-ed56-4057-aa42-6ef156b66393 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Espeholt, H
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 09668f0b-c48b-454b-8315-75fc0a8cb846 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Espeholt, R
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 65c5e111-101c-433e-b470-277b310c206f · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Hendrycks, C
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 50b91d8b-0a5f-45b3-913e-904d135b37b4 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Hilton, K
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ddd1dcba-6b8a-4ca0-9e75-122efaeeb268 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1a8ba6c0-7b7e-4f1c-a6ae-8e5caf4218e9 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d379ce49-8768-4e9f-824c-cb32e3bfe0d9 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 842f1a27-327f-4e04-9398-54f70b48c754 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 006a0879-9795-4d9f-a0a4-1c4ed76a61c7 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Kapturowski, G
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2bcfd707-7ecc-45ae-9639-e65d83ae8a57 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3afe3235-b742-4fdf-9ba2-163c2ef1226f · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a83167a8-ffea-4f7c-80bb-72c2b906a51b · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cd3d6ccb-9aac-4db8-8979-46e7f573bb14 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Liang, R
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 00c3bd1a-c9f2-4983-9911-f4b4947d08cd · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4c66119e-5d22-4e30-b656-2bc41e0ceaef · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5b2a3e38-cc70-4796-8b56-a8489d7d4d67 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc842cf3-3971-4022-b753-e1754b602810 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5187a572-3c6a-49dc-b575-ee6e3fe22140 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2fe33030-686b-40dd-a9d0-a8732103d5fa · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Specinfer: Accelerating large language model serving with tree-based speculative inference and verification
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9ac1dd7b-70c8-4f58-a4b3-041e34200782 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning URL https://openai.com/index/ learning-to-reason-with-llms/
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 22c0fc43-2c26-490e-8a5c-a489cb6e1b09 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning URL https://openai.com/index/ introducing-o3-and-o4-mini/
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 17be9c91-965d-422f-85f3-be94cfbb6859 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Ouyang, J
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 746cce49-6e19-44ed-a3ce-a0c2003f75ea · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec081b8d-129c-43c1-8ab9-c91d41defadf · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Paszke, S
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b34c7e4e-1a27-448a-892a-0bbd00897967 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Wiley Series in Probability and Statistics, Wiley (1994)
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4cfce521-ebf0-4a16-bfe9-efb03efa7f1c · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Generalized Slow Roll for Tensors
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 75a53cec-ed67-4faf-a111-953e497e4cbf · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Schulman, P
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1df6b9ad-1039-4be0-b252-ddcca0f8182d · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Proximal Policy Optimization Algorithms
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 80570ad7-d6ed-406a-8833-791c6d4c419a · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 55dbd829-c259-4871-b28d-53e033abd762 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Hybridflow: A flexible and efficient rlhf framework
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9b490396-40b7-48ba-8703-c5eb8ac69ba5 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf02d8bd-d30e-42e4-80c0-b9cd86e0f88e · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a9bb1a91-530b-4c64-91dc-330177927b5a · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac46e50e-a781-4a67-b569-1d9158477b2e · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Vaswani, N
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bb8f25bb-19ec-4261-b3a5-29ee5836815c · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Liar” ends the game, then both players reveal their dice. If the last bid is not satisfied, then the player who called “Liar
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 776aaf06-5ed0-4e76-b9fe-01a51c09984b · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b991c518-178b-426d-a6ca-76576e505494 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3628c109-2250-48f9-95a8-63f07031ee6a · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ed6d1c0e-4631-41c4-aa10-076c7defaa9f · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Job Scheduling Strategies for Parallel Processing pp 44–60
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 79e61d0b-a2c8-4729-9c60-20162a21aa35 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4fd6d89d-55ff-46da-b1b7-fae958d9c605 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e28f579a-43c0-4b89-bbaf-fd5e5bc8e433 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning I am EdgeRunner AI
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 52324d64-6c04-4d1f-801f-61360ca221ce · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Zheng, L
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 59953177-8abe-4da9-bdfe-30c73e0f5279 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ef40b99f-46ba-4556-bd29-cd07bb5f073d · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d73722ed-f188-4954-af71-82d1754581e9 · outbound
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning For most of the results, we use SGLang [63] v0.4.6 as generation backend and pytorch FSDP [ 62] as training backend
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 82ad9e0a-fd61-4bad-ad06-681f30b30f8b · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 191
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 87c5e038-0a6d-4e4b-b4a3-bd42cce82f05 · inbound
Reinforcement Learning from Human Feedback AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 134
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c27dc36d-5ebe-46fa-ab14-4e34213b8488 · inbound
Agent Lightning: Train ANY AI Agents with Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0525c134-88c7-4d1a-a5e2-4c201a03f947 · inbound
AWorld: Orchestrating the Training Recipe for Agentic AI AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8011cfc6-c1fd-46e9-a2f6-a47513364950 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 147
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d8d6583d-a411-4896-ac62-6ef44ab2e356 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 413d2afd-76e0-4a29-bcf6-2e2cd3ca6ef8 · inbound
Which Heads Matter for Reasoning? RL-Guided KV Cache Compression AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f02a984-a9be-4d4e-a564-8ec48c2572e0 · inbound
Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfc8bfee-2973-4ef5-ac6b-aeaf4c5df871 · inbound
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4bee1411-9875-4874-9e6f-8841b42ac7a5 · inbound
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aacc2d48-dda8-4a96-a6ab-3f900befb151 · inbound
HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 70d7ed3e-a225-46ab-9638-255e8247c743 · inbound
OpenTinker: Separating Concerns in Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d2d5ef5-381b-41af-b18b-44f0ae5fc23f · inbound
StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbaf49e3-f50a-4ba9-8bc2-a838fd1b49c6 · inbound
CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e97cb3-bbe2-4ffc-aa2d-49d925b892c4 · inbound
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eee55e6b-5056-48fa-a879-7ed6a6f470ef · inbound
Thinking Seeds: Leveraging Historical Diversity for Position-Aware RL in LLMs AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c0d5b6e-5517-449d-bb89-e4a7f943ed09 · inbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4c092ba1-561c-46a3-a115-ae396192aaa4 · inbound
ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation af19c22a-1aa9-4c5b-87fb-42df5799b123 · inbound
ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 084fafff-b067-49f9-a8bc-5ba0b9385e35 · inbound
RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 996aeed6-6acc-4490-8760-039426f88fec · inbound
When RL Meets Adaptive Speculative Training: A Unified Training-Serving System AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 082f7f1c-3122-49c3-9626-0c5d3de20b55 · inbound
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0167d83b-c497-4164-8fce-0954ae9341c9 · inbound
WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c0f8730c-77fb-48c0-ac05-e358dc8b2efb · inbound
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e196520-b2b4-49e4-9d2b-1e01e7157541 · inbound
OpenClaw-RL: Train Any Agent Simply by Talking AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9799de75-3d03-40f4-8300-dda427b53027 · inbound
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a80d0625-6c83-4771-ae17-5c4bf08115af · inbound
Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec567ed1-d217-4fe4-b075-28002f1fae0f · inbound
Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 15ec06f5-f9ea-4654-b1b6-b027fb62a2e5 · inbound
StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 28d5ebe2-0bb1-4463-abb6-4feedf122d17 · inbound
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3348572b-979b-420c-814a-95fb44087120 · inbound
AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8795d9fb-7462-4608-9133-2903e6a01b50 · inbound
DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6de70064-9924-4c14-83f4-7eebe1b6b2aa · inbound
Co-Evolving Policy Distillation AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bb3aa539-f7b4-4b52-b501-bf34ab776043 · inbound
T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0172fdfd-c0ec-48e9-b8f6-0b3c487c0e46 · inbound
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 60736719-82c0-4c9d-8720-03d335fe16cd · inbound
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 09e56093-512c-44af-9d7f-df841e883179 · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 25b28170-02ac-4885-9097-078af6c18838 · inbound
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d917e05-0452-4a23-8e23-8177d57a79de · inbound
BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b8030e71-5ff6-48c6-a20b-9ad790b99e42 · inbound
Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a2b86f5d-4b1c-425b-aac8-ef90ab2e10b3 · inbound
Position: Agentic AI System Is a Foreseeable Pathway to AGI AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 61174cbe-8304-48f5-95bc-82e7afc8816d · inbound
AIS: Adaptive Importance Sampling for Quantized RL AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2ca8497-d292-4e85-b552-d7a97237461b · inbound
Diagnosing Training Inference Mismatch in LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation da7d6f53-7e64-490d-be84-ca520310e3bc · inbound
ICRL: Learning to Internalize Self-Critique with Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 83b5152c-cd27-45c1-90b1-7541046ec6ff · inbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4f0699d7-5473-4da1-8a3b-44716b2409f2 · inbound
DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe810c69-4791-45df-a5f5-e6c151d4e47a · inbound
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 15a3adf2-8211-4f33-84b2-6124d127428b · inbound
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dd84ebee-5311-410b-a8cd-666e49a72249 · inbound
Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f6faf070-eb05-4ad1-8e18-b5aed790f443 · inbound
Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63abdda3-635c-4f93-b652-925b21489d17 · inbound
Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 21bb12d9-3fbd-4aec-a1a0-8f5b899e1174 · inbound
Trust Region On-Policy Distillation AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 237
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4c3d9624-528c-4dba-ba31-706b18038b2a · inbound
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 912f1d9e-9ba2-49c2-a818-18da27152940 · inbound
AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 00baaa07-cffa-4126-be70-b01ac06ab3a3 · inbound
AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 475c3c1b-5bec-4d66-b206-66aed008a67a · inbound
Rollout-Level Advantage-Prioritized Experience Replay for GRPO AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 43f1bb11-1658-4ba3-8889-5204cecca945 · inbound
When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 131
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf90f729-bc4c-4830-bad0-4c32d4ac6787 · inbound
Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e5c4347e-bbdf-42ec-b2c4-c64d284cef51 · inbound
Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aa490fe4-6ae8-453c-b334-d1b6aa5aa84f · inbound
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2dbde3c1-555d-477d-a89b-deda0b49132d · inbound
Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7b89e743-e38f-407b-b510-32c2ae4d4f9c · inbound
RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5149cbf7-dd78-4fc9-8aef-3cc9b5dbac56 · inbound
RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1d0c60b-fe75-4399-b3b0-a1d287f2b334 · inbound
Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 32bec269-cf7f-47ad-a24a-e0b2e9c40ea8 · inbound
Trees from Marginals: Autoregressive drafting with factorized priors AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d5e0de3a-0c46-403c-b27a-60b45309e067 · inbound
Trees from Marginals: Autoregressive drafting with factorized priors AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92f02aa6-2291-45cd-95a5-382a9e360cfa · inbound
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 55ea7ddb-1643-47ed-84ce-597ab67dd347 · inbound
Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0bd608c-2077-4856-a6e6-5c3f70d9e353 · inbound
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14b69ad0-a0c3-473e-83bb-8a8b6deaf7cf · inbound
WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f2db113-7f86-4a24-8845-7519a8acf1ce · inbound
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cd5dede-04e4-47e8-a00a-115009c85765 · inbound
QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a1c46ae-723a-478a-990e-bf9e9890cca4 · inbound
DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.