Pith. sign in

Paper Citation Record · LEDGER

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

As of 6 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 73 inbound Pith citation observations for arXiv:2505.24298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24298 v5

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 123 of 123 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 73 of 73 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:17:54.142833Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact15
  • verified fuzzy15
  • unresolved14
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch4

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation f1b5f65d-5401-46eb-a59d-e09f50458e72 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Dota 2 with Large Scale Deep Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.034522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:32e4612e96c3267b7027b46f22735648de523f555302aa3c318e053059680a3b

Observation 04c98360-1a8c-4d4b-9851-275de73d2520 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.024153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:b490e060ca412aebe3dcd0efef524046cd491ea2f14a50793495a1819d177b9a

Observation c68c561a-5cbb-46c8-b5b1-78b1f6320f9f · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.092893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:c423cd6d18249c40ca1f43274322add9570c379fa6c2ae360ef4f3ca3fcdef05

Observation f6726aac-2bc3-4aa1-867d-592e16976e21 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.061360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:f7da72ed3dacf2b4470efdcbd339aa59a1893faed683218a57be8f62cf9ec765

Observation 907e8104-ed56-4057-aa42-6ef156b66393 · outbound

This paper cites Espeholt, H.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Espeholt, H

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.095148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:65f7dd6734a4bbeac2f69748b3ac7dc18edb942f988095c44fb804a3b18e28c0

Observation 09668f0b-c48b-454b-8315-75fc0a8cb846 · outbound

This paper cites Espeholt, R.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Espeholt, R

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.097530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:58d7e49c57a8754546d0ee5ad951942abedff6b5184d35b279f4b23bd6323b75

Observation 65c5e111-101c-433e-b470-277b310c206f · outbound

This paper cites Hendrycks, C.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Hendrycks, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.099511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:3bced50e825faaaca3bfd4ff2550ba9f631d3bb21498baeb467bc8546eb84571

Observation 50b91d8b-0a5f-45b3-913e-904d135b37b4 · outbound

This paper cites Hilton, K.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Hilton, K

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.102020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:5ca40feeb2ead5958f4edf4b99d7dd54d4c11f345442a9b984010325351c1be9

Observation ddd1dcba-6b8a-4ca0-9e75-122efaeeb268 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.069097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:a964bf25a91b4ce2f314eadeaac6c4277b516fc1f01705829ddd66dda84a0907

Observation 1a8ba6c0-7b7e-4f1c-a6ae-8e5caf4218e9 · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.104566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:3b634016605f577b16572f116383296a7af29c75eb478f0800ebb00aeaafc548

Observation d379ce49-8768-4e9f-824c-cb32e3bfe0d9 · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.107087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:88904f50d3722953647bdc1f4ac07251d7eac1f9d6d054899e3d394219de28da

Observation 842f1a27-327f-4e04-9398-54f70b48c754 · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.109287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:a4155d738cc5840a92bc17d6bf44f224be6e816694c75f3c3792421c4004acc8

Observation 006a0879-9795-4d9f-a0a4-1c4ed76a61c7 · outbound

This paper cites Kapturowski, G.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Kapturowski, G

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.111431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:f0af8e133596029e8377c266815f4dcd2f64e3aea9e7ca1419c9c45be1e8d132

Observation 2bcfd707-7ecc-45ae-9639-e65d83ae8a57 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:24:21.973774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:3a02e216f868fb098b5538351875598f54785b4412394fec98f2dfaa1da3e5f4

Observation 3afe3235-b742-4fdf-9ba2-163c2ef1226f · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.113383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:a948137f8c07488696e6c0f4d63a86925b2727931d42626f00bc7f19dacd6b1b

Observation a83167a8-ffea-4f7c-80bb-72c2b906a51b · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.115514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:1f3a09c909abc9a92c7e06ebc392ebba7b95288ab0dd00b345eb49da198d80cf

Observation cd3d6ccb-9aac-4db8-8979-46e7f573bb14 · outbound

This paper cites Liang, R.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Liang, R

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.118256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:98bade433bd344a7ac49a9397019e0f5995af7c90aec9e8a9daddd8e69fb4587

Observation 00c3bd1a-c9f2-4983-9911-f4b4947d08cd · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.051843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:dfd11c3a11d120949dac5155a6750fcc751c65e262087f5da29224d58744ae7d

Observation 4c66119e-5d22-4e30-b656-2bc41e0ceaef · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.120199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:0fced4c2bc651132a03068247b2ee9da9a771031dab6c865b7a7642179ef0bac

Observation 5b2a3e38-cc70-4796-8b56-a8489d7d4d67 · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.122299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:7c972ec4f0ba6b41c9049761d3130251643519cff7bfa48be1a526b315cefeaf

Observation bc842cf3-3971-4022-b753-e1754b602810 · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.124185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:4c3fe65d4d94081afd02283c41aa276fb0198a431819dd8f4e987d2795b6b70e

Observation 5187a572-3c6a-49dc-b575-ee6e3fe22140 · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.126219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:0c347f4bd62118a27cdecdde82bc1b5edcac62891ccc79c8124bea1240c889c8

Observation 2fe33030-686b-40dd-a9d0-a8732103d5fa · outbound

This paper cites Specinfer: Accelerating large language model serving with tree-based speculative inference and verification.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Specinfer: Accelerating large language model serving with tree-based speculative inference and verification

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.009165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:25f69f48e14a153fec50491f002b374bc984e747fb90a07af1af764f62cf67a2

Observation 9ac1dd7b-70c8-4f58-a4b3-041e34200782 · outbound

This paper cites URL https://openai.com/index/ learning-to-reason-with-llms/.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning URL https://openai.com/index/ learning-to-reason-with-llms/

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.128959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:baa21481f450fd4006c0d1bf566a736a65efda3dc3ce4540d772f0e68786d43a

Observation 22c0fc43-2c26-490e-8a5c-a489cb6e1b09 · outbound

This paper cites URL https://openai.com/index/ introducing-o3-and-o4-mini/.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning URL https://openai.com/index/ introducing-o3-and-o4-mini/

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.131543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:f7a383df61816a175a26bdcfb848d0e08d755245a1d226788fee35685416ceed

Observation 17be9c91-965d-422f-85f3-be94cfbb6859 · outbound

This paper cites Ouyang, J.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Ouyang, J

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.134269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:b8329cb51f8aa1d2a2739d418906eefe75a61bca627482c3cd52bbfd3739c8ce

Observation 746cce49-6e19-44ed-a3ce-a0c2003f75ea · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.136472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:3447796f9b8e9b447d54a0a57c65d1846defbfa0601c7955931715f5fc737247

Observation ec081b8d-129c-43c1-8ab9-c91d41defadf · outbound

This paper cites Paszke, S.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Paszke, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.071955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:746374781eb2f39458134ce2da9afa4406df13d4046eb1170e6354d05e5467c3

Observation b34c7e4e-1a27-448a-892a-0bbd00897967 · outbound

This paper cites Wiley Series in Probability and Statistics, Wiley (1994).

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Wiley Series in Probability and Statistics, Wiley (1994)

Reference 39

Resolution
verified exact
doi, observed 2026-05-15T14:24:22.012386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:6806a86f664d96e670e1e1823fcd993f11f03fc6374528e44e4862797754bf79

Observation 4cfce521-ebf0-4a16-bfe9-efb03efa7f1c · outbound

This paper cites Generalized Slow Roll for Tensors.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Generalized Slow Roll for Tensors

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:24:22.016165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:5991c6b146bc0bc134cd6a7f492aa88e19db270b2480b05669df38ec9b71e4db

Observation 75a53cec-ed67-4faf-a111-953e497e4cbf · outbound

This paper cites Schulman, P.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Schulman, P

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.074665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:de61b2101e0743d2b1258ecca99aaf3d8c6d805b368478ce5d7ec3ff113d6fa6

Observation 1df6b9ad-1039-4be0-b252-ddcca0f8182d · outbound

This paper cites Proximal Policy Optimization Algorithms.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Proximal Policy Optimization Algorithms

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.038379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:351fda4b31483c9310e18c066321724ded3f4aca9ae4b69066e1482963717858

Observation 80570ad7-d6ed-406a-8833-791c6d4c419a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.042354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:4a1cbd02127f5eca35e4dbd80c02691e3f8ac94fc3697384ccf83b0c84060a04

Observation 55dbd829-c259-4871-b28d-53e033abd762 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Hybridflow: A flexible and efficient rlhf framework

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:24:22.003970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:0e61cca7b73b0cfa447bf813f7f6911ea27010bccbeb45a774ce02f4e237b2e8

Observation 9b490396-40b7-48ba-8703-c5eb8ac69ba5 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.046981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:da67304fa1219dceb4deb189908df6c7236e0d4b220b7063f3c4c10c4624a502

Observation bf02d8bd-d30e-42e4-80c0-b9cd86e0f88e · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.077208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:11d930af1c0f896aa47121f208e6f4efe06e6c81a834a8b0d9ef3502e19877b5

Observation a9bb1a91-530b-4c64-91dc-330177927b5a · outbound

This paper cites INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.057134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:5c746c4e0eb935867066b50dcd827390de192e4977600b1ee2f49b46e68f5042

Observation ac46e50e-a781-4a67-b569-1d9158477b2e · outbound

This paper cites Vaswani, N.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Vaswani, N

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.079530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:990e092aed12da3948ee6305011d6ece3f932ada6926afb46c545f358198c9c9

Observation bb8f25bb-19ec-4261-b3a5-29ee5836815c · outbound

This paper cites Liar” ends the game, then both players reveal their dice. If the last bid is not satisfied, then the player who called “Liar.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Liar” ends the game, then both players reveal their dice. If the last bid is not satisfied, then the player who called “Liar

Reference 55

Resolution
verified exact
doi, observed 2026-05-15T14:24:21.993334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:b88cf21460868a2d23ed1fc752db749507e9bcb3dcad8d8be8fcecb717658379

Observation 776aaf06-5ed0-4e76-b9fe-01a51c09984b · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.081758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:db64b551ad90cb33b5b475e618fc0ac8028f8c862df9ab2ffcae19e58665b27e

Observation b991c518-178b-426d-a6ca-76576e505494 · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.065868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:7ff499e43e7c129b31da518dffd959fd9e85eda9f117aa554699181e463fb343

Observation 3628c109-2250-48f9-95a8-63f07031ee6a · outbound

This paper cites an unresolved cited work.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-15T14:24:22.083916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:05fd2bf01588d19ca428ab68067f8768f027817c9108024480422dc556ae2c23

Observation ed6d1c0e-4631-41c4-aa10-076c7defaa9f · outbound

This paper cites Job Scheduling Strategies for Parallel Processing pp 44–60.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Job Scheduling Strategies for Parallel Processing pp 44–60

Reference 65

Resolution
metadata mismatch
doi, observed 2026-05-15T14:24:21.978165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:52805d68767f9aaeb1334671d2253532696478f4f929d815c55a756ecdc4a76b

Observation 79e61d0b-a2c8-4729-9c60-20162a21aa35 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 66

Resolution
malformed identifier
local_arxiv, observed 2026-05-15T14:24:21.986322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:911fe01cfdc4e7532b46c7a6f3570607c981cd7fd216e9af8083adf75bf4b3ce

Observation 4fd6d89d-55ff-46da-b1b7-fae958d9c605 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:24:22.020162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:7652604235456900a840288995498d50101d7bed2b86deba3d3e66209a3c701e

Observation e28f579a-43c0-4b89-bbaf-fd5e5bc8e433 · outbound

This paper cites I am EdgeRunner AI.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning I am EdgeRunner AI

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:21.999394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:77ac74efcd8ed8712f1fee6ecd8eb26181fc066884286663e4cf33d3ab1e3b4d

Observation 52324d64-6c04-4d1f-801f-61360ca221ce · outbound

This paper cites Zheng, L.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Zheng, L

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.086140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:f4fc955fc1cafc7e142b17e44d4337c061aa2278553b5e0558622119949e4629

Observation 59953177-8abe-4da9-bdfe-30c73e0f5279 · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 70

Resolution
malformed identifier
arxiv_id, observed 2026-05-15T14:24:22.029099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:0d2b2eef311273e68b9be6b664c60797a7c9c3e627e413463331452750db6090

Observation ef40b99f-46ba-4556-bd29-cd07bb5f073d · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.088384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:a3aa5ca3762e2314d10f7760355fd357daaf2430cc398ebf975dc10eda6e5ece

Observation d73722ed-f188-4954-af71-82d1754581e9 · outbound

This paper cites For most of the results, we use SGLang [63] v0.4.6 as generation backend and pytorch FSDP [ 62] as training backend.

AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning For most of the results, we use SGLang [63] v0.4.6 as generation backend and pytorch FSDP [ 62] as training backend

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T14:24:22.090551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:24:21.930934Z digest=sha256:ad5065aa48d8912cc94e92b2989099133d137ce806210e0c9d9794e63d75737c

Pith citing papers

Observation 82ad9e0a-fd61-4bad-ad06-681f30b30f8b · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 191

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:5c717dcd1983df277405e543f8201a38754f17fa534b4249ce9d4a26540055ed

Observation 87c5e038-0a6d-4e4b-b4a3-bd42cce82f05 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 134

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:32:01.288940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:0d895d98d8b974835eb2919cc7edcc67ac5c2da6b2b7b6454832e5b907bb7f72

Observation c27dc36d-5ebe-46fa-ab14-4e34213b8488 · inbound

Agent Lightning: Train ANY AI Agents with Reinforcement Learning cites this paper.

Agent Lightning: Train ANY AI Agents with Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T04:17:54.142833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:17:54.142833Z digest=sha256:b4ce3071fa34c8f45881cd0f742e36b4eccba0a64ceaedeaa84cd88890b0426d

Observation 0525c134-88c7-4d1a-a5e2-4c201a03f947 · inbound

AWorld: Orchestrating the Training Recipe for Agentic AI cites this paper.

AWorld: Orchestrating the Training Recipe for Agentic AI AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:17.779979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:17.779979Z digest=sha256:c46dcd9f2ac60afd8a56371cfd6587ca6bd8362f2b8c1a598eccab8f27b1b2a9

Observation 8011cfc6-c1fd-46e9-a2f6-a47513364950 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 147

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:25.401254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:400f43679b2ffddc75f7defe6c1aacd7bf81b447836cd2bc6cf1c8e7f780d7af

Observation d8d6583d-a411-4896-ac62-6ef44ab2e356 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.199481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.199481Z digest=sha256:744d9e828b12109bc7fec74ee08b7beef775360b5e8bd29df333e4fe144303aa

Observation 413d2afd-76e0-4a29-bcf6-2e2cd3ca6ef8 · inbound

Which Heads Matter for Reasoning? RL-Guided KV Cache Compression cites this paper.

Which Heads Matter for Reasoning? RL-Guided KV Cache Compression AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:50:09.254400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:50:09.254400Z digest=sha256:81aad742c0ecffb7246f5ed94f2be21d9abeefb4cbf83f11c51b7d5af16b7144

Observation 8f02a984-a9be-4d4e-a564-8ec48c2572e0 · inbound

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts cites this paper.

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:39:26.538300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:39:26.538300Z digest=sha256:1e775e79ccf4f9d1c1280eb67861e1e5734252173ca3218e9bca969da8d9e072

Observation dfc8bfee-2973-4ef5-ac6b-aeaf4c5df871 · inbound

RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs cites this paper.

RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.176765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T05:29:46.136115Z digest=sha256:17e2f68fedf1f2d29f3ef0175a790a6d82126814399731620db6f0ebb35bb7bd

Observation 4bee1411-9875-4874-9e6f-8841b42ac7a5 · inbound

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning cites this paper.

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:40:14.640754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:38:30.169363Z digest=sha256:af5d82666f5ca53a4432961b3db6d500aa07c8b1afdc2fac4103f74f38b2b837

Observation aacc2d48-dda8-4a96-a6ab-3f900befb151 · inbound

HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments cites this paper.

HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:23:36.646232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T22:21:26.271796Z digest=sha256:abb04677bf5c02a9fba5c9943675641fcec146f9c920e56293ce05772f930645

Observation 70d7ed3e-a225-46ab-9638-255e8247c743 · inbound

OpenTinker: Separating Concerns in Agentic Reinforcement Learning cites this paper.

OpenTinker: Separating Concerns in Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T11:08:09.185820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:08:09.185820Z digest=sha256:cd12060fe8cfeed91daebba95eba65ae81917de66f5094066eda13595f909838

Observation 7d2d5ef5-381b-41af-b18b-44f0ae5fc23f · inbound

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training cites this paper.

StaleFlow: Staleness-Aware Data Management for Mitigating Data Skewness in Fully Disaggregated RL Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T06:25:47.347729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:25:47.347729Z digest=sha256:b56047ab1159fb40824335012ec5a12f0f634236a805c68e4b7daa05f9df05b5

Observation bbaf49e3-f50a-4ba9-8bc2-a838fd1b49c6 · inbound

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning cites this paper.

CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:48.820681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:48.820681Z digest=sha256:c858b2e403333bd4b18a9a7a05f68523ae090bf7a83d44f751d4fbbea3ffd655

Observation 26e97cb3-bbe2-4ffc-aa2d-49d925b892c4 · inbound

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning cites this paper.

FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:07:47.092261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T11:07:18.839899Z digest=sha256:435f59681d9575195504060c1ac1974c73c4840b3b857027ea26c0779d24d0c2

Observation eee55e6b-5056-48fa-a879-7ed6a6f470ef · inbound

Thinking Seeds: Leveraging Historical Diversity for Position-Aware RL in LLMs cites this paper.

Thinking Seeds: Leveraging Historical Diversity for Position-Aware RL in LLMs AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T07:01:12.356383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:01:12.356383Z digest=sha256:73616575cfa77a0d7243c8bb4b15b22c5aaaced1e9fa86deba64696916de7b7e

Observation 7c0d5b6e-5517-449d-bb89-e4a7f943ed09 · inbound

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment cites this paper.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:40:49.102585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T09:37:57.120779Z digest=sha256:f2220e34eb835710ef891a35fe05eb85b77028221b4fb5d3ac909045db4c935b

Observation 4c092ba1-561c-46a3-a115-ae396192aaa4 · inbound

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment cites this paper.

ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T14:10:13.456057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T14:05:37.120262Z digest=sha256:2e222e126de88ff4124f489de0fb5a274a708370a09c5fb0bf70de1dc2bfbe7d

Observation af19c22a-1aa9-4c5b-87fb-42df5799b123 · inbound

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning cites this paper.

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T05:32:40.837833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:32:40.837833Z digest=sha256:15af94256811d60eb7875ad3e3960eb6f4b56d1d97f376342103705b4eacd5e4

Observation 084fafff-b067-49f9-a8bc-5ba0b9385e35 · inbound

RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training cites this paper.

RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:02:28.930758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T07:01:54.696390Z digest=sha256:d53c7c2a714c0803242dc968b867d5902ec7bfcaea9184dd8e47641575605c87

Observation 996aeed6-6acc-4490-8760-039426f88fec · inbound

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System cites this paper.

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:37:28.921689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T06:33:41.860803Z digest=sha256:d244ce0b7f39c62a9d6636169fad9aa7846aa259f46ed08e253a9366ecc8edb5

Observation 082f7f1c-3122-49c3-9626-0c5d3de20b55 · inbound

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System cites this paper.

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T23:32:11.710514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:32:11.710514Z digest=sha256:4dd77f1850f27be209c13bb0e534e5b16810967d4d0f930522470f5defbe7c85

Observation 0167d83b-c497-4164-8fce-0954ae9341c9 · inbound

WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces cites this paper.

WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T16:36:17.548558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T16:33:58.746944Z digest=sha256:0aaaaf0050f3607b73c1167be9b078a0533e5357650958022f8757c530b89b2f

Observation c0f8730c-77fb-48c0-ac05-e358dc8b2efb · inbound

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments cites this paper.

Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T18:45:20.779259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:45:20.779259Z digest=sha256:f6b1a31da7e36f89f53302a1396b417d8f27ec475a6d691c3260536e527dca7c

Observation 9e196520-b2b4-49e4-9d2b-1e01e7157541 · inbound

OpenClaw-RL: Train Any Agent Simply by Talking cites this paper.

OpenClaw-RL: Train Any Agent Simply by Talking AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T12:56:30.316823Z digest=sha256:0482ae03da87153d9bfa9dcf16afc32e2a900891dcbab7a0d1ecb02503471b95

Observation 9799de75-3d03-40f4-8300-dda427b53027 · inbound

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training cites this paper.

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:16:33.020610Z digest=sha256:8a1fe0bbabcfbe6cd18fdd5e34bd063c8060aa1ba2b519f442470f45e533f92d

Observation a80d0625-6c83-4771-ae17-5c4bf08115af · inbound

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale cites this paper.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:34d53465e898a69b128a21aad04b247a2a0ee8c248b1da0f8af32bf82b4a2c29

Observation ec567ed1-d217-4fe4-b075-28002f1fae0f · inbound

Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning cites this paper.

Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T07:27:15.270996Z digest=sha256:b06b1e6d08641d5670b9457f82563356788f2ac3328e69e5c8a17c6911a5c712

Observation 15ec06f5-f9ea-4654-b1b6-b027fb62a2e5 · inbound

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning cites this paper.

StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T04:26:59.781416Z digest=sha256:3b00174e1eed9fc35c3c8a32924b0845e90c216ed6baa7d700d41afe8b3cee59

Observation 28d5ebe2-0bb1-4463-abb6-4feedf122d17 · inbound

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training cites this paper.

JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:36:33.914193Z digest=sha256:f810cd311b5564ea43d34ea69656b0698136f87dfb46198e41d24776187ea570

Observation 3348572b-979b-420c-814a-95fb44087120 · inbound

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving cites this paper.

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T14:16:46.241126Z digest=sha256:1ebd81a3319d94c7e30c891e47d70725e85f2af67c948d0e993416bf6cbf41d7

Observation 8795d9fb-7462-4608-9133-2903e6a01b50 · inbound

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training cites this paper.

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T13:49:26.560459Z digest=sha256:347d47ac72b9a98487945d203e53ff6baabf9aab5dee9d03f1d3f8cb1f14d5de

Observation 6de70064-9924-4c14-83f4-7eebe1b6b2aa · inbound

Co-Evolving Policy Distillation cites this paper.

Co-Evolving Policy Distillation AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T08:23:41.819485Z digest=sha256:27c26eecbdb1020b2bd2911e12a20787376d2e26991d04abef6d585ac6a48c79

Observation bb3aa539-f7b4-4b52-b501-bf34ab776043 · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:60aa3342820f61ebca070d1245ea7124b581947763ca497b6980477bf832439c

Observation 0172fdfd-c0ec-48e9-b8f6-0b3c487c0e46 · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T10:12:58.421050Z digest=sha256:9bb876b7a3209c5dc7ef518caf3e6d4d4fea7a067e6aeb970cc96ba833d27cf4

Observation 60736719-82c0-4c9d-8720-03d335fe16cd · inbound

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence cites this paper.

Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T00:56:48.838028Z digest=sha256:b6954c7e5930b5bfb9c063e46f65d8309afc203ca57e86534cbd3f79379b87fe

Observation 09e56093-512c-44af-9d7f-df841e883179 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:ba35b31d7c731d8cef33074f2f8e653aa4bea5f3845ef1189235959855480659

Observation 25b28170-02ac-4885-9097-078af6c18838 · inbound

FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration cites this paper.

FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:46:59.400834Z digest=sha256:ecd5d35821dd24dce6c184175affbdf70d71c077ebfce86f485d52f30f409923

Observation 8d917e05-0452-4a23-8e23-8177d57a79de · inbound

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning cites this paper.

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:23:13.071107Z digest=sha256:c12dc71a6e4e4fc120b55cb8b3060826cef3a2060d5546fc24dc819a2c40ac9f

Observation b8030e71-5ff6-48c6-a20b-9ad790b99e42 · inbound

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search cites this paper.

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T02:58:46.728868Z digest=sha256:dbdf6c2a87397134ff2750094e6fd0ef218578cb03bf01bd7417e62575d35b55

Observation a2b86f5d-4b1c-425b-aac8-ef90ab2e10b3 · inbound

Position: Agentic AI System Is a Foreseeable Pathway to AGI cites this paper.

Position: Agentic AI System Is a Foreseeable Pathway to AGI AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T20:10:36.101426Z digest=sha256:cf36961212360a2608945c01ad879d7a65a2309c963fdc40eb53b42131b5fb8b

Observation 61174cbe-8304-48f5-95bc-82e7afc8816d · inbound

AIS: Adaptive Importance Sampling for Quantized RL cites this paper.

AIS: Adaptive Importance Sampling for Quantized RL AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T03:13:14.384567Z digest=sha256:8438202ae27a60eeee1b43cbcb989f7ba39ac9d8b8c6e7e681ae8657c24621dc

Observation f2ca8497-d292-4e85-b552-d7a97237461b · inbound

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning cites this paper.

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T02:05:44.813341Z digest=sha256:0dec56afd29ae379dcf854c23f2c930d170a741d0bcbcb2708b43ce4512754ba

Observation da7d6f53-7e64-490d-be84-ca520310e3bc · inbound

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning cites this paper.

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T18:02:42.448329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T17:58:05.817581Z digest=sha256:156117ffd2a03f6192b5a472f7effda1cfeac35a4a5ec77e7e3bbf849ed135bb

Observation 83b5152c-cd27-45c1-90b1-7541046ec6ff · inbound

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs cites this paper.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.758614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:8b4df24d5766faf8a80d821a6256ba4c3c338bd71662c0067007cb299a1f4fcf

Observation 4f0699d7-5473-4da1-8a3b-44716b2409f2 · inbound

DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training cites this paper.

DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:09:12.222518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T23:07:18.427357Z digest=sha256:d1cbff9be647b6ddf8b922a0f813f6db5d69049b3662318e50aa25507fc015d8

Observation fe810c69-4791-45df-a5f5-e6c151d4e47a · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:34:02.628850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:32:12.180233Z digest=sha256:b32f2bedfca8e8657692346e183872bd761b6039587321a2c6ed1f312ea402ea

Observation 15a3adf2-8211-4f33-84b2-6124d127428b · inbound

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning cites this paper.

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:01:19.399223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T08:59:28.405218Z digest=sha256:816ac3b0a464c799a02ec48ac6e0ab79dd9db521efb4608cec3136bd5bfeea79

Observation dd84ebee-5311-410b-a8cd-666e49a72249 · inbound

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor cites this paper.

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:50.046588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T07:59:43.755196Z digest=sha256:97da3cecf5333facbba2c11d285af2e8f04313038e46f0a434558695408ad494

Observation f6faf070-eb05-4ad1-8e18-b5aed790f443 · inbound

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor cites this paper.

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:50:23.726282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-25T05:49:08.484663Z digest=sha256:e88dca22ecf7d82c666308cc959c351327a537eb7ccf4853d02e8affb5000595

Observation 63abdda3-635c-4f93-b652-925b21489d17 · inbound

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor cites this paper.

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:04:57.966303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T18:01:38.509794Z digest=sha256:c7c8c69c4c7d2ac587aa2dbe12a8530c427ede5391a20caa8a9b1c76e91eef09

Observation 21bb12d9-3fbd-4aec-a1a0-8f5b899e1174 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 237

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.542325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:f9fa3300d4efed0252123d4de2bc4dd79c87294d37fc2a89d9b48e91958eb8ab

Observation 4c3d9624-528c-4dba-ba31-706b18038b2a · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T21:16:13.397428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:69e269000cefb4f2afc02d3e83f5a19e52687ffc0d8661cbe977e83602ce8372

Observation 912f1d9e-9ba2-49c2-a818-18da27152940 · inbound

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning cites this paper.

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:56:47.919524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T06:26:35.100859Z digest=sha256:b8295c47d9f57c9f6eaf74a3838bc1c47e3882e011383cd257faac3d42535ccb

Observation 00baaa07-cffa-4126-be70-b01ac06ab3a3 · inbound

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning cites this paper.

AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T12:27:08.279842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:27:08.279842Z digest=sha256:4a6d4ba41958cfeb37964599c3d54108c335db4d04393d4e4a6032db9167398a

Observation 475c3c1b-5bec-4d66-b206-66aed008a67a · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:06:41.093107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:804e1b91c0b9ebea3aeb80ce896b735fc4e09b0aa97e367d796729c3ef45a94e

Observation 43f1bb11-1658-4ba3-8889-5204cecca945 · inbound

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff cites this paper.

When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 131

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:27:26.415608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T18:49:47.876179Z digest=sha256:e16e7f3f3b8164fd76ce626ba8dc603abd4a7c56ff3b616fcff34f1e02798e70

Observation bf90f729-bc4c-4830-bad0-4c32d4ac6787 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T09:47:59.826024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:fa44a5ef42c58d4e6982a1cffaa42eab95cc4a7488336823eb3f8c8b9f7914e6

Observation e5c4347e-bbdf-42ec-b2c4-c64d284cef51 · inbound

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training cites this paper.

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-03T13:08:07.825297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T08:35:16.435272Z digest=sha256:e3e65398b9307369f78c9fd7226dbecce6e5c04bb24107415a09455c2e7cdff2

Observation aa490fe4-6ae8-453c-b334-d1b6aa5aa84f · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:17:25.513987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:985aecf16ca81d4c039805dda92b42f92764f644a6b2927aecce339cb42ff9a6

Observation 2dbde3c1-555d-477d-a89b-deda0b49132d · inbound

Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training cites this paper.

Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:24.703409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T19:06:16.100762Z digest=sha256:ad806f9b5dd4d07f302d5aa181d4f578e7e838cb357551086d887cfdae22e0d7

Observation 7b89e743-e38f-407b-b510-32c2ae4d4f9c · inbound

RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning cites this paper.

RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T14:39:58.209992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T02:56:47.102678Z digest=sha256:6c4f2a0f55395f96676b6aad2550847adb526e756943e551e439c711b097fa38

Observation 5149cbf7-dd78-4fc9-8aef-3cc9b5dbac56 · inbound

RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning cites this paper.

RolloutPipe: Overlapping Pipelined Rollout and Training in Disaggregated On-Policy LLM Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T11:52:35.159984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:52:35.159984Z digest=sha256:9a1eb72d287a3cb27509fee25ece31225f7d6526062bdc999785b2b6f0bdfd24

Observation c1d0c60b-fe75-4399-b3b0-a1d287f2b334 · inbound

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF cites this paper.

Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 7

Resolution
malformed identifier
local_arxiv, observed 2026-06-29T05:43:07.654879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T01:23:07.441218Z digest=sha256:0d09767f18f3c21bfcf7fd2f27439487df7cf54f3bf9453c4fe42c25756f2efa

Observation 32bec269-cf7f-47ad-a24a-e0b2e9c40ea8 · inbound

Trees from Marginals: Autoregressive drafting with factorized priors cites this paper.

Trees from Marginals: Autoregressive drafting with factorized priors AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T21:57:39.209960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-10T21:49:34.512370Z digest=sha256:ef4193d33d83685fd282160f671fdc894cb6abd8dd780fffbcfeb6254d42020d

Observation d5e0de3a-0c46-403c-b27a-60b45309e067 · inbound

Trees from Marginals: Autoregressive drafting with factorized priors cites this paper.

Trees from Marginals: Autoregressive drafting with factorized priors AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T15:58:05.356765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:58:05.356765Z digest=sha256:358e8865d71f81dd5d3330ec40b08cfb655377c00f645f7828e9037334290aa2

Observation 92f02aa6-2291-45cd-95a5-382a9e360cfa · inbound

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning cites this paper.

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-09T08:56:06.428422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T08:51:10.098370Z digest=sha256:b7e3655b6b96d955704c248e612c35d93de00971ab4bef08da4b96b3453bd883

Observation 55ea7ddb-1643-47ed-84ce-597ab67dd347 · inbound

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training cites this paper.

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
malformed identifier
no resolver link, observed 2026-07-13T04:42:35.589143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:42:35.589143Z digest=sha256:53f4c03ead11252c3dd087f6bd4e2849e93f6e4f99d2eabe45bb708f38a28156

Observation f0bd608c-2077-4856-a6e6-5c3f70d9e353 · inbound

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning cites this paper.

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T06:37:37.177652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:37:37.177652Z digest=sha256:a19628c3e2bf406625c5c1c1a12e138841627f2e6a9587ec2c2f4ce0a30d98c6

Observation 14b69ad0-a0c3-473e-83bb-8a8b6deaf7cf · inbound

WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning cites this paper.

WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:29:31.579850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:29:31.579850Z digest=sha256:fbbfd04df8369a8d2889e93de8e8ac6ed44dea087772eb4c9e3d59976ec6a12e

Observation 7f2db113-7f86-4a24-8845-7519a8acf1ce · inbound

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning cites this paper.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:28.853734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:28.853734Z digest=sha256:01ca4851e0c86a594614d73546a861ffaaa810179f78a160b34e064fdcb73434

Observation 4cd5dede-04e4-47e8-a00a-115009c85765 · inbound

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization cites this paper.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:17.311342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:17.311342Z digest=sha256:26d7bd681dd5c4b416db48799b5701ae48e8d042d46c36d74389a8ae04ad8a92

Observation 4a1c46ae-723a-478a-990e-bf9e9890cca4 · inbound

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training cites this paper.

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T11:13:59.560465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:13:59.560465Z digest=sha256:0ce7c1586c592ce834556cdf2277f724c899308f5003efc2927564bcfb6bad80