Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T20:10:32.300423Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2605.15565.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T20:10:32.300423Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e5ac3e37-e40a-49d5-95be-b33b108e4b03 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs arXiv preprint arXiv:2511.16108(2025)
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 64a59955-792c-4894-8561-9c2230e39d4a · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Why Do Multi-Agent LLM Systems Fail?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 21b2c46e-5089-4ddb-afaa-704d575f4f6e · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems , author=
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6156e740-7730-4aaf-b7d0-6935a859492c · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Frontier rl is cheaper than you think.https://fireworks.ai/blog/frontier-rl-is-cheaper-than-you-think , March 2026
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 83b5152c-cd27-45c1-90b1-7541046ec6ff · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cc460b16-ab70-4c03-8ca0-427c5b8225b2 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2f9e8ab8-1dd4-49fa-a234-e81503f366ec · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 968fb06a-b677-4af5-9dc9-75ea096ead2a · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs LLM Multi-Agent Systems: Challenges and Open Problems
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5140c882-4a16-4283-816c-bda3d05ccc19 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 69499fa6-ac48-494e-8c0b-ac469511e626 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 62d2a544-9f76-4179-a953-6367120a43cd · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9046c92b-22ca-4f4f-a040-2679e90e6c4d · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Art: Agent reinforcement trainer.https://github.com/openpipe/art
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 810f253f-aad9-4d69-bf90-4eca9a565b99 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3fef6357-5d32-41eb-a601-f2c997b7c9b8 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Prime-rl
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0059c9ed-692a-4e1b-9034-1108bcf710a8 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6bb570e2-8c71-432d-a870-14b40df9836e · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 11430c34-f4b3-4ac8-a656-fa7690e67b00 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3c225aa8-a2dd-4d7e-b057-297c8f30a4dc · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Computerrl: Scaling end-to-end online reinforcement learning for computer use agents
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e36dea94-3082-4d10-88d6-ee702e8fd02b · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Deepcoder: A fully open-source 14b coder at o3-mini level
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c508ca35-1289-498d-95ae-0e8577b480b4 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 62b51473-9f4e-44f7-93fa-3a1685713e0c · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 19336e7b-b9c9-482b-9bd4-ac708c7ee35a · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Real: Efficient rlhf training of large language models with parameter reallocation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1a6ddd1f-f4ff-4e66-9639-4b98d2145ec3 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f0aac165-e892-4322-beda-7603a055b7f8 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 83b30e56-ffc8-40d7-b135-5d8b6f2970b1 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Training language models to follow instructions with human feedback
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 20260cb5-267c-4991-ba5f-5b8eb6c00367 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Proximal Policy Optimization Algorithms
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1801ca8e-c178-4c98-996e-213fbce79f34 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f275f9a5-465a-4df6-90d9-fb8f31b817b3 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bc652e75-f322-4d52-9d89-accc02eeafa4 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs HybridFlow: A Flexible and Efficient RLHF Framework
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c5855360-b852-4146-a87f-bdbf1630bfc7 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Improving data efficiency for llm reinforcement fine-tuning through difficulty-targeted online data selection and rollout replay
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation abce3304-3ba0-4a1c-bc56-d1841b7bd503 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1e8cabbe-a2bd-4812-be4f-c4ce7ea66b79 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Marti-mars2: Scaling multi-agent self-search via reinforcement learning for code generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6efcf597-9f36-4c42-a2b2-993d3c4866ac · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Rlanything: Forge environment, policy, and reward model in completely dynamic rl system
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e54ac121-514a-4c13-9bb0-2a88f1b536e7 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 591e94f5-fa5c-4e94-ad96-9e2b90393888 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Autogen: Enabling next-gen llm applications via multi-agent conversations
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d4c13c69-fe25-43b0-a6fd-5dc55bdb2550 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation baca4235-a3ab-4962-b775-976bd4d23c9f · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Less: selecting influential data for targeted instruction tuning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f0910be4-e456-42bb-8fd2-ea4a3144cab1 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9495583-66d6-4193-92ef-947a9c18d515 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Areal-hex: Accommodating asynchronous rl training over heterogeneous gpus
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7959dc1c-335d-4f0f-b6f6-a2b6ee583d59 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 04a651c2-4e2f-422b-8f0f-a6280f66e224 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Qwen3 Technical Report
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b90d3184-8539-48d4-9d1e-b5469aae3006 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d207b28e-0eb3-451d-b031-739af31a52ac · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 276b815a-5661-4bb8-8f3f-9010f045d6fc · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Agentrl: Scaling agentic reinforcement learning with a multi-turn, multi-task framework
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fa48a93b-bf3d-4aec-8f05-28af6dfba6da · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Stronger-mas: Multi-agent reinforcement learning for collaborative llms
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff5148d9-995c-4f9a-b5b5-c0c504056ea1 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Group Sequence Policy Optimization
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 72467af6-8d03-4f29-8f8f-1424b2767578 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a1073676-9eb1-49cd-a3d0-2c0b44902166 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Prosperity before collapse: How far can off-policy rl reach with stale data on llms? InICLR
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2727de08-137d-4805-96ef-5bd097032a20 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Deepresearcher: Scaling deep research via reinforcement learning in real-world environments
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d060c8be-386b-4cbb-be5f-fad2cb4be93d · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e6ca8ac4-376e-4d69-8f72-b591ed56edbf · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Optimizing {RLHF} training for large language models with stage fusion
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f129269a-ebc8-4592-b110-51497ffb580c · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Let’s think step by step . . . \boxed{}
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 97000aee-cd75-4152-b2b0-26c0963fba8e · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Any actions except provided available actions will be regarded as illegal
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ec4340f9-926c-4185-8846-9b8ccddb3085 · outbound
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
No inbound Pith citation observations are available.