Pith. sign in

Paper Citation Record · LEDGER

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs

As of 4 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2605.15565.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.15565 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T20:10:32.300423Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact37
  • verified fuzzy15
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e5ac3e37-e40a-49d5-95be-b33b108e4b03 · outbound

This paper cites arXiv preprint arXiv:2511.16108(2025).

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs arXiv preprint arXiv:2511.16108(2025)

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.816406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:54f2a261144346a52c23d23443d26331119e13653e1cf6168bc18ace8a799ad8

Observation 64a59955-792c-4894-8561-9c2230e39d4a · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Why Do Multi-Agent LLM Systems Fail?

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.821503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:fd4b3375da4acb8b38b4e28fa8af0ac8d6129880e7cfeb505d2badc15acc8938

Observation 21b2c46e-5089-4ddb-afaa-704d575f4f6e · outbound

This paper cites MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems , author=.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems , author=

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.840692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:43d05693b2eafebdc072e9c4d3556aaffc7a229ad3f4322ef505b14f206b1a01

Observation 6156e740-7730-4aaf-b7d0-6935a859492c · outbound

This paper cites Frontier rl is cheaper than you think.https://fireworks.ai/blog/frontier-rl-is-cheaper-than-you-think , March 2026.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Frontier rl is cheaper than you think.https://fireworks.ai/blog/frontier-rl-is-cheaper-than-you-think , March 2026

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.831613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:b1afc84f7e2210ba1077c0ff45eb9c0d6bf298fdd0193250b5aa89a6a0a0d67c

Observation 83b5152c-cd27-45c1-90b1-7541046ec6ff · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.758614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:ce60f41f628e83bcccd6676c6af856da97c95d1a5637769f84038a6fd6c1b0c6

Observation cc460b16-ab70-4c03-8ca0-427c5b8225b2 · outbound

This paper cites Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.812045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:3c7f1a743617dc5d362c8d9e2dfc10c44348e465af5640a4e360e40622b39342

Observation 2f9e8ab8-1dd4-49fa-a234-e81503f366ec · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.793215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:d0ebb28c7f712f38c3516e856c0ec4026ab2583a4f5686029ae17a4d085fda68

Observation 968fb06a-b677-4af5-9dc9-75ea096ead2a · outbound

This paper cites LLM Multi-Agent Systems: Challenges and Open Problems.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs LLM Multi-Agent Systems: Challenges and Open Problems

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.797792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:f5f2409f79c2061208068ce0cb94b941bf17a673dc47b8740d686a2dbca0ed3b

Observation 5140c882-4a16-4283-816c-bda3d05ccc19 · outbound

This paper cites AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.849930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:f09c319ae9288bd49e9f274283069afee39797d7a499dc4d7c11e21188d9473b

Observation 69499fa6-ac48-494e-8c0b-ac469511e626 · outbound

This paper cites History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.860135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:80a3ef0d3e471df89ac6c626fe2fb0564f7df50369c23685c1596308b4a7cee1

Observation 62d2a544-9f76-4179-a953-6367120a43cd · outbound

This paper cites HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.801716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:23b1548b871d49adc6d24320627203c3db9b80d138d3381a78076f42d23d740d

Observation 9046c92b-22ca-4f4f-a040-2679e90e6c4d · outbound

This paper cites Art: Agent reinforcement trainer.https://github.com/openpipe/art.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Art: Agent reinforcement trainer.https://github.com/openpipe/art

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.826771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:9526f13d92994fb434e213a4cd125c595692b05e2f5a008221c7366377880b67

Observation 810f253f-aad9-4d69-bf90-4eca9a565b99 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.798594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:0e3058d6d023b555dff42a6e8421ab47784415c68ef83923a41db7f111b92cfb

Observation 3fef6357-5d32-41eb-a601-f2c997b7c9b8 · outbound

This paper cites Prime-rl.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Prime-rl

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.819615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:3fb0938456b389ec9d80ffae63f4bf8d6904b7e2159a7a8a2e77f498737c6d52

Observation 0059c9ed-692a-4e1b-9034-1108bcf710a8 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.711778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:e0fe2de76a84a91a51b9f0fc672086751b971f74af37894b31de070023c06594

Observation 6bb570e2-8c71-432d-a870-14b40df9836e · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.718261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:51e9f50ea2d83cba4fe4d860f56d2700bd18a806a2f90cb138ee04094f34734b

Observation 11430c34-f4b3-4ac8-a656-fa7690e67b00 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.737537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:a8b398fb4a89b34a664d7f436692702a4c59c89d9111e24bc864a144358168e8

Observation 3c225aa8-a2dd-4d7e-b057-297c8f30a4dc · outbound

This paper cites Computerrl: Scaling end-to-end online reinforcement learning for computer use agents.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Computerrl: Scaling end-to-end online reinforcement learning for computer use agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.871508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:aa623ffcbab2dc907715c8658bc85fea6c5a1d07f873b7209d079d8ff88938ca

Observation e36dea94-3082-4d10-88d6-ee702e8fd02b · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Deepcoder: A fully open-source 14b coder at o3-mini level

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.810298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:473bbf3d801c714431580df537e52c4e75441815989f7e72c5d422a5489252b1

Observation c508ca35-1289-498d-95ae-0e8577b480b4 · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.813510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:adafa096e79fa333afa7174d91414362b59e1d13d705d6e397f48a5b4881f54c

Observation 62b51473-9f4e-44f7-93fa-3a1685713e0c · outbound

This paper cites ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.884991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:a5477ee9ae2cfc7c0dd62ee7741ba274fbd6d50d145345f578b038c039d2db1b

Observation 19336e7b-b9c9-482b-9bd4-ac708c7ee35a · outbound

This paper cites Real: Efficient rlhf training of large language models with parameter reallocation.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Real: Efficient rlhf training of large language models with parameter reallocation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.806799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:37d0830bf16b1ee4560853b46661919b9c45c58599f92333b085063bd9c8a697

Observation 1a6ddd1f-f4ff-4e66-9639-4b98d2145ec3 · outbound

This paper cites Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.730105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:9bfb9f9d650153f0679eba0ebd81f5724ea2166ee11dc6e6ff9bb8ddfc86a365

Observation f0aac165-e892-4322-beda-7603a055b7f8 · outbound

This paper cites Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.696571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:9f91d47e340f9fbd8196b8998dac312831af501bdbb8349480d1c95df4662954

Observation 83b30e56-ffc8-40d7-b135-5d8b6f2970b1 · outbound

This paper cites Training language models to follow instructions with human feedback.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Training language models to follow instructions with human feedback

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.816717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:26d3ac663ee76a1cf219081949f98c6ec0f5c2f3966915d8a3367abd15880dc8

Observation 20260cb5-267c-4991-ba5f-5b8eb6c00367 · outbound

This paper cites Proximal Policy Optimization Algorithms.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Proximal Policy Optimization Algorithms

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.692460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:738896cae4d14c6bcb70c2a3bd3c14a2e4ec29b0a1fcc9ca69a2671777bd92bd

Observation 1801ca8e-c178-4c98-996e-213fbce79f34 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.773103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:f89870c17a7fb2f5241b3b9d0f03c7dfeab2939ead659e0bdb1a6a7431ac27b9

Observation f275f9a5-465a-4df6-90d9-fb8f31b817b3 · outbound

This paper cites NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.733269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:12a2ab7e1a523ec7d31d2acb9551e479201bd78694e469c0a8b2181252821bf2

Observation bc652e75-f322-4d52-9d89-accc02eeafa4 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs HybridFlow: A Flexible and Efficient RLHF Framework

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.788293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:2a261603d387a191b6a9ff53e493e594b645607db95c0a7d33b315569a558013

Observation c5855360-b852-4146-a87f-bdbf1630bfc7 · outbound

This paper cites Improving data efficiency for llm reinforcement fine-tuning through difficulty-targeted online data selection and rollout replay.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Improving data efficiency for llm reinforcement fine-tuning through difficulty-targeted online data selection and rollout replay

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.823009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:906223f5d60a3d2fba32c9f72633fd25dc2e7b7243dd3313197784b8686b6bd7

Observation abce3304-3ba0-4a1c-bc56-d1841b7bd503 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.767751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:4e60c704e4e9dc7cabc364170993882fc8e9737222b9b78fd757d842b5532e7a

Observation 1e8cabbe-a2bd-4812-be4f-c4ce7ea66b79 · outbound

This paper cites Marti-mars2: Scaling multi-agent self-search via reinforcement learning for code generation.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Marti-mars2: Scaling multi-agent self-search via reinforcement learning for code generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.748091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:b6facc4c8b56a500f965465e71dc2b1deda6c97ccbd5c65d2888c3ac356a4999

Observation 6efcf597-9f36-4c42-a2b2-993d3c4866ac · outbound

This paper cites Rlanything: Forge environment, policy, and reward model in completely dynamic rl system.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Rlanything: Forge environment, policy, and reward model in completely dynamic rl system

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.835972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:e03cc272805f40872df74ed13c118bc0133d8e336486cf45966b1eb983faedee

Observation e54ac121-514a-4c13-9bb0-2a88f1b536e7 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.865329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:799e7f34fd018d9782f57cb2ab2a02b07ab688711fc26d5cb86e37ae13423e04

Observation 591e94f5-fa5c-4e94-ad96-9e2b90393888 · outbound

This paper cites Autogen: Enabling next-gen llm applications via multi-agent conversations.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Autogen: Enabling next-gen llm applications via multi-agent conversations

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:59.194486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:ed1a4ec6e64f985bb721923fefa4e75060b9e6a8d4b45c8a3dd341c291af0a73

Observation d4c13c69-fe25-43b0-a6fd-5dc55bdb2550 · outbound

This paper cites RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.876213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:9f8d253f23a7a2ff515af28837aec77de77f20fec91a924d15efcd828bca5896

Observation baca4235-a3ab-4962-b775-976bd4d23c9f · outbound

This paper cites Less: selecting influential data for targeted instruction tuning.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Less: selecting influential data for targeted instruction tuning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.752068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:ea0537ed02f3171e3700ebff871e0198e03625f1f49d4a9ad0a4a4eb719abffa

Observation f0910be4-e456-42bb-8fd2-ea4a3144cab1 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.844864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:b2c12a45e6b46bba72f7d1c31a57dc7086d6537ecf90b98a9293e4d1a458dd15

Observation b9495583-66d6-4193-92ef-947a9c18d515 · outbound

This paper cites Areal-hex: Accommodating asynchronous rl training over heterogeneous gpus.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Areal-hex: Accommodating asynchronous rl training over heterogeneous gpus

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.831396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:5820d5cb3953851131eab63c28038de3e56a2c89109d17cc6620b229d2b21463

Observation 7959dc1c-335d-4f0f-b6f6-a2b6ee583d59 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.826382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:82ef23c8ff39d8297b10ae3aecb3236370a848b3f6998a57bb6ed85b2a6ed7a6

Observation 04a651c2-4e2f-422b-8f0f-a6280f66e224 · outbound

This paper cites Qwen3 Technical Report.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Qwen3 Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.854447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:9499c7fec4f5630bee896cc109ef57215c6219414f054faf89831a78890ba515

Observation b90d3184-8539-48d4-9d1e-b5469aae3006 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.880509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:0293bd83c89f9c82301ca20791b07261b953c312733e1cdc9fb615f7b2f82fd7

Observation d207b28e-0eb3-451d-b031-739af31a52ac · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.806981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:6149be844e1161b341d66967ba713360bd8abd8aea4e1bf1bb554504523b32cf

Observation 276b815a-5661-4bb8-8f3f-9010f045d6fc · outbound

This paper cites Agentrl: Scaling agentic reinforcement learning with a multi-turn, multi-task framework.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Agentrl: Scaling agentic reinforcement learning with a multi-turn, multi-task framework

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.793461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:dfc92febc984a74f81f3705eb9333aecce2a2f190054d1a60e85e16daadad454

Observation fa48a93b-bf3d-4aec-8f05-28af6dfba6da · outbound

This paper cites Stronger-mas: Multi-agent reinforcement learning for collaborative llms.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Stronger-mas: Multi-agent reinforcement learning for collaborative llms

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.803157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:e0027bed121f87df10ded973d85d9ccb330e5271c7f4cc2a30635c448f108b5a

Observation ff5148d9-995c-4f9a-b5b5-c0c504056ea1 · outbound

This paper cites Group Sequence Policy Optimization.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Group Sequence Policy Optimization

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:13:43.783001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:11fea40d75e243c9c30495df077f8b350be744895155c8138d635a5f35be9120

Observation 72467af6-8d03-4f29-8f8f-1424b2767578 · outbound

This paper cites Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:13:43.788110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:d671d5ad2071cb3f95033c68a6b84080f25aa961fdd6d74e3e3b6283c20362ff

Observation a1073676-9eb1-49cd-a3d0-2c0b44902166 · outbound

This paper cites Prosperity before collapse: How far can off-policy rl reach with stale data on llms? InICLR.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Prosperity before collapse: How far can off-policy rl reach with stale data on llms? InICLR

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.848514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:521a0c74909ca4c5330c6fc6b0f912e9c1ccf8ea35484417dcfd1de434ab9969

Observation 2727de08-137d-4805-96ef-5bd097032a20 · outbound

This paper cites Deepresearcher: Scaling deep research via reinforcement learning in real-world environments.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Deepresearcher: Scaling deep research via reinforcement learning in real-world environments

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.839196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:d9400595a22f997a05fcb325d95a8c04924becc3d74879a9daf8c6becea0f58f

Observation d060c8be-386b-4cbb-be5f-fad2cb4be93d · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.778925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:8acb4bcd9f2078464fa178dfe92badd0ffee5157b470da84c4833ac759d0ecb7

Observation e6ca8ac4-376e-4d69-8f72-b591ed56edbf · outbound

This paper cites Optimizing {RLHF} training for large language models with stage fusion.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Optimizing {RLHF} training for large language models with stage fusion

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:59.182924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:c05c1b2b596f56b6387f0c729e03b2222116c29fbb7f891e386d9c716ef4588a

Observation f129269a-ebc8-4592-b110-51497ffb580c · outbound

This paper cites Let’s think step by step . . . \boxed{}.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Let’s think step by step . . . \boxed{}

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.842356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:422ac0041eb9150fa38ac0a0541baf5d6f7934ed58e327410c69f0c7d1544298

Observation 97000aee-cd75-4152-b2b0-26c0963fba8e · outbound

This paper cites Any actions except provided available actions will be regarded as illegal.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Any actions except provided available actions will be regarded as illegal

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T20:13:44.834814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:707a48c480d2dee821b6136ac8a1106b7d8bccf806c35e2f51a339d43aa64201

Observation ec4340f9-926c-4185-8846-9b8ccddb3085 · outbound

This paper cites an unresolved cited work.

AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-20T20:13:44.769678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T20:10:32.300423Z digest=sha256:175c188fae6b0c25d194e193677d03c40bd840ef11be64a933cb7833b13f27b5

Pith citing papers

No inbound Pith citation observations are available.