Pith. sign in

Paper Citation Record · LEDGER

AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2509.08755.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08755 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:17:44.662944Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ab033493-353e-4c3c-86e5-a01590349bc5 · inbound

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails cites this paper.

From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:44:22.032163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T20:42:40.823721Z digest=sha256:e543b5a60e58594c4ef3902e8e5fd04c58934e94a752ad34d01cb29475d7c065

Observation 1bac4949-4614-4099-84cb-bf272862a396 · inbound

Graph-Enhanced Policy Optimization in LLM Agent Training cites this paper.

Graph-Enhanced Policy Optimization in LLM Agent Training AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T07:17:44.662944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:17:44.662944Z digest=sha256:2ef246daca982e4122bcf8d95c2f9812913ee67882df451d0710cc4b57d46de4

Observation 65ecc095-b3c4-4c04-8fd5-45bd71d0c2f1 · inbound

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning cites this paper.

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:03:04.350165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T16:01:48.789986Z digest=sha256:84e9fa145e0e635c276ff8ea1232a09cdcac8eaf1429c14d4147aaed27dd1425

Observation 521c8bd2-31d2-4ddf-bc5f-0bfa840077e2 · inbound

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents cites this paper.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:44.971735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:44.971735Z digest=sha256:b9d029b7ea1a80e7994be4641bc7d61430e9f73ac2852eeecc4701ffe186f356

Observation 62076d23-5f87-47e3-b40e-a118ffdc1d62 · inbound

HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents cites this paper.

HiMAC: Hierarchical Macro-Micro Learning for Long-Horizon LLM Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:36:28.302124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:35:45.900606Z digest=sha256:49fd60b6df7212340240dbcc4bca2efc2fa5a7b098e513b47de61581813a54f8

Observation 80e94209-9660-4f03-9d68-3a17f887eee7 · inbound

Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures cites this paper.

Multi-Agent Systems: From Classical Paradigms to Large Foundation Model-Enabled Futures AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:51:04.003863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T04:31:28.242097Z digest=sha256:e1e1d281134d490fefa1e91f3575cdb25be90c971f3c9f1e099d7a646ef467c6

Observation a0a00fb7-0848-4b85-af89-ff144e01af2a · inbound

Gated Coordination for Efficient Multi-Agent Collaboration in Minecraft Game cites this paper.

Gated Coordination for Efficient Multi-Agent Collaboration in Minecraft Game AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:16:10.668460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:03:01.533652Z digest=sha256:08993411290b197bb5b5c8985055087bcfeafe1c82380468d941f20662ca9841

Observation 3f2fe3ee-5043-4fd1-86a5-ce7f9e1e0052 · inbound

Ask Only When Needed: Proactive Retrieval from Memory and Skills for Experience-Driven Lifelong Agents cites this paper.

Ask Only When Needed: Proactive Retrieval from Memory and Skills for Experience-Driven Lifelong Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.266164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:07:50.798934Z digest=sha256:85647d99e5f7c9aab8e66e73ecd6f1636213dbfb4291cbe6e86d1b084b492811

Observation d56ae31f-1127-4d63-8d06-b6f79feefa8a · inbound

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length cites this paper.

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:40:40.777758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T18:13:25.735085Z digest=sha256:7d30853f19a4e3ff4586bc3e30f099d4358dbcc103b8bbbcfb84726ee8cbd46c

Observation 62f24da8-3e94-423f-9baa-ad1e2f5b43ee · inbound

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction cites this paper.

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:11:12.218032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T09:59:48.604813Z digest=sha256:962df6656db0ce9c0d140b4577960beffe6debe12466ad90941175540b1565fc

Observation 1e85163d-8550-442e-9751-e64bd6830b91 · inbound

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents cites this paper.

Weblica: Scalable and Reproducible Training Environments for Visual Web Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:25:56.957645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:25:44.578007Z digest=sha256:043b34eb25bf803bd24b7224e7a5242d2b87b4c36118e2afd00a82c55a90f5ee

Observation 8088afa4-dd0d-4fa3-8c33-a1292a4c46b5 · inbound

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems cites this paper.

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:27.547132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:19:49.062330Z digest=sha256:486b495da60b960d67596cd46de9d6b334be631350961eff6d983b3ea67dfa01

Observation a6a7d221-1648-47b6-aefe-2964ec339729 · inbound

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems cites this paper.

AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:25:03.885299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T05:24:54.265411Z digest=sha256:e81111257db23529d891fbb37c2aee1ce93d09a4f544d06a6ad9ec0c8ef5a78e

Observation f9d5f345-c44a-4e88-8985-0fe8803df340 · inbound

GRAFT: Graph-Tokenized LLMs for Tool Planning cites this paper.

GRAFT: Graph-Tokenized LLMs for Tool Planning AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:22:28.269254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:22:03.560420Z digest=sha256:303c9e7122081f9f7e3e482d32857a2bc74904b24e55e50fac7c20d3cccacf8f

Observation d857d5c2-67c7-46cb-a55c-bc15d944de6b · inbound

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control cites this paper.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:32:28.836944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T07:30:41.399083Z digest=sha256:db1042cdd4ba01f468336a89692c559d5c459e3a50a23bb3414f733eda171042

Observation 199a2e9b-eb9a-4a3a-940b-4e6a27353c78 · inbound

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control cites this paper.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:05.669079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T06:04:32.640299Z digest=sha256:1c6a91656bbbe995dd3e0732b0627310db4589465e6187683fef1be34833e1cc

Observation 72827b3b-f976-4a54-9e0e-970c19d50e4d · inbound

PriorZero: Bridging Language Priors and World Models for Decision Making cites this paper.

PriorZero: Bridging Language Priors and World Models for Decision Making AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:27:18.624132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T05:25:05.907123Z digest=sha256:bfd354f536f1854e50b13cb27646146cd1d6f19466cb31eb258f3152f0cfefe1

Observation 468889c8-af3e-4213-baa3-d9b4c468334a · inbound

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy cites this paper.

Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:48:28.618421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T01:46:24.724553Z digest=sha256:7e6f00f93914de0eb899054a7a032c1a50a2b0fd372e127e90dea854f1b35258

Observation e1ac94f0-db25-4541-9ebc-593bb43dfd4e · inbound

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning cites this paper.

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:02:42.395542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T17:58:05.817581Z digest=sha256:2b99eec480867d738dd1b494236c51ff1b115f684f80e4c84d5bf625ff2eeaec

Observation b59c1e80-b311-4087-a119-9dbbe090a57a · inbound

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents cites this paper.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.470067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:0a25e2248c4ce1b698a6861a570e914d8c51fc526a7fa5d905f81b7b571ee0bf

Observation 8d583fbf-e3c7-403d-b117-3b59177e3305 · inbound

SkillGrad: Optimizing Agent Skills Like Gradient Descent cites this paper.

SkillGrad: Optimizing Agent Skills Like Gradient Descent AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 1

Resolution
malformed identifier
arxiv_id, observed 2026-06-29T16:53:41.304135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T16:46:09.930635Z digest=sha256:1d282591a47984be03e9d892a6b90112b9400734423661e40871a3180595d8e3

Observation 6b1e739c-bb86-4664-9402-992f048bad50 · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:00af689f31138a7173242721fb6508a6bb69b140efe95e0f0d32af1b84545498

Observation 763e6f28-3062-4056-96fa-b9052d9b42fb · inbound

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training cites this paper.

From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T09:45:49.609534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:45:49.609534Z digest=sha256:b662e366ddd78adfcde7a70ce6c205f9bf9ce01c198220c08ab59f7e1c406a50

Observation 1a45e017-0649-429c-b45c-3054a1e94154 · inbound

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation cites this paper.

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:56:35.257327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:56:35.257327Z digest=sha256:5a35a66a7b662ade81d12d1af9a269f54f82578a840c9de11dc577cf005eedeb

Observation 1934517f-cac9-4a64-9ac3-74851e6dce68 · inbound

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play cites this paper.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:12.420962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:12.420962Z digest=sha256:7db67e1a1cdd7eda7e9d75f8e695b2d26d67c51ca1927557efa265ff40f885f5

Observation 4e2c48f9-e951-48b3-8d4a-71238d638059 · inbound

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution cites this paper.

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-30T21:12:59.428554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:12:59.428554Z digest=sha256:07e3c24107a0f9fed94e5ecf1e3a05b118998b85aab304219fddf97924dfae2b