Pith. sign in

Paper Citation Record · LEDGER

ToRL: Scaling Tool-Integrated RL

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 85 inbound Pith citation observations for arXiv:2503.23383.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.23383 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 85 of 85 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:43.322456Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c0ec26f3-934f-475c-a272-d028e2921326 · inbound

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs cites this paper.

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs ToRL: Scaling Tool-Integrated RL

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:42:39.069463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T18:42:39.023650Z digest=sha256:1865189d880f963aaeaf09f9c3323722989182d4e381d03b68ada7c3935dd08d

Observation 25e6e5a0-2022-4752-86d8-c3a998fcf750 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering ToRL: Scaling Tool-Integrated RL

Reference 181

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:43.322456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:43.322456Z digest=sha256:e266a5ee9f3e925f58e23534065cd4645087ace8c10eeb9a68e8110e2e74ffcf

Observation 8f1ff404-6ca0-451a-b0be-ea32b7af48b7 · inbound

AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset cites this paper.

AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset ToRL: Scaling Tool-Integrated RL

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:58:59.031839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:58:59.031839Z digest=sha256:99677d3e6991411d2eebd835d4c1370041e5774ceab7740840f4bf6df4c13a5b

Observation 7df66722-4f73-4a07-886d-78b479392da2 · inbound

WebThinker: Empowering Large Reasoning Models with Deep Research Capability cites this paper.

WebThinker: Empowering Large Reasoning Models with Deep Research Capability ToRL: Scaling Tool-Integrated RL

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:14:25.328698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T19:14:25.283645Z digest=sha256:a1e7c5677589fa839b01eef7653d8b9b6ea8c89d6c1c472e168d9b65123da290

Observation 6c980726-5d3a-4f9d-9940-6fce536c2387 · inbound

Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving cites this paper.

Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving ToRL: Scaling Tool-Integrated RL

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:54.092671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:54.092671Z digest=sha256:efb84a2475539c926d57f284db7b55cdf18e50a93135c0a67ef56bf785292603

Observation b2062e6d-11ff-4ab5-bf6d-f0fe946c1bf1 · inbound

Time-R1: Towards Comprehensive Temporal Reasoning in LLMs cites this paper.

Time-R1: Towards Comprehensive Temporal Reasoning in LLMs ToRL: Scaling Tool-Integrated RL

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:13.059774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:13.059774Z digest=sha256:52c1b572b053aa02aa02b241495979fd76cdab13df6c48d07f473184621e3ce6

Observation 242dc923-cbeb-4a5f-93d5-edf32159a6d5 · inbound

Visual Agentic Reinforcement Fine-Tuning cites this paper.

Visual Agentic Reinforcement Fine-Tuning ToRL: Scaling Tool-Integrated RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:27.826306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:27.826306Z digest=sha256:27781c3390d026af6e645bffd25387d2c77b90f44d345876c6761b81b0ee6641

Observation 34f7b3a3-7b5c-4681-948c-188b4fc6f61e · inbound

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning cites this paper.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.394679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.394679Z digest=sha256:b37c06378c58cee9f21d2086ab84cb6bedbf34e1f739de8af8d36f97f13ee4e0

Observation f88a0921-bfd6-4689-b5af-88624a1ad566 · inbound

DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation cites this paper.

DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation ToRL: Scaling Tool-Integrated RL

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:16.110667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:16.110667Z digest=sha256:58433bf5cc6e28c5afce9aa085f5fd615cb94326809050c089549c6f139cccee

Observation 79512292-bbd6-4bc9-af93-958343527375 · inbound

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking cites this paper.

SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking ToRL: Scaling Tool-Integrated RL

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:19.031069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:19.031069Z digest=sha256:fe743607c143681975392d7320d9ec0c6d21894fccc6994edecf429682c19184

Observation 5ce54d20-9228-4fce-bc41-0c6183dc2f59 · inbound

Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers cites this paper.

Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers ToRL: Scaling Tool-Integrated RL

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:20.079725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:20.079725Z digest=sha256:f1c754b777dc0e8e30f0162ce10cc91aae9d567baee1e7e5b0fb3cc7304b9fbc

Observation 7dd91fe7-fa36-4dd1-b7cf-d4b0f6f84d79 · inbound

Reasoning LLMs are Wandering Solution Explorers cites this paper.

Reasoning LLMs are Wandering Solution Explorers ToRL: Scaling Tool-Integrated RL

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:12.497316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:12.497316Z digest=sha256:93af17be1729027028f68ecd80eeb1b1700623cfe9a6aa3ad39a318cad4aa586

Observation 91035334-31c4-4bf9-8900-7c926ce8639f · inbound

Towards Effective Code-Integrated Reasoning cites this paper.

Towards Effective Code-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:25:26.808955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:25:26.808955Z digest=sha256:61f3720b4b21d9c8f79cba919e06384b7206e988f340b3a23eb10926c7d70043

Observation 0b82f9e9-4aab-4d41-aa95-17b48f455999 · inbound

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning cites this paper.

Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:37.503027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:37.503027Z digest=sha256:d29e503b2d051c51293b6ca27dbde90e113ebd2fc24f3dc45ab1e433a8e234dd

Observation f7e5c7f2-04d0-407f-88fc-59ac5af16301 · inbound

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents cites this paper.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.665196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.665196Z digest=sha256:35c819358caf86494f908ffdb8b932d195102af0ff176ae3850af81cd5b529b1

Observation aca1611a-9a7e-42cc-aac2-3e1b73269c53 · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning ToRL: Scaling Tool-Integrated RL

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:47.327873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:47.327873Z digest=sha256:3631bc5a051a4907a82eea256a80c3b64032003d3a6f8f446e206214922b2635

Observation 215f307f-beb2-4e89-8e06-91b84257bee5 · inbound

VerIF: Verification Engineering for Reinforcement Learning in Instruction Following cites this paper.

VerIF: Verification Engineering for Reinforcement Learning in Instruction Following ToRL: Scaling Tool-Integrated RL

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:18.797764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:18.797764Z digest=sha256:ca6d6dde3dd82695211ff1543ef6844ebf531565ab7a9b2ad1ea0d75fd044ca2

Observation 601b0de8-2bed-4708-a7fe-8102af2e7d01 · inbound

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications cites this paper.

A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications ToRL: Scaling Tool-Integrated RL

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:20.031722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:20.031722Z digest=sha256:6864db7df55075ddcb0af413ec3e4a8fc1bcab0f567f695c817d639a8b4479bd

Observation be98b4be-dba1-44c8-bb83-0f574e20da63 · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning ToRL: Scaling Tool-Integrated RL

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:31.307164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:31.307164Z digest=sha256:31a48dd040846850befc095e0e64b828c286b846af90a8788a182b6598041833

Observation 7c9e36bc-fb98-4b7f-90c5-0b5834acc2fa · inbound

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges cites this paper.

Measuring and Augmenting Large Language Models for Solving Capture-the-Flag Challenges ToRL: Scaling Tool-Integrated RL

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:04.518537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:04.518537Z digest=sha256:6845a4a04eab75510052caec414c4ad37e607811b761ec68740c3961f651b1c0

Observation 5d1b38b3-89b2-4ed8-a746-8d265c087767 · inbound

Distilling Tool Knowledge into Language Models via Back-Translated Traces cites this paper.

Distilling Tool Knowledge into Language Models via Back-Translated Traces ToRL: Scaling Tool-Integrated RL

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:58.308371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:58.308371Z digest=sha256:b2dcc8cf7a94555dd9afb62243169c9e4b90a46d54ba5872c466588c8c060a67

Observation 380fa6b6-651c-49c2-94f3-e4a6ef083491 · inbound

PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization cites this paper.

PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization ToRL: Scaling Tool-Integrated RL

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:16:12.930533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:16:12.930533Z digest=sha256:97c3e02413ad77f61f62561f4c1d822a6995402e47e93bac3aae3c4b8b6fd590

Observation 05729999-5f89-4d04-a2cb-2cda13c47ab4 · inbound

StepFun-Prover Preview: Let's Think and Verify Step by Step cites this paper.

StepFun-Prover Preview: Let's Think and Verify Step by Step ToRL: Scaling Tool-Integrated RL

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T13:47:38.217893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:47:38.217893Z digest=sha256:8def0eb21756fa2ebe808860d6ed4144b7409b4eab16779f971cb98975e4e1dd

Observation 4c924188-6d0e-4605-9afb-89eb632f15d7 · inbound

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning cites this paper.

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T12:23:55.175839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:23:55.175839Z digest=sha256:030b96612676f99a91cfc073b6cbcbcd8c97ea4c3716146ee3c9ece89f06c0f3

Observation 77514045-5d44-4f40-be4c-22d915a50560 · inbound

UserBench: An Interactive Gym Environment for User-Centric Agents cites this paper.

UserBench: An Interactive Gym Environment for User-Centric Agents ToRL: Scaling Tool-Integrated RL

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:26.376258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:26.376258Z digest=sha256:c7ef82009e23caf4a66059739c22bd4818a12fdad6927419d34b692685651215

Observation 7a0b5e96-f4a5-4364-a336-25872a9514fc · inbound

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection cites this paper.

Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection ToRL: Scaling Tool-Integrated RL

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T00:03:29.999533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:03:29.999533Z digest=sha256:b1967279fc0a3d47fb528aa7b761a4de48033b1b50fd83973367cfb9e323fc2a

Observation 243c78a3-5ade-4f3d-b10a-0bff989019c9 · inbound

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance cites this paper.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance ToRL: Scaling Tool-Integrated RL

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.597081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.597081Z digest=sha256:2f9eb6ea8de67b58c0d3cd83af8fa0b0b3e07db40fec187ca74c8011088c0cdd

Observation 0824de45-a673-4ba2-a2a3-3a113de407a4 · inbound

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL cites this paper.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.485908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.485908Z digest=sha256:2ae277ca222d42ee528075048adfc39804529ec702fcfc43b66c51bd8bb058d1

Observation a7d18c52-045b-4390-a996-eb97c0cad2ef · inbound

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use cites this paper.

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use ToRL: Scaling Tool-Integrated RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:22:51.584383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:22:51.584383Z digest=sha256:4de485abb61cc75b0cc6d85fb52e7ba45743203f75c736061b437c9401a264cd

Observation 6a3e08fd-da83-46f4-8a76-dfd42cea623b · inbound

Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction cites this paper.

Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction ToRL: Scaling Tool-Integrated RL

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:16.817954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:17:16.817954Z digest=sha256:4b894d4c1f9b4ee0f6423d6318c70df33de7e92f87319ddfa4ff87b218e8d9bc

Observation f5cc4b08-122f-4dbe-8bac-4bcd6feb5e3f · inbound

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning cites this paper.

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning ToRL: Scaling Tool-Integrated RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:14.848125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:44:14.848125Z digest=sha256:2da7b0cfd0bc9d8204aaf797408a38475c9fc234961a695993f72c3b05f96cb8

Observation 0b4a37c3-6a47-4f09-a775-3a90ffedb885 · inbound

rStar2-Agent: Agentic Reasoning Technical Report cites this paper.

rStar2-Agent: Agentic Reasoning Technical Report ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.427791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.427791Z digest=sha256:bfb19538786b93d48c9b91003e5865070edc9bb013fb59a0f4642e81cf3ef9f3

Observation aa12c784-d0c1-47d4-9e2a-41d1e08b1339 · inbound

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning cites this paper.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:22.953751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:22.953751Z digest=sha256:5d80cd430f81253aa70f055c2b0c61d73859f31c0d58e0ef7a577457eda8a04e

Observation b9f49a87-c1ee-4848-adad-06a39be94915 · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:13:59.036699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:5395cfafcca43fb4d70b505632d8e209158b6576c6931574fdb431b49b83ac1c

Observation 7d4cf452-a9e5-41dc-835e-67e5a5324a0f · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey ToRL: Scaling Tool-Integrated RL

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.647828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:f2db5eda921f2335cd5facd9f9fb205edbe0c03eb6512623a3417ab8d8ffa016

Observation e44d5100-5f37-478c-8a4d-474f38a0c806 · inbound

SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents cites this paper.

SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents ToRL: Scaling Tool-Integrated RL

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T23:55:38.101554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:55:38.101554Z digest=sha256:6d79257bbcf4ee299ef654c9ddf72809535390058ac1c3481fc0b7555ca343e5

Observation 70fe5237-d565-4d4d-8b25-e3e90c7d63e2 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models ToRL: Scaling Tool-Integrated RL

Reference 290

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:24.780909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:cbf06f6e7e41aac29be4cdfe514cae9944594bd3de523414fe15648605314b83

Observation ac6d0e7b-4034-494d-b8cf-bf0226db0283 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle ToRL: Scaling Tool-Integrated RL

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:33.433285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:33.433285Z digest=sha256:d2190a6d78783804362da4ae488b1d02cff591066f9ffd2a20884af1dbb9968b

Observation 4017163c-30ac-4cba-9d59-a1bcd7e95c4e · inbound

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions cites this paper.

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions ToRL: Scaling Tool-Integrated RL

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:01:31.394429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T15:00:51.162221Z digest=sha256:55cc1f9bf9d32e75617fe79722d308a484ef7889a64fe912058230ff18f7e63d

Observation f3054fae-3073-4eae-82ed-0095c656cfff · inbound

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination cites this paper.

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination ToRL: Scaling Tool-Integrated RL

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:10:51.356802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T04:09:50.183494Z digest=sha256:82310b1581291435c02592d4fe055f688ca49621bb4e8e3225af79219dfdb38b

Observation d543d618-6245-4541-a094-87a016794b19 · inbound

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving cites this paper.

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving ToRL: Scaling Tool-Integrated RL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T06:40:24.762093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:40:24.762093Z digest=sha256:db3091e7d5dc9d7e8d0740d4c30e5f5eddd88ff7d186b8a7d3237d4bd99a0ef0

Observation e8fec37e-6cc5-425d-950a-b976230eb3af · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation ToRL: Scaling Tool-Integrated RL

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.934885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.934885Z digest=sha256:beb7c00cefa1be60c1a5cf62d54440eddfbba18632cdac587a585037764c06a8

Observation 4fb04403-fcde-49e9-9366-19131b72d040 · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation ToRL: Scaling Tool-Integrated RL

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:51:40.765848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:d7900bc6edcb8e32d42b4b0fa9bf2940f29d15df2c7294b2a144df1d3727e60b

Observation 8ef8c2e9-57e1-49e3-8c4c-3e37788322b3 · inbound

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding cites this paper.

CharTool: Tool-Integrated Visual Reasoning for Chart Understanding ToRL: Scaling Tool-Integrated RL

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:08:12.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T20:07:23.153064Z digest=sha256:b892ab6c13e6a75b72bc1d4e5b718f63e3500efd95b4fff6253a3e23f9d7fa82

Observation 706f8afb-266d-4440-9298-f4ad4fa4182c · inbound

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning cites this paper.

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning ToRL: Scaling Tool-Integrated RL

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:10:51.855593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:40:04.944348Z digest=sha256:65c3c21ae016acc8b9bfc4b766867631c664f1e445d55582c78c8006717298d8

Observation 218d5a03-420e-463c-9b2f-0f0cf5ad4642 · inbound

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning cites this paper.

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:59.826281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T18:08:18.056525Z digest=sha256:aa58a1930fa27ec24b45440b2248d4c112da223dfb8023b2d5dd7f8642bc4bee

Observation 187e6633-ab0d-4b30-9cb5-7edc87a06037 · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence ToRL: Scaling Tool-Integrated RL

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:54.216722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:58f89fc78608c820d39864ddc48987b00b4635989eb44f0bb9760beb7dd063fd

Observation bcc06785-14f6-4835-beb0-663c880a8336 · inbound

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling cites this paper.

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling ToRL: Scaling Tool-Integrated RL

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:10.366674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-09T19:32:57.054584Z digest=sha256:726dea1e02b14888dc4d1167f53b14201bcd200535158a46e6cbfc7cae608878

Observation 2001a2f7-cab5-4e06-993e-153caaf1c8bb · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:30:58.761107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:2506b829dc4f5fa2ddf8e88d40a33d7e180e5dde7b2a25c053ca2076513d2abe

Observation 7a1abce0-7186-407a-84da-7f16f046f19a · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:45.032153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:45.032153Z digest=sha256:77caafefc9cad3eac9744adc76c58588209321764106fdcbdc45930410288664

Observation 8d5585b9-95a3-409d-81b0-3bb2dd92a7b2 · inbound

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning cites this paper.

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:23.221704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T04:42:49.165066Z digest=sha256:42207fd4cb33488e0ea4408d8704130dfff1cd9df20688fd01e035f7bea78895

Observation fa50e7bb-92cb-4f0a-9e52-486162a28127 · inbound

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox cites this paper.

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:23.444039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T04:42:47.496496Z digest=sha256:b91d1952b6f5e610e0cda3e8dc7f1dc397f6df57b739ea87073f037cf90d8bf7

Observation e38e9970-45c9-415f-8c2f-b14b8d94550d · inbound

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox cites this paper.

ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:09:51.260787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-21T08:08:04.544564Z digest=sha256:df2c27f5297420893de4142ff7450dac52ff6fcbfe879972938cb1e45e791cc5

Observation 160a545c-9bd6-4ab7-ba7d-e17491a1ca7a · inbound

Harnessing LLM Agents with Skill Programs cites this paper.

Harnessing LLM Agents with Skill Programs ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:28:14.372406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T11:26:19.382463Z digest=sha256:c4790a42915af087d3da49db66d2ddfa974cdfaff61638c231b587052bd4a48a

Observation d07a6f08-b72c-426f-944b-2693cdfc2f61 · inbound

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use cites this paper.

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use ToRL: Scaling Tool-Integrated RL

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.597051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T13:49:58.677299Z digest=sha256:e3cf56d6f81261039958c426d7efc1df0c6a7cc848d05aedb4f78fafb34baaaa

Observation 0cd270cd-f03b-4a05-b28a-9734d0724f15 · inbound

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning cites this paper.

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning ToRL: Scaling Tool-Integrated RL

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:23:24.097141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T12:22:39.655615Z digest=sha256:5e4279b881ec7835f6081b5b15b1c2c38e558fd7eeb4a1a565a4f35c6ca3fb60

Observation c528edd2-8f50-4d0b-aee0-6209ce4db83c · inbound

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning cites this paper.

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning ToRL: Scaling Tool-Integrated RL

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:23:23.709428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T12:22:39.655615Z digest=sha256:dbfbe5f434852be0074088d1ba51a45864d4c80393c1810328cbe16cdf4182ec

Observation 89728c42-fa89-46bf-94cc-d12ce2f5c9bb · inbound

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning cites this paper.

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:26.298738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T11:19:31.702516Z digest=sha256:b3127a799140a6f2e01f4536b66984d8e01844b497c80efd6eb0f2b9ae52cad5

Observation 3385bd01-37da-4c78-8fe7-45aac4978f0c · inbound

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating cites this paper.

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating ToRL: Scaling Tool-Integrated RL

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:17:08.740016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T22:54:44.329613Z digest=sha256:e8aebc460eef59d488ca700efe0153fab0bd2eaf7a62736a0f264495708fb52f

Observation 527b852c-249b-406e-bffd-26a8d7b36d08 · inbound

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs cites this paper.

Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:30.699154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T16:43:00.259139Z digest=sha256:658c43966e12181005e142dbe8c725e9edde102ac908f47eced8dfb47da2b58f

Observation 8813d5a7-129f-4fa4-8f6f-300c3a5cb865 · inbound

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation cites this paper.

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation ToRL: Scaling Tool-Integrated RL

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:39.261630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T13:23:42.745788Z digest=sha256:eba2557d25f2e2c8a86b104986ba40e6b2bb79550226fb5cd7b9e916cdea8147

Observation 89052348-3a9c-4647-871f-62b814bbd5d7 · inbound

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents cites this paper.

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents ToRL: Scaling Tool-Integrated RL

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.064660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T10:57:55.875707Z digest=sha256:a68c3e2807d149a0ea2c14c351cfebef5ebb84afc002c1851efd57c3db1e639c

Observation 8a2e1419-c0c5-4db1-9992-b40160bc9eb4 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization ToRL: Scaling Tool-Integrated RL

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:37:49.403443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:a6f24167392394f3ac0e645c81731936e4bab4417e145068043b4540a39d4091

Observation 73e0d108-c07f-4b76-95e9-cbeb0f34d7fa · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization ToRL: Scaling Tool-Integrated RL

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:28.554279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:28.554279Z digest=sha256:0e37e514b3b0841dfe3f0a26920772850f464861c2c2f22e5a347b9e6e1b366b

Observation f3056329-2a84-46e7-a6ac-8d65e3d85169 · inbound

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents cites this paper.

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents ToRL: Scaling Tool-Integrated RL

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:58:33.209084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T06:49:12.070481Z digest=sha256:8a74bd5d59ecb4264e43271955de3dd5172b28f344cb6e7e0e642547a00f58d6

Observation 206a8438-71f8-4e3a-8575-8313c113c327 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ToRL: Scaling Tool-Integrated RL

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.636889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:989def5e6bf51fcfee691ddc52d890e17df2a8c2fa651b2b0b9f2d874f6f635a

Observation 23264379-35df-4bd4-bf00-2e6d31f696a2 · inbound

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It cites this paper.

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It ToRL: Scaling Tool-Integrated RL

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:12.369980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-25T19:28:36.499352Z digest=sha256:5ab13ca286e1068c988a27c914263887a5767613603c4970bfdcdadd1a421305

Observation baa37537-78f8-422b-b35c-ec386e6386c2 · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:40.244054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T06:13:54.960194Z digest=sha256:2d4ac5618c66c2f4fd2ef5d9b8ea754137cc7bf115d87aa6626c078e876c66dc

Observation 6ebedfcc-fc11-4d90-97d7-91509a9fdb01 · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T07:12:34.663753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:12:34.663753Z digest=sha256:fcefa03ca806993115a66cb969c8f251b3ce4bbaa3483d373ab4ded024dbf3f6

Observation e3b45eb5-c77a-4597-bb91-d81577474950 · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:12.937563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:12.937563Z digest=sha256:3c8adba42bc0166a9e3aff81f6e9c278c11e2a4b52f7ee47daf393e1bac78b3b

Observation 85bdd02c-b90a-4538-b798-a0eae5c34095 · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T02:05:44.909805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:05:44.909805Z digest=sha256:cb2cdf9dfa736b403bf367132abc8965c9d1ad231c45ad64bf5e05e74ab27dd9

Observation bcee6077-162f-49d4-b9ec-4b8b61e6d5cd · inbound

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL cites this paper.

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL ToRL: Scaling Tool-Integrated RL

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T04:36:25.292766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:36:25.292766Z digest=sha256:5c2cac8b5298834246b010b9cf67db0eef3edccc55def4382c18f2285101c47d

Observation cbe8b569-abbf-4213-be6e-ee19d8d30bbf · inbound

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use cites this paper.

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use ToRL: Scaling Tool-Integrated RL

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.326575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-02T12:16:47.299349Z digest=sha256:a5a7892ea26bc769c083da8563ccf3fde66fd6e39661310cd788ae337ead7b12

Observation b57a85d3-e8bc-44e0-b596-f4f14bc7aa84 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception ToRL: Scaling Tool-Integrated RL

Reference 134

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:449017f3b90dee9ec6093a2be15d9734969a3a41723226aef05f206e9d3055ee

Observation bc658aa5-335e-456c-bbc8-569f751eb622 · inbound

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents cites this paper.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:dc6d8fc3da9975e6949d7ca328587b0dcbea19fc5cd0a3b72f95ff4b0b00ab51

Observation 012ef096-fe13-4c84-b804-db4c759c726e · inbound

Knowledge-Centric Agents for Workflow Generation in ComfyUI cites this paper.

Knowledge-Centric Agents for Workflow Generation in ComfyUI ToRL: Scaling Tool-Integrated RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T22:12:53.833936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:12:53.833936Z digest=sha256:70f3d2ace1a357523cd90f6d309456493aab2c5b31615585db620a7b12dad794

Observation e8d2fc97-d1e1-45c2-985b-edf2ec06025b · inbound

H$^2$SD: Hybrid Hindsight Self-Distillation cites this paper.

H$^2$SD: Hybrid Hindsight Self-Distillation ToRL: Scaling Tool-Integrated RL

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:49.091047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:49.091047Z digest=sha256:7ec231ec5f7058014aa6f6e8d8e6d5890f4295b7b21a5af80ca5837442a6141b

Observation f134a709-de0c-4f9d-9fbf-e0f22b86e1b4 · inbound

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents cites this paper.

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T22:56:53.271564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:56:53.271564Z digest=sha256:d25629b454d51207b5538343a8f445b44df57bb81f37e3f70b280bed26620a44

Observation 8b721a9c-415e-4ccb-a10e-af5f6890dec9 · inbound

TCPO: Turn-Level Credit Policy Optimization cites this paper.

TCPO: Turn-Level Credit Policy Optimization ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.645199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.645199Z digest=sha256:bc7bbcfd776d0d4fab2777ef1488facc47443a9dfebf7d7b21c3f6e7370abd73

Observation 05877cd5-5ffe-4f1a-9b12-286a45b46ca7 · inbound

Progressive Agent Skill Generation via Reinforcement Learning cites this paper.

Progressive Agent Skill Generation via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T23:05:27.937398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:05:27.937398Z digest=sha256:77ef6316fb31bf6d8c28cee169ded5a488d85ea21441a10543416dacc77bc430

Observation 927b7980-bf7f-4d01-b64b-5900562bc05e · inbound

CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents cites this paper.

CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents ToRL: Scaling Tool-Integrated RL

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T21:49:39.554840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:49:39.554840Z digest=sha256:8f3e30c042865bddd004a6042eed58c0643af1dd88b36d7b9098fde2cb0d1636

Observation ffc6c40f-2636-40c7-a5d1-2b2f38aca1c8 · inbound

Contextual Information Policy Optimization for Search Agents cites this paper.

Contextual Information Policy Optimization for Search Agents ToRL: Scaling Tool-Integrated RL

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:44.357730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:27:44.357730Z digest=sha256:e6f6d57b49038992707b7566d06fc104481837e181b36ef70f4e1a53cb7c7f28

Observation 8661dab7-fd78-44f8-99d9-6c81b7dbc4ff · inbound

Contextual Information Policy Optimization for Search Agents cites this paper.

Contextual Information Policy Optimization for Search Agents ToRL: Scaling Tool-Integrated RL

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T00:54:15.445890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:54:15.445890Z digest=sha256:288e3170777d2382acbf2646efb106392f28d4ccb383f050a8fdae91ee36c591

Observation 3c61b79c-030c-42a0-88ff-cbac19d0b4fb · inbound

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing cites this paper.

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing ToRL: Scaling Tool-Integrated RL

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T04:39:03.633077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:39:03.633077Z digest=sha256:d6688cfc75f421e019ea97db7f14056e2c9bd1d833f76cf0cccf388babb7251c

Observation cfeb87f9-c46e-4edb-8262-dfd60b8b77e8 · inbound

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents cites this paper.

Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T15:03:54.120450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:03:54.120450Z digest=sha256:2ae528acbddde96f42032ac76d1604ed6978a845ea5a6baf7510970d29f6efdc