Pith. sign in

Paper Citation Record · LEDGER

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

As of 21 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 34 inbound Pith citation observations for arXiv:2509.02479.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02479 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:39:25.500635Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:54:44.263730Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b5aff72c-2ec4-4548-be35-d4ad6a0273f8 · outbound

This paper cites Towards Effective Code-Integrated Reasoning.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Towards Effective Code-Integrated Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:21.563335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:21.563335Z digest=sha256:36b64f556c5005266fe16c4efbe63d7e5f196da377d11ba3c000012c3bb759b3

Observation a67d6f6b-1a0a-446e-89b8-17b1e5b55bcf · outbound

This paper cites Multi-Turn RL Training for CUDA Kernel Generation.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Multi-Turn RL Training for CUDA Kernel Generation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:39:27.606728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T11:39:21.654132Z digest=sha256:e59525728e9ecaad99d8222b367768a6ffa60fce081ce5f9b51e5c09a555cc1c

Observation 59660ec3-2511-4b83-9db5-747d7c676b81 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:21.804526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:21.804526Z digest=sha256:37ced916322b3399efbc36d29c56c98f245754285268dc7782790cf92104b580

Observation 566ac665-b736-4d69-bee2-e56aa708acc7 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:21.969591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:21.969591Z digest=sha256:c61a653cdf0a946fb48d262e9d320044c3701328c9ab5ca87fb0a89692add3d2

Observation 99fb6fa1-2182-4737-b173-c0a7fc7bb149 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:22.034107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:22.034107Z digest=sha256:8021a493cf1895d31140470b0dc2d3faff933025ae3b9df42bfb9e6dc8049a7b

Observation 062dd9b9-b253-4590-8115-89169a1bba7f · outbound

This paper cites Agentic Reinforced Policy Optimization.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Agentic Reinforced Policy Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:22.139956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:22.139956Z digest=sha256:440131f1cff569239477baca8302525e09b1ee614004647480d94626f126df32

Observation aca43fa9-b2ba-43ba-8dd0-fa865dff7efd · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:22.278687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:22.278687Z digest=sha256:f84fd251cb4830ae20c80ceff76bccffe8c5de863abd65e0badfb3736cacfa12

Observation bb62df75-3fee-44be-a48d-9e4b13103466 · outbound

This paper cites Dean, and Craig Boutilier.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Dean, and Craig Boutilier

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:39:27.431990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T11:39:22.431007Z digest=sha256:8dd8038ebea0a28c40c89723085660259918ffb487350df411c52da8ebf95fc2

Observation e6b85788-a3ea-44f3-b61c-e93e64009369 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:39:27.219744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T11:39:22.510306Z digest=sha256:87dd85444d697044dc0b4f60a758f139fa4a8720662cd9d2bc8e4d7abadf9c16

Observation 16893c04-35ac-4ca2-8f93-79a6d4101784 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:22.593363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:22.593363Z digest=sha256:a33ab12f43dfbba260f1c96e2107d9af6c23c03a15d5afdbb6cfc6b529a8fae8

Observation 2d60a3c3-8f4a-46da-9d69-0c34540f9b59 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:22.710331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:22.710331Z digest=sha256:3f928fd4fc8a9242ff2f33f586a5060271213321160d67003d152f3f9a8c9418

Observation f129fb19-4881-43e6-9880-950cf7834823 · outbound

This paper cites CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:22.819067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:22.819067Z digest=sha256:2c50eb73bdad8bfa867611a334d553d333611d1f3ddc539656893f2bba1c0a58

Observation aa12c784-d0c1-47d4-9e2a-41d1e08b1339 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ToRL: Scaling Tool-Integrated RL

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:22.953751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:22.953751Z digest=sha256:5d80cd430f81253aa70f055c2b0c61d73859f31c0d58e0ef7a577457eda8a04e

Observation 95da2a5e-6f43-4d90-bc06-f783947f2a9b · outbound

This paper cites Logit Dynamics in Softmax Policy Gradient Methods.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Logit Dynamics in Softmax Policy Gradient Methods

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-05T11:39:25.983870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T11:39:23.104253Z digest=sha256:58798ddbdbec1240410f1fc875b8fa16247fc8a211562629a1cc26a8413d7988

Observation 68007f4e-83cc-476b-a620-3e5e402546d5 · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:39:26.978676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T11:39:23.251376Z digest=sha256:8090e0262a3e330983384cbac42dde37dfcea5e2ef3418ab0ed2fec7d7eff509

Observation 656d61f4-b0e2-4ca8-a945-5f8df29c399b · outbound

This paper cites Understanding Tool-Integrated Reasoning.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Understanding Tool-Integrated Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:23.377485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:23.377485Z digest=sha256:086596f9abdfb124197f1581b6f650320015dbbf908e4e068523dc84e3015be0

Observation 185b89bf-3df2-45a3-ada4-b3de776752b5 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:23.530235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:23.530235Z digest=sha256:37ad020901266df8e2a37a772ced3a107d46641fc3265fc700f4b1319b13c350

Observation bc9c65c1-49ea-4e09-a14e-86f4fcd3e5df · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:23.676782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:23.676782Z digest=sha256:0c83d5eaa7c389c7ff3eddb19cb560f1501b0c9074198e8f4e7299590cce7016

Observation fa14734a-4921-4aa2-a9d5-daef1525768d · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:39:26.799315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T11:39:23.885010Z digest=sha256:f50058b0f109cc76f2e5078703fd50f91ed3e81c1b1163cedbb826d9578d9273

Observation 8bb1be83-9bc9-4c92-9ae7-b325b209dc4b · outbound

This paper cites Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:24.061038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:24.061038Z digest=sha256:1a325aae1f623d8579bc8a25f86c73fbdd207c5f193fc9cc65dd12f154fd21e4

Observation 27094400-20a4-4f5d-8c11-137218c4e3c5 · outbound

This paper cites Kimi-Researcher: End-to-End RL Training for Emerging Agentic Capabilities.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Kimi-Researcher: End-to-End RL Training for Emerging Agentic Capabilities

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:39:26.623796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T11:39:24.223199Z digest=sha256:3669f9bd382f3c1f80d7cba302d7e3dbad199172431a2fdac6d0695814a1c9b9

Observation 8b1e3065-ce6f-481a-a813-47345c27fdeb · outbound

This paper cites Proximal Policy Optimization Algorithms.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Proximal Policy Optimization Algorithms

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:24.330070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:24.330070Z digest=sha256:20fef9025e67b41fa1ea5a15d49846f48d0910c0b9a9c6e18621bdb63907b124

Observation 2e244270-f369-4138-946c-852d0afebb79 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:24.436677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:24.436677Z digest=sha256:d270c7614e3523877e6ca08ce724b09ccd4b28403d45dbd92279aaca7cfd3a3a

Observation 0900ce65-3114-406d-b631-cf22dd7b6e5f · outbound

This paper cites Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:24.525311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:24.525311Z digest=sha256:b3348844dfbec0a0672ad2b3bab2d153c7d963f5b5fe5083724084105de39af0

Observation a308f37a-c074-4895-97cd-db536de8835a · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:24.627187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:24.627187Z digest=sha256:c5daabab714eadd76aeaaf0ec50962551743ed8aab3761b7685a040d487a777f

Observation bf7ddfff-1716-4aef-81f3-9a389bb5ebfb · outbound

This paper cites Divergence-augmented policy optimization.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Divergence-augmented policy optimization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:39:26.414168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T11:39:24.700391Z digest=sha256:8440688863f3930aafe8f186dc7ecfc565f216cfa4966b1fd49923ca2d6117a3

Observation 2beff123-71bc-4bc3-bd4c-76dbac3f7860 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:24.770094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:24.770094Z digest=sha256:972010e6f52e4eeea04f0991a5c42543c67d0be46287a235cb11bda9de7ea2b5

Observation e8576bfa-7340-4826-bebb-5d6786874288 · outbound

This paper cites Your Efficient RL Framework Secretly Brings You Off-Policy RL Training.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Your Efficient RL Framework Secretly Brings You Off-Policy RL Training

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:39:26.233721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-05T11:39:24.867017Z digest=sha256:d16aaf73deb088f47ac09d8e92329be6d909cc02b466e24010ad0d8f6a36b5b2

Observation c9111b61-3ccb-4206-a04b-fe2f9936071c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:24.939131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:24.939131Z digest=sha256:8e670b0e0539b2ba74b36617d7a519c90cd0fb5c9975f15f72085c41d14e330d

Observation 741430aa-f2dd-4067-afb3-fb19b8582aa8 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:25.032006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:25.032006Z digest=sha256:32715bf0c31688f1e89063b876365510346337c6d78066a3f9aa82428d7e3930

Observation c7f2a4b8-a266-44fe-b9f7-48d2ba1cb883 · outbound

This paper cites R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:25.091160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:25.091160Z digest=sha256:4e8a82525320954e65139abc08df16f8bd1b5682eabd681ac8963f1f084bdd20

Observation 9fd22fb0-2e4d-4989-8d55-aae81c84c9ca · outbound

This paper cites Geometric-Mean Policy Optimization.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Geometric-Mean Policy Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:25.219497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:25.219497Z digest=sha256:ff2061ca1b8868093410d1788a7447abea5475d2218127f59da1f25109a0c382

Observation e524eb84-a13d-49ce-86c1-ab3c70b72080 · outbound

This paper cites Group Sequence Policy Optimization.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Group Sequence Policy Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:25.423245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:25.423245Z digest=sha256:62a9441d03bd55d35e638de8972464116b85a3b6053376bd06ac72673571a07d

Observation adb6998a-e01d-4627-9cd3-e36056005c2b · outbound

This paper cites The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:25.500635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:25.500635Z digest=sha256:d33030c906173868498e6b86e9b48c84f97764d46cb899debd3e000e60267552

Pith citing papers

Observation 3b1a0ca3-49ad-4662-8895-f46c2393af3f · inbound

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search cites this paper.

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:17:55.589474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:17:55.500268Z digest=sha256:9ecacb70f26be6e7fcb94f377560d8a5b03f759b1310357f16f3980578844624

Observation 96af294e-368c-48d0-903a-3429d47b5ec8 · inbound

Training Multi-Image Vision Agents via End2End Reinforcement Learning cites this paper.

Training Multi-Image Vision Agents via End2End Reinforcement Learning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:01:24.304910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T00:59:28.618477Z digest=sha256:504eb586da7c475262a9620212d51fbd41af7458636a13a186d3365be384f44b

Observation a612c8e4-589d-4e9a-8735-08555f0ef9db · inbound

SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models cites this paper.

SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:23:09.869455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T17:21:54.516597Z digest=sha256:54726ade1386a32baf5673fb67bb7cbbf27a782c740f78dea5f77d4454da4d5c

Observation fe72e3d9-1493-4d5e-92f7-5cd4f18923dd · inbound

SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models cites this paper.

SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:18:10.591575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:18:10.591575Z digest=sha256:1054a2bb73243d353fb11203475018d81721d632d6ddcf98f5986ddacfe93392

Observation ccd8f0f6-98b2-4483-af25-c7093a70a153 · inbound

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning cites this paper.

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T16:40:22.776109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T16:35:24.557809Z digest=sha256:c165975f0a14ac28a2b4718f195daed92493344f9ffa7cb20166441a328e26f4

Observation 7b0580b7-7b2e-4ae4-be77-75e1a5414e09 · inbound

MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning cites this paper.

MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:20:17.554708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T15:15:22.055616Z digest=sha256:078c9b5eca6e9d2db37e8655aab2e76460d0b757fe9186cc2db126211954eb2a

Observation 91ff23a9-3d9f-4891-9a74-03c8c64502b9 · inbound

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization cites this paper.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.171881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:5f4e0f89460de7e512dd9952deede29cc21488b37fc31e80fd079ce393e7f423

Observation f1517022-504c-424c-bc17-2bc433818364 · inbound

AnomalyAgent: Agentic Industrial Anomaly Synthesis via Tool-Augmented Reinforcement Learning cites this paper.

AnomalyAgent: Agentic Industrial Anomaly Synthesis via Tool-Augmented Reinforcement Learning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:35:42.804115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:32:31.955785Z digest=sha256:85e9d074424bdd9c03efb7c234cf702312727d8cc0bdf1274affeb805a8a47af

Observation 560baf11-4602-4904-9981-a6ab9380935c · inbound

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning cites this paper.

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:51.983542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:40:04.944348Z digest=sha256:9eb2f2e526176a2487c4b4d7a08d6662d28d85b3c2e350bad4a545b70e3df8fc

Observation a69f55b3-9496-44f2-b475-685ca459cb6a · inbound

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning cites this paper.

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:59.842752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T18:08:18.056525Z digest=sha256:86c0cb005af58109130542f13c29dd3a15b30c275905a873e5b84b43a781ace0

Observation 6ddc7489-e2cc-4559-827b-17057b6bd84f · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:45:22.194748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:05dfa6580657e9625ebd9d54244ad5381a789178c503bbf54fdbb9df7fbdba99

Observation ac9ea6eb-31f3-4f26-b14a-8460bc5b76b7 · inbound

Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning cites this paper.

Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:01:10.534979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T10:30:36.330305Z digest=sha256:73099842770eb274bfa068b7f7f76190c54a94d02cd93be8d9fe681661081249

Observation f30c080e-4697-4ce1-bf3d-af7bb0848535 · inbound

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective cites this paper.

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:55.365079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T01:19:58.448344Z digest=sha256:651493d0af98d01cd6d829012ffb62f9c0971b8b122b2a127e031f3da96c4754

Observation 31f4b51b-df37-4077-ba13-c1f579eaefd2 · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:30:58.767379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T02:25:59.056181Z digest=sha256:998ccfdfe05e513efb185f9c73b8a3802033a73862d3ffed2e0ae394d1807aeb

Observation 2bc2a4a6-fece-4205-bff3-57e46fc074bc · inbound

SOD: Step-wise On-policy Distillation for Small Language Model Agents cites this paper.

SOD: Step-wise On-policy Distillation for Small Language Model Agents SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T05:20:45.024277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:20:45.024277Z digest=sha256:c7673beac395d8a55c66692e2f465d53e4eefc539348bc9b97f2f412ed0a4e6f

Observation f2561b43-d11b-4d1f-9b53-3dcd18f5c7c6 · inbound

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits cites this paper.

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:28.831917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T01:24:03.186413Z digest=sha256:1cfb84433ad819fbf9c3aec70fe854d5e4c2dfca710752ffc8f4c1ba6b6a3689

Observation 7a237f99-e4d5-46ed-ab9a-7f6ec9deeed2 · inbound

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning cites this paper.

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:51:18.092249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T02:47:31.830410Z digest=sha256:ede6f5db8e4b27de9253cc113361fc42ecfccd77e854e3bc1786fde28b170837

Observation deb6548f-7b70-4c7d-94d1-dba84e7c51ce · inbound

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning cites this paper.

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:01:23.336361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T04:42:49.165066Z digest=sha256:ea3e4bdc44eb383f95f14f9c395d995edff36d35d5c8fe1f63920f1410015193

Observation 8ac226f4-f045-465f-9c53-00abf4f637b6 · inbound

The Agent Use of Agent Beings: Agent Cybernetics Is the Missing Science of Foundation Agents cites this paper.

The Agent Use of Agent Beings: Agent Cybernetics Is the Missing Science of Foundation Agents SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:25.931625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T05:03:58.419364Z digest=sha256:deb0b64eb4a201f9fa3244e0b0f5521005e7433a056448336afd67e36372dd14

Observation cb0d95a3-6838-485d-8318-373958214abb · inbound

Harnessing LLM Agents with Skill Programs cites this paper.

Harnessing LLM Agents with Skill Programs SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:28:14.399273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T11:26:19.382463Z digest=sha256:1a9e380ce32d891f67dcf03142653d1594ccb84a887a1b99ef206dbc31f28a01

Observation 9d65937d-eaaa-40ce-970b-180b2e575a03 · inbound

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning cites this paper.

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:34:40.940904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T06:33:36.846345Z digest=sha256:c1a5d3d1f1d4f4b7e24fc6d19084fb9eb43d6bd22f22fef883de42010d0e5ad9

Observation 940e9d91-fd9b-43b7-853c-a6c43e595780 · inbound

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL cites this paper.

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:43:30.596549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-29T14:41:13.191919Z digest=sha256:9c224c35ccdeba8eedd03ff87b75861bf72fad3f3d0d5c20c9ef0e40c1a75012

Observation b717725a-31f2-495d-82de-d44406034bd6 · inbound

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning cites this paper.

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:20.272365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T14:59:57.983338Z digest=sha256:5a5e5e1d66e4dbc1e19c88f5f200b21fd229ffaeae799db2dd9c33c4d0d062c1

Observation db400056-07f6-443a-8d66-87109d368a53 · inbound

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions cites this paper.

Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:36:26.618510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:51:30.546134Z digest=sha256:57190463ea517934c3817fcccdfc87c1773e128fb28a90de5e62cfafb8322313

Observation ee508e64-b31f-46f0-a3e3-2603a54f33af · inbound

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation cites this paper.

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:07:39.263609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T13:23:42.745788Z digest=sha256:5ef32750a6a01355021aadaa65eda6777b2c6ff4120387628966550b5d5e1c19

Observation 3489bb9e-b3f5-41e4-bbad-1e1de868f8f4 · inbound

AIR: Adaptive Interleaved Reasoning with Code in MLLMs cites this paper.

AIR: Adaptive Interleaved Reasoning with Code in MLLMs SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:45.008651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T09:06:38.001604Z digest=sha256:1b27177a91f709171cc83beba5957b6afb16805ad50b9cca3d93ec124f0e2baa

Observation fb20a44c-432a-4c90-98ee-f3ee07a3d67d · inbound

Latent Visual States for Efficient Multimodal Reasoning cites this paper.

Latent Visual States for Efficient Multimodal Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:57.215761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T00:38:11.619574Z digest=sha256:86cd54f880253b4ba3356fc9e5c02ce50d58045cad4d1c634fb2a4f969db057b

Observation 46ee27e6-2fc5-4a18-9928-240853ad722b · inbound

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It cites this paper.

Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:50:12.363329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-25T19:28:36.499352Z digest=sha256:853eacd1c576e4275cc3d028b509f04617c439777c3dcd412c5ec9a2b2590bff

Observation aaa91678-462f-4546-be5c-1b43ee95d308 · inbound

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation cites this paper.

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-09T21:06:34.844654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-09T21:05:08.742517Z digest=sha256:b2fcf2fb275906fd325ff18e47d76ddf5daeae4b7b0c977cc8d9c97c2d84e4b1

Observation 94fd6642-92ba-4341-9089-0f5ea2b83c00 · inbound

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation cites this paper.

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T08:11:32.658488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:11:32.658488Z digest=sha256:a2cd06aaa814273c03397b9c78f6efa040699ada6bc2ba2ed8365ebd9ce07365

Observation d27f17f7-783b-4aa4-aa49-d12de1ce277a · inbound

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation cites this paper.

When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:28.548139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:31:28.548139Z digest=sha256:e73860f7ebd4f3f6b5bff970c50c09f0da80a953fb89b75c01d41cb2eacc6db9

Observation 513c8760-7796-4bf8-8b78-cc2040292517 · inbound

ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability cites this paper.

ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T06:36:47.880415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:36:47.880415Z digest=sha256:636a5e66f5555e78740c503c66108f91f0be672885c52876459cb5d0aea0e325

Observation 62823d65-5a71-4f60-9b0f-88342c65e2d2 · inbound

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL cites this paper.

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-02T14:50:17.781800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:50:17.781800Z digest=sha256:cd3e4beddebf26153109730ed2784f772cf649502b083a3ba58b4e85005e0de4

Observation 0b978fe9-44cc-42d7-996e-074b439cce26 · inbound

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning cites this paper.

Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T21:54:44.263730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:54:44.263730Z digest=sha256:5fd79e8485066f3741e61195112338bcd2aca827e5c20af7fe86d3ef2670787f