Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T11:39:25.500635Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 34 inbound Pith citation observations for arXiv:2509.02479.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T11:39:25.500635Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T21:54:44.263730Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
34 of 34 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation b5aff72c-2ec4-4548-be35-d4ad6a0273f8 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Towards Effective Code-Integrated Reasoning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a67d6f6b-1a0a-446e-89b8-17b1e5b55bcf · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Multi-Turn RL Training for CUDA Kernel Generation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 59660ec3-2511-4b83-9db5-747d7c676b81 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 566ac665-b736-4d69-bee2-e56aa708acc7 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99fb6fa1-2182-4737-b173-c0a7fc7bb149 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 062dd9b9-b253-4590-8115-89169a1bba7f · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Agentic Reinforced Policy Optimization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aca43fa9-b2ba-43ba-8dd0-fa865dff7efd · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb62df75-3fee-44be-a48d-9e4b13103466 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Dean, and Craig Boutilier
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e6b85788-a3ea-44f3-b61c-e93e64009369 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 16893c04-35ac-4ca2-8f93-79a6d4101784 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d60a3c3-8f4a-46da-9d69-0c34540f9b59 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f129fb19-4881-43e6-9880-950cf7834823 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa12c784-d0c1-47d4-9e2a-41d1e08b1339 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ToRL: Scaling Tool-Integrated RL
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95da2a5e-6f43-4d90-bc06-f783947f2a9b · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Logit Dynamics in Softmax Policy Gradient Methods
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 68007f4e-83cc-476b-a620-3e5e402546d5 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 656d61f4-b0e2-4ca8-a945-5f8df29c399b · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Understanding Tool-Integrated Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 185b89bf-3df2-45a3-ada4-b3de776752b5 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc9c65c1-49ea-4e09-a14e-86f4fcd3e5df · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa14734a-4921-4aa2-a9d5-daef1525768d · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8bb1be83-9bc9-4c92-9ae7-b325b209dc4b · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27094400-20a4-4f5d-8c11-137218c4e3c5 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Kimi-Researcher: End-to-End RL Training for Emerging Agentic Capabilities
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8b1e3065-ce6f-481a-a813-47345c27fdeb · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Proximal Policy Optimization Algorithms
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e244270-f369-4138-946c-852d0afebb79 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0900ce65-3114-406d-b631-cf22dd7b6e5f · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a308f37a-c074-4895-97cd-db536de8835a · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf7ddfff-1716-4aef-81f3-9a389bb5ebfb · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Divergence-augmented policy optimization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2beff123-71bc-4bc3-bd4c-76dbac3f7860 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8576bfa-7340-4826-bebb-5d6786874288 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Your Efficient RL Framework Secretly Brings You Off-Policy RL Training
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c9111b61-3ccb-4206-a04b-fe2f9936071c · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 741430aa-f2dd-4067-afb3-fb19b8582aa8 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7f2a4b8-a266-44fe-b9f7-48d2ba1cb883 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fd22fb0-2e4d-4989-8d55-aae81c84c9ca · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Geometric-Mean Policy Optimization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e524eb84-a13d-49ce-86c1-ab3c70b72080 · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Group Sequence Policy Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adb6998a-e01d-4627-9cd3-e36056005c2b · outbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b1a0ca3-49ad-4662-8895-f46c2393af3f · inbound
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 96af294e-368c-48d0-903a-3429d47b5ec8 · inbound
Training Multi-Image Vision Agents via End2End Reinforcement Learning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a612c8e4-589d-4e9a-8735-08555f0ef9db · inbound
SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fe72e3d9-1493-4d5e-92f7-5cd4f18923dd · inbound
SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccd8f0f6-98b2-4483-af25-c7093a70a153 · inbound
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7b0580b7-7b2e-4ae4-be77-75e1a5414e09 · inbound
MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 91ff23a9-3d9f-4891-9a74-03c8c64502b9 · inbound
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f1517022-504c-424c-bc17-2bc433818364 · inbound
AnomalyAgent: Agentic Industrial Anomaly Synthesis via Tool-Augmented Reinforcement Learning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 560baf11-4602-4904-9981-a6ab9380935c · inbound
When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a69f55b3-9496-44f2-b475-685ca459cb6a · inbound
E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6ddc7489-e2cc-4559-827b-17057b6bd84f · inbound
T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ac9ea6eb-31f3-4f26-b14a-8460bc5b76b7 · inbound
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f30c080e-4697-4ce1-bf3d-af7bb0848535 · inbound
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 31f4b51b-df37-4077-ba13-c1f579eaefd2 · inbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2bc2a4a6-fece-4205-bff3-57e46fc074bc · inbound
SOD: Step-wise On-policy Distillation for Small Language Model Agents SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2561b43-d11b-4d1f-9b53-3dcd18f5c7c6 · inbound
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7a237f99-e4d5-46ed-ab9a-7f6ec9deeed2 · inbound
TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation deb6548f-7b70-4c7d-94d1-dba84e7c51ce · inbound
PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8ac226f4-f045-465f-9c53-00abf4f637b6 · inbound
The Agent Use of Agent Beings: Agent Cybernetics Is the Missing Science of Foundation Agents SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cb0d95a3-6838-485d-8318-373958214abb · inbound
Harnessing LLM Agents with Skill Programs SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9d65937d-eaaa-40ce-970b-180b2e575a03 · inbound
Efficient Agentic Reasoning Through Self-Regulated Simulative Planning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 940e9d91-fd9b-43b7-853c-a6c43e595780 · inbound
Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b717725a-31f2-495d-82de-d44406034bd6 · inbound
Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation db400056-07f6-443a-8d66-87109d368a53 · inbound
Local Guidance, Global Impact: Gaussian-Reshaped Trust Region Unlocks Behavior Transitions SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ee508e64-b31f-46f0-a3e3-2603a54f33af · inbound
Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3489bb9e-b3f5-41e4-bbad-1e1de868f8f4 · inbound
AIR: Adaptive Interleaved Reasoning with Code in MLLMs SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fb20a44c-432a-4c90-98ee-f3ee07a3d67d · inbound
Latent Visual States for Efficient Multimodal Reasoning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 46ee27e6-2fc5-4a18-9928-240853ad722b · inbound
Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation aaa91678-462f-4546-be5c-1b43ee95d308 · inbound
When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 94fd6642-92ba-4341-9089-0f5ea2b83c00 · inbound
When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d27f17f7-783b-4aa4-aa49-d12de1ce277a · inbound
When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 513c8760-7796-4bf8-8b78-cc2040292517 · inbound
ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62823d65-5a71-4f60-9b0f-88342c65e2d2 · inbound
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b978fe9-44cc-42d7-996e-074b439cce26 · inbound
Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.