Pith. sign in

Paper Citation Record · LEDGER

Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2505.07773.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07773 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:23:55.198073Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:29:57.263901Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 180f9b04-4a4c-4050-8278-f640fa91051b · inbound

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning cites this paper.

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:23:55.198073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:23:55.198073Z digest=sha256:ae777bb8aa3e7e5c4352ece952637f44ab18f1967b7e2d4747061f7af10d1a7d

Observation 82b74efe-b908-4be7-8a9d-f20ce2219763 · inbound

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL cites this paper.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:33.803769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:33.803769Z digest=sha256:d4ba6ebb09ebcb8b6aedd63226cc3442e19c8b53ff4d18d87a5359a57ddb352d

Observation 4e85ad13-c4e7-40ce-8ed5-8bdbda4a1f96 · inbound

rStar2-Agent: Agentic Reasoning Technical Report cites this paper.

rStar2-Agent: Agentic Reasoning Technical Report Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:39.441530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:57:39.441530Z digest=sha256:95619cbd03cf0cf6b63112ba227000569bc6a9b9a5e0a990fa0d6a0c9ce25b6f

Observation 8bb1be83-9bc9-4c92-9ae7-b325b209dc4b · inbound

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning cites this paper.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:24.061038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:24.061038Z digest=sha256:4de2db1b51c5e37a8bab79b68757f0dcfdd059e8406435eb635b3d28d02c6a6a

Observation 7140552f-ec9d-448e-97bf-b4347eeb36f3 · inbound

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search cites this paper.

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:17:55.696222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T01:17:55.500268Z digest=sha256:5c6b2cf308bfe92ba28a32e1e790a1e030367b52bc83c907dcee3358be91bcfa

Observation 06159f94-f92e-435e-92db-076035518def · inbound

ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards cites this paper.

ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:21:18.739433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T11:20:02.118579Z digest=sha256:3b341ff4ec693cebba53fbaa1dd1c9a59e295be29c701f4672603e3ea8816122

Observation cae61186-19c1-47b3-a9c8-511c1c15a7e5 · inbound

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents cites this paper.

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:12:15.363234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T01:10:50.325484Z digest=sha256:df4f9ee156ade64663e5a3d1be0855e71f93dd162e1470781ef1f589ef465dff

Observation 988f592a-a2e1-4807-a6cf-a8a1c1f34491 · inbound

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents cites this paper.

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T00:12:31.672833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:12:31.672833Z digest=sha256:ca4bc0d8ba23a06acab9bc835c57dc105331184a32c8dcfc1ee069381b7926e1

Observation 7dcb72e0-1c50-4f99-8526-881157608874 · inbound

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving cites this paper.

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T06:40:25.051327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:40:25.051327Z digest=sha256:4360dbc305b2d5b85c989e1a3881121d05f7e895fbe5de107cce278f1b2310ac

Observation 19d6a823-d3f4-43b7-b473-f5f0fbd076c8 · inbound

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning cites this paper.

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.500962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:30:39.500962Z digest=sha256:dbeda6a2f5c0987d81381ea657cf56bd6f2031d6687c4ff592c40fe5d0480a33

Observation b980d36e-f6be-41c2-9d23-73eeefa72421 · inbound

Making Expert Reasoning Learnable with Self-Distillation cites this paper.

Making Expert Reasoning Learnable with Self-Distillation Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T05:25:32.878045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:25:32.878045Z digest=sha256:bb108b7ebba69388d971d29c18db6fe9668c1045ebf3fcb3d80218332d9c7831

Observation a63896ad-167b-4433-a866-9d3e277308e7 · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:51:40.699405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:ddec8b55863f84b07df2c42cf0500b9ffac80d8732f67ca6273f10ddfae96e4f

Observation fb5d1019-2f23-4e62-9b69-c82919ee8576 · inbound

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization cites this paper.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:45:50.167268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:ba49049cb3b940b3ec7cd0118453b773170eb12ff5cfd42da7cb2060f1fc0ae9

Observation b733163e-6620-46d7-815e-d53f3b665ac6 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:29.709212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:2a064df6f9b4282602ab6c7ec1bfbe8f797fffef1159b78bee4fdb223cf4d55d

Observation 410b48f9-75f4-4f9a-bc58-6b4240c966b2 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:17.002277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:f9cfefcb89dd16e5670ac17b592232352f5f05478a4eb95b4994771d6cb473f9

Observation fb9f6857-d435-4436-8ddb-fc5d03233f04 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:02:40.823135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:891ede0c61a988b77d4bbbe24a6acc1ace82bb7ccd71b8d4835ab777d7266dcc

Observation 4c808a61-188d-4e55-9e86-c5dd66b4d398 · inbound

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning cites this paper.

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:01:23.195048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:42:49.165066Z digest=sha256:32f592173c819e47f7f3ada52c5395cdcb776e83a7b89ae638bc1ff167cee6ef

Observation e74073e5-3fd5-4790-919a-d63fab83ad43 · inbound

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents cites this paper.

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:17:45.059514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T10:57:55.875707Z digest=sha256:4cc99c5391e83e1e7f8503037099d01424894ba4dc48f7420747fc2c5f0caa8e

Observation f13c40ec-3a25-460e-8d45-abf23c9f8f8a · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.394598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T06:39:34.199607Z digest=sha256:b54bd8e356ea0eaea72be7d56fca0aa5c091813b186a35dec438ea617ac00e2f

Observation 8f44afc6-8f45-4f69-b31c-6d109f32a5c8 · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:21.353101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:21.353101Z digest=sha256:90ece30cab411773de8ec342250839991b93b8342f72a5f042292075743bd646

Observation 41f5a501-6ba6-4947-a74c-a9677759db78 · inbound

STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability cites this paper.

STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:29:16.433527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T21:09:30.899636Z digest=sha256:cb154646655ebe6e3c1747c6ef33e14c53e1af021c248a0b50ae3af51d7d6f36

Observation 9dd4607a-78b4-4fc7-8563-ef806333cb08 · inbound

Latent Visual States for Efficient Multimodal Reasoning cites this paper.

Latent Visual States for Efficient Multimodal Reasoning Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:29:57.265562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T00:38:11.619574Z digest=sha256:a7a00d49f3c4e372cd49224ea914e8be14fb7c8f320deabc131908bd6b4351ab

Observation 65a78ef7-3e5b-49be-aeda-577874f59027 · inbound

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language cites this paper.

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T14:04:50.220784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:04:50.220784Z digest=sha256:c99bc4c647c944fa42f368dd4115edb4407da5876ff272d939ced17f646daf9c