Pith. sign in

Paper Citation Record · LEDGER

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

As of 22 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 38 inbound Pith citation observations for arXiv:2505.16410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16410 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:06:08.486342Z

measured 120 of 120 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:26:56.575865Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

82 of 82 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved65
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a768c1a7-f587-45d9-afb8-9244aff5719c · outbound

This paper cites Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:02.921197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:02.921197Z digest=sha256:d685c3a5533fa1925f30a863ab56e20e926b1824b9eb880d578de01af27ffa23

Observation 9a52689b-1077-4ebc-ab31-0003e6de246c · outbound

This paper cites an unresolved cited work.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:02.965213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:02.965213Z digest=sha256:c18cfce1360dc5b37ab2f30bcedebb5d2568a55c7c8627047c571d4c7acdcfa9

Observation c5f52244-f3d1-4478-92df-621630b3f6d6 · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.057637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.057637Z digest=sha256:887bd9cc334609f86cb7344664a147a9fad5eb8adcecbccae92d5df395ab423a

Observation 08d50b1f-5230-48c6-9bf5-b7e9303ae65f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.258543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.258543Z digest=sha256:635f66d77976adc33324225c81bb09ce5aa89050360e3f7a2867152bcd41602d

Observation 6e0db75b-ba54-40fd-a9c3-04d5116bb108 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.338887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.338887Z digest=sha256:9e5c0b20bb8e5cb54fd87642ff68d33845c042c8c96d4ba52981e7a9a663d431

Observation 3f2a9b67-4572-412c-9ed6-5b8da5fb3627 · outbound

This paper cites Reinforcement learning for reasoning in small llms: What works and what doesn’t.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Reinforcement learning for reasoning in small llms: What works and what doesn’t

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.410193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.410193Z digest=sha256:fd4a1999c8ba13a2b1dd5386dbeaefb0d7d80c12fa7d6cd1d2530f118b6ef44e

Observation 9db8e6d8-9076-4853-be05-13d988fca364 · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.377945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:03.487236Z digest=sha256:993b1887e888afebebf4926e7970fae289a02bc187d5e21558c6d8060324c362

Observation 8a01c378-6eea-464c-945e-8b952fd37e6d · outbound

This paper cites an unresolved cited work.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.575092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.575092Z digest=sha256:99ab95bf076847c20c2966eb9523f05cb607cf53ac9f19f28d70b15d76dfae7d

Observation 463ec17e-ac30-4bdb-8ea2-81a9d406ba19 · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.644180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.644180Z digest=sha256:2d926dd6dfe73e4a44f5b5773f0829f298ff3b5293bd6b2cdbd55ec9844b5157

Observation 59afab32-de7d-400d-9f4e-bda9cdd99418 · outbound

This paper cites How abilities in large language models are affected by supervised fine-tuning data composition.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning How abilities in large language models are affected by supervised fine-tuning data composition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.722976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.722976Z digest=sha256:9526d28338ab3eb1a041163b5a83a13085e754df6a714262931d6dcb8c346e6c

Observation 534947a6-a771-46e7-89c1-5c0506bb8b8a · outbound

This paper cites Progressive Multimodal Reasoning via Active Retrieval.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Progressive Multimodal Reasoning via Active Retrieval

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.782011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.782011Z digest=sha256:26b7db99e8368938f0b570ed1910fc358f2ee5787370cbbe923c470f79a375b1

Observation 05b45d04-ff58-4c3c-9d94-db92a37592e5 · outbound

This paper cites Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.853422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.853422Z digest=sha256:ae5c04b5c4ca683e6b91828cb5c02775d20c7ce0f9c2b8f3b9be338669969496

Observation 3a8ccf65-cd6a-4971-93df-8ea0870efad9 · outbound

This paper cites The Llama 3 Herd of Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.926256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.926256Z digest=sha256:c6b06f07c394ecacc655394aaf2deb984f54a6530d8a2c3b3b3cc07447bfed1d

Observation 913923bf-d874-4927-8775-ad0923d608cb · outbound

This paper cites Concise reasoning via reinforcement learning, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Concise reasoning via reinforcement learning, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.356157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:03.998962Z digest=sha256:9eee895c0b5981210cca93d81cca874170f256d1f712c92bd0a2f165802a6d2b

Observation f7aea587-87f7-4ccd-9cd7-cf89a904e6da · outbound

This paper cites Retool: Reinforcement learning for strategic tool use in llms, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Retool: Reinforcement learning for strategic tool use in llms, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.079207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.079207Z digest=sha256:b5f37d5c2b51a65395cf356d0f79be213787c7a2a630978f320e67117b2fc9bf

Observation 23cafa84-7db2-4b24-bf89-2225134fdcf0 · outbound

This paper cites Tora: A tool-integrated reasoning agent for mathematical problem solving.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Tora: A tool-integrated reasoning agent for mathematical problem solving

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.339342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:04.133757Z digest=sha256:31650ffd121f21ac84ad71b3d93a939772f3d396bc575f05bf301873454783f3

Observation f0ec14f4-adee-4938-ae80-09f21d6f692b · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Measuring mathematical problem solving with the MATH dataset

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.244504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.244504Z digest=sha256:57d9e0b3eaa7356348b44d73fc178e893b0402f1683944f6ae67efd4071f3bb3

Observation 8458cacf-0f84-45b9-8de3-68df07e66012 · outbound

This paper cites Constructing A multi- hop QA dataset for comprehensive evaluation of reasoning steps.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Constructing A multi- hop QA dataset for comprehensive evaluation of reasoning steps

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.304466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.304466Z digest=sha256:36208069c607c7b494d6c03d0fbca79bd0e1c69119e92b711c3516cc8de1ac1d

Observation f8588b63-1591-4eb6-bc10-acab3f7332ea · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.343006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.343006Z digest=sha256:d762b91e5003121b6f1f8aeeddd2fb293589c4c6a9e16da6d619f36676e108be

Observation e9160051-d954-4e11-b638-b95197f87cbb · outbound

This paper cites Towards reasoning in large language models: A survey.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Towards reasoning in large language models: A survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.459705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.459705Z digest=sha256:545620587edb4f6aa4086308f2570d1d4a2dfc9402942b90ae8695295b6bc78b

Observation 3cddfc57-598a-4628-8a98-68a78fd6e4d9 · outbound

This paper cites RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.555452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.555452Z digest=sha256:087ec1866e0ec5cf18f488f2c77179d86a74a3e1bf5afd2e5dc56c459cea5802

Observation bffe5425-3732-4e3f-9c2d-fde963dae4b9 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.638038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.638038Z digest=sha256:2b5ed87afd95e2c8a7e6594c36f60d84563821643c5c715025ac1df80ed04478

Observation 83ed383c-4019-413e-a777-e91fd6877bd9 · outbound

This paper cites FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.709451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.709451Z digest=sha256:c0452662f1db5eb53ac1661270af1237f6558dd79792880491ed42c44568c06a

Observation 200f13a6-785c-47f8-b91a-5eff891f837a · outbound

This paper cites InstructERC: Reforming Emotion Recognition in Conversation with Multi-task Retrieval-Augmented Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning InstructERC: Reforming Emotion Recognition in Conversation with Multi-task Retrieval-Augmented Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.761722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.761722Z digest=sha256:dcd0a5a578b2be862a09d23022412a6aca33f0b4edc1c23a97fda1d68a4d898b

Observation 8416c664-e224-46e3-a43d-0bdebc16618d · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive NLP tasks.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Retrieval-augmented generation for knowledge-intensive NLP tasks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.311394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:04.856927Z digest=sha256:cff25ffc1a59c0653faefe5c702a27d102edc56a5fdecbf91509e10a9729afee

Observation daa241b4-b288-4055-81b0-54ead7f9de7c · outbound

This paper cites DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:04.931191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:04.931191Z digest=sha256:acf93bf7edab318ce689d1f1d682e09410878dd4e250ed38f4d2611e9b038168

Observation 25f7b39c-0cb8-4f13-a9ed-05c22fd440b6 · outbound

This paper cites START: Self-taught Reasoner with Tools.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning START: Self-taught Reasoner with Tools

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.048539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.048539Z digest=sha256:c3449d1ad6f707ee84761984a2b97a481e98ba217383645477bb112148aca5ee

Observation 283ff077-586d-43e1-8ac4-58ba3cee46c4 · outbound

This paper cites Chain of code: Reasoning with a language model-augmented code emulator.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Chain of code: Reasoning with a language model-augmented code emulator

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.302008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:05.120040Z digest=sha256:8b485b8bde326c4c287d48c922a10739f3ac281e404458984c51ce7c7a277962

Observation 57c7b695-3dee-4172-bfcc-48b6d2e5a6fe · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.190265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.190265Z digest=sha256:29e1a2744e1c01a395a452c13f3fe873468dc4300b721394d20aa542f862d249

Observation 04539f72-ae22-48f5-b753-e1a7d0010b5e · outbound

This paper cites WebThinker: Empowering Large Reasoning Models with Deep Research Capability.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning WebThinker: Empowering Large Reasoning Models with Deep Research Capability

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.232033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.232033Z digest=sha256:b126191901b88232c69b8cf41a810da8b8c806108bfd697b9a4c4db0d6b5d7d9

Observation 3575f901-c584-4012-9934-210449d77720 · outbound

This paper cites RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.273622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.273622Z digest=sha256:dfdeb9d35c199d093b449f65232829ce133be8015e8d84006b55fb64be133f84

Observation d0dcaede-b418-4ca1-974f-bb7ca676e1a3 · outbound

This paper cites LIMR: Less is More for RL Scaling.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LIMR: Less is More for RL Scaling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.361350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.361350Z digest=sha256:471db1c7fea2392c7f8285e8d8b3a3f9c882f237b581664431741a6c9ec667d1

Observation 34f7b3a3-7b5c-4681-948c-188b4fc6f61e · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ToRL: Scaling Tool-Integrated RL

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.394679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.394679Z digest=sha256:b37c06378c58cee9f21d2086ab84cb6bedbf34e1f739de8af8d36f97f13ee4e0

Observation 745a06c8-a473-4d00-abe6-0858cc0119fd · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.495850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.495850Z digest=sha256:a598a4a628783fa54fb071890189fc2a093ace687a34190144890ebaa67af7a3

Observation aea5fb85-94ca-4a23-9392-ebdc726e520c · outbound

This paper cites Let’s verify step by step.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Let’s verify step by step

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.292687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:05.585309Z digest=sha256:f535a56f7188110db97fbd26fe839f3e9952e92d7341c0ac8eb06cb2cc99c9a0

Observation 6b7fd29a-5863-4e42-96ef-e1de6797f489 · outbound

This paper cites OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.698406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.698406Z digest=sha256:93ead1e719964ca0aa3e70b15fad0eb873e85b2ad9b08c36d377dc4d1b55b15b

Observation 445726c9-8530-4733-8eca-b261283df9e9 · outbound

This paper cites GAIA: a benchmark for general AI assistants.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning GAIA: a benchmark for general AI assistants

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.280556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:05.804990Z digest=sha256:04072a90da54a2406f61d96b518d2bdb0433f84ddfd479c083a8c53620882d2f

Observation 93a9665b-d25d-40fa-bd45-651c6f9aef19 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.887519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.887519Z digest=sha256:3032b87aeb4541818d9bacf2094f7a893e2d5795f5afb7af32c74fdb70307cf1

Observation cf84693b-90e0-48cd-a54a-a6034290062d · outbound

This paper cites Learning to reason with llms, September 2024.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Learning to reason with llms, September 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:05.988114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:05.988114Z digest=sha256:af50e3c08bd478eb919bf75bf725e37a042e2b46877a66fcef658af96affd950

Observation 3144e645-50d1-48ce-885a-71094a43d092 · outbound

This paper cites ART: Automatic multi-step reasoning and tool-use for large language models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ART: Automatic multi-step reasoning and tool-use for large language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.095753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.095753Z digest=sha256:eebb625ca28c5eb84115706a5f5440c1165fe692ac940e5d301df3582c8580d1

Observation 5f3d20e1-635b-4c5d-a32f-33d56633289f · outbound

This paper cites Humanity's Last Exam.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Humanity's Last Exam

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.199430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.199430Z digest=sha256:bb695fb9181cced06b140f3b439d3449079e97505d5d4bf864ace5581fc524e6

Observation b8539a7f-3075-48f5-bbce-6105aff6cb86 · outbound

This paper cites Smith, and Mike Lewis.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Smith, and Mike Lewis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.237798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.237798Z digest=sha256:2d543372af12b6a1355d0247b8060c90ee144941f48d726ddccaa55a0edc2b05

Observation b5233d9b-b8cd-4656-88ee-e72349f1e852 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning ToolRL: Reward is All Tool Learning Needs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.338068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.338068Z digest=sha256:9236a4e7151fe0e9280add3f6b7cce7dd3aa3e9c3be965f498ba07c82a6fd040

Observation a7298ef0-0a11-46c6-a798-9ec61ac14a85 · outbound

This paper cites Toolrl: Reward is all tool learning needs, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Toolrl: Reward is all tool learning needs, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.259059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:06.394366Z digest=sha256:f5872389aca63524d8e1fe06b358360f232c7dd32e42f2c00e81fbd99f5d7f4c

Observation 0a372f50-8758-4b4f-b9b3-60ce4a8d6509 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.397967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.397967Z digest=sha256:483b418f18ab308da37c60e1380b12d950d201ef5b712bbfba8f1158d5d7c6c2

Observation 0adb45c0-3add-4d97-9c99-26536730064d · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.440420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.440420Z digest=sha256:db1d1618417b103b6553c03155213fb33000a5d1d453f9b4caf18ff1e675e407

Observation 2997ecbc-83d1-4c34-bf61-4ee3967d1677 · outbound

This paper cites Qwen2.5 technical report, 2024.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwen2.5 technical report, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.491523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.491523Z digest=sha256:9704e61c6132f2ef046fc054630fcbb50ed9dca893cfe5d231485db2370dd16b

Observation f7e10529-5f86-489a-8250-4815c38c2293 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.244120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:06.528602Z digest=sha256:a9d2bd00fbe43af78700a24f946e328802762ca727097f2a509816bd431bd43b

Observation c0e5f2f0-0796-41dc-aa5e-5b70c7d0bdbd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.594525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.594525Z digest=sha256:58e2895baf1bf2c7c448a6e60ec41352b05fd9ef8d8199b8ad162ff2c4385859

Observation 00a5d9b7-a885-40fe-b01c-681f3acdd004 · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.664381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.664381Z digest=sha256:b2c7bab28447390446475888183b81a6f551264b885350418321113e28220e47

Observation 93374031-0617-423a-82a1-ccf6d4d99499 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.718929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.718929Z digest=sha256:9fee945692bd48e832890ae09f861c312c946e40219913aa123bef1540191aeb

Observation 3fc324cf-9944-49e5-aea7-6b77dd04f0ab · outbound

This paper cites Curriculum learning: A survey.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Curriculum learning: A survey

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.233430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:06.770243Z digest=sha256:c3cfe9faf0fd4cd3c050c95777feb4046a8cc9238bd5972e6a5ec4f8ef7c34d7

Observation a1453fd0-e7d7-43f6-806d-6f5e890cfc88 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.794401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.794401Z digest=sha256:ca840055a13056dffe8531ca338d721ac247a531695d78ca77b3f881d9fbecee

Observation c7c6dfe8-d0b4-4b8e-8f90-86b78c232e7b · outbound

This paper cites Zerosearch: Incentivize the search capability of llms without searching, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Zerosearch: Incentivize the search capability of llms without searching, 2025

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.835005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.835005Z digest=sha256:1e1a999f81150338bbb05d946f4b47e9c9dfd5c07462435492a41c01134edcd6

Observation c6e6190f-f8ce-4ad7-b7b1-f53ae29ec45d · outbound

This paper cites A Survey of Reasoning with Foundation Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning A Survey of Reasoning with Foundation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:06.912225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:06.912225Z digest=sha256:166ea243045aa4eaf98976f735f9c515edb28eb3693263f493c68877a4ae899c

Observation 98a456f2-71ba-4ab3-acf7-9470f5701e26 · outbound

This paper cites Simpledeepsearcher: Deep information seeking via web-powered reasoning trajectory synthesis.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Simpledeepsearcher: Deep information seeking via web-powered reasoning trajectory synthesis

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.217964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:06.939970Z digest=sha256:2f68c91350fcb8733376ad57b5cfd5b9174a065aa3cf1649a0b88332498a15a4

Observation 2c61a8ef-fdc0-4fec-bd37-f51574ecfb16 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.075956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.075956Z digest=sha256:bdb77affc4abea72961d622b66b7c563fac20546a8b5111e5539d5dad1cedb2e

Observation c4530ae1-730f-487c-b9de-cd35b72a215c · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwq: Reflect deeply on the boundaries of the unknown

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.190061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.190061Z digest=sha256:8325ecf1ad974a7a6bdbd7925930c1221585ba576f32fd928663f4e6948fac11

Observation 8e086096-b12f-4f26-82b3-691746961c7b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LLaMA: Open and Efficient Foundation Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.292322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.292322Z digest=sha256:b254a492e18c2d0bf4a549aa5e9927fb71fe6820d5bc36ab46b802aa2dfcb6d1

Observation 41c61fc1-3617-4a5e-a0f4-9210c9366594 · outbound

This paper cites Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.353123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.353123Z digest=sha256:15ed40975dd07d234fca0944b916c66063c3834f06f388b658887dbe8b96407e

Observation 291d80f0-80dc-4ebd-a067-e19d68876a68 · outbound

This paper cites musique: Multihop questions via single-hop question composition.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning musique: Multihop questions via single-hop question composition

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.492226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.492226Z digest=sha256:f088ef323fb4770d1ac6fab07d712a67c05cfd6a8b201be2e31f73a072b8b72a

Observation e3927b1f-8cf6-41e1-a29e-5b85e26a419c · outbound

This paper cites Acting Less is Reasoning More! Teaching Model to Act Efficiently.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Acting Less is Reasoning More! Teaching Model to Act Efficiently

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.633692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.633692Z digest=sha256:1c75beb377bd0287c7b743e089a815f5cb027045d04f4579982c9a7974fbde3a

Observation abd5b097-0a92-431b-afce-308df4f64937 · outbound

This paper cites Text embeddings by weakly-supervised contrastive pre-training, 2024.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Text embeddings by weakly-supervised contrastive pre-training, 2024

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.793407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.793407Z digest=sha256:9394d44cb92c64da5671e133a62e587393d545d0f74b6dac15a68b3f6cffbfd4

Observation 6865c8fb-0fe2-458f-82c7-17d3d5577937 · outbound

This paper cites Reinforcement learning for reasoning in large language models with one training example, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Reinforcement learning for reasoning in large language models with one training example, 2025

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:07.888830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:07.888830Z digest=sha256:39ee6149b58cad4fee99fee3e800f93cec806db05b4f576af71bb5351502b539

Observation 1c6d98e7-73d1-4a9f-823c-a9a0497ebfc8 · outbound

This paper cites Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning, 2025.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Ragen: Understanding self-evolution in llm agents via multi-turn reinforcement learning, 2025

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.183390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:08.048505Z digest=sha256:6bedfc587334897b60e9c78fb3f84604768bda830aeac3635100fca26c763001

Observation ce961658-b063-4860-8900-c03507313a72 · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning WebWalker: Benchmarking LLMs in Web Traversal

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.136388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.136388Z digest=sha256:16e0459e7687f5313ec97a30acd8a13c4dfe8be5c8f2f1e2ac880fd150ac03f1

Observation fc269afc-78f1-4380-933b-e35f5fe7a0d2 · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.224735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.224735Z digest=sha256:0f78b0b997f5191f52a3aaf13d04030a28599d0367f405e5c62b5f0c42adc6ce

Observation 2fc8f876-fd04-4022-a5c7-510118aae02a · outbound

This paper cites Qwen2 Technical Report.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwen2 Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.283937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.283937Z digest=sha256:10c9474e91008f0534b104011393fd901278e80e940b3f8fa6f676a537f8514c

Observation b498f759-afe6-4708-9ac1-e06713c18d97 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.324488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.324488Z digest=sha256:2dd005374d92cff7d5df0d6517e68bff9605074cc9676bf2552f5b58b602b14f

Observation 307a74a1-7a9a-47b0-b6d7-5d55cb513d6a · outbound

This paper cites Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.363789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.363789Z digest=sha256:e105abfb6e71787fa70ea8f5b1f7c6970b05156ff014238197793acd3bb17b42

Observation e548cbf4-4dc3-4a94-8cd7-2d6e1c820185 · outbound

This paper cites an unresolved cited work.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.408176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.408176Z digest=sha256:a959bd7d4d5ff7a6bf8841bb7f65d0c0a897e768959669136dde71446cc97ab6

Observation 81248e37-f9ac-4e80-a775-969a6c60d58d · outbound

This paper cites LIMO: Less is More for Reasoning.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning LIMO: Less is More for Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.440907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.440907Z digest=sha256:8e3646094391d238412157eb62c042ba93e53371451a25a80cab41f3dca1a508

Observation d4d0aa70-6d7d-4dfd-a556-9709db7a630e · outbound

This paper cites Reinforcement Learning with Knowledge Representation and Reasoning: A Brief Survey.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Reinforcement Learning with Knowledge Representation and Reasoning: A Brief Survey

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.455553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.455553Z digest=sha256:1ca545816c6e817114ca3fac9a17fe3c57a04e9a0df8b3d4a70e109cbe88cd1a

Observation c60b9b13-4423-4230-8314-19ace93877e5 · outbound

This paper cites SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:06:08.651537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:08.459251Z digest=sha256:4cfe60bd0f085cded7212235f61c3fd057bfa4fd9e427aa6b6d915327c73f899

Observation 4ee74034-adfa-4c42-8c6b-7a8371bd8dbd · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.462641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.462641Z digest=sha256:8109e581c16863e106e1ac55584a2d8af17b07f71da5a7dc8a18498e5f79d15f

Observation 981b328b-5cad-4846-a3ff-9e4652bff27c · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.466277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.466277Z digest=sha256:f144ea2872825241cbb03b6bfcaef58ad31dcbfb2c7c5f19fa07c2c28fb685c0

Observation d58cba6d-1eda-493c-91ec-c37fb3e07075 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.469560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.469560Z digest=sha256:17f2156a5184c298aae4a55dd86f63938b4ff60d7c620af5bbcc52d86d41e0de

Observation ba4b92eb-7b3e-4d32-9517-c76a3d75a398 · outbound

This paper cites Agent models: Internalizing Chain-of-Action Generation into Reasoning models.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Agent models: Internalizing Chain-of-Action Generation into Reasoning models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:08.472839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.472839Z digest=sha256:42a0bf6a41cf8a930435451f1118289a623ac4803181b678f11b320a27dab5c0

Observation 23835f8d-1b01-4f55-8421-2fea2308e4d4 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 81

Resolution
malformed identifier
no resolver link, observed 2026-08-07T15:06:08.476018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:08.476018Z digest=sha256:d384968c0f13bff2292e11d5fdaccb3ee8022fe22a1ef89f87ab7eddc44ede8b

Observation 98931d4b-53cc-4d13-815c-8d322e3a3bf0 · outbound

This paper cites - t: Time spent in the coffee shop in minutes (which needs to be converted to hours since the other times are in hours).

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning - t: Time spent in the coffee shop in minutes (which needs to be converted to hours since the other times are in hours)

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.167010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:08.479995Z digest=sha256:ccf6c0b05c867a18ebe9e7e780ae4f20a9a4e4ef07c61f742e6a46e9aa843101

Observation be2b3f86-a863-4906-a4a4-aec64f33799c · outbound

This paper cites Converting t minutes to hours, we get t 60.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning Converting t minutes to hours, we get t 60

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:09.155451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:08.483160Z digest=sha256:f906d111490d19da9b8e5ac1a7a8815906752cc7e33552c1b46d1fe045b65e85

Observation e10bd72b-7cc0-464e-9a4c-f2c71d769f71 · outbound

This paper cites {numerator}/{denominator}.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning {numerator}/{denominator}

Reference 84

Resolution
verified exact
raw_fallback, observed 2026-08-07T15:06:08.583497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T15:06:08.486342Z digest=sha256:5ebdfd74ba702a6f22c24a1eb790bba6b4a171aadfd15e8596b8aa993f4322c1

Pith citing papers

Observation 26238ddb-55ef-4b2a-9f10-2735e6547f87 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.575865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.575865Z digest=sha256:7555fec650d07aa54194eef783001a914a1bb5b3305fddb6abe64c61c926a645

Observation 57099b37-beae-4b05-a5c4-9e9ca5be68b5 · inbound

Leveraging LLM-Assisted Query Understanding for Live Retrieval-Augmented Generation cites this paper.

Leveraging LLM-Assisted Query Understanding for Live Retrieval-Augmented Generation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:36.509403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:36.509403Z digest=sha256:c31ff76c68836d22beaa33fa014fa43e18c8bf8819d608c2b1b80bd01872e78f

Observation 8249ae1b-0f50-430e-a455-2913c4bae6bd · inbound

A Survey of Context Engineering for Large Language Models cites this paper.

A Survey of Context Engineering for Large Language Models Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 231

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:45.317119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T20:58:45.060041Z digest=sha256:9eb93ed96061182f712d77caf59a819eaa296d2e5ecb2f2e00296b001eaa55a7

Observation 61e64d85-0ef9-4e7b-8d37-94c49eb4807c · inbound

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning cites this paper.

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:23:55.066336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:23:55.066336Z digest=sha256:3dd60a77aac55ded0b4adc3e290064d6ba0b09fe962105e932ea6d9fb736f06c

Observation 1ec53e26-e08c-4eaa-b6f8-4349080788a5 · inbound

MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning cites this paper.

MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T10:17:13.551033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:17:13.551033Z digest=sha256:192b0157186f71fd3f0c96c9688751f562cb2b5a3fc00e1d10ab90e83d804296

Observation 5c8e51c0-4b72-44f0-baa2-9ff8928beed1 · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:21:42.329827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:46cea709dec6940684407ad84772a673b32bf92de5e32e49a476e273245eb7b4

Observation 6d12bb68-7528-4280-9e78-e7c7d7d238ba · inbound

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL cites this paper.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:31.860107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:31.860107Z digest=sha256:2f352e2d13995ea3d31350c3ac2eb23ad41b76bbb6f45e09d75125462667e355

Observation 8c997f8d-da85-4462-888f-99a8a655b5c8 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.278117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:6aa3cac9ce11b5b70b954364927802a9d56c9f15ef9bfc6f9bcbcb20dfa842fc

Observation 4cf73df6-8d84-4a8c-9504-70b53a7eb141 · inbound

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs cites this paper.

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:52.334039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:52.334039Z digest=sha256:e74b035ba298fc8326da0d1a957713152b789db26e6520d6a0a2bc206d8a189c

Observation c39a617e-8a38-456c-a1b9-3b25bf595665 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:28.753936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:28.753936Z digest=sha256:c0f52e05f82099f622e5fff9aa22a6836faf685d4cc9ac6674b73584e7a6f8c3

Observation d53bfd2c-ef7d-496d-9924-bbd31e9e9dc6 · inbound

Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning cites this paper.

Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T14:49:51.555062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:49:51.555062Z digest=sha256:a456ec4563984146cda606eb3f5b52bee3487a4bcd369bae5237c6990497bf16

Observation 72f33c36-2fab-44a6-93a7-49a1d4ea5b14 · inbound

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory cites this paper.

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T18:03:36.415695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:03:36.415695Z digest=sha256:04f7023fb34befd9136498a063f591426541447ece18e91c51a5b71721672986

Observation cca8698f-4717-4727-b186-7865c2f30559 · inbound

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning cites this paper.

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T05:56:57.071345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:56:57.071345Z digest=sha256:6062000a6565c8d69d08bf4ba7e2f23a8677f362556cc3ea743a0848e294d7c1

Observation 2eee2387-aef1-4049-93d4-754602ccb4fc · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:51:40.747261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:fede4eeed6318f9d1e46f0e7818e4d0320d35923f8b10b93138255b5dcdd35f7

Observation 782c5788-7aec-4c3c-ac47-15f9e354c63a · inbound

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments cites this paper.

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:23:27.288717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T01:20:03.181903Z digest=sha256:8fef357b22f36f7ce3bf6de33344320dc44ac1bc08dff84b5c950ba02de6f2c1

Observation 084cdba8-6810-472e-b25e-6bcfd79a999d · inbound

Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA cites this paper.

Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:49.508970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T19:25:26.811390Z digest=sha256:4aaa6f0d0d462af4181391b30fedd53a35e6d1cbe7154509407a0022eb9d2bcf

Observation 31bd0abe-80a2-4ab8-9381-690740a6df9a · inbound

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning cites this paper.

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:59.937831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T18:08:18.056525Z digest=sha256:e71afe0041e86ab83c17a7d8c4494c39b5c869ca7d91fac2b83408abe4d147c9

Observation 3bee09b4-f1f4-4348-9163-07973725b8ee · inbound

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning cites this paper.

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:58.424047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:42:57.596073Z digest=sha256:310a743d63c567d42e9a229c6c2828613b699f7d124b05492ba6fdd94b081891

Observation e9ac6733-1ad4-4aa6-b9d9-15bdb1d04c7b · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:54.223482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:c2c4bc7b9fb4e9657407c0ebc860753aa4d1f8cac677af0beb40a65e7cd23f09

Observation 407fad62-21f3-4e68-b6b4-9bf344089a81 · inbound

Teaching Language Models to Think in Code cites this paper.

Teaching Language Models to Think in Code Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:56.498224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T02:33:02.995360Z digest=sha256:b1fd966f2117247cfcb5c8d5cf8d2fa2ce80f2060ec1d2691a5d5d705ce212f7

Observation c63639f7-9229-49a5-8728-2f23fb64545b · inbound

Teaching Language Models to Think in Code cites this paper.

Teaching Language Models to Think in Code Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:16:28.390643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:26:11.265781Z digest=sha256:1642525f822cae86f04f23aebf4a5b583f0a251ec9e8ba72847cddc3b6918b59

Observation bab25d3d-c1ec-460f-8e49-fc9827819e64 · inbound

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning cites this paper.

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:31:23.986138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T02:47:31.830410Z digest=sha256:fed755cd6b5bf1f710cfddbf06740e5f8e3f43b9e40ec4f7e5f107ccb0861793

Observation 2154c720-8d36-4496-9cab-3c05b7e3268e · inbound

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning cites this paper.

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:01:23.278748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T04:42:49.165066Z digest=sha256:6f90440d499ee658f679967ebf204daacfdb5298c2b86ecdebfe6ead806db11d

Observation 751d4fcb-e620-4ec5-abcc-dc6947058c96 · inbound

Learning Agentic Policy from Action Guidance cites this paper.

Learning Agentic Policy from Action Guidance Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.727198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:3e682094a930b3e32c2cc4634eded372c4a3e4942da76a21066746120a3218fb

Observation 0a0d90ec-834b-4ada-b8af-32f410603583 · inbound

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents cites this paper.

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.303055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T03:51:54.401200Z digest=sha256:04c0f202806acec2c36daa13db57d819f368695d9615898921da716cb188e22a

Observation 707e0454-297d-4895-a8f1-af98616163da · inbound

Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction cites this paper.

Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:09:41.517513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T06:05:45.046156Z digest=sha256:073ad5c6b806155be420268cd3aedefd45e213ee2b9198d422e1fb3d76e7c7a6

Observation 898960cb-a66d-4f3e-8cb9-5dd058afa2b4 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.553725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:de3c223e878602e16f56f4eacf00c88b70c717603801a13fa05ffca2ed8ae7dc

Observation 639eeddd-856a-470d-8cbe-1fd351a1493e · inbound

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning cites this paper.

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-28T15:02:18.700268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T14:59:57.983338Z digest=sha256:f1110306e5165d0ad2005a02dd2e93305c6d320a0549c7314cdd04c8dd00f406

Observation 80e8188d-8c77-4e2a-b6bc-3d98bda5218c · inbound

ToolFG: Towards Well-Grounded Fine-Grained Image Classification cites this paper.

ToolFG: Towards Well-Grounded Fine-Grained Image Classification Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.338611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T14:52:06.186594Z digest=sha256:5fb627f61bb0af246fb152df9eb75fb3fff17ecb115d6efb754fbbe6e076fd54

Observation 27969eb6-8c88-47cd-8b7e-643059f52c85 · inbound

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning cites this paper.

Tool-Aware Optimization with Entropy Guidance for Efficient Agentic Reinforcement Learning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:26.289106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T11:19:31.702516Z digest=sha256:82062ae93c731d72c3e612ee9afb0f0bff2e141c27daab86ef9739bbf906aa96

Observation 5d04b888-c649-42c2-b7ca-ae6f7837d636 · inbound

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training cites this paper.

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:56.884050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T02:17:32.324432Z digest=sha256:bc8806f37ce4c666c7e2bf7ce206f1bba7e128f2067c26826a4cf5e0435d6c8f

Observation 0e47b6b6-929a-44e0-b716-146985896a3f · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:37:49.295388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T10:21:55.485624Z digest=sha256:09c28a54355ddc13bddd8431b16db81466632dd43b9f7099f6412a2d8d979764

Observation 44a3994b-2af1-4c3e-9f49-65b04b8794c1 · inbound

APPO: Agentic Procedural Policy Optimization cites this paper.

APPO: Agentic Procedural Policy Optimization Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:26.150859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:26.150859Z digest=sha256:d7e2676f2b2eb2549a274b6b51993c2aad2744fe334f18325ed490ef4df19d62

Observation a6de5c47-ed61-469e-a188-460913426f03 · inbound

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents cites this paper.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.509228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.509228Z digest=sha256:871bfa365e45e0be7934e4adcdafed73a013d2e14d7fd582fe9a8403f453db77

Observation e01b077e-76b7-4348-b052-0bcf74aafbde · inbound

AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs cites this paper.

AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T11:06:38.022901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:06:38.022901Z digest=sha256:db607732df84b9af8c9da95250871455cecc190435071c18c2b9ecffaab01a2f

Observation 032d3af7-9043-4df6-94e3-156e805cafde · inbound

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation cites this paper.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:10.315097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:10.315097Z digest=sha256:e794f95c1b8814aa07155d9f9861e8f1031440e1b579e6eef90cc0cf6a059a89

Observation b6ccbbc4-bf13-4618-b828-bd37700535d2 · inbound

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation cites this paper.

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-31T20:06:18.144305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T20:06:18.144305Z digest=sha256:a02bb7118ad0507f71e47787126c7949d4bf67a4b3ce4b7da1c0ed49083ea92d

Observation f9f149f2-ea60-4886-ad4d-d4362bbd38fb · inbound

ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning cites this paper.

ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T18:36:23.555649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:36:23.555649Z digest=sha256:5751895c5b9d3356d83d6667687df8c46395978475cd190fa8b24cea3f237692