Pith. sign in

Paper Citation Record · LEDGER

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 6 inbound Pith citation observations for arXiv:2507.01489.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01489 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:55:09.996272Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:55:55.496339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:53:28.611729Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2980a235-2774-4187-bc1b-f97f589fd2ef · outbound

This paper cites Brown, B.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Brown, B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:55:10.276010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:55:08.809952Z digest=sha256:704c55ae2e123b45705ef13d929b879877f933cf20348a8429938ab57edd70aa

Observation 741dd78a-e701-4532-96fa-7b09cd18633c · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.044949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.044949Z digest=sha256:dbe4137a0730e42f89684d6011d01fcef3f4bde712f92e17b2f0e1a2ed080834

Observation b56633a1-1bb9-4945-9047-497a5ffb5378 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.251681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.251681Z digest=sha256:901ef804ef1ac4717ae0c21747f2bbf773b7f474ff19e18d2ea1b73b2d5c822c

Observation cd883f29-6f39-45e4-a322-8bb6646e2d57 · outbound

This paper cites Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.340912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.340912Z digest=sha256:b6e2147c2da72e69f8e64b55bfb1b2380143982ad556a78bdca902b6a88b1d9a

Observation 579394ad-e4ce-4fc6-8220-841b831c109b · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Measuring and Narrowing the Compositionality Gap in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.438805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.438805Z digest=sha256:e1e63c46ef5b969b7c819490887fb7ba3406e2cd53a0a339594f6240889597b6

Observation e4058c2a-53ed-4ea3-a043-9ebed6af1788 · outbound

This paper cites Qwen2.5 Technical Report.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Qwen2.5 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.563729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.563729Z digest=sha256:65a100aac75626d965d2d8d3e803da1ddd2be53248d27eafd57e0e0265b25a6b

Observation 65766e1d-9f9c-413a-b2a4-8a57bdb89f94 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.825758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.825758Z digest=sha256:2a5b739678d9ecac4ea4bf62f70d61d1a316f2c2aafbd1931a56b4e64da0ea39

Observation 5a78b56a-609d-4642-bbdd-b527cbff4f49 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning ReAct: Synergizing Reasoning and Acting in Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.879680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.879680Z digest=sha256:1c63e78f9b788ce6a7fd248b0c935173a9535b09b471204b5f66f2f625338836

Observation cc12c887-380f-4ca6-a429-67fadbedd70f · outbound

This paper cites OpenResearcher: Unleashing AI for Accelerated Scientific Research.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning OpenResearcher: Unleashing AI for Accelerated Scientific Research

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.948075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.948075Z digest=sha256:a900065b0cbda21c280306b6c06496e42e15ec3fb2cccadbbd33b007221b9b10

Observation 3779de78-c0be-4671-ac6a-69f58a4549d4 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.996272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.996272Z digest=sha256:6d767b460f497c4bb4db71c3be1e10cd1dabdb37046ed20326be973ff2af362d

Observation bb878f04-f70a-467a-a495-7e8d242f156b · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.108495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.108495Z digest=sha256:c44680fc39f5371d2b9334235e3602537f54c46217c28cb2939d04e3921add6f

Observation e7748746-bc8d-4743-82b9-64a1f7610247 · outbound

This paper cites MuSiQue: Multihop Questions via Single-hop Question Composition.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning MuSiQue: Multihop Questions via Single-hop Question Composition

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.770719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.770719Z digest=sha256:94712f09baef995b11cbd1f648dc0f9b836ce535c68ad08888f0c878ea2f4a72

Observation cc514193-ee30-4d29-9e02-81507033a662 · outbound

This paper cites GPT-4o System Card.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning GPT-4o System Card

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.165242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.165242Z digest=sha256:c7c6535e1f7e3c775d69e5f4f159473eccbbe46b8cd55f238c6171b4c07a22fd

Observation 44c22c11-5e34-44c1-a8b0-410e413b38c8 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.695260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.695260Z digest=sha256:ebdf5755b239e38c1d6b700be28a70f0cae16cf42b5f1b112da6574d25867306

Observation 4b64dbb8-5014-4415-9e30-908677aa8367 · outbound

This paper cites Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:08.936392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:08.936392Z digest=sha256:968bdaa954b9939a971b94137041b5d4653991417c1528928434a9a21815e409

Pith citing papers

Observation 88b47baf-aff7-4e85-bc8e-def515eb4e0b · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 228

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:13:15.635577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:d407e987777eb47fb52dcae6f792f511ad2741563ffb4882b4c2c2e748c91474

Observation 56960e50-3558-4410-8bca-ed3bd3392f08 · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:26.597990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:44:30.069444Z digest=sha256:bd7fd5a958e6a14d78b0f0a416d511331301e4b11cb4c36d624709ac000d39f9

Observation 91881407-ef2d-4594-a886-d4bd1467fc30 · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:59:48.707183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T05:59:23.949124Z digest=sha256:8eed8fb8247fd4eb06541690549ce0cc27f940e3ee460350164b12a2036216c5

Observation 2d126fbc-da39-44de-b05f-8ef1dc014e1f · inbound

IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents cites this paper.

IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:24:40.531043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T06:21:56.971204Z digest=sha256:a09ede52048a0475873b8e9dee70b861c9e8b0cc4fddc0368b510c296574f949

Observation 3ae45dc3-c1ce-487e-9d3b-aa30ec874642 · inbound

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use cites this paper.

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.613407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T13:49:58.677299Z digest=sha256:c1e27cbc9ce03e328126ae20d45ed5f3471f4b3a740b80a994a8e07891510a56

Observation 4efbff9a-c058-4e2b-9631-ae4d821c1538 · inbound

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems cites this paper.

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:55.496339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:55.496339Z digest=sha256:977bd03608b5ddb6785e43daebf78a319590a633962b0162748f462e06867598