Pith. sign in

Paper Citation Record · LEDGER

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 6 inbound Pith citation observations for arXiv:2507.01489.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01489 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:55:09.996272Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:55:55.496339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:53:28.611729Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2980a235-2774-4187-bc1b-f97f589fd2ef · outbound

This paper cites Brown, B.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Brown, B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:55:10.276010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:55:08.809952Z digest=sha256:079d0c6d06a02352bee19cc8bea9df9d127d3afa8f20cf3a387c773dbb641d08

Observation 741dd78a-e701-4532-96fa-7b09cd18633c · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.044949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.044949Z digest=sha256:5cb14a3a5ad4bdeb24b2f8e8273fffc9e3dd56edc176ab6645540795b7eeed16

Observation b56633a1-1bb9-4945-9047-497a5ffb5378 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.251681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.251681Z digest=sha256:5f4193ddb5a0b1d45cac4f05cdef3ff8c9f737f43cfc950cc6889e7134061657

Observation cd883f29-6f39-45e4-a322-8bb6646e2d57 · outbound

This paper cites Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.340912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.340912Z digest=sha256:adc75bde5e2ff30e6a7ee731f66e345db79e7efc53d56405ec3cfe6ec064287a

Observation 579394ad-e4ce-4fc6-8220-841b831c109b · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Measuring and Narrowing the Compositionality Gap in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.438805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.438805Z digest=sha256:ccc310a9eae5fe391c62911d80c48af865dec88d83cc89959628afe1f280808b

Observation e4058c2a-53ed-4ea3-a043-9ebed6af1788 · outbound

This paper cites Qwen2.5 Technical Report.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Qwen2.5 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.563729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.563729Z digest=sha256:3219e45fa5026017843991102260d66e068bdf8ee48dd62f81f299fd8878daee

Observation 65766e1d-9f9c-413a-b2a4-8a57bdb89f94 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.825758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.825758Z digest=sha256:521c7ff700f241a3a9b9462e492f2bf295c509e4837140ceabe036a36e3da2ac

Observation 5a78b56a-609d-4642-bbdd-b527cbff4f49 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning ReAct: Synergizing Reasoning and Acting in Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.879680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.879680Z digest=sha256:31c9a385e587259b21d160817b594a2f760f24d9a9cd6e8da1d4d59dafab9a75

Observation cc12c887-380f-4ca6-a429-67fadbedd70f · outbound

This paper cites OpenResearcher: Unleashing AI for Accelerated Scientific Research.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning OpenResearcher: Unleashing AI for Accelerated Scientific Research

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.948075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.948075Z digest=sha256:dba96592348038e8bbe68fd967acf9323f8d25c7c0af142d93a9caa9d5fdcab0

Observation 3779de78-c0be-4671-ac6a-69f58a4549d4 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.996272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.996272Z digest=sha256:81f38aec81ba8cd5dba9e78c6257a907effb93ad9ee66fbd32992c570e133635

Observation bb878f04-f70a-467a-a495-7e8d242f156b · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.108495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.108495Z digest=sha256:5caa68a3f3e8415051711b988c197651925991d2549c9716e20cbb43d62fa10c

Observation e7748746-bc8d-4743-82b9-64a1f7610247 · outbound

This paper cites MuSiQue: Multihop Questions via Single-hop Question Composition.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning MuSiQue: Multihop Questions via Single-hop Question Composition

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.770719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.770719Z digest=sha256:308effa360cee3fd6f3c253d62509c060570a13ba0f27b14751047a5f4e3c117

Observation cc514193-ee30-4d29-9e02-81507033a662 · outbound

This paper cites GPT-4o System Card.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning GPT-4o System Card

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.165242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.165242Z digest=sha256:201e736c2d498753321a5991cc41feabed7322185e96aab98d03c5c6b8bee4e4

Observation 44c22c11-5e34-44c1-a8b0-410e413b38c8 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.695260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.695260Z digest=sha256:4dc372adfb39aa0dc4b88c05be050b2d57018d57c815c95a1699bddada2ccb16

Observation 4b64dbb8-5014-4415-9e30-908677aa8367 · outbound

This paper cites Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:08.936392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:08.936392Z digest=sha256:f828b2fae92553b9bb09c3d59e5bb4a9ad39f94fd47cb55f944841a67934d6d9

Pith citing papers

Observation 88b47baf-aff7-4e85-bc8e-def515eb4e0b · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 228

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:13:15.635577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:7a93f5a3fe26e78a9b4bb44152f21ccf3a12fbee4d0d4b28764c87e9eb45092d

Observation 56960e50-3558-4410-8bca-ed3bd3392f08 · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:26.597990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T06:44:30.069444Z digest=sha256:f7ba8f05a197a5933de45d725d664f9aaafbb5ea93839eebf5cc73bb71dc2e1b

Observation 91881407-ef2d-4594-a886-d4bd1467fc30 · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:59:48.707183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T05:59:23.949124Z digest=sha256:6cbdadae5e0aef67417599375bac3d677569c9cc8aff72a0e38ddb1a03eb1a35

Observation 2d126fbc-da39-44de-b05f-8ef1dc014e1f · inbound

IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents cites this paper.

IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:24:40.531043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T06:21:56.971204Z digest=sha256:113e8da4b9c30607c2ad13b905fc6b989ee23338ea13b90a7c71aa3a7b6e0643

Observation 3ae45dc3-c1ce-487e-9d3b-aa30ec874642 · inbound

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use cites this paper.

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.613407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T13:49:58.677299Z digest=sha256:bade2016a74e901023e69e0cca87dea5fcec5fd32c8c5a7004bec88beda73ba8

Observation 4efbff9a-c058-4e2b-9631-ae4d821c1538 · inbound

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems cites this paper.

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:55.496339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:55.496339Z digest=sha256:7f377c7070a41e53bc1bc135669f700d807fdd278f382961364f8608c6c43e63