Pith. sign in

Paper Citation Record · LEDGER

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

As of 6 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2605.29697.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.29697 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T07:58:55.548382Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T15:12:55.587685Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact34
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7b817a55-7a00-43c4-987f-c72140818004 · outbound

This paper cites xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.308144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:d6ad6962a2591e2a0e6c326b772b0d6c169492626552b3a4a81b369386b4b2f0

Observation dc32f3e3-071d-4696-a079-8e242b35ae44 · outbound

This paper cites Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.260619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:bb16411c48483fb9c7520a913ae43145717c4474fa0e34e0e2951b70443045fd

Observation d4166a03-44b5-4587-8f17-255f64d4260f · outbound

This paper cites Agentic entropy-balanced policy optimization.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Agentic entropy-balanced policy optimization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.263798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:566d83b3e95ef11cd96b292e86ab81df73fe28736e17923fb07ceca39a902bb3

Observation 80b367dc-3b22-480d-88ab-3bad878c848f · outbound

This paper cites Agentic Reinforced Policy Optimization.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Agentic Reinforced Policy Optimization

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.276093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:1756010ab8cffbd29aec1136c0b09284d1cc688dd68b857925011bf8c2f20694

Observation a1ece557-5e33-492e-bd90-82107d9b4733 · outbound

This paper cites Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.283420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:2bcbc285eba397bfb79e30afdded6eeb35c2525b739cd2abd5494d279fc8800d

Observation 42ec8530-48b9-439f-bbf3-73416b70dfc0 · outbound

This paper cites Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.270067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:4648ddb95dce688ef82dcbbb7e500d3869bb1059395e14f1cb7750a94361e2f2

Observation fdf6f91f-04bc-4300-aa13-a456713746d3 · outbound

This paper cites Dynasearcher: Dynamic knowl- edge graph augmented search agent via multi-reward reinforcement learning.arXiv preprint arXiv:2507.17365.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Dynasearcher: Dynamic knowl- edge graph augmented search agent via multi-reward reinforcement learning.arXiv preprint arXiv:2507.17365

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.254457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:7732161fd38b82bd8c4200a849410627fd3475d683642ee118556922fc24cef9

Observation d66b4fc2-8352-4580-91ae-06aaccc2e7c2 · outbound

This paper cites TreeRL: LLM Reinforcement Learning with On-Policy Tree Search.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling TreeRL: LLM Reinforcement Learning with On-Policy Tree Search

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.272895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:bb3ee73d5a3167063cee376beec1555e78485ee064ecbeec176a264d79f6ceba

Observation f5f31f6a-ce65-45c5-801a-975b2c08a6ff · outbound

This paper cites Tree search for llm agent reinforcement learning.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Tree search for llm agent reinforcement learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.286116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:38374b444b47f1456d8666dbfc72064d8ecc41f472d929fceacb48fb06669246

Observation 879b6e96-9067-4bcd-8650-668666258e28 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.318817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:953b7ad693e060d4e700f41a796636ff6bebe2a7d811dce1fd4a52d347c43a35

Observation 9a0f0a16-8580-4da1-83d4-d88bcb579490 · outbound

This paper cites WebSailor-V2: Bridgingthechasmtoproprietaryagentsviasyntheticdataandscalable reinforcementlearning.arXivpreprint.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling WebSailor-V2: Bridgingthechasmtoproprietaryagentsviasyntheticdataandscalable reinforcementlearning.arXivpreprint

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.239225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:6067b270f14f88dc0cda9a1d0e6b1ecb98316955975456211f5ff596223ebfde

Observation d9978c48-f546-4c7e-9286-f988a7dccc87 · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.255321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:95daac4d04a24aa1aa43bce00e337aa1e40a559d4a9cf831ef58bba51f0acba3

Observation 7a96e49d-6bf3-4359-9f33-65a07be13d5a · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.291784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:1c4f4563d84ceaedb6a78684ddfae24a178556fa89a69d2b972f2dfc5c48a820

Observation 4c322c72-3308-43a1-a22c-9c6ae5ee641b · outbound

This paper cites Webexplorer: Exploreandevolvefortraininglong-horizonwebagents.arXivpreprint.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Webexplorer: Exploreandevolvefortraininglong-horizonwebagents.arXivpreprint

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.267520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:bc74d92882cbe1d1c8a03e51d4c41a962bde6fcee44560e654ff279ea1a3b704

Observation 76415858-d599-4028-84d9-55888fcc69dd · outbound

This paper cites Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.275661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:c05ca190ed4ab9611f0471bde77921ec042d7bcb2b8969915b8ce8bc1d11de37

Observation f56e9f0f-98d6-4990-82ab-71bf0281071d · outbound

This paper cites an unresolved cited work.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:0607b1e5360aa2b3119a72fbbb5f1c9eb3d0d9107707eec74c291e266ea8d188

Observation 151a1e16-529e-46a2-9ef3-13c5a411277d · outbound

This paper cites an unresolved cited work.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:395875e0a73376df3c71c29b6aad93ca59a8499117fc69f790ad18da43d13256

Observation a8593af6-2164-4dc2-b455-8a6ce0bd781c · outbound

This paper cites an unresolved cited work.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:e8ce2af46252284c68d44ab37dd0b239066b64964235b8552a33027b95579965

Observation ae93f585-3973-4684-9b9f-d9ced1423fd6 · outbound

This paper cites an unresolved cited work.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:6eb2280170db14a3c86386a4c8b75e53e51379253dbf0030a2ec97a1f31c141c

Observation 0f41ac71-f1e7-40a1-98d1-ee709d3e9eee · outbound

This paper cites an unresolved cited work.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:c6c96ef3928ff2964d036815fb1b690039f0051a76e753274345593d3d96f217

Observation 2cbd5282-9b6c-416b-82c8-d4c3dbf02bc8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.263792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:9fc518140b020d758652a00c1edc02107203800e601362c074a23ffc8440c696

Observation 18c2caa5-c661-43de-8d3a-47db8d1a0964 · outbound

This paper cites an unresolved cited work.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:d77049fdc556fe3cf0c4e4cf200c75f4db6cf9783ca367b0b9a2f7dc00ba2788

Observation bbb2ac45-8346-429e-b00d-935f134bb0ea · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.278322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:5a7d798e5b65dd12f4ba4e05356b1577b9c9027fe6ffb19d6fc3ac3ba78c0794

Observation 16a2bd18-21b9-4845-a22c-54197f94a5b6 · outbound

This paper cites Webleaper: Empowering efficiency and efficacy in webagent via enabling info-rich seeking.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Webleaper: Empowering efficiency and efficacy in webagent via enabling info-rich seeking

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.257602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:c153d0f21738eebabd990a1219155c2a46a46cde61029772a349f233fd6922d7

Observation b321a08b-9c8f-4288-823b-9d76ec4f1bcb · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Kimi K2: Open Agentic Intelligence

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.299693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:539cba34a8dba5f5d6bbb64dd2d10de12587ff269f90a02ddc1df4d5328f4439

Observation 33b6ab0a-0fa6-44c0-8078-9432b1919293 · outbound

This paper cites MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.233391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:9ffec884f3f116768875604f5a408f8f7cded073ef6050284feedf7e927e2607

Observation 0b47dbcc-de77-4b33-94bd-f8a6804ec892 · outbound

This paper cites Tongyi DeepResearch Technical Report.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Tongyi DeepResearch Technical Report

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.250181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:11e06df615a46daaeed844f9cdf119bbc31301b6a3ab9f9f038e1885916d34d5

Observation 1bbb14fa-0f20-4455-ba0c-3ba2a2688f58 · outbound

This paper cites WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.302996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:d5ec1a746eefcdf17d8425c5501a16f9e60a0e334cf82590d35a642d7f79bf20

Observation e920c3d0-f433-49c0-b0a4-43eefbca219a · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.287902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:140d985e538290bb85b6f507239d895e976f315d1731f3d9d96713569ea76610

Observation f128c70e-03f9-4f48-b480-86714e441ecb · outbound

This paper cites an unresolved cited work.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:389ebde38aafef49ce77be6cb939e9b086e80c6552db8cb94d5c3eed50ad7636

Observation cc0b7253-a762-4415-be16-5809958259ca · outbound

This paper cites Reinforcing multi-turn reasoning in llm agents via turn-level credit assignment.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Reinforcing multi-turn reasoning in llm agents via turn-level credit assignment

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.239956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:f71479e809eb04670585d46ece24efa2c7b1e09fa94f268074eb5af32d5d8146

Observation bc87a14a-d656-4f65-9f93-c601cb9a3d82 · outbound

This paper cites WebDancer: Towards Autonomous Information Seeking Agency.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling WebDancer: Towards Autonomous Information Seeking Agency

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.247506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:bd92bcbc5246999ad05d60d013ff23524c7fe0c2ee414344a066f8102c967690

Observation ded4a9ee-fb55-4456-ae80-430575425e3f · outbound

This paper cites Qwen3 Technical Report.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Qwen3 Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.305567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:567b5f84a19f12bd9797a4cf7992cfc9bacc8b5c5e781c60a3e10f09ecbfe2b2

Observation 8f92fc6d-fa29-45b4-a62d-647ef2e1fe93 · outbound

This paper cites Treerpo: Tree relative policy optimization.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Treerpo: Tree relative policy optimization

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.316363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:5f71069c8c7cc1d69c3db472c240f2a1a541d0524b8d47a57f8030de04cbd5cc

Observation 1e4102e9-d317-478e-9fd3-b3d9e33a8742 · outbound

This paper cites an unresolved cited work.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:dfe2ce945573fd29d573495f5a80fedb511c2dc7dc16169eb2836767214d565c

Observation 1d03fe18-37d5-47bd-8db7-482b1e88f803 · outbound

This paper cites Process-Supervised Reinforcement Learning for Code Generation.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Process-Supervised Reinforcement Learning for Code Generation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.288936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:cc7f833c986fad89f16900e61fa039999a7d41312165f5138802ece8d5e1c2a9

Observation 0d40df21-ee9d-4796-b9bd-c1e2f5d7d8fa · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.280667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:23c661b6586efea480f7e2af4870a6a07b669a78419c112559e9e5ff0adab40e

Observation 6fccf41f-ad9f-4079-8f76-8d91bc0ebaee · outbound

This paper cites an unresolved cited work.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:97e8907a8e373f2a13efadca28ab24f3e56d59ea2f26d75f4727e6d7d428d365

Observation 21de9f86-7d70-4bd5-a42e-b000aa59ea22 · outbound

This paper cites an unresolved cited work.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:1023dcfefa4f60e84a63b654a3c9702bbe0a735ad23a0e9f5046d92ce90415be

Observation e3e1dae7-6099-4437-8040-68a3af042211 · outbound

This paper cites Chaining the evidence: Robust reinforcement learning for deep search agents with citation-aware rubric rewards.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Chaining the evidence: Robust reinforcement learning for deep search agents with citation-aware rubric rewards

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.297033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:7695ec9d4d5ea6ed43fc15dcfbbcf552f2879711f3f0ab618be8c9c22cfecf59

Observation 653e8ddd-a8cf-45f8-8a45-109615f69b6f · outbound

This paper cites an unresolved cited work.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Unresolved cited work

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.266459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:37048ec56e6300ca2f1c7f4a2780d931252191c8edbd4c3b3db432ff3a272497

Observation f7d6d69f-96bd-4c89-8d14-c5cd76566bce · outbound

This paper cites Group Sequence Policy Optimization.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Group Sequence Policy Optimization

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.294048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:1122d07836e8efbf3449dc39b2435f23950e1c9d264c3618eba5d23a5299e511

Observation e60b15ae-b6c3-42b6-9654-755a87b60b16 · outbound

This paper cites an unresolved cited work.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:5db70de916ef6fa0ccaa212aae6fcb348dfa6b91bbcc9eea256437b21f73637c

Observation bcd3b997-807a-454c-9b71-30c5c2f650fc · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.290658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:adfb78e73dc85c781feabbe91698d75fb8480c9dd3564f6447887086e5dfed57

Observation 96300555-28ba-4565-b59c-c24e00eda21e · outbound

This paper cites BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.311800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:12649785e91c0569bd570bc404fd05493db1b68c9e65f609e762d36493ddad85

Observation f08e5b1f-3fc2-4fa2-bcab-a6fa3b497cbf · outbound

This paper cites online" 'onlinestring :=.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling online" 'onlinestring :=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:e1d3e1a6484fbbd57a574903a0fceaf8299ab0b0e010242c0650df5e6ecc7d78

Observation 1fb06325-5626-4a26-a44a-bdcc9db532d5 · outbound

This paper cites write newline.

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling write newline

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-29T07:58:55.548382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-29T07:58:55.548382Z digest=sha256:91e2f89c10a97cbed2a49a1faa02940002e8cabf6bf702c3a12b03e86db738f4

Pith citing papers

Observation 5ca19be1-64de-4342-8ce8-60c8f729ab0d · inbound

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents cites this paper.

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T15:12:55.587685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:12:55.587685Z digest=sha256:de8147a7a0f3b1b34897a74253e0b97318ffeb9124fce643cc90a0113ba18f32