Pith. sign in

Paper Citation Record · LEDGER

StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2505.15107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15107 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:20:40.072107Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T13:04:56.727934Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5f9244b7-87c3-4e89-b3db-e8c95144cc37 · inbound

Group-in-Group Policy Optimization for LLM Agent Training cites this paper.

Group-in-Group Policy Optimization for LLM Agent Training StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:15:08.622111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T09:15:08.193357Z digest=sha256:3a5ef3d0bcc095099ed7b9173974f343e166c7fe9f85f21dbdad03f7701c6084

Observation 82c3f128-cd6b-4ca1-b174-dfc382795218 · inbound

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning cites this paper.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:12.451731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:12.451731Z digest=sha256:44381358f26acafb6aad9e34a1131f7d4fbf4c0c2642afc9015bf4e8a90e8730

Observation 47a1e444-233c-42af-bd0b-a92e31b51857 · inbound

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL cites this paper.

Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T23:57:35.446112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:57:35.446112Z digest=sha256:0624a70bfe5e3aecce4043639d6ae5d208d72a1c3acf6b23038ecea17f46566d

Observation 5e1666c2-d673-4616-a540-ed43729ee018 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 280

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.572696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:d54cf2f9f48995f74807a2c574937a0f0dd5d0a77ab38acb4474358c982eb8f1

Observation 8e9fd6de-99b2-4e1f-b9e1-580c3c687b17 · inbound

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs cites this paper.

Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMs StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:11:18.260446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T11:06:20.058342Z digest=sha256:5970e746d7f80b0ab99a28798287bd4cf2cae766b4ba00e8c530376f0816e624

Observation bdd4d1db-c793-41da-ab9e-65725839cfbe · inbound

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation cites this paper.

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T09:50:50.408334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:50:50.408334Z digest=sha256:3617c10453f1cb4bef7751b2d125ba926a4068392dbe5c512fabcebdbe6c31b5

Observation 77e100f0-2904-4f0b-aea0-c0055fab618a · inbound

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework cites this paper.

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:50.356852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:33:50.356852Z digest=sha256:d250ebd747d7152719c04dea585cde7e9d627ab5ae33ee8e50d9822eb19234b4

Observation dda22789-8f60-40e1-b75b-0e9db858154c · inbound

Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning cites this paper.

Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T20:25:01.404692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:25:01.404692Z digest=sha256:cc346000b71f5604f671485964e136969124e006bd8a88bbb6a2693335e56ecc

Observation d89acafe-afeb-498e-9828-b297d9155d1f · inbound

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning cites this paper.

E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:25:59.776769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T18:08:18.056525Z digest=sha256:379501140fbaaf0843c936482055bbfe85cbcb08d98d30f26ee359302cc07da1

Observation 90f0cf76-e5e0-4776-9f45-f82b1adae920 · inbound

$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data cites this paper.

$\pi$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:27.840116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:43:09.581054Z digest=sha256:22fee17ff9fe8b3975ba2b580c6a3d4b2c5bda33670e979eea24fd034fefb63a

Observation a715c2fd-ff47-4ab7-94d8-a02b8f4ac245 · inbound

PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning cites this paper.

PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:26.662716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:35:21.469086Z digest=sha256:25d66e14b2bd7ef3b85631d890a00245904c76284f35e84d0b033a27a4db3d83

Observation 62eb5c5d-e137-4c90-b007-1781216907f4 · inbound

PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning cites this paper.

PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:37:29.679737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T07:34:32.166480Z digest=sha256:b0a8f4bd3c3b2676968f8b743bfef8a5f9dc3199e338b8d990a9335d841b8be0

Observation 89303f49-90db-40e4-9daf-e8468738297e · inbound

Harnessing LLM Agents with Skill Programs cites this paper.

Harnessing LLM Agents with Skill Programs StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:28:14.387703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T11:26:19.382463Z digest=sha256:02330cdf9db37d30505074162dbd3b908f4bad64a1f392b69c710a85b3b8ed3f

Observation 5554b0c3-3278-4092-a2dc-63a5476a0f77 · inbound

SD-Search: On-Policy Hindsight Self-Distillation for Search-Augmented Reasoning cites this paper.

SD-Search: On-Policy Hindsight Self-Distillation for Search-Augmented Reasoning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:03:14.662474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T09:59:17.346548Z digest=sha256:ae4d8018ce74117040d643af36e217d012167bc5d7d007abb085848c1cc5dce8

Observation 04b16e29-d8c8-413a-88da-6f308ad9e551 · inbound

Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning cites this paper.

Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:11:09.155728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:08:22.252524Z digest=sha256:0d9f3f877e3c8a52f1b2834b1f9135137e600810ad72d785d7e8eac3b32e444f

Observation f5a9e757-4c28-4535-bde2-fbadbae3ab0f · inbound

Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning cites this paper.

Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:56.743506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T17:24:26.923687Z digest=sha256:c8fdbc7d9326dd917559bbbc4b8fb8147581e9b2c23c014173f257d457efc225

Observation 67705cae-0a56-41ee-9152-ac4e29548a0d · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.462913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:5b6e0da02e2c36fd217c80b8792d0cf18954590944ad0f24088ecd2a6595cce4

Observation 80832a9b-e435-4a00-afdf-f06b668aa714 · inbound

Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses cites this paper.

Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:26:22.099072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:25:08.052988Z digest=sha256:338da1288d2ba7a1cacac4c769be8245c10ffba6cce503cba41ccfe2599cbef1

Observation ae6b84d7-f4b5-4510-9421-cecfa4c3c4cb · inbound

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training cites this paper.

When Denser Credit Is Not Enough: Evidence-Calibrated Policy Optimization for Long-Horizon LLM Agent Training StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.871475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T02:17:32.324432Z digest=sha256:b4a9e62b06d4828b88e3bd7511741c03cb552f039159b56a50e629a05c56c9ad

Observation ffaea00e-6581-4880-a5f1-68dd30ab359c · inbound

Co-Evolving Skill Generation and Policy Optimization cites this paper.

Co-Evolving Skill Generation and Policy Optimization StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:25.881525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T18:37:00.015083Z digest=sha256:61c4d76bed11246736400091fd3f7fa56bae605ea50c17358acb66ae4d1826eb

Observation 538e4a03-48c6-4322-8bc7-350f39348636 · inbound

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment cites this paper.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:28.917625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T17:22:16.055502Z digest=sha256:155a5d610be2f6dd5044e095eb6394765357170f84802dbe0c9a4d2d31338cf5

Observation d52464a9-a930-4b0c-8ca9-24415466a351 · inbound

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment cites this paper.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:aee4123ffe9ce6a842b0353b3fd2dba5afa00f1cda0ccbb4ba72d7850fa4c29f

Observation 91115b55-54b5-4efc-b863-5981c9d9aaba · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.357678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T06:39:34.199607Z digest=sha256:c9dae966d5ed714403fcd9cfd30a85d46e92ea81fb2d3273e20a95aa78172c61

Observation a1b21114-9d2a-4b4f-aed3-ec6ed1c30fda · inbound

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning cites this paper.

ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T02:12:23.414741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:12:23.414741Z digest=sha256:21ec6fcccc38f3f74f7d4984248e93bad32ddc926bd111a7071dd83a12d03ac7

Observation 2409857c-a3f5-4f08-aab0-b5a6b817a4cd · inbound

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning cites this paper.

Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.757145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:44:23.085858Z digest=sha256:54b106bb68a090baf737f483453e484afb59f0ebdf2af56b1ff77fb49c9fffbc

Observation 733fe686-4d61-4648-8494-88de7018c776 · inbound

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents cites this paper.

Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:04:56.729725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-08T13:01:45.049174Z digest=sha256:15482972cf05154df880a4502a853a000400527c2742257c7ffed49f10043aff

Observation 10c737c3-645d-49ab-9e8a-d5b20ac9ffb9 · inbound

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration cites this paper.

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T23:46:17.626014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:46:17.626014Z digest=sha256:25511864ec7b29e2a462b9e2a07b448126e527dfdcf1f66cd884481cd5423d45

Observation ab23ebee-b632-4e80-ae75-1775ba2bac1d · inbound

EviSD: Evidence-Conditioned Self-Distillation for Search-Augmented Agents cites this paper.

EviSD: Evidence-Conditioned Self-Distillation for Search-Augmented Agents StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:20:40.072107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:20:40.072107Z digest=sha256:e4fc63c45f530d587034c8c0a9e56533ba5fd1ba627a9773f014db28fc07ff9f

Observation 16c9d078-39ed-4e86-877e-921490f8de89 · inbound

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents cites this paper.

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-04T15:12:55.832980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:12:55.832980Z digest=sha256:511fb76cfe02eafd70bb92376663e2b7fe6cec556abf403cd0ef5366244e8f1d