Pith. sign in

Paper Citation Record · LEDGER

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent

As of 10 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2608.00969.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00969 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:37:25.304636Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:37:22.965089Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T00:37:26.010355Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54373557-c971-46ed-9353-69796bbe3beb · outbound

This paper cites PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:37:26.163501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:22.965089Z digest=sha256:f1bfbcb25080b3b847eb45f0604b74f7e09fbe8091276ebc8c93a0177bbf6679

Observation 3fb9add0-3062-4433-bdf0-1c54e8dc20ea · outbound

This paper cites It has enabled open- domain question answering [3] by LLMs through the retrieval of evidence from large corpora and the generation of answers based on the retrieved context.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent It has enabled open- domain question answering [3] by LLMs through the retrieval of evidence from large corpora and the generation of answers based on the retrieved context

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:30.975049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:23.016570Z digest=sha256:d5fb4642f4b0ef611266f04acc9e543c0b16b9abeb51f046d65e04aab0096a32

Observation 435cc06d-50a3-40dd-ac41-ede8b629232d · outbound

This paper cites Then we provide details on our proposed coverage reward with teacher-guided training framework.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Then we provide details on our proposed coverage reward with teacher-guided training framework

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T00:37:30.617126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:23.134193Z digest=sha256:49b46321d8da4318924c12a1e4be999344870e5378911431c48f255266154a31

Observation a0463496-3f85-4695-8e62-ed67c9fcb8c1 · outbound

This paper cites an unresolved cited work.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:37:30.426338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:23.199205Z digest=sha256:f6aee811ca42c407b8f5d7f5b44e681f8a3c74d36b72ac8cf55c319e0f61d968

Observation d1fca82b-1a51-448a-a05d-20c18d89ff59 · outbound

This paper cites an unresolved cited work.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:37:30.233566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:23.291445Z digest=sha256:be295c63575628dc2d1b01a67eeb06c1f40d6c1fc1887ec07b1482750df34dc2

Observation 0456a6a4-4db1-4e42-b685-43be9bbeb345 · outbound

This paper cites an unresolved cited work.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:37:29.877431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:23.342030Z digest=sha256:c680d29e5f5488c030ff7ba626dbf236bea6e09f6e3e46eb21c2bf5e28c033c3

Observation dc596767-7e36-40ec-9e0f-5150a5e5f9cd · outbound

This paper cites React: Synergizing reasoning and acting in lan- guage models,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent React: Synergizing reasoning and acting in lan- guage models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:28.345951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:23.941253Z digest=sha256:79cea213cfe0538e0f4236877176095ed0e744c0c60e835c46a478b648fb5c7a

Observation 9b1eb094-aace-4fff-a20c-40390a6b0152 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T00:37:23.402058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:37:23.402058Z digest=sha256:d9bf13be79111cc4966d5044f2048f2ca1b1cdf95d298473a4572e93e2d14639

Observation d7ffc96a-a5c4-4c25-95c9-7060a7b281c3 · outbound

This paper cites Deep- researcher: Scaling deep research via reinforcement learning in real-world environments,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Deep- researcher: Scaling deep research via reinforcement learning in real-world environments,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:29.626970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:23.501648Z digest=sha256:e82703b89f0f0f7c7b725693f843fd4ffe0eaad0398273cab593f86dcef45646

Observation 64ff9fca-eef5-48d8-918e-39fcd395f3de · outbound

This paper cites Reading wikipedia to answer open-domain questions,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Reading wikipedia to answer open-domain questions,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:29.376053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:23.560798Z digest=sha256:ec7649190da2e247d8c821265396efd01a1e59804b34ff080b10fc23c0cdf7fb

Observation bcccccfc-dc1b-4079-b3b5-a75eeba4d86c · outbound

This paper cites Natural questions: a benchmark for question answering re- search,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Natural questions: a benchmark for question answering re- search,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:29.051272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:23.644884Z digest=sha256:848c1a50ad0da7df87437152975fcaad9e6ae8570b156eadf30dfb732e0df1c6

Observation 8f4edd40-8e4e-440b-9d8e-7bf609785314 · outbound

This paper cites Hotpotqa: A dataset for diverse, explain- able multi-hop question answering,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Hotpotqa: A dataset for diverse, explain- able multi-hop question answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:28.804074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:23.727908Z digest=sha256:5cf65e3ae95a442b988f8e3dd70fa7261d81853499d5b009d94702b8b39b3812

Observation 664e026b-47a3-4c85-9345-a7b6aab7a057 · outbound

This paper cites Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:28.613955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:23.847679Z digest=sha256:02db6b5c6f100257b7b3c0550e05c78fc13214511bec8777e5a63759f8a22965

Observation 42d91b26-5ed7-48e4-ab77-65c6c60556c6 · outbound

This paper cites In parallel, Ze- roSearch [16] addresses the high cost and instability of RL by training with simulated retrieval during RL.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent In parallel, Ze- roSearch [16] addresses the high cost and instability of RL by training with simulated retrieval during RL

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:30.798418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:23.079209Z digest=sha256:06b46d07503d9bc3528b751b37cd65fec806e3193e415b8fa092872274e592be

Observation dde2cfa6-3482-4b1b-9008-de6eb1cd687d · outbound

This paper cites Chain-of-retrieval augmented generation,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Chain-of-retrieval augmented generation,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:37:24.008663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:37:24.008663Z digest=sha256:af83c8cde2cd86cf58c7475aa8f51d5c012b1e31366111b9acc65dcd5bd41ef6

Observation a09846f1-9385-43e4-9eb6-7e153ec07bd4 · outbound

This paper cites DeepRAG: Thinking to Retrieve Step by Step for Large Language Models.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent DeepRAG: Thinking to Retrieve Step by Step for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T00:37:24.065196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:37:24.065196Z digest=sha256:dedf0d97073b91e0c324aab2ecb9629f25dabda7b64222ce925d8994a532fba1

Observation 12b88403-6e6b-49d1-a9fa-03a354bda57c · outbound

This paper cites Self- rag: Learning to retrieve, generate, and critique through self- reflection,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Self- rag: Learning to retrieve, generate, and critique through self- reflection,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:28.108699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:24.150643Z digest=sha256:7206d15157ed00c57ced1e63832785b34234c080e83d40a17844b17f5fee0f25

Observation 7c021b4d-6adc-46d6-8dcd-28b0f7a9678c · outbound

This paper cites HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:37:24.239405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:37:24.239405Z digest=sha256:c7819de9f86806767c1d23f8b6d22f88bba78902dd7b245ad3f92c35fcc90dbf

Observation cbf7f50f-2bf9-4d17-9be9-06154b315522 · outbound

This paper cites Beyond the limitation of a single query: Train your llm for query expansion with reinforcement learning,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Beyond the limitation of a single query: Train your llm for query expansion with reinforcement learning,

Reference 19

Resolution
verified exact
raw_fallback, observed 2026-08-06T00:37:25.657837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:24.332702Z digest=sha256:d342e432e88a77c9df9f050859312fc6b99546251e6e94ce19c427de0cc53deb

Observation c39cdff7-14a1-448e-923c-1e6dbc1b3a2c · outbound

This paper cites Frugalrag: Less is more in rl finetuning for multi-hop question answering,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Frugalrag: Less is more in rl finetuning for multi-hop question answering,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:27.877953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:24.427769Z digest=sha256:826f71bfd9addaa87ae372c7b668bdab368f3b49b8eb026a7e379aad02b12054

Observation 2381ac15-4ae2-4160-822a-36d73801b58d · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T00:37:24.521869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:37:24.521869Z digest=sha256:734a8fa7246f822b6d58a64f14b5d4bf19b9157ae5fc2c367850cc30ec4738e6

Observation a0edeed5-dc2f-4829-b781-6d3531d018e5 · outbound

This paper cites Search wisely: Mitigating sub-optimal agentic searches by reducing uncertainty,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Search wisely: Mitigating sub-optimal agentic searches by reducing uncertainty,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:27.546859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:24.612533Z digest=sha256:d0f6f92a20b23870a525c8ea130bc3873ecb77ee50e208d73e23159f5428e5ee

Observation 2c1bf993-9c97-46b7-a545-892b5d0e4681 · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:37:24.709002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:37:24.709002Z digest=sha256:79a01a0d32e38052c5839304b2db78042ec9c33f2b4618e03522791db191d790

Observation bc26b024-8e89-41d8-a5a1-cd82da4b0c6e · outbound

This paper cites Proximal Policy Optimization Algorithms.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Proximal Policy Optimization Algorithms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T00:37:24.786950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:37:24.786950Z digest=sha256:6e1315f903c666f3e7d9f6f58f6ac515a2cfe95bf8142f086fc7df158c6e2dda

Observation e0c35b74-3571-4c96-848e-98407ef02022 · outbound

This paper cites Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T00:37:24.848150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:37:24.848150Z digest=sha256:b530999f09caf93d6cd57f547ca7b15f60209bfe686cb284558b1edacd6fb9de

Observation e4355908-5b51-401c-9410-c6ab1fbab9d4 · outbound

This paper cites When not to trust language models: Investigating effec- tiveness of parametric and non-parametric memories,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent When not to trust language models: Investigating effec- tiveness of parametric and non-parametric memories,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:27.210753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:24.909995Z digest=sha256:5990bcac46e7bf45522a1d08f4dab9f6702cdc3301f86bc46f7faed5aad7ee48

Observation b61f4cc7-fb6b-47d7-b93b-0d4dd29fb68f · outbound

This paper cites Con- structing a multi-hop qa dataset for comprehensive evaluation of reasoning steps,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Con- structing a multi-hop qa dataset for comprehensive evaluation of reasoning steps,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:26.915806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:25.005957Z digest=sha256:c30a0221267771ec1c64f6ecf3ad9b099e2cc13e138b7eccd490442d27680c92

Observation 83f4963e-bad6-4ddf-af4b-6b0a7d284845 · outbound

This paper cites Musique: Multihop questions via single-hop question composi- tion,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Musique: Multihop questions via single-hop question composi- tion,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:37:26.453902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:25.088278Z digest=sha256:697f2c2d8c1ee721d06a5254598c9b45a9aec7d3f4d9fe040d6d6d3923b7aab0

Observation fb31c455-c342-4e19-8fb5-5bc8dfe86381 · outbound

This paper cites Dense passage retrieval for open-domain question answering,.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Dense passage retrieval for open-domain question answering,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T00:37:25.196726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:37:25.196726Z digest=sha256:d24f669d66035516953c426cca9188e36f69e8c7912676ca52d3fd9f5e623f20

Observation 5187a253-79c8-4d32-a87c-c7495542da9e · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T00:37:25.304636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:37:25.304636Z digest=sha256:3f728f67158e4187b3cbe8b85b62c59d06a33187786f1b86a1f02ece899e9660

Pith citing papers

Observation 54373557-c971-46ed-9353-69796bbe3beb · inbound

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent cites this paper.

PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:37:26.163501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:37:22.965089Z digest=sha256:f1bfbcb25080b3b847eb45f0604b74f7e09fbe8091276ebc8c93a0177bbf6679