Pith. sign in

Paper Citation Record · LEDGER

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification

As of 4 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 3 inbound Pith citation observations for arXiv:2601.15808.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.15808 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T12:29:06.229857Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T03:36:57.168246Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:45:55.587604Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact12
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b4d81aeb-84e0-4f0d-a949-39facba0d072 · outbound

This paper cites TapeAgents: a Holistic Framework for Agent Development and Optimization.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification TapeAgents: a Holistic Framework for Agent Development and Optimization

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:30:53.857510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:ad97a50905f8f22ae7ee08d62bb7aa0d2262f037c5f67a096e4d67fdfe503ee4

Observation 8d64e2b4-bdb0-42d8-bfb6-0611144d6125 · outbound

This paper cites xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:30:53.853442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:93c52ccf0d5b71da2d58da164b7402b3261a564f7c58837a7402d6dcf9496152

Observation 71ebb346-b115-4363-8325-d79ca0f34d83 · outbound

This paper cites KCTS: knowledge-constrained tree search decoding with token-level hallucination detection.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification KCTS: knowledge-constrained tree search decoding with token-level hallucination detection

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T12:30:54.310156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:614a58d2e2fb0036cae2eaf6fb91656c3c9b0a467427c7832cc7106ade27570d

Observation 31f93883-e168-42d1-9d4b-d7876a02e6d5 · outbound

This paper cites URL https://doi.org/10.18653/v1/2023.emnlp-main.867.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification URL https://doi.org/10.18653/v1/2023.emnlp-main.867

Reference 4

Resolution
verified exact
doi, observed 2026-05-16T12:30:53.793557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:282ab237001855562a24b4c679cff9f1c017db8991ba062cace1c8904b990e43

Observation b96ad2a7-3a35-4532-953a-5050d2e65b44 · outbound

This paper cites WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:30:53.875593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:85cbade9b5a7c2c82d64f94f68a43ff40b2ccd20c4e8b54ebf5c7327df16f118

Observation be25c1d3-fc9a-4a65-aa86-687ff91db3e9 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:30:53.898939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:a73990f2588a9b88acfb197054a54ff29c0603fe61573fcea693d65ef0fe628b

Observation 026dae31-73ab-495c-945c-42a9fff7faef · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:30:53.881950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:3ddefb4cf3c5eab38fabe01ed611e13d400cacddfa21d46890c36a036f59a91d

Observation a8ef525d-b8c8-46fd-8b57-92e6c755c1ae · outbound

This paper cites WebSuite: Systematically Evaluating Why Web Agents Fail.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification WebSuite: Systematically Evaluating Why Web Agents Fail

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:30:53.902823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:04b8d5d5c9d3438208592e44fa275bb2a64b6460dee88d1f1d8dcbbf76ce5e8b

Observation b90de6d2-0ee8-41f7-8d01-98fe6dd19626 · outbound

This paper cites WebSuite: Systematically Evaluating Why Web Agents Fail.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification WebSuite: Systematically Evaluating Why Web Agents Fail

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:30:53.786370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:59a4173e96829f8ed751fbb92ebc37ebcbc09eb424fdf7fd0934f7153732af07

Observation 27b55389-9037-44f3-8b5c-0034a8739e23 · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:37:09.773663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:746d258f7b0298a270003dfc4a7a998a46a0b99ad061ed90a532eb83da75a135

Observation 0a67d5ae-d413-403d-894d-6673788b514a · outbound

This paper cites Agentrewardbench: Evaluating automatic evaluations of web agent trajectories.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification Agentrewardbench: Evaluating automatic evaluations of web agent trajectories

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:30:53.891278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:5c78aaa83b0003189f0c23b868579ed5bed962370e0f143c3fd3b762c5a2258d

Observation 90cb41df-f7d5-42c5-be03-73a1f6f0d849 · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification Autonomous Evaluation and Refinement of Digital Agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:30:53.849781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:b731624d6dbac4a0b70990b70aff9ec738b5f4be8ee4ffe07a678836c9ce4989

Observation 13226e46-a925-406b-81a8-b4fe06df8540 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:30:53.906553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:c864fe33f59329954872953f17825c4d77b64d1b84e9518c859226870e36c061

Observation 5fc8b545-4daa-4caf-9018-d1ad5b35e4dd · outbound

This paper cites Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:30:53.870904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:e492c1c46a62fcc48474391dfe0f6d8cbd7e6af3ccf228711e015fc555a3d74f

Observation d7d99522-aeb0-4016-b5c3-b86392a3243b · outbound

This paper cites Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:30:53.791034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:7c0896690ca872699738bf66c6a86ae5e8c4d4d5a7ffed10181449dfa5ab3afa

Observation 83764175-b33f-4820-a6c4-22956ca5265d · outbound

This paper cites WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:30:53.879074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:29c0612419983bfc6add8ed5d2a7d60a62fa9a4702f0b5b9d90b95015d5789b9

Observation b885cc4b-c826-4328-8b02-ac9eea9dfe58 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:30:53.895063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:309fd49c6f8b6ebd143003f7d18bdddbbf3d0595b0e9ce5e98f0b3c97f3d9849

Observation a9769eec-10b3-4eef-b38c-56cc115b91e1 · outbound

This paper cites TextGrad: Automatic "Differentiation" via Text.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification TextGrad: Automatic "Differentiation" via Text

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:30:53.864528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:0547ce31501597fe4ed51475aa9b6c811844dffc8910f42a0cc4fb8f695e3534

Observation ee259c2e-f85b-4828-af71-34fc4bfc00c2 · outbound

This paper cites How far are we from genuinely useful deep research agents?arXiv preprint arXiv:2512.01948, 2025a.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification How far are we from genuinely useful deep research agents?arXiv preprint arXiv:2512.01948, 2025a

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:30:53.873857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:57ae1ae93145e8483a76bb7f2cda5365e685e4eadf2fc6089575970da3f0079f

Observation 034f1ad8-b452-4d29-a483-52e8d4704339 · outbound

This paper cites Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment

Reference 23

Resolution
malformed identifier
doi_truncated, observed 2026-05-16T12:30:53.796093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:d5868ecd1d5f73734804910707c025a1b277cef42b0012f92c91a1344c3b0ad0

Observation 692c93f6-dd87-4ead-8762-64128bc144a2 · outbound

This paper cites Agent-as-a-Judge: Evaluate Agents with Agents.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification Agent-as-a-Judge: Evaluate Agents with Agents

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:30:53.885575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:f085cf3458e7b3c7cf57163746a2f0c5e956bb9d3608697669b99596f35a8d30

Observation a4a8d047-23bb-449f-89f3-51f18f922c05 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification TTRL: Test-Time Reinforcement Learning

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T12:30:53.780583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T12:29:06.229857Z digest=sha256:3361b6c2f6b0c5240961aa30ee75552a214c178d8fa9766f582746083b1d43ac

Pith citing papers

Observation f72d902f-f82e-445b-88b4-4500d7bcc4df · inbound

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration cites this paper.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:10:23.484189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:d36b5ddcd343aa646e277d62940619db3a4871c97025ee6540866b87164b20d6

Observation bc690305-63b0-4174-98fa-5d095c2462be · inbound

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning cites this paper.

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T17:01:08.148937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T14:18:14.048230Z digest=sha256:509c35fc47a1cac616469569143888a565e47b97fc6f0d8f971a73de575993bd

Observation c513d312-eab5-4b6a-9e8b-b56cbce3139b · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:45:55.589146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:39c54181743868b7341e9ff1bb006bd1de5d9bafcbcd76e673c98b618a7056a5