Pith. sign in

Paper Citation Record · LEDGER

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 61 inbound Pith citation observations for arXiv:2508.05748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.05748 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T18:56:23.817544Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 61 of 61 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:26:56.603067Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact20
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch9

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation d4bd7a48-51bf-435e-ba07-6ad6cb3d9cd1 · outbound

This paper cites Qwen2.5-VL Technical Report.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Qwen2.5-VL Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T18:56:23.915900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:41d8a0af36e9d40c5d5bb8d1929e1aa5135f8a64fcd75e49862b246e5dcc2da3

Observation a15cbc98-4ab1-4f8e-beb2-7e60657564d3 · outbound

This paper cites Why reasoning matters? a survey of advancements in multimodal reasoning (v1).

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Why reasoning matters? a survey of advancements in multimodal reasoning (v1)

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.857412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:74d02339c9782b722716c91a52b905aa212b6307e3794b805c14822bfd76e8e3

Observation e4c2ebe9-4399-4cfe-ae47-0b775b4bcd14 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:56:23.864934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:c3125525483e64db9ddfc433c14b21fbed056a69836af8d3586617d267b98f21

Observation 499edf28-5844-4674-ab3d-bcb80ccd7474 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.872966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:7192771237dfcb0a775ba69b056bc96d4476b976180e348977a4cea3454d3a71

Observation 76c4a3da-e9a7-4fa5-b398-b019731279ff · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.880172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:f4437709ece1df643cec5f952642ad5d6c09fb316ba6d3a203e649a2a57a8337

Observation a1f07c42-8af2-451e-b620-24358f1f71ef · outbound

This paper cites Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inference.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inference

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.887970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:73d42106914fcae9adc7d6ae963cdf07bed77f9602b18218401fa1a7f0e4256a

Observation 08bc7117-0e4f-42c3-be50-d03b44bf8b8c · outbound

This paper cites SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.895091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:c99058b7d2b1f709beeb006c6b68ce0fe2996b49d790b0cdd77fba04b898ab23

Observation 06a31577-b3b6-41d8-a96b-58f6ee127d9c · outbound

This paper cites FullStack Bench: Evaluating LLMs as Full Stack Coders.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.902839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:5976128c9c7d766375e91ab1c3e20cfd96fa663c23548b4999c5da327016cdba

Observation 1c89c951-07b8-4fb4-857b-98a3de03a057 · outbound

This paper cites Yuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang, Winston Hu, Yongming Rao, and Ziwei Liu.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Yuhao Dong, Zuyan Liu, Hai-Long Sun, Jingkang Yang, Winston Hu, Yongming Rao, and Ziwei Liu

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:56:24.048466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:2455b84be821cd1d825a0a487e942a2262441c54d5f58149c86dc366a96484f5

Observation 9bf24133-f692-496d-b86a-d713551450a5 · outbound

This paper cites Seeking and Updating with Live Visual Knowledge.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Seeking and Updating with Live Visual Knowledge

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.922842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:fd1bab2baac44e05665ee2a175f2c225e444f2d99f58f7a6e78fcf19a361743e

Observation d5fc73ca-9b82-466a-8a88-91440c9e06d0 · outbound

This paper cites Toward Structured Knowledge Reasoning: Contrastive Retrieval-Augmented Generation on Experience.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Toward Structured Knowledge Reasoning: Contrastive Retrieval-Augmented Generation on Experience

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.929624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:cc63e7e27127d9b85b7f47801b8d091e4cb7d76eaf0e96a0ddeb7d1c6b4c3059

Observation 28812c4e-aa19-49d7-b707-b9c88ffe31e0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:56:23.936880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:0c93611220d99e55e465c0e8764af3f3acaf87f06adf1af4af32c46e8c5fc08b

Observation 89b0f690-1830-4613-b407-6e923d38572a · outbound

This paper cites OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:56:23.944685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:1a42ddc16e65de59c44607815a9511ef02f07d40e9f075a3ccc8f0bf431ed903

Observation b7e4f7ab-7720-4187-ae82-22e5d575534c · outbound

This paper cites OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-27T02:04:28.324233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:f55315392e18b311a343ad98b4de779350935780c076e28e00d79c75bca47825

Observation 6320c711-898e-4ac0-90cc-f39425bca9d7 · outbound

This paper cites MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:56:23.959929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:4296b70dda038d8e737dc6c6cd334ef5f8ef840df4e6ae3171a43e42c21e099f

Observation bef75ba1-73ff-4487-bdcc-a4ae44e0fb8b · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T15:37:09.773663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:887fe7ae46c6c8c8bb347c93e55b2a9ead54c314715833f511a064ab20a0ed0d

Observation 22521279-05b7-4ede-be3c-007bf6381880 · outbound

This paper cites Thomas Mensink, Jasper Uijlings, Lluis Castrejon, Arushi Goel, Felipe Cadar, Howard Zhou, Fei Sha, André Araujo, and Vittorio Ferrari.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Thomas Mensink, Jasper Uijlings, Lluis Castrejon, Arushi Goel, Felipe Cadar, Howard Zhou, Fei Sha, André Araujo, and Vittorio Ferrari

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:56:24.052950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:a3bc1eb06e05b82fdbbd1412a012b87d85b31d2491822a9f36ec90c9cbb63b12

Observation 357233fc-e9f1-40f4-a029-ff02ee2e2435 · outbound

This paper cites Humanity's Last Exam.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Humanity's Last Exam

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T18:56:23.980780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:b4b1161b2685434b4c4cef2eb715ff48fa3f07aadf7ebcc29fa29973210b88be

Observation 1d9134d9-2d67-4ef6-8865-6454a3103638 · outbound

This paper cites Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.987203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:67e0f753314ecd14b588251c160e8022f21fbd2758c7b64fb3e58d4ccb7b7511

Observation 8014801f-cc3f-4541-a341-24145d26a7c2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:56:23.992073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:18f6f140d940923a1f13c841c48ce48511cf73eac2ac166f4cb5e1af4455ad9c

Observation ee46ac06-1f6f-40ea-932f-c54783cf6fa3 · outbound

This paper cites Assessing "Implicit" Retrieval Robustness of Large Language Models.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Assessing "Implicit" Retrieval Robustness of Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.997401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:9fcfbc483f2baacc4a50bb229d9147b30854df39ae5ff15bea34cb9d27d1ed44

Observation 0418bf23-1913-45d5-9e3a-b51bc9548e61 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:56:24.002556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:259e421ab86fc8583b0132df390857d4743126c33e357a0f163c90f2c13ee152

Observation a8097119-183a-4e24-b0ff-e5c84960d97a · outbound

This paper cites OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:13:13.788897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:d6d3e446cbcfd287ff71bb055d58f468e0de963c7c807a025bdd8df49778db83

Observation 0e52505f-8ea4-4214-8a84-daf5d6ce139b · outbound

This paper cites WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.014124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:28575a1f9d1c3718be3d7f896a7a51d6737f184b282822f3df61f67e477c13d7

Observation 90c1f86a-9076-421d-9b75-43c0b280e8b1 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T18:56:24.021133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:fc30be9180b121eae27980328b5601be945f003a63b239df06ef8383fdc71e9e

Observation c3c255b3-2864-451a-b5a5-b941577cd2e1 · outbound

This paper cites A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:56:24.027581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:183f96850596fef6941ee6308adf2470df08cc35b27ab29bf0f6623000652ab2

Observation 623233a9-efc2-4d9b-826d-7f857dfd8e81 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.034271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:00737e8b370e6a77c8dc0f96c4d6afa2fd7bd66b048b3cc5c1b1f4f334f9b321

Observation ab0a14ae-da34-4a13-8fb7-81a7de51354c · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:56:24.039726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:bc08b6d7e9312cb00dccc444e481b5876fb9a388d3fe810ecb1769f5a64938e8

Observation af2c8f61-5ca4-4d2f-a1c4-f5ebfe77fa78 · outbound

This paper cites PyVision: Agentic Vision with Dynamic Tooling.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent PyVision: Agentic Vision with Dynamic Tooling

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:56:24.044382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:74c7d08defc4340acc10c2eb4f8bcc9179fda3a2045bc8f51da883e64ebf5daf

Observation 2afbc784-1962-4708-a775-7b80ee1d7563 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:58:59.023888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:231c09b93ee5c5e52520dca80d1ab7f55915186583bd1fc314346a5b98a2c6bc

Observation 703b22d4-151c-4059-879e-d92854343ae8 · outbound

This paper cites OAgents: An Empirical Study of Building Effective Agents.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent OAgents: An Empirical Study of Building Effective Agents

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.848489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:89e6d73fcc76a6488e826dc3fb0ee18a643f931a9858ca8b03ff95cb571ebad5

Pith citing papers

Observation c5665ad9-e340-4a1e-9b73-9b821722e60d · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.603067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.603067Z digest=sha256:1f23ec2f0d0d3f080582f34547567cba8798b90bbe9413f2c13d22c0972c096e

Observation e06f674b-d6b4-4622-9983-709648e89d1c · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 283

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.586305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:4e6ed6425b925a085ddf3625125281f43cfa20069d4d43d2a2158fb40715775b

Observation cf5555a3-91d0-472c-b320-e5c9e72f4d97 · inbound

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search cites this paper.

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:17:55.644770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:17:55.500268Z digest=sha256:44552681f054f54d11be32e4baafcb79aeec42e29fd4841118ee6d2b86b48634

Observation ba2045a0-37ff-4d05-a19b-a233e566c4a0 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 159

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:25.446636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:9c445a99647fa6813575949921b738aa93524e9879447b6415f4d810dc28f9c5

Observation 2917a802-1ca4-4263-8138-551b4280f645 · inbound

Latent Visual Reasoning cites this paper.

Latent Visual Reasoning WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:41:30.307521Z digest=sha256:1e2e0be55f6d9cc6bfe4d0d86793c6875707088285d2c534daff512cc1300e79

Observation 12d06612-af7d-4170-9488-945c6a25002d · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:32:29.566253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:69dc96519c61e49a2750b9ab7641a3e5fc130180677a81ba15cb899635378f78

Observation 8275f9e2-84ef-4b02-b282-00c1add04fb5 · inbound

DynaWeb: Model-Based Reinforcement Learning of Web Agents cites this paper.

DynaWeb: Model-Based Reinforcement Learning of Web Agents WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:37:41.514211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:33:30.444057Z digest=sha256:1aae4808309ce5f56277ce842ca4b6092169780671d13479d7f539072c4ff8b2

Observation 7f4433c1-af72-4bb9-b18c-8b06d1776dcc · inbound

Imagination Helps Visual Reasoning, But Not Yet in Latent Space cites this paper.

Imagination Helps Visual Reasoning, But Not Yet in Latent Space WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T20:40:30.373169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:40:30.373169Z digest=sha256:45870c3a86585da82040ce9c8f08a373bbe1a9cea28c0709f8e0ed0a59fa9981

Observation 36d15297-3d43-4bd5-b8f9-6c0922e912e5 · inbound

Evaluating the Search Agent in a Parallel World cites this paper.

Evaluating the Search Agent in a Parallel World WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:02:10.963839Z digest=sha256:6e830dc07a65bd741c96241c9e893cc99bbeb142ef827ce1dce06f7310fdc97f

Observation 820af3fc-8672-493b-b050-4b2ec880093a · inbound

Gen-Searcher: Reinforcing Agentic Search for Image Generation cites this paper.

Gen-Searcher: Reinforcing Agentic Search for Image Generation WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:18:10.087258Z digest=sha256:69247bf27dbd697061edf1c9a0380e52d5771a45f0c234d1dd875782e78dbf33

Observation dd1b7442-fd09-4dd2-80df-f7763f3c36a3 · inbound

Gen-Searcher: Reinforcing Agentic Search for Image Generation cites this paper.

Gen-Searcher: Reinforcing Agentic Search for Image Generation WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:35:25.175264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T06:31:13.232880Z digest=sha256:bee4a561ba1098e52fc49ada0159193e5901ab62685f66a84cb0a9517ef762d1

Observation 6bccbf98-d5c7-46d3-ad74-040726d283e9 · inbound

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces cites this paper.

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T17:31:08.575993Z digest=sha256:face3c9f27d25b36d6ac706ab078f9f3f4b8925b3b69861c72322e3ddfd9b769

Observation 7b074a6a-42d9-4f4d-b6d9-ebadc9e61e0c · inbound

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization cites this paper.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:90ae85153c3ea208e0c25dade7b71312266f18523127856e68a32af615447f51

Observation 4fe2c27f-2b7c-41d2-a98a-eae3194eb7df · inbound

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning cites this paper.

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:45:21.088097Z digest=sha256:4e119ebab03d2191e4793d6d7dca1b28cd1cbdf568ab3b87dc01a975414afec9

Observation 503034a4-bb44-4dc9-89e5-530b27442ec9 · inbound

Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation cites this paper.

Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:42:35.824427Z digest=sha256:5b13aac9310c6abc5e0960cc4ca33da72310ea8681e3225a0ef68cc57ce62f62

Observation 9c833cf5-f319-4c0b-865c-1a4e9b1e6aac · inbound

PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers cites this paper.

PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:51:39.424754Z digest=sha256:d270bdb9b5ddf7809b8d5375ca18fb27bee6ac7f4c84f38e03168a14f7a4a9ac

Observation 1a554a16-a74c-4636-a155-85442316779b · inbound

Towards Long-horizon Agentic Multimodal Search cites this paper.

Towards Long-horizon Agentic Multimodal Search WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:40:32.137708Z digest=sha256:c257992324766740a5b17de7b505dc73de9b255e6bc845bb8709041f3850dece

Observation 6dd5d7df-1fb5-404e-8fe0-4b0dfd57dea1 · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:09:24.304696Z digest=sha256:9956f9e10b77f1e0c5a914b2b87f648762b09034c4030bc3720cf9736a182184

Observation ffa279fd-48a8-4c9a-946c-f5dc401d0c82 · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T16:18:20.625700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:18:20.625700Z digest=sha256:c6dd7febf3ccbc038261b8e8fa6d424a499d0f36c41e96f3c8cf169c4fedc454

Observation e45160d7-6dc6-4449-9820-7952aa799f3d · inbound

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents cites this paper.

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T03:21:30.732925Z digest=sha256:63cbb7145e338705273b6489f5f0226c28f4a6a48a2a5700212856306a8d71a4

Observation 402544bc-25d2-4de5-b99b-466770a08906 · inbound

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards cites this paper.

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:12:17.469552Z digest=sha256:73ec286d02fc714490ecb4c9d5a0544cd91b42ff08e0d052701c7d174f207f20

Observation 5b9efe9e-9ddd-4c7e-9836-5fada0fa4c1e · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:a6b07261345c1e2d421032fef08d86d2df39fad89d27c8ab32c8e26bee991bf7

Observation 7b83d99e-7dc7-4b98-a7f6-6f1ce5e02178 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:ce333b2bfa10de38294f373e5c49230c610f75a276c376bfd69c5530d4ba0e4e

Observation b1a61f13-3a14-4d27-88c0-2a2cb97f9bba · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:02:40.734170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:8984ffd72af29cf521ef07be505dca1542214b1a33c4d846679c9962719e5c4c

Observation 2b7fa124-5252-4e77-90ff-3fd61ebbe583 · inbound

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents cites this paper.

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:28:36.266167Z digest=sha256:d4a933a2b3bea404caa3e1caacdd2978f0869e48a830d4f11f4563fc220afb69

Observation d0d608c4-2264-4227-8fb3-4d911c741437 · inbound

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents cites this paper.

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:18:01.006274Z digest=sha256:0abf7b25a69519c165cc3a5a2d590ff35bfb70a21d0ccca64611a3aed12ba41e

Observation a655cd2f-6599-4831-89e8-f433ac23cea9 · inbound

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search cites this paper.

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:21:52.972107Z digest=sha256:18f7a76c18cda9da62e9d382f11f89a7b99bb14c360b2bce42e01d048a068832

Observation 2f1efc30-0f67-4ba7-bd69-be5f89ac5d50 · inbound

From Web to Pixels: Bringing Agentic Search into Visual Perception cites this paper.

From Web to Pixels: Bringing Agentic Search into Visual Perception WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:47:43.959052Z digest=sha256:ce189a6aa685b937f8066eacb9175a0aade7308b7ef0d8d040df466fca10dea1

Observation af59cca7-70f7-4c3a-9247-d0d14a7aacb9 · inbound

ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence cites this paper.

ViDR: Grounding Multimodal Deep Research Reports in Source Visual Evidence WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:18:52.801531Z digest=sha256:036b0d4c5c5d84286f067222baa0a6da0ffacf55dc649533418c445a98a7ffdc

Observation 0655a808-fe76-4928-9461-b2593ad1df10 · inbound

FIKA-Bench: From Fine-grained Recognition to Fine-Grained Knowledge Acquisition cites this paper.

FIKA-Bench: From Fine-grained Recognition to Fine-Grained Knowledge Acquisition WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:25:36.746335Z digest=sha256:cde01e2680d03b9168aef42f5edd933fd25f5d3433000f2db4c3b2f3f34da4ea

Observation a736928e-a736-4e52-bc6c-443ec8e3549a · inbound

FIKA-Bench: From Fine-grained Recognition to Fine-Grained Knowledge Acquisition cites this paper.

FIKA-Bench: From Fine-grained Recognition to Fine-Grained Knowledge Acquisition WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:03:47.201538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T22:02:19.717335Z digest=sha256:71dc629906c1a27c4e8e8779231c7cd4c5f53651536965dda6b8c9c221b1b954

Observation 7b0e288d-a7f0-4f43-8c8b-7f1a43b2a071 · inbound

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context cites this paper.

Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:16:07.851098Z digest=sha256:7d57245386090d6a51680501fbfb48616ed1cfe905677175ff3a6b74d5ee16eb

Observation a2ff6fbc-a5e2-47f2-b78f-d6ebcc5accb2 · inbound

Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification cites this paper.

Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:43:19.757653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T13:38:21.492236Z digest=sha256:8e296f70c1b6d58e70831c27d868fb386d89a29c7c558d65a34595ca6be639ff

Observation 7249b038-2dac-4b4f-9a18-70814771cd9e · inbound

Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification cites this paper.

Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:15:01.021796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T19:07:05.208780Z digest=sha256:0505360193a8a4ea84a6d6016eebe733d18b67c82a64f7cf812fdc796f95dc2a

Observation 1dec591c-7eb6-4e14-978f-d87fec19beef · inbound

SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain cites this paper.

SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:13:12.017523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T10:08:56.397296Z digest=sha256:917a7a03aa3b6e4205f96371b6e1e4cf11cd5c1c8374593cbed29ef9929eb6c3

Observation 7f895b10-fbab-40c6-b091-7b5edb9fc2a2 · inbound

SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain cites this paper.

SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T08:49:53.419752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T08:49:32.503461Z digest=sha256:bb8b954c8cdea65673d025908cc7f2fbbb435df528f93e4b0e7b44cae5279fe8

Observation e5e8b6ac-6de9-4a93-996e-17d8e643d459 · inbound

Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models cites this paper.

Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T18:33:50.544221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T18:29:26.731925Z digest=sha256:83e3b0e9620b5ee1cdc79c23c792b2ea302091ff1d6e6fab8b041564ba23a976

Observation e50ccc53-2b6b-4b3d-8d8c-a69cb767320c · inbound

DeepLatent: Think with Images via Parallel Latent Visual Reasoning cites this paper.

DeepLatent: Think with Images via Parallel Latent Visual Reasoning WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T20:22:37.711207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T18:44:39.545911Z digest=sha256:d926338c2b1a50374b90b824013637835a31c191617bafd63edac117d24aec8d

Observation c9d5eefa-354f-4890-ab14-ba686bfd7a51 · inbound

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents cites this paper.

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:16:57.776433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T02:11:11.638029Z digest=sha256:c4ea71e8171f3b90d7152e6bcda5e14e8bd6bd9788ac8c9eba0006a70a31b3a9

Observation d16a255a-209f-4411-92b5-b8d847858a47 · inbound

Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking cites this paper.

Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T16:47:10.186792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T22:20:53.187362Z digest=sha256:48f0dd5b7c4c085e494c44dd45531820a6aa2735a5da9e936d3e43ccee21a3b5

Observation 4e8908a4-db3e-4159-b517-d05ba0ef7b14 · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 270

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T09:50:48.464619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:35877f5648437078a31d088546ce7a945b3211bfc43b45a79e2f8cde8a2e068e

Observation 144f6d7d-5000-4250-8505-b70cbed5271e · inbound

ChartWalker: Benchmarking the Cross-Chart RAG Task with Hierarchical Knowledge Graphs cites this paper.

ChartWalker: Benchmarking the Cross-Chart RAG Task with Hierarchical Knowledge Graphs WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:49:52.168131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T06:02:40.752332Z digest=sha256:11c0ea6da625734806e707f7909e397f0c7358f559d0fa6518e4afcc11190db8

Observation 70d8f59e-a107-4500-b06a-1a1c4eea4dda · inbound

Latent Visual States for Efficient Multimodal Reasoning cites this paper.

Latent Visual States for Efficient Multimodal Reasoning WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:29:57.263044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T00:38:11.619574Z digest=sha256:72d8ac289019c51e2b07e346b208d0dbf3ce3320fefa61d0407e6e170acd3afa

Observation 5d9c1944-0cc8-4bbc-b61a-27e1eb4dd461 · inbound

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning cites this paper.

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:20:06.665253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T21:23:10.051805Z digest=sha256:fbeaa03079d70aea665187ef0b5b7df636e87983d725765996c1a813be8f8ca3

Observation 47a7c605-d050-4f56-b92a-60970bc6e0dd · inbound

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning cites this paper.

Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T14:29:53.361671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T03:51:51.827622Z digest=sha256:9064a6c7dbce8768ac366139cd74a4cf0b9b779cef195f767d6bfa26a35e0f61

Observation 9c69a688-107c-4c3e-abfd-a5f5fa40a24d · inbound

Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents cites this paper.

Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T10:05:41.384777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:49:52.979376Z digest=sha256:ea9679718c474ffec5958a7daac01c7fc1599cd6a7f6952387894eb8649da3b5

Observation b684a661-c7a1-4e4e-9a76-2df8839e7ce6 · inbound

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search cites this paper.

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:55:41.130831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T06:02:48.532478Z digest=sha256:78a20cdb4041d7531bca8def3ec81db2924b31cb240d50d6ec4af27ddda806a8

Observation 6cc76311-02e7-4c67-b456-e86ee37c2194 · inbound

When RAG Meets Query Planning: Logical Query Trees for Resolving Exploratory Reasoning Problems cites this paper.

When RAG Meets Query Planning: Logical Query Trees for Resolving Exploratory Reasoning Problems WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:56:43.883868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T06:47:57.679984Z digest=sha256:fefac00efb57fad2c6b1196e0ad575ecddaf01412ca9549988f0f728023c9e21

Observation 633fe60b-42bb-4ede-9f82-35a552225014 · inbound

When RAG Meets Query Planning: Logical Query Trees for Resolving Exploratory Reasoning Problems cites this paper.

When RAG Meets Query Planning: Logical Query Trees for Resolving Exploratory Reasoning Problems WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-03T19:08:49.360613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T19:05:04.593481Z digest=sha256:20eaaabadc136870c90b9bf38b0997c67a778b90979578f9f27803bbbfbbcf2f

Observation 7bf7bec7-5459-495d-869c-3fd177d991bb · inbound

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning cites this paper.

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T06:04:16.637378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:04:16.637378Z digest=sha256:873a7af45927432417f6562bdd661e56518ba05c1cc1fe267c03c3278523d032

Observation 7c282df7-e02c-403c-bbea-5721de79ddd4 · inbound

SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation cites this paper.

SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-08T20:05:34.061221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-08T20:04:30.942032Z digest=sha256:b6aaa13c515cf189a9b988bb9b14310671bcdb61fd88fa78c1206abdd8a48db0

Observation dd67ed7e-517d-47d8-9f84-5e275e848281 · inbound

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery cites this paper.

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-08T08:04:48.031528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-08T08:00:31.873224Z digest=sha256:22e071736d0b9b1182fc5474af935b80952d34a8875fd0fa640f8dafc4b787ef

Observation 6575478c-4069-42f1-830d-a2a24755cb4e · inbound

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment cites this paper.

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T17:27:26.697652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T17:25:25.161103Z digest=sha256:8ec3287eeda26584dd79d5664a1f17a73c971bc5d9e3f205b52753b128b0fb89

Observation 6e821f7b-1ac5-47b0-99a3-46c50c1f83e9 · inbound

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment cites this paper.

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T15:41:28.727344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:41:28.727344Z digest=sha256:dc2de0e8230abbe7eb50b7b158ec2a4be64154325cf62a81c89d572720af503c

Observation 47fb2a3f-d067-4cca-b29b-07da7e062e2f · inbound

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents cites this paper.

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-10T07:26:54.432175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-10T07:18:57.823444Z digest=sha256:7d0d9e939bb0f815fabd64ef5e132d3b5ef286224fc119549299c67f00380caf

Observation a100912a-4bb8-4824-822e-d1489526ba12 · inbound

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents cites this paper.

Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T07:56:56.599749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:56:56.599749Z digest=sha256:03ca55cf2c45ce9c1ec71dd4a6f9fe4b36af2a6cf109f02757c1197e66f61e69

Observation 3307efb3-a7f1-4fee-9362-b6166c78c5f1 · inbound

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp cites this paper.

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-14T10:51:16.019022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T10:51:16.019022Z digest=sha256:3682f54fdafae36db6d336b58963ab48e35a61044189d24982e51df0fb6edbfa

Observation d8113168-6b42-40be-80c6-509dd82f840f · inbound

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger cites this paper.

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T09:24:20.416024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T09:24:20.416024Z digest=sha256:d575d59b6a1e88b2e8f8ab841d6dac7e73fd2bf7f50dbe0bc75723ef9049f6b1

Observation 2d81ff00-6fa5-4cca-96d8-eb93c7f3c713 · inbound

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents cites this paper.

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T20:12:42.009030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:12:42.009030Z digest=sha256:21b40f4d3ee625754437affe332ae8472426e13adffd1a13b664a840b72f26db

Observation 4f6d35e9-fb6f-426c-a75b-6bbbdf13e80f · inbound

VC-Tooler: Learning Compositional and Adaptive Visual Tool Use cites this paper.

VC-Tooler: Learning Compositional and Adaptive Visual Tool Use WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:21:31.750266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:21:31.750266Z digest=sha256:fcb6bf221dabd2277ec66dab1f4cbb316fb5690d726291f5e2c9fa368fc87009

Observation 93872b99-4b08-4ebc-bb94-d7dcbdde654e · inbound

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent cites this paper.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.969281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.969281Z digest=sha256:63a224fdc973072f5e5bc43ebdcf1b9ce113344862524cb4fe127ad2b34b4cc3