Pith. sign in

Paper Citation Record · LEDGER

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

As of 13 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2608.03979.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03979 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:44:32.473958Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 07c32dab-92fe-471c-a56d-9a84f9aad946 · outbound

This paper cites VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.203082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.203082Z digest=sha256:d30e84cd712b36ace33ec5e8312c406fddd9591149054a993c7ea942db47b12f

Observation a50a00fc-ade3-4222-aaa8-9188f8a7fb9d · outbound

This paper cites Qwen3.5-Omni Technical Report.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Qwen3.5-Omni Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.295574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.295574Z digest=sha256:8181365701902930a4b714bcd4d209634a3afe361f990720dedf52bd6d6979a4

Observation 3b8658af-e80d-41ee-954c-72225ab03aae · outbound

This paper cites VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.371880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.371880Z digest=sha256:0ddd7b9cb54628526e0317e1e6d7405d3c71fda188609f50e1c68532b170a057

Observation 72a4bef8-678f-484b-91b9-3bd47f45f266 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.468763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.468763Z digest=sha256:eb54a76161b2c934c51263252c4e763724b2cfa325044ab00bbe48d156860027

Observation 24579356-993c-4206-8feb-9216a76b1c50 · outbound

This paper cites an unresolved cited work.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T04:44:36.137174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-05T04:44:27.491048Z digest=sha256:8bea0c6e354b3a6b0462379b88193cbc983d4c84745a655282a22774209fcab9

Observation 8497ac67-a285-40e7-acc5-e1452a4a9c0e · outbound

This paper cites arXiv preprint arXiv:2601.22060 , year=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent arXiv preprint arXiv:2601.22060 , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.545939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.545939Z digest=sha256:5189f34956741326e7ee51104c062954b65a4911b952c87413c05212d747578c

Observation 066cc371-bfe7-41c3-a0a0-cb2a65ff824a · outbound

This paper cites Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.675736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.675736Z digest=sha256:0e7db64143eecbc34361568dcf274217483eb1f979052a71c9ba2bd4502251af

Observation 0053b198-a6bf-4a4a-bfa7-edbc5a969a16 · outbound

This paper cites OpenAI GPT-5 System Card.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent OpenAI GPT-5 System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.712772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.712772Z digest=sha256:90d60eed737a08f7cac356ea96aa3a935e2a960504ee997151137791904c5483

Observation 3f93d5c3-1b82-428c-9f88-138f0981786b · outbound

This paper cites arXiv preprint arXiv:2503.17736 , year=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent arXiv preprint arXiv:2503.17736 , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.756752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.756752Z digest=sha256:ccbe14800cc126e106332ba15bd92a72b34fe89d4c03216488b567d7abebad98

Observation 4501895f-dd9e-4dbe-825e-6fae3ace9fae · outbound

This paper cites arXiv preprint arXiv:2510.01304 , year=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent arXiv preprint arXiv:2510.01304 , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.789044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.789044Z digest=sha256:9f20c86f6c29f580d87e7659d668774082baf36c9a8b1479e8e7bc653dd27bde

Observation fe893917-5459-49c5-b0a6-37e3b86c925b · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.880473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.880473Z digest=sha256:f1dc4bea646d9c826f52184032bf4dbb0e25db35fca4319cc3e373c9423a6d82

Observation 944333b8-061f-453f-8787-94df744891f6 · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:27.956682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:27.956682Z digest=sha256:a62cbf2a331e9a1b5232129affb62120d47077daaa329dedd8a917f648f97c32

Observation 89d57bf1-71b3-4fc6-ab0d-f91f05759ae7 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.010688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.010688Z digest=sha256:92fde14cfd91c969b3e03fa3198c88a858b60beaac5d953f54c69b424d0a8178

Observation d364a7f1-1aee-41cc-914f-30a67aab5a2b · outbound

This paper cites European Conference on Computer Vision , pages=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent European Conference on Computer Vision , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.072160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.072160Z digest=sha256:f40459c2986d8b4de03d5ef0a5a596ea126183b5516fd06c1c9bb1152273915a

Observation 030cf6ed-183c-4cb8-8ab7-b2d7f1330bff · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Advances in Neural Information Processing Systems , volume=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.135146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.135146Z digest=sha256:a67e5a143e377f5251f6a1bf35010124e32302eca318011f3a0d7d2072a59c73

Observation 0fce546a-b5be-478b-99ee-a3188c212469 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Advances in Neural Information Processing Systems , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.197665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.197665Z digest=sha256:0b8f481200ff60398e3f2c77fb0b2c3ff73b5807f40a90b5331882615602d4d6

Observation 7f3e6fec-a6c0-4ea9-94c1-9ee706bf5810 · outbound

This paper cites an unresolved cited work.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.258543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.258543Z digest=sha256:7352ac8a2a00309a333bb8e7f612731e40da1426bbd23803730d3e935ee3df40

Observation 6263cb7b-7c51-44cd-98ef-84d851c2d80d · outbound

This paper cites arXiv preprint arXiv:2601.03193 , year=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent arXiv preprint arXiv:2601.03193 , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.318851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.318851Z digest=sha256:4c48d18adeca85bcf8ba6baab27299981c5949865980480e876681591ca8cd3c

Observation e3544784-064c-427d-8fdb-2f8c6f9a08c1 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.379916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.379916Z digest=sha256:fec695d6e5a8661285051b95d8c523778cfc1b9f5688162b18b74d2a9e7b8616

Observation 7b605e53-04d2-4b19-91bb-36281c8f2109 · outbound

This paper cites arXiv preprint arXiv:2511.22134 , year=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent arXiv preprint arXiv:2511.22134 , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.406524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.406524Z digest=sha256:fafa45f1dfd7bc8d75d73a3ed8670ac041b8e0d82dc70990d09e81fccfeeb00d

Observation da4b10f7-2315-4881-8542-3bc14fb3835f · outbound

This paper cites CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.494145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.494145Z digest=sha256:1c7d4e743f778b2fac4d2f02be25cd875af3c10c8196d608de3e6ce3e245376a

Observation 097f2bcd-f015-46ea-b082-5db26a881451 · outbound

This paper cites Interleaving Reasoning for Better Text-to-Image Generation.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Interleaving Reasoning for Better Text-to-Image Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.524092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.524092Z digest=sha256:e2413a99e290ee0bce5d5736c2c94a718d72ebc12ce4505fbe2d4de21646d6fd

Observation 9075839e-afea-4ee5-babc-20563e17991c · outbound

This paper cites Seeking and Updating with Live Visual Knowledge.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Seeking and Updating with Live Visual Knowledge

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.589607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.589607Z digest=sha256:16b31d1de1656299ef21859b8924c6ff36d178a2a916a549b870d3530112d0e5

Observation 637e9dc8-9f97-41b9-8b13-f92c671435e6 · outbound

This paper cites MMSearch-R1: Incentivizing LMMs to Search.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent MMSearch-R1: Incentivizing LMMs to Search

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.704735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.704735Z digest=sha256:caffea40157cee7f94923e142cf3cf14d81037d9e66abce8918c8c2ae2dd4be3

Observation 4a65b9b0-cea0-4aaa-9adf-de440fdf79f1 · outbound

This paper cites arXiv preprint arXiv:2510.12801 , year=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent arXiv preprint arXiv:2510.12801 , year=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.859609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.859609Z digest=sha256:6aa2f080632af6be6ed6cf446609f138cc6591b2e1ae000d8593f46bd32c6b03

Observation 93872b99-4b08-4ebc-bb94-d7dcbdde654e · outbound

This paper cites WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.969281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.969281Z digest=sha256:4fe0e3d5ac2dc5caf0f3e8e27b940e8c5406decf35589555357a2641a229ce73

Observation 8fcf68b2-ada1-4deb-8469-b9e11ef52053 · outbound

This paper cites DeepEyesV2: Toward Agentic Multimodal Model.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent DeepEyesV2: Toward Agentic Multimodal Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:29.094287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:29.094287Z digest=sha256:ad12914b7671317e5039e404b2987167908bdd6133e91d6697cfcd71d2b9bb79

Observation 635dcb0e-469e-4277-af42-37fa7bfef9cc · outbound

This paper cites Tongyi DeepResearch Technical Report.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Tongyi DeepResearch Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:29.189555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:29.189555Z digest=sha256:53eaddb6f55c4dfb5e171205a7321885004bfbaf06e22a447bc455feba859746

Observation 4b2f2681-3270-418e-bf12-105ca9a541a8 · outbound

This paper cites WebDancer: Towards Autonomous Information Seeking Agency.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent WebDancer: Towards Autonomous Information Seeking Agency

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:29.293979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:29.293979Z digest=sha256:d45764737a8c908daed92c995af0847ecec2f0059433b38ae4469dbae5ff893c

Observation 88dfc456-8fce-4535-9a97-876e6a2a419d · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:29.403241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:29.403241Z digest=sha256:9dd85300cec713ce82ad04c33a1ef34bd81d4d011e7e991c5c70ff1ea5f64ddd

Observation a40ec824-5d1b-417f-b688-9dec1a1b28c4 · outbound

This paper cites WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:29.547768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:29.547768Z digest=sha256:a9f0f50f286d9c5caf2f16810ef1b2e56dd13808f82fb769b422129c9c4ffa63

Observation 5372b521-88d0-4cdf-94c5-6ad3f667ed21 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:29.646003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:29.646003Z digest=sha256:ba2425418a74e1142f877312ddee162892cf6081b27247736f2f008fd44da783

Observation a7cce302-aee3-40b2-ac08-93fdff9de299 · outbound

This paper cites IEEE transactions on pattern analysis and machine intelligence , volume=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent IEEE transactions on pattern analysis and machine intelligence , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:29.746477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:29.746477Z digest=sha256:6cc9b6e7881f426b23577f601bd43985c57787474ac88bc51647d795f60e74f4

Observation 5bf954f1-ee33-40a1-9b60-d5801e39a2c0 · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:29.820822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:29.820822Z digest=sha256:d6fc16eb2e108d8ecd31daed494ae059ae5a662ffef58036dcb3461d962da2b0

Observation 0cc449e3-67e1-4360-a6be-bf242213d6d8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:29.935629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:29.935629Z digest=sha256:1e34a0ee02a09bb772c1209a1fe69084186122bb8f5e72ebbcef27bcd0fcdda2

Observation 52cf3f54-6486-4036-8185-8ffa131cf1a7 · outbound

This paper cites Qwen3-VL Technical Report.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Qwen3-VL Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:30.041611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:30.041611Z digest=sha256:2d4ee455131fc65b2740bc15b2f7881d8559bd199bad04f3e8c67cacf2f721b5

Observation 75e65da6-332f-40f3-9a91-415fc5092505 · outbound

This paper cites Proceedings of the 2024 conference on empirical methods in natural language processing , pages=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Proceedings of the 2024 conference on empirical methods in natural language processing , pages=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:44:35.881911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-05T04:44:30.150386Z digest=sha256:806974ecf4d8e1cfbc4ca1da0672247c4382eda81053abdda823dfeb042d7b90

Observation f174c949-a9a0-4d90-a449-9bd883251944 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:30.255721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:30.255721Z digest=sha256:fd37f45024ce3faefa15bad83a59ebab3f9af33056ddc7623557339fa86c7898

Observation d2bc9930-dabe-4aae-ae8f-ac747acc8859 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:30.356821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:30.356821Z digest=sha256:e6f73f9197d7a5d57f4c6ccdaecda88c330afb52328f9b4c9784dc81011078fc

Observation fba83264-40dc-4253-a0ee-d9b25e798be5 · outbound

This paper cites LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:30.446320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:30.446320Z digest=sha256:e3e3c84036280b03e8e75cd074279976a2b68aba5eaa95777644054d23cc2179

Observation b943b3f8-bd6f-475f-b17a-530aae2b75e3 · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:30.525201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:30.525201Z digest=sha256:1792ab6a91c06166e3963c9aab0706d98c5849b8885bec6a9038932283d743e8

Observation 2c336cee-0198-4967-87ef-4c11cc1ca7a0 · outbound

This paper cites Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:30.600425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:30.600425Z digest=sha256:6caf5cc5b7ca7130300595165bdbc92daf77a234323a914766d2ebcbe7105378

Observation 58f20515-b1ce-457b-a62e-9552faccfd82 · outbound

This paper cites arXiv preprint arXiv:2512.14870 , year=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent arXiv preprint arXiv:2512.14870 , year=

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-08-05T04:44:33.476174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-05T04:44:30.684319Z digest=sha256:603769273fca6ffbb733bac1a26a1377da0fdbc116484043e02508c3f2cf0310

Observation d2f5b0b7-75a7-4c28-89ae-467e4091ebeb · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:44:35.651100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-05T04:44:30.761599Z digest=sha256:34cf17a7ce7b394dfcb7af0324383be37988e8b56cdb771573b4777d622493a5

Observation c648c6b9-d2b0-44d3-a52c-d7e55d67019e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Advances in Neural Information Processing Systems , volume=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:30.866744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:30.866744Z digest=sha256:0bfa25fe4d8f626a6e7caa1f5750941172f9a9132f6300ceff21d109a646ff66

Observation 5a1213c8-f9e5-4a53-9fa0-98921ecad209 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:30.939009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:30.939009Z digest=sha256:5bc976af845cb9fcce3d1b6a8959dd44f241442b2bcba7d1a90ffd30369dd7cc

Observation c3e4ad40-9018-4eea-9bcb-d36c6c3cc83a · outbound

This paper cites arXiv preprint arXiv:2603.19217 , year=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent arXiv preprint arXiv:2603.19217 , year=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:31.001168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:31.001168Z digest=sha256:96a33b3758f487e258f4be9e31c324d129242346aa13f38475c894669c986d57

Observation 164e7250-c2f0-4bca-8069-fbb09c9e92a2 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:31.067191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:31.067191Z digest=sha256:62a9a558922c9f860b3cc7245eaca8e1868934741b28be93893b3900994b121d

Observation 3983520e-04dc-436f-a665-faa8d643b5f2 · outbound

This paper cites arXiv preprint arXiv:2510.10689 , year=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent arXiv preprint arXiv:2510.10689 , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:31.184609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:31.184609Z digest=sha256:55c1d95ed5cb62bd6ec7a92af285e9d70dfa2efd40278ebb2d436853b55dd3d7

Observation 16bdf053-cf44-4b4f-9257-a8077c613024 · outbound

This paper cites MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:31.241489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:31.241489Z digest=sha256:3188c2683bbd4e8489ad5a4f11ae01465d829ed2dc01cfe34fc1e8850a15bab0

Observation ded02916-e121-45fd-940a-b0ac7d8bd37d · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Advances in Neural Information Processing Systems , volume=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:44:35.465510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-05T04:44:31.286651Z digest=sha256:3ca9d15996a756c9146fd3fe7bc45f4e4e9e26fbb4e6a4cb50433b3bbb44170d

Observation 7458f876-fe13-4bb4-995b-6e8cfd35982f · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:31.371937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:31.371937Z digest=sha256:89fecbc0a507809065715d5722a07cf524585ceaf17ec880754d703f840c4737

Observation 62a1bb0c-ce06-41eb-a7e1-6a210e69a66e · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:31.444260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:31.444260Z digest=sha256:b23aca8d968afec8dd013fb765cf54ea16ee6d2d6596a416ce561d3041b93c36

Observation fd6709aa-23b5-4ed4-8d69-c5906b3ea488 · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:31.521258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:31.521258Z digest=sha256:b192088a142f270890b8f3d9aff8cc6e3974344546b340996999c9fe81e5b09e

Observation 6b64bbf6-ff09-462c-9140-ce71861d595f · outbound

This paper cites and Han, Rilyn and Fei-Fei, Li and Xie, Saining , title =.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent and Han, Rilyn and Fei-Fei, Li and Xie, Saining , title =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:44:35.297932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-05T04:44:31.597799Z digest=sha256:fdac7a50b9defe716d6f06541a3b162216337cbc294cc22b89a29e5ec92199e6

Observation ddfc1082-a5f5-4cf3-8aae-e386e27d7e0c · outbound

This paper cites 2026 , eprint=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent 2026 , eprint=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:44:35.120998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-05T04:44:31.649344Z digest=sha256:915cbb41d05791a28cd583d5eab22d3ad2adf2396c236d71218135c4a572ec4d

Observation 60e642cf-e2e9-49b2-a522-ff4045044394 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Advances in Neural Information Processing Systems , volume=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:44:34.922307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-05T04:44:31.727354Z digest=sha256:c94ca24f0d54e22689e58965dce99c56c12c4873dc21c8c1e1086d1c14480dc8

Observation c62a29f0-cc7a-4563-8e9a-40e7fbf97594 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent WebGPT: Browser-assisted question-answering with human feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:31.811427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:31.811427Z digest=sha256:334426824b3a96a7ca394f13cb503977272b079923a7ffa032c88e1e31e510dc

Observation 9a4b58f0-e5f0-49c6-8756-72d8140e0c8e · outbound

This paper cites Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:31.890534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:31.890534Z digest=sha256:eb29771d898e3659584fdbacfc2ac873b9de76ec5e7a93a62d1074ae708aa10e

Observation aa25b629-98da-4b97-908c-a78538470cb3 · outbound

This paper cites OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:31.997680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:31.997680Z digest=sha256:2fff1d4952bfac0f643db4df7733055413ab485dee49efa63b81f4e96d47a6f0

Observation 694c7a00-1928-4c6c-b21c-cf47442a0eeb · outbound

This paper cites arXiv preprint arXiv:2602.14234 , year=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent arXiv preprint arXiv:2602.14234 , year=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:32.078052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:32.078052Z digest=sha256:38b80b6ba344a4ecf42a23fd0f827cb15805b897f4857c80deba6520bb42f754

Observation d773791b-43ce-4c1a-b8c8-42b5bfb8a1c2 · outbound

This paper cites arXiv preprint arXiv:2509.25027 , year=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent arXiv preprint arXiv:2509.25027 , year=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:32.134345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:32.134345Z digest=sha256:40ae02a5ef857e1e3aa8c2276f0d056a98ed6eb30cc59e1c660ed8ae720760a1

Observation 731cdedc-1297-4081-8a5a-4d7d4122a14a · outbound

This paper cites Gen-Searcher: Reinforcing Agentic Search for Image Generation.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Gen-Searcher: Reinforcing Agentic Search for Image Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:32.205411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:32.205411Z digest=sha256:be03f7e349c58d34df10db85ce917f9db8706b1561f7283a24de75bfd1b0a46c

Observation 9a810b8e-ed71-4deb-9371-cac9960f024d · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Kimi K2.5: Visual Agentic Intelligence

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:32.269924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:32.269924Z digest=sha256:22c8da5d01980cc6b69c848706db8d103a23a8fd0c539b7e21192d63dac68594

Observation b1f64d43-be90-44d5-8b2c-8f0dd8dd4960 · outbound

This paper cites Qwen3-VL Technical Report.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Qwen3-VL Technical Report

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:32.384273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:32.384273Z digest=sha256:992f311beda9027fba166ac760b1917213b5db326b97c8ddf8231f2bf5df3b1a

Observation 4bfe723a-b2cb-4c3d-87be-94251c02948e · outbound

This paper cites International conference on machine learning , pages=.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent International conference on machine learning , pages=

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:32.473958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:32.473958Z digest=sha256:e753e2d054f4e81d95a0244fded511a8245121ba07d75c6ca963a4cbca17f04d

Pith citing papers

No inbound Pith citation observations are available.