Pith. sign in

Paper Citation Record · LEDGER

Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2302.11713.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.11713 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:22:10.964965Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:59:46.832017Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b6bfed60-da5a-4134-b7a6-b87f524b9d9b · inbound

PaLI-X: On Scaling up a Multilingual Vision and Language Model cites this paper.

PaLI-X: On Scaling up a Multilingual Vision and Language Model Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:36:10.023151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T14:36:09.971825Z digest=sha256:42d7363fcaf7ea074c455ded24d4f77c79947dc59fad80d9538c37a412921390

Observation a4f59d38-5b45-41f7-af56-f14df06e3ed3 · inbound

VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation cites this paper.

VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:22:10.964965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:22:10.964965Z digest=sha256:e661584d6055354a80c6d066a104eaabbe75d007cf3cd51453f309c5922c2698

Observation 3109636d-67a5-415a-baff-0e63cb21a129 · inbound

UniCoRN: Unified Commented Retrieval Network with LMMs cites this paper.

UniCoRN: Unified Commented Retrieval Network with LMMs Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:15.864357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:15.864357Z digest=sha256:fa3d386dd01f4a6b617d6e66a1ad50b391f7059cedb5e67fa00daa7ae9d4d1d4

Observation 16a8baad-8059-40d3-af0d-6ecc6d78427f · inbound

Towards General Continuous Memory for Vision-Language Models cites this paper.

Towards General Continuous Memory for Vision-Language Models Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:45.236100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:45.236100Z digest=sha256:e6dcb267750a0ceeb1ee548465a79711eebfc576b988d8dd12c0a54b86b8ccbd

Observation c1780cfb-b22e-4278-807b-4b7dcdd9248a · inbound

Benchmarking Poisoning Attacks against Retrieval-Augmented Generation cites this paper.

Benchmarking Poisoning Attacks against Retrieval-Augmented Generation Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:55.697518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:55.697518Z digest=sha256:0fb979727d33960cfec1bba1f1f49c14200f034522e1053d455f05a7e3d37c96

Observation e8b593f0-fa3a-43dd-82f9-2874a2ab121a · inbound

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding cites this paper.

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:17.778638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:17.778638Z digest=sha256:adf66313742a7c163d7ba850a6adc43cd3a24920952fa1929231eb23e42231aa

Observation bfa67e94-d0bc-4c11-8444-04a73de81995 · inbound

Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation cites this paper.

Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:52:19.894702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T13:50:30.090068Z digest=sha256:c97f0bf6188bf6955dc25216eb5c4caa9b6b372ea80206d49f76d7211d326641

Observation 261c7fb3-9da8-48e0-95c6-3f92a1fafaf9 · inbound

Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM cites this paper.

Spa-VLM: Stealthy Poisoning Attacks on RAG-based VLM Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:40.614803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:40.614803Z digest=sha256:d77beeb9849df7af17f24b3e610ade187672024d20b694886ca5d003775dfe6c

Observation 0eed1381-9216-4444-9d52-308e0adc1ca9 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:21.914960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:21.914960Z digest=sha256:5b7aa01c36be83fc7876dedb1aab93b823d761e6d45f10f0f78366f418719ce3

Observation 17b0c8a1-b73e-4bf0-bf58-3901c577b7e7 · inbound

MMSearch-R1: Incentivizing LMMs to Search cites this paper.

MMSearch-R1: Incentivizing LMMs to Search Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:27:04.427821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T15:27:04.228144Z digest=sha256:ca40607d16ef5dfc1f7cc4c79c6dc632c97760afba1832d6b46a9c7394f939ed

Observation b2b48a63-6335-45e0-b0b6-35b2fc97d4c6 · inbound

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning cites this paper.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:32.356535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:32.356535Z digest=sha256:af8b992a666a4086a925566a2b29c2554d945bc48275adc2d34e0eb223498d48

Observation a7d9639e-f645-4088-80e7-25eeee52864d · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:31.536273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:31.536273Z digest=sha256:912de2ca1b2d5a5204e6edd89d9038118a760b05fe68da4fe5d360c9dbf66aeb

Observation 76c4a3da-e9a7-4fa5-b398-b019731279ff · inbound

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent cites this paper.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:23.880172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:19e678156d30443f21bb487416311112c8420d7ce93fabaa256e61fb7f908e4a

Observation ea7054b4-df8a-41ef-a8c0-4066d635aa22 · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.538412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:1d96d4eedb9109005c7693a1ab2614e3bb2f53a0926a8879de198c53bcc5400a

Observation e880ec27-3ea7-45b3-8290-d63e329b92e4 · inbound

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation cites this paper.

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:47:45.512121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T10:46:17.411843Z digest=sha256:7725bde86da55ea22ec93e6c1e425c1c9b637443da725e503bcc26d37cab1f0b

Observation 0f8b9170-1c4c-4886-bf4a-bcaa4ed1c4c7 · inbound

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation cites this paper.

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T08:13:52.772526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:13:52.772526Z digest=sha256:cf1aba7f30760530d4de0e74c1a978a7b1fcca299f4525e273b652c8d2757ac9

Observation bd1c256f-8a67-45d1-bb0d-622d9aac99ed · inbound

Evaluating the Search Agent in a Parallel World cites this paper.

Evaluating the Search Agent in a Parallel World Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:06:19.398333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T17:02:10.963839Z digest=sha256:db04296cea7999152aa9eb9f74d6edcfd5e151e0bb4f1298d37da1aaa36e9c4a

Observation 4fd4f9ae-c567-4359-a94a-2355379060a8 · inbound

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum cites this paper.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:22a1ca5cc79c5d1452556680f409a3f01327aec6072d7ca97201f030ab9d8ab0

Observation 2df65d05-b063-4535-bd42-fb830a64fa75 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:15:50.617495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:cea601b90597df63db8e244f4cb5cb24430981ad4ec13c35a223e4546f7236c8

Observation 3c589ea0-4333-4bed-9a40-68f66b2496c2 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:90ef78cda6b803414ed47ea0837aedd1726c01c6144b75cb92a3d85c58e5ffb1

Observation 04d5088e-69fd-4c27-a7c0-95a147f5c88f · inbound

Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering cites this paper.

Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:55:59.357888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:52:32.352399Z digest=sha256:fd8450efc612e8c892383b11aa2eee9f0cf160aafd37478bb3b0ecaa8a8af4d9

Observation f65682d5-c688-4328-a206-d66109f75ec6 · inbound

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents cites this paper.

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:08.003540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T03:21:30.732925Z digest=sha256:aaaf3716d4e98f79065ec1ec36b4fe283793d8e52eb83a8a8d5177cec787411d

Observation 17253701-54e0-494f-99cf-108285154a07 · inbound

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards cites this paper.

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:09.791197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T01:12:17.469552Z digest=sha256:02539a01530c11c608c430f37c526824eefba3dbbeab32b25c4520219e87eb43

Observation 4dfc9e15-e7ef-4c77-9ef1-356eb94b06a6 · inbound

Delineating Knowledge Boundaries for Honest Large Vision-Language Models cites this paper.

Delineating Knowledge Boundaries for Honest Large Vision-Language Models Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:25.309052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T13:56:19.635005Z digest=sha256:a70bfca0f2673a800678bc002c6b6b931b133e3c545d65ddf29bc6a9fe4cad1d

Observation 6fabff02-2f74-430d-995d-3bcac17718dc · inbound

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models cites this paper.

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.194601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T21:00:25.664841Z digest=sha256:ead384ec542b4d46edfa472d3fa88324c8ca564cfa13a5f5c7c952d0b564b2e8

Observation c7219241-79d1-4bb5-9a7d-668821257ec8 · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:57.778198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:495e8fc4625f7247884dd87b4df7f61eb099720b10eac72afe848f3b4c5f4009

Observation cdc4ff57-90a9-4fd6-938f-fe148251fffa · inbound

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification cites this paper.

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:59:46.833645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T08:12:14.829556Z digest=sha256:52fd29bc738fd62e2ea5676e024fca6043ed905de1dd27fc6f19cb69d2a05749

Observation fcf7514f-5791-4cb5-b6b6-aacdc5ce8d7c · inbound

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search cites this paper.

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:55:41.073138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-01T06:02:48.532478Z digest=sha256:e8acee63d19b59c49c58effdf2532ed7fff18c7eb5905fd7c721f61dfb89ce7a

Observation a1c92dfb-b7d4-4e0d-8081-6a6b85515075 · inbound

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting cites this paper.

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:17:17.746673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T19:13:44.015782Z digest=sha256:f5e14450d1feca5b247da4bf9822e93e80b08fb76dc2e9d8e0696e8cdc4456d5

Observation bfdd7a1d-5754-44bc-b4ee-de7fe08b7503 · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:52.536689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:52.536689Z digest=sha256:c01657dc6e337ad6105ef3f1c2386acadee43a39fc478beed83281a2875b9101

Observation a9ee170d-edfb-4ad2-8703-e2033d4bd1f9 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:14.890550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:14.890550Z digest=sha256:7ec7c646183d931f9bbe659d6f015f305936e55d3a74138c9b1769fb04696021

Observation 10cad550-eafc-443e-b7b9-4b4853ee7ce6 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:50.895716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:50.895716Z digest=sha256:2ae2623fb2a30482df34000e040ae3f3d4366a0429519298901bc5b83ec28e1a

Observation 5bf954f1-ee33-40a1-9b60-d5801e39a2c0 · inbound

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent cites this paper.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:29.820822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:29.820822Z digest=sha256:d3e888bdb5d1eb3c355c0f8ef670a3f673d8c2b0e4799c36e26ada57c97123b4