Pith. sign in

Paper Citation Record · LEDGER

SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2404.16790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.16790 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:48:30.185030Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-05T04:30:40.718071Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eb1b67c9-cb8f-4cf3-9bed-c91a0b6ad9eb · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.782957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:13b1e9ad58c2941693ea5dfcc8f0851c8ddc6b802cf5b346177f115e9d6d4bb2

Observation bf601f4f-cc32-4b4f-884f-34104f92a460 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.074856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:c8fe8c11813baaf9216904034a233bff0e26b8215a8d1a569b2a5963e241d7d2

Observation 7fed4a12-5eaf-40e0-bb7f-58a14df4529b · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.706356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:d252c995f2984c1f1cec932c9f7475e77d49ae659b71f218d0c61c741884a362

Observation 6ffa1d7f-aaee-47f1-9efa-8d2f643abf52 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.337965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:7658e2ca67cd19a5132d507e0ab098d8789eaeb0be82052bfd5a430bdb34da2a

Observation d0599d4f-1320-4427-957d-93ceb82ad4d3 · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.185030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.185030Z digest=sha256:5969b3dc98801eaae5ff725d536279bb1dbb9cf4e1b3c141915383c9e28d7a98

Observation 3e46d437-d4fa-4afc-800a-916d564b420e · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:23.943515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:23.943515Z digest=sha256:6237c437594e4e468284d2df606316db570d7334d0d379a06c2f5431ecacd09d

Observation 0f0d8dc1-58fc-44ab-9453-12b949975ae9 · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:13.998784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:13.998784Z digest=sha256:b600eff929149556ab5b5f4ea13ba63c2e22a0c55ff03180e37823e23387e6d0

Observation e446c60f-acc8-4f68-9a7d-af9aa250ad29 · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:08.173192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:08.173192Z digest=sha256:113a73c2417dfc7cb680bbbf846a0ce7e1b099023e86896f70eac4bdb340e71e

Observation 26a51221-4347-4ac5-88fe-8b51366e9d8b · inbound

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents cites this paper.

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:18:54.074131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:18:54.074131Z digest=sha256:ae818f25a397e9ccd1d4fab543ac59056a7d08da93af76d11075d98faa4effd3

Observation 92bbbf6b-91b7-40b3-8fd0-8423d9c3780a · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.159584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:7a4594a3c04f81c843d17c5e1afcfc71a047f0fc1fbf3fd995469360385d359e

Observation 93a3a951-c912-49f2-8512-585b1fc62165 · inbound

BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning cites this paper.

BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:29:24.990122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:29:24.990122Z digest=sha256:84323e67d10df4154f99d12d6fd170dd2df21e7a06cf9425b1ef3c89187a66a3

Observation a1b56a09-3792-4a48-af5c-26eeedbf0746 · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.335668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:1d48194154e39f1e60ddea5c2e1a8207e2886e03a68b288eb6bc71dd65f15728

Observation 8b799e30-2925-4d39-a890-2e38cc6736f3 · inbound

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR cites this paper.

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:20:11.514676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T20:19:52.701262Z digest=sha256:f8f09af643fb3ea76f21c8e98f00f15ba31863db5a84010b6186166c4abcdc29

Observation d2f8dcbc-db99-45ae-9962-7e9141e36ba9 · inbound

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning cites this paper.

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:01:05.178856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:58:03.654150Z digest=sha256:62b8af464ad63b224cd525e9853535b8eadc9a895e619b77b20ed1c726c7b270

Observation 9ed6f8f7-af64-4e22-8ed1-c90afa19f5ab · inbound

Learning More from Less: Unlocking Internal Representations for Benchmark Compression cites this paper.

Learning More from Less: Unlocking Internal Representations for Benchmark Compression SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T06:01:57.751567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:01:57.751567Z digest=sha256:f4425da1cba651d5d6de0ddf642045b7870c84916090341c5d8be4c14199cb1a

Observation 87b8e29e-dfbe-4f72-81af-26fe4762df7f · inbound

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion cites this paper.

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-15T13:43:40.241796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:43:40.241796Z digest=sha256:18f782e0f2ee3861739184615a0dc2d98600239d49b5dd5f60dbd3236693e318

Observation e40648cf-11bb-41eb-af3e-203b45d85a7d · inbound

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models cites this paper.

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:53.805058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:35:21.514502Z digest=sha256:97f7ffce836ad6612dec6389aba7a7ce14c8de85458de1f9a5047e12385f6dff

Observation 3e929891-5d7f-4ccf-9fc2-81d28e7affe8 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:04.284210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:46:36.010169Z digest=sha256:fc5c1382a9a1ef077d09853084065e6e4a44cf9f46ab111843e1ee00c79595fd

Observation 4adad9f9-fad9-43a7-8d8e-15396f9d3750 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:26.854067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:38:06.774877Z digest=sha256:baa179670bcf71581a2275e9f93019d1fc7d9031944d0aaa5ceb07ea5b31789c

Observation 4b66fcbb-5d88-4ff7-abdd-cff4dad2ba57 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:27:29.383248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:26:59.917150Z digest=sha256:a77e29282e7381c780661863679e114c0e7674061e53bad2ae377cd49abda3f2

Observation f451cf29-e6b8-4dd3-a0e5-1b08896683eb · inbound

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow cites this paper.

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:24.877582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T09:02:09.075097Z digest=sha256:86e9c06040d459082a0a92d67497e29e5cdb6c7c360c526f110ef24be4b95948

Observation 75727995-6101-4240-89b7-3e9c522d0f1c · inbound

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization cites this paper.

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:41:04.248345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:17:33.668451Z digest=sha256:b78f86107a8f37f0885f4c5ff2d1fc9a38f8713a66677a21c23c59c64a554391

Observation b0fe877f-5bf5-449c-9e76-3a59ace3047f · inbound

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization cites this paper.

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-05T04:30:40.720043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-05T04:29:52.480873Z digest=sha256:4fb094955991c8c5320f50d6ef989332e5d7cd76917204c13ed0adc8d6e791b1

Observation abb05a70-1fbc-41c8-a357-413b306d3a7e · inbound

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging cites this paper.

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:58.909922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:24:30.679219Z digest=sha256:3f3db10cb3ddd123122eb4e2764cbb2064d2edafd99c13fe06e1b431f8418128

Observation 16914475-6daa-40bb-ad3e-c59290015fe3 · inbound

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging cites this paper.

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.730394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T09:12:22.712240Z digest=sha256:a2472bbb2394c6c3e2635593a9f0a6ccd820efb6d4e6ebae37baf1e089a2865e

Observation 579364dd-983f-4c76-bfd3-ebb0c2944b74 · inbound

Deep Pre-Alignment for VLMs cites this paper.

Deep Pre-Alignment for VLMs SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 146

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:27:39.577374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T16:26:41.094936Z digest=sha256:dfe7619b4d3d1a6d6c85bb9ff361058c98b1311aba9874b6ec1ecf975c193e16

Observation efcc01c1-daba-4837-90d9-1bdce7d19ce4 · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:43:38.878107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:ee40bf3a566b99c470b7a8d2a5ea6a811130f3a06e65b0ab591346b9e0a4e8b1

Observation 46758ddf-3d42-469b-872d-32c4c5cb52da · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.447465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:d111d1922821109f2de8c2f93327ab83f77e9378284a968f215357e0335e4300

Observation 0ecaf041-da59-4b17-bf3e-18d4eb31c305 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.039683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:0a7011bb43b65d0d4cdd1f1f054032faa114d96f9d32f161792fa271556b5e3e

Observation 929228b3-17f2-4f93-a933-9619e559e750 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 171

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.641668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:71db6420e312c9a89e68ce6ced43d51fa1f3b440d8492bc76354add7a5542cb4

Observation 7c4026ce-6994-4e23-8df0-612519f719b0 · inbound

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering cites this paper.

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.614558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T00:24:20.208132Z digest=sha256:db5b269ccc94c77770b669deb0cfd5aadeeb5dcbf7612d41b0ffda258bf265da

Observation 238a0bfb-d25a-41b8-90ea-c3601466b7c3 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.731258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:de8d8c00c6f4973393780912363f6b84fae59839865d07af5960e6fef39efd4e

Observation 8e6216d3-e66c-4f9a-9e92-8eb022bc2b9a · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.162441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:e273dc1b0154b516e6ea27bfbae563989ce7bc8e361e4f652407e9035c8be7c2

Observation 73ee3721-3cd8-460a-a73f-6e71170c2f86 · inbound

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning cites this paper.

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:07:03.920797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T15:04:18.880016Z digest=sha256:4092f62d59abccc7c627b845781952af7dd12c45d7db430e1cbfbf44685825b7

Observation 6c79f30b-a7c0-4f88-904a-6c8d6a29cdf5 · inbound

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions cites this paper.

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:58:01.879688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:58:01.879688Z digest=sha256:f7a447cacfb0fbe8bc4a8c0ea86d530054a967990bd9048b6ba7f3884fb3c72d

Observation b8adfbc4-0697-4892-8d1e-9bce6b39468f · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 143

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.200481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.200481Z digest=sha256:b6c07bc12f30cf1b812cc6114f56a69c23aec95f4698dccff73011e956de534b

Observation d6899fd5-23ac-494e-848d-f0574a043eac · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.312878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.312878Z digest=sha256:a4266a394c9e91a6d861ecd4f580314b6307b5a46311b1a2b151a80453e7d8ef