Pith. sign in

Paper Citation Record · LEDGER

SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2404.16790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.16790 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:48:30.185030Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-05T04:30:40.718071Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eb1b67c9-cb8f-4cf3-9bed-c91a0b6ad9eb · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.782957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:ef0c7ddf08679fb247096227c712c128fe71609f208517a451bf08e74e272d8d

Observation bf601f4f-cc32-4b4f-884f-34104f92a460 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.074856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:ba3ce1c0d2b67e296b4e94f84edef02e279d7520391723979ae3ca8535f74728

Observation 7fed4a12-5eaf-40e0-bb7f-58a14df4529b · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.706356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:5fa8ca02ebdb93aecf18a81588d10fab2f7b0aaf522051596365477fb0d99ec3

Observation 6ffa1d7f-aaee-47f1-9efa-8d2f643abf52 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.337965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:a91016970f9b9f48d622c9214dddd1d913523b733ed2b1a4960444f385a189f3

Observation d0599d4f-1320-4427-957d-93ceb82ad4d3 · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.185030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.185030Z digest=sha256:5969b3dc98801eaae5ff725d536279bb1dbb9cf4e1b3c141915383c9e28d7a98

Observation 3e46d437-d4fa-4afc-800a-916d564b420e · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:23.943515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:23.943515Z digest=sha256:6237c437594e4e468284d2df606316db570d7334d0d379a06c2f5431ecacd09d

Observation 0f0d8dc1-58fc-44ab-9453-12b949975ae9 · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:13.998784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:13.998784Z digest=sha256:b600eff929149556ab5b5f4ea13ba63c2e22a0c55ff03180e37823e23387e6d0

Observation e446c60f-acc8-4f68-9a7d-af9aa250ad29 · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:08.173192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:08.173192Z digest=sha256:113a73c2417dfc7cb680bbbf846a0ce7e1b099023e86896f70eac4bdb340e71e

Observation 26a51221-4347-4ac5-88fe-8b51366e9d8b · inbound

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents cites this paper.

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:18:54.074131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:18:54.074131Z digest=sha256:64b8cc8b4748ad03e9e2019db635f4aecd43c10be5540f396695f5bb01a5dd78

Observation 92bbbf6b-91b7-40b3-8fd0-8423d9c3780a · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.159584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:b781c4127504d1b130decc44c716406c0aae7249ef1d931c0264c4e2b940b7aa

Observation 93a3a951-c912-49f2-8512-585b1fc62165 · inbound

BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning cites this paper.

BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:29:24.990122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:29:24.990122Z digest=sha256:84323e67d10df4154f99d12d6fd170dd2df21e7a06cf9425b1ef3c89187a66a3

Observation a1b56a09-3792-4a48-af5c-26eeedbf0746 · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.335668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:11088643b12790ca54730967d2511f673e430ff2f0a035698247855baef4a1b2

Observation 8b799e30-2925-4d39-a890-2e38cc6736f3 · inbound

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR cites this paper.

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:20:11.514676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T20:19:52.701262Z digest=sha256:fc804e6a2483eb49ccd0a585a77f56972743496f8d0209c89e74446b293d0d84

Observation d2f8dcbc-db99-45ae-9962-7e9141e36ba9 · inbound

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning cites this paper.

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:01:05.178856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T15:58:03.654150Z digest=sha256:f1e0db487b1e0919b8b6d50a30eab70603a37e56225ce5e7ed1e8e48d27e87c0

Observation 9ed6f8f7-af64-4e22-8ed1-c90afa19f5ab · inbound

Learning More from Less: Unlocking Internal Representations for Benchmark Compression cites this paper.

Learning More from Less: Unlocking Internal Representations for Benchmark Compression SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T06:01:57.751567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:01:57.751567Z digest=sha256:f4425da1cba651d5d6de0ddf642045b7870c84916090341c5d8be4c14199cb1a

Observation 87b8e29e-dfbe-4f72-81af-26fe4762df7f · inbound

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion cites this paper.

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-15T13:43:40.241796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:43:40.241796Z digest=sha256:18f782e0f2ee3861739184615a0dc2d98600239d49b5dd5f60dbd3236693e318

Observation e40648cf-11bb-41eb-af3e-203b45d85a7d · inbound

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models cites this paper.

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:53.805058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:35:21.514502Z digest=sha256:325de84e175f946978e21d88865cca54d1f48415d6960949e09b8f18ddec4db3

Observation 3e929891-5d7f-4ccf-9fc2-81d28e7affe8 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:04.284210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:46:36.010169Z digest=sha256:ebaa880c184c5a6b77b9653a93aed80c978782358df036dbb07c8277b1679762

Observation 4adad9f9-fad9-43a7-8d8e-15396f9d3750 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:26.854067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:38:06.774877Z digest=sha256:edf5dcf8d98fc8cfc51939f270f8c06c11ae5d22257f619ec5ce9d8dba734e6b

Observation 4b66fcbb-5d88-4ff7-abdd-cff4dad2ba57 · inbound

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning cites this paper.

Visual Enhanced Depth Scaling for Multimodal Latent Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:27:29.383248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T07:26:59.917150Z digest=sha256:b32d27fdc80f148abc1404ebeeadbf436da0a89383593c032410d5f0eeb5aaec

Observation f451cf29-e6b8-4dd3-a0e5-1b08896683eb · inbound

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow cites this paper.

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:24.877582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T09:02:09.075097Z digest=sha256:ae337930f01bd189241ec8718eb4db8ef498776e10bec096ef4fa22ae21ee630

Observation 75727995-6101-4240-89b7-3e9c522d0f1c · inbound

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization cites this paper.

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:41:04.248345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:17:33.668451Z digest=sha256:187d28dd06e4700935bb20cec66a7ff459f32d98860f64c191eba340c8b549b0

Observation b0fe877f-5bf5-449c-9e76-3a59ace3047f · inbound

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization cites this paper.

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-05T04:30:40.720043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-05T04:29:52.480873Z digest=sha256:b957918987b3fdbc32f9c235fdd0ed597bb4af6215b3f852d46faab211241cdd

Observation abb05a70-1fbc-41c8-a357-413b306d3a7e · inbound

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging cites this paper.

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:58.909922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T20:24:30.679219Z digest=sha256:464fd308cbb95e1bd3c2b6d829ea087ed0153ae4d9fae1d07caa8d9a770b8f94

Observation 16914475-6daa-40bb-ad3e-c59290015fe3 · inbound

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging cites this paper.

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.730394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T09:12:22.712240Z digest=sha256:ef253643c1f249b43327ee2088e82c105b2e630e810dbb389cf86f0d62eb90fe

Observation 579364dd-983f-4c76-bfd3-ebb0c2944b74 · inbound

Deep Pre-Alignment for VLMs cites this paper.

Deep Pre-Alignment for VLMs SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 146

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:27:39.577374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T16:26:41.094936Z digest=sha256:8b3184244b97508ef9766404794c052400182a67592a8aaa6da3b33cf6c2c502

Observation efcc01c1-daba-4837-90d9-1bdce7d19ce4 · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:43:38.878107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:7ccdd0598c6cecfe2381ccabd4bfbb03fafb42f9ee6d21697506d9d1417ac749

Observation 46758ddf-3d42-469b-872d-32c4c5cb52da · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.447465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:7bad87ff01b4b6dc343b9e44ba3ab2062b263c3da2c73a22d4147b8eebfa83f9

Observation 0ecaf041-da59-4b17-bf3e-18d4eb31c305 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.039683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:05dc5cb1983fcd3bd3b43ebb040beaba725faa34d9445726e0cf3b5300419456

Observation 929228b3-17f2-4f93-a933-9619e559e750 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 171

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.641668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:09b194b2c5d189bcf6ff6569d2e7ea83f6b0ee8a3b484b284dcccd6e3f6f7993

Observation 7c4026ce-6994-4e23-8df0-612519f719b0 · inbound

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering cites this paper.

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.614558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T00:24:20.208132Z digest=sha256:d34e343928a588fcae089d6d3b353e94eeb9400cd8573b966b13ba383edba2d7

Observation 238a0bfb-d25a-41b8-90ea-c3601466b7c3 · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.731258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:14ea13565f190fc7371474a4f320c4da696d766995fa71ea524d8b42a46a8e8d

Observation 8e6216d3-e66c-4f9a-9e92-8eb022bc2b9a · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.162441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:2182f0a863e9b8b1af776b9ae66f3276c75d1fe29f2cc93b66ad78a2712bcb30

Observation 73ee3721-3cd8-460a-a73f-6e71170c2f86 · inbound

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning cites this paper.

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:07:03.920797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T15:04:18.880016Z digest=sha256:25a6b611624417ce3bf4ed0e01cdce895ade67522ec68bca5270038679817ccf

Observation 6c79f30b-a7c0-4f88-904a-6c8d6a29cdf5 · inbound

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions cites this paper.

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:58:01.879688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:58:01.879688Z digest=sha256:f7a447cacfb0fbe8bc4a8c0ea86d530054a967990bd9048b6ba7f3884fb3c72d

Observation b8adfbc4-0697-4892-8d1e-9bce6b39468f · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 143

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:14.200481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:14.200481Z digest=sha256:3fe1b8212f4c5da2ca844c0cc9ca55c374f19186fcc4a010d48139ac02ae9533

Observation d6899fd5-23ac-494e-848d-f0574a043eac · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.312878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.312878Z digest=sha256:31f554e96d0961f2b80be9ff53e32fee7d40ea0a42f4ee02cf195c91e7a09117