Pith. sign in

Paper Citation Record · LEDGER

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2404.16006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.16006 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:12:00.710731Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:10.187396Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f07e53ec-0581-49c0-b1d4-c97c9b444c5e · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.300055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:f4218ec1b9c588a7d424f9dfc3e933cc437487112b798e43b747811ea4e17f0f

Observation 2203c694-e3c4-49ff-9ae5-f353746bb08b · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.282180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:c636cfceaeb3260d4dd23a2d4be2fedec5e227312b983511900c784ea943c460

Observation 8cee43a7-b33a-4385-a8e3-8b2429966526 · inbound

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions cites this paper.

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:53:40.596039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-24T00:52:52.056076Z digest=sha256:88ffb3a83c186bc86260d4b832eb7dc6713d5431705f7b3a1ded4f5b833ae473

Observation 1cf8737c-5e06-4ea3-a4af-69cf81b378a6 · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:09:30.450662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:251730588ee5a3b091c903ad71656483d25d93af6a8f6e6b8d525c7ac7f32985

Observation 64c8d833-788c-4e0b-8f5d-9d19661b3933 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:36.790918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:36.790918Z digest=sha256:591703b229d4bb79d331a6dfb60dc1730a21ec1302fbbbba020ce34ee9c3cfbd

Observation 2fb0e1f1-6f62-4fc8-846b-6f1d09200c36 · inbound

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation cites this paper.

OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T11:10:54.885006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:10:54.885006Z digest=sha256:fcea096c494adc8e599bded1a685afeeab5abcefa29233892563c5915ca15625

Observation 2b1ed913-0ec8-46fb-bc60-0014c828bf2a · inbound

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios cites this paper.

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:00.204585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:00.204585Z digest=sha256:b8d6dc4f1d056752a4fe401afd5ec5cb2ceb874896327fe0ee8ff8dbedab57e8

Observation 41f9a92b-b410-4268-9701-a2246ee958eb · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 278

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.241936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:ea91327db6e761c95f7faea01dec1af1d1ffdf1df0371c9aada2ecd215a240b6

Observation 3f4071c6-3d95-48d8-91fc-6fcd2733ea34 · inbound

Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models cites this paper.

Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T13:58:43.757280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:58:43.757280Z digest=sha256:1e629e850ec931a7a454fc04f996d03f287a36c28114a7fba806fd210677962a

Observation 9b89acf0-6d5f-4b6d-9e8a-432c6991a3c8 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.733088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:89fe117588f4bb655a508066028569e30a666d4939929fd81965131535c70267

Observation 94ee971f-2195-4d7b-ab91-d332cb320c83 · inbound

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark cites this paper.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.777581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.777581Z digest=sha256:8d4c34bd0d9ffe0816963b73fdd9d973abb3da451ba233917f5dda068188d1d0

Observation 90c41d2a-8e4a-4132-bf31-981baf459a15 · inbound

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding cites this paper.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.313329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.313329Z digest=sha256:4772905c53319925a225e5c46b34fbb3d51c65e7c1156ebecc6eca006afef12b

Observation 0e128430-72d1-4109-8572-04960d51c3ba · inbound

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs cites this paper.

From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T22:43:59.488259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:43:59.488259Z digest=sha256:2f35bb723c5d58263ba23b150bb8c5ade08928772263d6ad647858a2d14ba5a2

Observation 122975f6-eed8-4776-8518-f9f0845635f9 · inbound

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? cites this paper.

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning? MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:42:13.621900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T22:38:35.969273Z digest=sha256:cd55391772676c2bb289e42a6213fb5697b3c707051309abf2b948490b37927f

Observation 61ff7e0a-45e4-481a-840e-ac8901063e05 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.153948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:eba37b40089b4a4fdacda45b383fd8dc814d8a755963722942c0b75328ec01a6

Observation 3b940da0-3cc7-4cf0-a4d7-ebc30e705242 · inbound

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks cites this paper.

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:00.710731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:12:00.710731Z digest=sha256:5b0579dfbd7c4d4131096ad9afbecb6f02a2fa1e4f4d7bb60faaa12f57c94f4b

Observation 67fd94ee-9da5-4dbd-b54e-bf48ab174532 · inbound

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization cites this paper.

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:07:59.809431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:07:59.809431Z digest=sha256:6e3de2591462a7a6d05c34f8d0b0bdd2d28fb19c3f9350f9098bdebbf188e6b7

Observation 05d5a28d-78ef-46e4-87e7-9f579b2fc602 · inbound

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind cites this paper.

Can Large Multimodal Models Understand Agricultural Scenes? Benchmarking with AgroMind MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:43.699366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:43.699366Z digest=sha256:f821adb8fde7a3ec6bb55d1ee42221c393accf9cbae0dc2243c6149f55f0a330

Observation dd1247fd-ca59-4c6e-8d84-ab8666766be7 · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:25.406739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:25.406739Z digest=sha256:d55cfd0e49e1ed0c2c6bfab8c04dfb6eec96ba33d0a4f39eb927a8177bd559fa

Observation 0552af4a-974a-47da-b9ed-a3540cec5b83 · inbound

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos cites this paper.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.642571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.642571Z digest=sha256:1745f2ee0f2d18c95b4f9fba1854774e1491cda660c045d95d3d53e12a4cdf22

Observation 359cf8d7-ad72-45e5-8231-774b361010e1 · inbound

CoMemo: LVLMs Need Image Context with Image Memory cites this paper.

CoMemo: LVLMs Need Image Context with Image Memory MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.488263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.488263Z digest=sha256:1fc48f35bd3396c84fe9a9e4600fb04452658e5b82563f0eb19d377ae71f1a73

Observation b975bd57-1ac1-4c2c-9a4b-a220e9bdc4d1 · inbound

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? cites this paper.

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:19:05.062097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:19:05.062097Z digest=sha256:e9498c5621445d5dbf17e9c33c059a1ba956542475dcf37e93088cd99dbbd673

Observation e302ac2a-9625-464c-9a83-d3a02c2d75d8 · inbound

Taming Vision-Language Models for Medical Image Analysis: A Comprehensive Review cites this paper.

Taming Vision-Language Models for Medical Image Analysis: A Comprehensive Review MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 231

Resolution
unresolved
no resolver link, observed 2026-08-15T18:52:30.947865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:52:30.947865Z digest=sha256:80e2ce6328cdf4c16ccbe0e6f69c7105fcb956fee007ae3b08a38cead0609ed2

Observation 9ea9b439-fd49-40a5-9db9-5e1f2a8f0b91 · inbound

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI cites this paper.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:40.493129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:40.493129Z digest=sha256:6384c4cc60afab88c143e93501e202168303b50dea5c59244bda29b95cd0eec2

Observation 194404aa-8c2c-4111-bd6d-595658e045ed · inbound

Large Multi-modal Model Cartographic Map Comprehension for Textual Locality Georeferencing cites this paper.

Large Multi-modal Model Cartographic Map Comprehension for Textual Locality Georeferencing MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.092552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.092552Z digest=sha256:14ab6f1a5c9edec2aefa24ab3c5ecc6e243c06615feaca9474bdf741e768c30f

Observation 5dfb38a4-4956-47d4-9875-0bea4f982ff9 · inbound

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models cites this paper.

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:34.927360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:49:34.927360Z digest=sha256:0b1979c6a71802eac93730a799d9389fc4bb7140bbe6be84d2be2348342bb970

Observation 49b1b18b-a872-47ff-8655-2e2f8a63c8c1 · inbound

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark cites this paper.

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:13.357316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:13.357316Z digest=sha256:9ed959ad745c6eff2caf40de33c135d65fc39bb994085f2ede02d6bc2509712b

Observation 26a3a3dc-d9ba-4060-ba7c-8bd3e211a70c · inbound

MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models cites this paper.

MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:20.657760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:20.657760Z digest=sha256:7fc0233e45e75198b0ac66a797a7f2535b225970edefbe755c6ae1691996a5a0

Observation 623233a9-efc2-4d9b-826d-7f857dfd8e81 · inbound

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent cites this paper.

WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.034271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T18:56:23.817544Z digest=sha256:42b6b94597e4f1a7a2d8098c6a741cd661bd1877064d200d03cc80dc5b70d513

Observation 7b3eee89-68d3-4e4b-8a3b-26b972ad55da · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:05.388212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:05.388212Z digest=sha256:613e4ca058e1869ca1c8b956a9023f5c06d6c24218577f094fd526a346c52c1d

Observation 6cfdf7c3-1f09-4e39-8537-ae6415cfb2a5 · inbound

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models cites this paper.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:43.360928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:43.360928Z digest=sha256:3c856ea715ec5f60f0a209b4c743cabfd56c09b7ecfe609494c9eec992c0ede0

Observation 9d72ab86-6f58-411f-83cb-727847e57973 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.921223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:9a4e6013b3c02e3d545b32666cbca53783f4d30e48fa536e20108a6d9ecd9dee

Observation d08a92f7-3106-4809-bb68-07ac28c4ed8c · inbound

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis cites this paper.

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:40.845414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:40.845414Z digest=sha256:75c86bbb6d0ae32fa87080b28fe57601b8aadbf15fa82708c38019b9abed84ad

Observation 1bde5717-df73-4b65-8f18-d240ca650786 · inbound

SPHINX: A Synthetic Environment for Visual Perception and Reasoning cites this paper.

SPHINX: A Synthetic Environment for Visual Perception and Reasoning MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:21:30.726988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T04:19:26.808804Z digest=sha256:0be4b3c2bc2b5e5dacdbfaabfd74a75f6e5a829984ded18c10cbd6a7ad89ee29

Observation a7e25596-8751-4631-ad05-a1ae6b969843 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:26.465826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:82b67d37e8597fbc756444b198553642b4a4081443b989aa99b435df0855ee12

Observation 1e326ccb-4340-43a7-99de-9e432e5af207 · inbound

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents cites this paper.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:11:29.864333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T03:09:31.161760Z digest=sha256:6d13fd9a0f8f4ee6d42cb26238e856db9584dcbf97d2b610847f6b3455cb7879

Observation e9f9dc04-1acd-416f-ac6a-8f94a592e8ba · inbound

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents cites this paper.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:48.676604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:48.676604Z digest=sha256:0b303ca1ea3a46353c9023a4ce96a7601734ff95377767aec6385754589a740e

Observation 2a854ec4-9137-4869-94f1-f5370a6ef96f · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:52:59.949730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T16:51:48.705876Z digest=sha256:8e8e488f48b65e7e48096cd126c85a76291408d3991b32b4813a24a97c1bfedf

Observation 2f087083-3736-4a58-96f2-e94383d27b47 · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-13T12:10:53.720348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:10:53.720348Z digest=sha256:8cb8d94bf49c82fa26ad7057f8f76a86936251d6d7dbfc71c369e04b14550834

Observation b62ba793-183d-4028-87e8-a70d8dab483c · inbound

HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing cites this paper.

HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:57.859380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:06:49.114269Z digest=sha256:35c64fdf2da26cc1375072e6f3cdbbb0ac1b60cf394d7b39dc4c4ed2bba059c5

Observation 17f99df8-7dbb-49af-bc06-2fcabbc55bbd · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:22.671101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T18:53:06.494640Z digest=sha256:c3f1051f329a5e29194c0327a6c623bde6a51d5f6259d02ca5b503604ae50f4d

Observation 739c674d-a30c-48f2-8f91-e0179c88281d · inbound

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs cites this paper.

Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:50:51.555655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T01:49:15.136031Z digest=sha256:be4989925ec0615b719cfbb7b407c3bfb11c5231dc5b44b61c390a5679a319cd

Observation 575ae2a5-bd24-453a-abba-a6d0cea6c8a8 · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.318295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:b457c6acfaeff8be4813429bf1b7e8ffbdb324670b45e3ba1f25c75ece90b02c

Observation dac1914c-acbc-439b-b6b9-538dd6b00fbd · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 108

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:34.595588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:9ee7fbd747fb03d10ce71e5c5545f44803d6784a8d72b21ff5d4758b5363052b

Observation ccd1fb11-eedc-46fc-a1d7-f1fed4e22ceb · inbound

TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding cites this paper.

TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:27:25.893702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T18:54:34.353940Z digest=sha256:ae50cc0db7b5f297ec5f150abf84c6cfca581e00c754b49e47f7fc54c3a489cd

Observation 03d97ab5-429c-4341-9b82-2c8f44b93dad · inbound

C3-Bench: A Context-Aware Change Captioning Benchmark cites this paper.

C3-Bench: A Context-Aware Change Captioning Benchmark MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:50:10.188929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-25T21:02:52.529391Z digest=sha256:aa5591c332ac61044458ca9a86814baca5483750781ab47529040fae268fbe08

Observation 479473d3-2b3b-4790-aa44-1eceee4481a3 · inbound

CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts cites this paper.

CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:15:44.513584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T05:41:20.492139Z digest=sha256:889ba5486abaab41d78156dc23cf5fa675d319e7a36cb1afc769bbdf14d834cf

Observation 9b5d014b-bf73-4425-b7b5-cf1bcdd2f25a · inbound

CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts cites this paper.

CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T09:59:21.899377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:59:21.899377Z digest=sha256:2171be4cd33aac42ffd26bf965669c53ad6b4983c85d89d3a5128e9aecb92bc9

Observation a36d6b0f-1480-48ef-bb41-49ab19ee5561 · inbound

CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts cites this paper.

CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T04:38:23.452380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:38:23.452380Z digest=sha256:447ff91cc7841f59a5b95df0bc77049550db7aa98ba0b8467527211450ee4d8b

Observation 51cafaa1-e006-4bd0-a2f0-c01b90c8e58b · inbound

Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs? cites this paper.

Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs? MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T22:11:21.555191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:11:21.555191Z digest=sha256:d548e153926aaa60638c359dcd95e85db5b0bad3cba26699e2b8325973cabb0d