Pith. sign in

Paper Citation Record · LEDGER

A Survey on Benchmarks of Multimodal Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2408.08632.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.08632 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:30:22.093509Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:49:41.635040Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6df3758a-82c1-4555-967f-f0116940e192 · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning A Survey on Benchmarks of Multimodal Large Language Models

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:33.237633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:3be4cc7604144a089453ed523fb0ce29a3038d09cef6d1b354cc60a2814e3f78

Observation c4541fd2-edff-4e88-8ec7-b424820fe260 · inbound

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization cites this paper.

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization A Survey on Benchmarks of Multimodal Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:47:13.058347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T22:47:09.500229Z digest=sha256:5ee6cc7b935abd5f7004a5de90a050d7c872242f875bc4bd584952592c9e7cd8

Observation 7e2f3ed1-726d-445c-be3e-c0fba0554758 · inbound

MLLMs are Deeply Affected by Modality Bias cites this paper.

MLLMs are Deeply Affected by Modality Bias A Survey on Benchmarks of Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:22.093509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:22.093509Z digest=sha256:24526d528de0daa6c76bd8d170a4ebd5057fbddc7e076df473cf8830e829e9e2

Observation 792c852f-1151-4710-90d4-5a20d32e9ccb · inbound

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models cites this paper.

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:13.867592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:13.867592Z digest=sha256:3751dae960931e0c90a7675c6573b68eb678c7aeb1ae44027527480aaedcf7d4

Observation 93700d1e-7cf0-4a5d-8199-fbe04144cfe1 · inbound

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests cites this paper.

Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests A Survey on Benchmarks of Multimodal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:57.018980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:57.018980Z digest=sha256:d0c6c51c980e82ad962486ddc04bf9367c933ff09de537217678f4df80f4597c

Observation f51502c7-a87f-45f3-a521-4b4b84b61990 · inbound

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models cites this paper.

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:59.994525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:59.994525Z digest=sha256:6f0101f134ecae09a29e1b29e49cd7eed74e1fd38b278b1b60e699340090ad11

Observation b3b66b49-a8e3-45ae-99c4-5c84ace87e7b · inbound

Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? cites this paper.

Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? A Survey on Benchmarks of Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:36:56.960671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:36:56.960671Z digest=sha256:d81cbc0ee419a62220a72eca6e5b3affd0a1f6ba50d075085beda3741011b547

Observation 2bd0f30d-8816-4898-adaf-fc8413a2e385 · inbound

Unified Multimodal Understanding via Byte-Pair Visual Encoding cites this paper.

Unified Multimodal Understanding via Byte-Pair Visual Encoding A Survey on Benchmarks of Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:55.436309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:55.436309Z digest=sha256:5185ed5f6f05900f44e1d78b82fce3397e64fd139e61a3c67c9bb25f7939fdc7

Observation ab6c4cd2-5071-4449-802b-8455bbee0d98 · inbound

Re:Verse -- Can Your VLM Read a Manga? cites this paper.

Re:Verse -- Can Your VLM Read a Manga? A Survey on Benchmarks of Multimodal Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:16:54.122059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T23:14:25.088197Z digest=sha256:ed27b495dfcd46cf88e61308414470d419a5bc7a6685f9b9f3a901665226ca17

Observation c3afe12d-cd83-4cc8-9109-5f02e39d92c9 · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM A Survey on Benchmarks of Multimodal Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:48.481837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:48.481837Z digest=sha256:4d83a9a323f6ad26f655be56e13620c081f08562a1dbc23e9cdf246ef93c67f6

Observation df52c407-f409-4ba5-a79a-a4dbcea8c987 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding A Survey on Benchmarks of Multimodal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.843114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.843114Z digest=sha256:42a6f1830e783d432a48a52973c450e842dcc827f891a82ceb8715c0f9a08a29

Observation 9e7b8396-5a13-45bb-8dc1-0c787d59870d · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks A Survey on Benchmarks of Multimodal Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:10:01.934373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:09:51.816554Z digest=sha256:4d1f4bd03576ab0e6daacdaef1eba869f027af5c7ee39eafef4ae2a2a3196b5b

Observation 50d669ae-960f-4164-a420-1135b47ce2a0 · inbound

MARINER: A 3E-Driven Benchmark for Fine-Grained Perception and Complex Reasoning in Open-Water Environments cites this paper.

MARINER: A 3E-Driven Benchmark for Fine-Grained Perception and Complex Reasoning in Open-Water Environments A Survey on Benchmarks of Multimodal Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:11:00.913053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:46:15.107175Z digest=sha256:ca3a57dd52ce520853d4b6b4744172d5b656e8237d384761ac28103ea001e5ca

Observation 37c82ed5-a722-435d-9946-cd0650edd23c · inbound

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models cites this paper.

Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:25:18.559769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:23:46.371799Z digest=sha256:4180b6f056f687a72647dd76c579a611645f433fc5cd75608c4f85ad7afcdcc8

Observation b1bf0ca9-97bf-4be4-86a5-21bfcd4fa148 · inbound

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety cites this paper.

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety A Survey on Benchmarks of Multimodal Large Language Models

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:05.630683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T03:00:34.862711Z digest=sha256:e38fd5a390860d1bc985d710186560375259385193dc4073c4e6de99a7ff1b62

Observation 2c46244f-d023-4efa-9798-d884440b9e7b · inbound

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing cites this paper.

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing A Survey on Benchmarks of Multimodal Large Language Models

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:46:46.595108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-07T16:18:35.484800Z digest=sha256:6595dea7cc69501fe3eec91630d9151d8d2218468be8996e6feea0ba47e7c31b

Observation 7d41d828-6c75-4b97-a4f0-ffa9c0d0d6a4 · inbound

SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows? cites this paper.

SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows? A Survey on Benchmarks of Multimodal Large Language Models

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:03:39.464904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T19:03:29.501299Z digest=sha256:b90fd8fdedbc4d7488ed46b44438b43c5b2b3609b06f56548631a6a5d5f2d4eb

Observation 601467ae-c5fe-497f-8db2-94132ff06ce8 · inbound

MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models cites this paper.

MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models A Survey on Benchmarks of Multimodal Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:56:44.630245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T07:16:45.665440Z digest=sha256:973946864aa4e71b16cb22d69eae472155bfd03583c450f80982475b903bd7bd

Observation 3f26418d-e546-4b78-9e59-d466642eec04 · inbound

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments cites this paper.

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments A Survey on Benchmarks of Multimodal Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:55.944986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:36:11.721143Z digest=sha256:0c8b099ddd00ebc5df9f4c47c1b0ad4727dddbdf96f22621dc157f67e854d0f7

Observation 80be2a3e-6e86-4a65-bfd2-480ed0f1d975 · inbound

MMGist: A Comprehensive Multimodal Benchmark for 2027 cites this paper.

MMGist: A Comprehensive Multimodal Benchmark for 2027 A Survey on Benchmarks of Multimodal Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:41.637190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T11:05:14.573386Z digest=sha256:f25a5307b1bdc39a7a5ac5ad4e864cfef7921e90be9dc95a3a6dbd1495171e70

Observation f3148e2d-14da-4ba9-b0e5-0a7ddd658545 · inbound

PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement cites this paper.

PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement A Survey on Benchmarks of Multimodal Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T06:18:42.939955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T06:18:42.939955Z digest=sha256:d6dfb75e4d5445ea607779a0ca4f7ca610445d0c28cf6a3c287143c7721edd1c