Pith. sign in

Paper Citation Record · LEDGER

A Survey on Evaluation of Multimodal Large Language Models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2408.15769.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15769 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:12.097192Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T10:34:07.256051Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 04abcfd6-1a15-454f-a924-b65aa7dcd855 · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey A Survey on Evaluation of Multimodal Large Language Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:28.036970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:d7e1d2b241d7695ec0efa5ed0c442efcb3cfd1eaf03d0e51b1e98b01c42e900e

Observation ec607f9a-a940-4aa7-91db-7c1eca2d044a · inbound

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey cites this paper.

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey A Survey on Evaluation of Multimodal Large Language Models

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-11T23:54:23.655726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:54:23.655726Z digest=sha256:3369872486b7d0f8cec43327022141623e30219b381dfe5618687844bdcc114a

Observation 37a6830a-9e8e-43ea-953f-aec62da1ffe3 · inbound

Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search cites this paper.

Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search A Survey on Evaluation of Multimodal Large Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T04:53:11.781191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:53:11.781191Z digest=sha256:0d3bca6330796e039daeba7d4080aaee0325eab930a58d2c5b80e93df8592af7

Observation d4471af5-eab1-4499-b9f6-648073ee7d76 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey A Survey on Evaluation of Multimodal Large Language Models

Reference 168

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.859085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.859085Z digest=sha256:2f42d68e04772b8e91a51967f5730bcd9fccb1c4bf5c774c2f6465f810a63848

Observation 00933da5-cdf5-4a31-b37a-4f39eab45692 · inbound

Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering cites this paper.

Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering A Survey on Evaluation of Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T23:10:33.015365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:10:33.015365Z digest=sha256:029ebd846f3ab9064e436589399cf622ce55e9893a9cc2b8534ae66b0c134007

Observation 7e1cb9d6-6066-465c-97fb-6a20d02cbeac · inbound

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! cites this paper.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! A Survey on Evaluation of Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.194763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.194763Z digest=sha256:2327c1f7a7c7acb5246ff4bb0d700af74f195b238897e6ee2c290bba2add599c

Observation 73dd6bd7-988e-42e1-a932-9b5a3e87f154 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models A Survey on Evaluation of Multimodal Large Language Models

Reference 289

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.688765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:0b54c671285246e9ae88d133a45d48f4f2d7c2c0bf3474d3f5df4edad936f73b

Observation 3bd41662-074e-4ff0-ac5a-ffa8f1d94484 · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment A Survey on Evaluation of Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:12.097192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:12.097192Z digest=sha256:077117d76705f0df8dea9fb07cb5839ba033fec3c1a9c44ddc012d9afc78157a

Observation bbe1d679-9c90-4203-9523-e2d6fdc4b450 · inbound

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks cites this paper.

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks A Survey on Evaluation of Multimodal Large Language Models

Reference 143

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:00.700076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:12:00.700076Z digest=sha256:6ef24372a554ef4c05d24d755fe58c0050e1ad3b7b56ac7b8442e11ec3ddf646

Observation 6023184c-710a-4066-a4cb-332ab1951308 · inbound

Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models cites this paper.

Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models A Survey on Evaluation of Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:55.906329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:55.906329Z digest=sha256:0c2a532ec356950437f74cc588dd53b7e4b3cdb3e30d6651455a1266fcac7732

Observation c2252e6d-8f3a-4ad4-81d6-87a0f853c88b · inbound

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models cites this paper.

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models A Survey on Evaluation of Multimodal Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:15.568736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:36:15.568736Z digest=sha256:90e5bdc6f078a72e98847da11983a1a91dcd6ad0ce7466ade810b2dc2f03a056

Observation f09d3e57-92b7-4da0-9496-38c4d707aeac · inbound

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models cites this paper.

AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models A Survey on Evaluation of Multimodal Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:13:02.831115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T11:12:41.130806Z digest=sha256:626cd60b180161e5bc517a2ff8b940aed27ad0da36d33798e16e52244f53631f

Observation 3b6839a1-2a9a-4720-84af-974319007cc1 · inbound

From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem cites this paper.

From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem A Survey on Evaluation of Multimodal Large Language Models

Reference 138

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:10.154494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:45:10.154494Z digest=sha256:ed6ea1121748f0b44ca11b69795e417d3d899bfc028860b1ace052d5a802b6e6

Observation 6d1b235c-0761-47ee-8af3-9719f156e79b · inbound

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation cites this paper.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation A Survey on Evaluation of Multimodal Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:46.842259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:46.842259Z digest=sha256:373523314a51db38949dfbd7c9c2dca71065f02ec5d7ac07b4832c7082db60d2

Observation 7b2b85b1-491e-4086-bd9c-b0819c6da1b1 · inbound

Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation cites this paper.

Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation A Survey on Evaluation of Multimodal Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:34:07.258882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:32:41.447253Z digest=sha256:99f0d89a9858d04575cb98769ef3d0ec4d9955749c88d8b3c1424b505ae933d7

Observation ff959bb2-be03-4605-ab1a-9684f071899b · inbound

QoS-QoE Translation with Large Language Model cites this paper.

QoS-QoE Translation with Large Language Model A Survey on Evaluation of Multimodal Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:41:01.175195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:01:57.703684Z digest=sha256:936442cc5c34823f7d9cd97496221c593618de72552568e2844dc31f60c5bb3b

Observation a26a0f95-4b7b-49c8-ba2d-a90c8f330a53 · inbound

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs cites this paper.

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs A Survey on Evaluation of Multimodal Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:01:08.880958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T10:34:33.399684Z digest=sha256:d1d62bcda03eaf75fe044c269a249ebc80ba3bd1d54da91f23eb01793216d0d8

Observation 66c826d2-5bb5-46b3-a91a-e361fe3713d3 · inbound

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs cites this paper.

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs A Survey on Evaluation of Multimodal Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:49.680393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T00:49:23.895938Z digest=sha256:c5fcb718bdc6d3a9b41c8eacb46c7b13598ac4028047c91531718b5a9ba47e06

Observation fdb1882c-9298-4079-854e-465c923a66c0 · inbound

KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation cites this paper.

KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation A Survey on Evaluation of Multimodal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T12:40:15.292794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:40:15.292794Z digest=sha256:5f16ee19d12cf8b8917734b93a5824483d2c227d7ed450c169a2f4b7cdc54003

Observation efed6e94-41bb-432a-b7d6-2927cea06f9f · inbound

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences cites this paper.

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences A Survey on Evaluation of Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:29.980951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:29.980951Z digest=sha256:32739635bfd428a60253d85113e8b74034462620fb8b5a999c9358cdf661ad4a