Pith. sign in

Paper Citation Record · LEDGER

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation

As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2506.07202.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07202 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:43:44.008852Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T08:36:21.863358Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:43:15.762081Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca4abf32-e965-4329-8ddb-2d501f7288dd · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:49.228008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:38.190734Z digest=sha256:ef003ba79a4d39a5c66332e1cb4e2260a306411bc71008cb98b7be9a29aaf8bb

Observation 461a6f64-d357-4671-838d-54ffc7894c6f · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:38.257889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:38.257889Z digest=sha256:055523f73ad07547cb6507e727820141aa1092d4248d49f40922f499f3785aa3

Observation ff8bb995-86a5-47ab-a9af-6e3f4a808755 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:38.419260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:38.419260Z digest=sha256:3846f377a5bb198c22e20a4daadb66b5312a1ac95915c01846a3a6aa5298e8d2

Observation c7bd01bc-355d-4378-8430-4d52309701ae · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:38.541820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:38.541820Z digest=sha256:61b206ce70750e72dbd982f62916b83015e5851fb0c5bb312e77e5e2b22649c8

Observation 6e3b7325-fb3c-403b-9538-584a770ebc6b · outbound

This paper cites Le, Sergey Levine, and Yi Ma.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Le, Sergey Levine, and Yi Ma

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:38.711427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:38.711427Z digest=sha256:e261dd7037bd01d7af15b6b4c4e19a261660e873860e5c8f94113dfb70b7d42a

Observation a5dce37d-d9c2-4d0b-888a-cbfc4ac61195 · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:38.885301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:38.885301Z digest=sha256:fc54b8c5046eb983abeb3e48d68fc8db7d35462751371b0c3b864bab38bd1f21

Observation b6d79c24-7fe5-41ca-af19-30b895ee88de · outbound

This paper cites Complex Video Reasoning and Robustness Evaluation Suite (CVRR-ES).

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Complex Video Reasoning and Robustness Evaluation Suite (CVRR-ES)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:48.912141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:39.042400Z digest=sha256:f2d74f920a619566965c3d2a34a9194314625a39c83f2006fbabdbdc28597798

Observation 145fe6b8-60a8-4bd3-a933-6aa0a5eaad46 · outbound

This paper cites NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:39.255176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:39.255176Z digest=sha256:dc417cb8cc260f23c2092e54a4b64df9ef9430955f4acf2ab96588102210f0b7

Observation 18484495-cf17-4f71-b15a-4214b86012b5 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:39.375817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:39.375817Z digest=sha256:5bbd48309cb30d709954e39f704e64ceb33ba1f97fdb4362a3cb7837fea7faf3

Observation 7f864f67-399d-4d94-8d52-7712ddc73784 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:39.449895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:39.449895Z digest=sha256:caa27ae8ef5193f45c5c08d8de92d1d23a81c2651c2c501e92a0df2c5d7ef306

Observation dc258f49-a426-4b23-bc29-558dde71b26f · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:39.543588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:39.543588Z digest=sha256:d6fb117a3c353d2c6ad9548ce56133a8ea4551de9873383c4aac01397dc39202

Observation 373f9932-70b9-42e9-8684-9ed355346882 · outbound

This paper cites Time travel in LLMs: Tracing data contamination in large language models.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Time travel in LLMs: Tracing data contamination in large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:48.603675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:39.598573Z digest=sha256:9f0bb2d779991a4a97e0eab65f0d82e45503cf0cd5ab68590efae8bd7d67a112

Observation 6dad6892-dcbe-4f3c-8327-94cc60b97c7b · outbound

This paper cites Goodfellow, Jonathon Shlens, and Christian Szegedy.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Goodfellow, Jonathon Shlens, and Christian Szegedy

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:48.325245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:39.737219Z digest=sha256:54b8be96b1a2d9c27adc313ef49808bf7c74a9043a3ae64addf81e5543e9e2f0

Observation 03daa9c5-f96a-43e8-944a-775293501894 · outbound

This paper cites Flat minima.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Flat minima

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:48.054042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:39.855755Z digest=sha256:61f77628ca7c927c5ba4b81da82b85ac4bda386682b090aca942127fa47270c3

Observation 296a0b41-3fd2-40f9-ba56-9bbb5d0eb0b1 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.028594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.028594Z digest=sha256:033d6e11ea91351e599983e50ec42a95c6262449ea482738ecf6273b8599390e

Observation 16b0f9da-ef7e-455e-a53e-c4a85c7e0b23 · outbound

This paper cites On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.205852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.205852Z digest=sha256:0fb0f027b380c3081ae17f2b2c8a138f537c5125095ece1934d076cf46b92522

Observation d940628d-e01b-4cad-acd2-5bf23370c18f · outbound

This paper cites LLaV A-OneVision: Easy visual task transfer.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation LLaV A-OneVision: Easy visual task transfer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:47.780360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:40.352290Z digest=sha256:5972ec57078f803bb961193534e87259d5d6a221e57114818623902807a0cab3

Observation c7cc41dc-b9de-4638-aa37-b8f895cd19a0 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.501360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.501360Z digest=sha256:65ee6611d90a218cc636f54450fd2471bc96052466e0d3a336a2f1ff006e08ca

Observation 56be8d9d-0980-44d9-8a08-da5276953aa3 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.648654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.648654Z digest=sha256:9cc8cba95158e5a803823d54edef66a912ab14ac2b1cf9f68c35f3ec29121706

Observation 954ca561-503e-4eec-9ed3-dd3d1fe530e5 · outbound

This paper cites Visual Instruction Tuning.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Visual Instruction Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.744698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.744698Z digest=sha256:319a2bc88b769962cf9d3c7ea979805fdc968bd82434cabe4ff86ec538304239

Observation a1214211-d31e-42b3-811c-b63afed3d312 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.944687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.944687Z digest=sha256:b54f13c501c4e1f7794098f230b909632ffbcb8ac96e5119b5183e7d5904826a

Observation 655f1235-d44d-4cb3-89c3-0231e9e6d2cd · outbound

This paper cites On the robustness of multimodal language model towards distractions, 2025.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation On the robustness of multimodal language model towards distractions, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:47.469322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:41.115259Z digest=sha256:7810f6b202f4a7ae50c775a2247c866cace07e41325d16d2b5823c009850089a

Observation 9bc5acdd-12dc-415d-97c4-5eeadbcac686 · outbound

This paper cites Is your video language model a reliable judge? In The Thirteenth International Conference on Learning Representations, 2025.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Is your video language model a reliable judge? In The Thirteenth International Conference on Learning Representations, 2025

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:47.201742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:41.231453Z digest=sha256:7c282c0f7a22c6d4e0658c6c3a922df76e4bb51de3789210b9a37b8770f29479

Observation 786f99bd-9ae8-4afc-a830-939e58122de1 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation MMBench: Is Your Multi-modal Model an All-around Player?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:41.387231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:41.387231Z digest=sha256:d55afc35272c9a45296c75babb4c78994847c27110670b39931d57a1ab15ac81

Observation 3002a829-5539-413f-8145-f53522016021 · outbound

This paper cites The Llama 3 herd of models.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation The Llama 3 herd of models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:46.871877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:41.529067Z digest=sha256:fb7170bf5c740e59f7a30763848176dd694e569c6b90fdd52b941a4f9a6dccee

Observation 719807d8-f564-4bf8-84ec-85ffd002dfec · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:41.682613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:41.682613Z digest=sha256:2f07ef00de870ee67c1d46d5bbf8d71759930d64c38334a2dd205313a056684d

Observation 7c71a775-0782-4af8-98a8-9bb109f52d3c · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:41.824761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:41.824761Z digest=sha256:dfd0c4f12fa4c1fc526d85d4ab1f78351dafb5cf934b5685c97b87c2b7eaa2ff

Observation 5588a6fb-1bff-444c-affa-7cb2e6634c9a · outbound

This paper cites Introducing GPT-4.1.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Introducing GPT-4.1

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:46.605726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:42.007117Z digest=sha256:c0e5e5a72e0e39c1c636ccc7c467bf4443716e0b6dc98013a2d1841a982249b2

Observation e6db3cb7-00f5-467c-a8e7-ea121436aac4 · outbound

This paper cites Introducing o3 and o4-mini: Our smartest models yet.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Introducing o3 and o4-mini: Our smartest models yet

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:46.340798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:42.126854Z digest=sha256:badd898c23bed9c372007d1a8f08203cf1569d9e20f97e80ceaf1296ae40141e

Observation b7293655-e8c3-47b1-beba-d9e66f85cc81 · outbound

This paper cites Chatterji, Faisal Ladhak, and Tatsunori Hashimoto.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Chatterji, Faisal Ladhak, and Tatsunori Hashimoto

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:45.997943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:42.312172Z digest=sha256:efc77ce815f2e310fd13ade2cfb0af899923ccfc2aa5371fe9290fb1fa45995a

Observation e33612d5-bdef-4b89-8fd5-1966ba17a24f · outbound

This paper cites Qwen2.5-VL Technical Report.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Qwen2.5-VL Technical Report

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:45.644329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:42.449448Z digest=sha256:1cb0f2771bc467aed19578b5c6c539e94b31f1ad4b223a636bc7b6b55ff48eba

Observation 7fdd844e-eb3c-416b-9c9b-7bf7065b460d · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation LLaMA: Open and Efficient Foundation Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.595992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.595992Z digest=sha256:00e63ea4273060c0414aaea29c2118dc074ddc3f2ffb36dc5da60676c35578ac

Observation 5d741959-b5bc-45b1-a7a8-54dbea542666 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.737932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.737932Z digest=sha256:72eb257762808ee39682e8100e35ff4298116fccc47d32f1db50c9e7916ecb8e

Observation 83b791cf-4563-4bac-b69b-e823938961cd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.884189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.884189Z digest=sha256:0f7529963637e137ea3412524db42e4a98eb6caeda8560b79ccbc5767a78d57b

Observation fc628e12-ea92-4ca2-8b7b-a2916806fa5f · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.067246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.067246Z digest=sha256:ca7d3922f163b6439671d21ad4aed4603be99d26f158a157f42155f7d38b8662

Observation 49db2cd3-1be6-43e4-9ca3-4059d3f4df2b · outbound

This paper cites Realworldqa.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Realworldqa

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:45.404107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:43.206948Z digest=sha256:8b244228cd27368acaa9ddbc9ee16c962154e4d504a494674a40191dae6cdd39

Observation f5299504-2ffd-4225-86e8-3561e2128509 · outbound

This paper cites Dynamic multimodal evaluation with flexible complexity by vision-language bootstrapping.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Dynamic multimodal evaluation with flexible complexity by vision-language bootstrapping

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:45.090141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:43.306522Z digest=sha256:21ae92fc1c6ea6c5ddc5b70ee89d5b49f360de2562ca8174df527b433ddabb9a

Observation 170e434f-ad70-4ad2-8b5b-37a9c533797f · outbound

This paper cites A Survey on Multimodal Large Language Models.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation A Survey on Multimodal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.489140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.489140Z digest=sha256:3dd9393f85b2e0b70fde80b617fa6398b56cd1620fdffdaf9062432437fe4a97

Observation e027e542-f702-40f6-9e38-0c9c8015b887 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.659170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.659170Z digest=sha256:7b23758d8a341f157e2a0bc96d89f636224ca966f5adb3db9086b28550af0e36

Observation 5879e056-8321-4c4d-9c80-1ed495a4e8ba · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.761515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.761515Z digest=sha256:6752de04b9c9270459f514a86b3e1b9f5b4a4e66d550bbd268c38e3e209730d0

Observation bcc4f788-8ab4-443c-b4bf-217c5acd20de · outbound

This paper cites DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.859281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.859281Z digest=sha256:6599ed2252f7156831425003b146ed6c2a5af373e1684f05fb5d3a17b48fc98f

Observation 678d5dee-677b-490a-9f62-21926e94ddac · outbound

This paper cites reasoning MLLMs,.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation reasoning MLLMs,

Reference 42

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:43:44.803931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:43:44.008852Z digest=sha256:fc5bd7e0c568f9ece30e2da182a0843591a856d2e81cd01db8c7818eaf51690f

Pith citing papers

Observation 87a9aca1-b6d3-44d1-a386-294b3273e171 · inbound

DMC-CF: Dynamic Multimodal CounterFactual QA benchmark for Causal Reasoning cites this paper.

DMC-CF: Dynamic Multimodal CounterFactual QA benchmark for Causal Reasoning Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:43:15.763580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:36:21.863358Z digest=sha256:c872df0211be970deee0735dc023fc158b4ce5ae52b78c43c1e82dc429ae6d37