Pith. sign in

Paper Citation Record · LEDGER

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

As of 23 August 2026, this Paper Citation Record lists 100 of 113 outbound references and 45 inbound Pith citation observations for arXiv:2501.05444.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05444 v1

Coverage vector

measured 100 of 113 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:18:59.867228Z

measured 145 of 145 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:47.134304Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:50.233935Z

Reference resolution

100 of 113 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved79
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 46b02530-cd70-4fc8-9b73-00e9ce62e98e · outbound

This paper cites Greg Lan- druwu2024plot2codem, 8(31.10):5281, 2013.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Greg Lan- druwu2024plot2codem, 8(31.10):5281, 2013

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.449559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.449559Z digest=sha256:b10e9c550d4f375cfc704e8b58a95fc148223347ed10ac91f451255c23d7a3a1

Observation 14fa4d6e-2547-4c55-a821-15c2d7a109f8 · outbound

This paper cites Khan academy.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Khan academy

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.454610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.454610Z digest=sha256:ab7fdc8f67e69a928d4832d6ceeae4f2137c88320becd5004d07a2b367194b58

Observation 3d093895-4ab0-4d83-a0e0-bde7659ab4e2 · outbound

This paper cites GPT-4 Technical Report.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.458930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.458930Z digest=sha256:316bf24d5c20ecb4d7719f9c9bdd24b43087a468cb695db8a4bdf3a481f9228d

Observation 3211d22e-a5e8-4f81-b9e5-b6479ac530e1 · outbound

This paper cites VISREAS: Complex Visual Reasoning with Unanswerable Questions.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark VISREAS: Complex Visual Reasoning with Unanswerable Questions

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:19:00.635641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.464128Z digest=sha256:d44926e348d165deacf802a471e72842871cf5de4d64d524e2fdae22e59d0fbd

Observation 076f321a-ad4a-425a-9f09-a46c3a3634a1 · outbound

This paper cites Claude 3.5 sonnet.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Claude 3.5 sonnet

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.468573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.468573Z digest=sha256:fbe3e60020b997f6d6362a7905303e1a3e2b54f5318eaece4e8099699432e01e

Observation 82627291-00dd-40e2-b146-0a320bfa6a21 · outbound

This paper cites Vqa: Visual question answering.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Vqa: Visual question answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.472998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.472998Z digest=sha256:49bf60b295ddb7ee31bb6f41b1b386c1216a3ed3dbcd480381be4a733942a4d3

Observation f3cb3538-a059-4694-859f-d1f529e922b2 · outbound

This paper cites Qwen Technical Report.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Qwen Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.478229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.478229Z digest=sha256:83820d3c97195c7a2146ae3b50bd563679f6dc95b5457b6306e05ab383824174

Observation d8f4ba97-6a51-4967-a794-a984b2a9ef25 · outbound

This paper cites Viseval: A benchmark for data visualization in the era of large language models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Viseval: A benchmark for data visualization in the era of large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.487005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.487005Z digest=sha256:693419d1e76a50eb785be858d71e5107f62bc26f5567d5998f446033589a84ff

Observation 70951517-0338-4286-9c48-afd52e19ba1d · outbound

This paper cites Uniter: Universal image-text representation learning.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Uniter: Universal image-text representation learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.491249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.491249Z digest=sha256:383f8a823fa321884266cf375f98bde2daf1290fefde731a15b5c7c96691d50f

Observation e14bc7da-57d3-437f-930e-d95c1e01be9e · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.495233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.495233Z digest=sha256:e54b5acf81e408437b7cc8ed755ba773bcb38940074271ec8aaac7883342911f

Observation ee5f873c-6dee-4731-8241-a898d9499189 · outbound

This paper cites CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.500206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.500206Z digest=sha256:0b4e7a8f65f3dadf49de9cdecf86535080a209fbfeb65afb3ccb2c5aac1ec2d8

Observation eba7d77c-7753-4477-ac84-c127ba18b377 · outbound

This paper cites On the Measure of Intelligence.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark On the Measure of Intelligence

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.505071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.505071Z digest=sha256:33abe675e9feb589745a565ffe8bf08033de01712c6c46d093c330428b4906b2

Observation 1a1c06c1-4593-4752-bc9f-0dbf6cca4292 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Training Verifiers to Solve Math Word Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.509911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.509911Z digest=sha256:67fbb17485be3535e82a8f102264f25aa927382ed23f33a8cb45c0285af071e8

Observation ba0b953c-972e-4b30-9c11-fa618e9f2640 · outbound

This paper cites Codeforces.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Codeforces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.514094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.514094Z digest=sha256:e2e0f05878325982d5e0d725bda952a2f0b055b2bf67b4dc994a2af088c7b2e4

Observation c7e84cf4-f6d5-4005-b2e4-6fd89c07ae6e · outbound

This paper cites EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.518452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.518452Z digest=sha256:69398dc4336108d586b8ac0631945fb892526cd19df122bf077640e4f55e6b45

Observation 6cd4b9f5-91e0-4d7b-90c5-96dd9a1cae71 · outbound

This paper cites Gemini 2.0 flash thinking mode.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Gemini 2.0 flash thinking mode

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.522338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.522338Z digest=sha256:626ac29cfc712573e911284e5998ca1ad8211f1b3c1c39a55981718fa253d9d9

Observation 90b3e9d7-14ba-4eae-a20e-36798e4c72fe · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Introducing gemini 2.0: our new ai model for the agentic era

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.526922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.526922Z digest=sha256:d7b4d6b85b6d256e157c7a288de5d47d732c531b2de02bd928cba333014550a0

Observation aa282e4b-3422-47e8-bf01-f17ed1e52da6 · outbound

This paper cites Deepseek-r1-lite-preview is now live: un- leashing supercharged reasoning power! https://api- docs.deepseek.com/news/news1120, 2024.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Deepseek-r1-lite-preview is now live: un- leashing supercharged reasoning power! https://api- docs.deepseek.com/news/news1120, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.530484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.530484Z digest=sha256:a2e564789d8c929f6fde65eebf7d85dc9d96f136d55bd4813217e04563eec6c2

Observation 2565ee91-18a7-4971-9f5d-691e5bb64fa7 · outbound

This paper cites The Llama 3 Herd of Models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.535228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.535228Z digest=sha256:701fcb72628d2aefd7517fadea4f3740a972ea40f6a92e3966b7249afe8f671d

Observation babb6af9-8727-4a53-b713-e0b26a5d3279 · outbound

This paper cites IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.538886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.538886Z digest=sha256:e3ca97e2491166c21b726479197482e37bcbb7ba6b0554b2807087a9816d3a65

Observation 70a7997a-fe80-416a-bdb6-13e5673290e2 · outbound

This paper cites The epistemology of visual thinking in mathematics, Feb 2020.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark The epistemology of visual thinking in mathematics, Feb 2020

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.542806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.542806Z digest=sha256:984ef9125a8623bec73cc6850c8c9ef8c1bf3c3d339b7b6577e28f5c941ef1cb

Observation 37cda82d-31bc-42e2-942a-d063efc0c57c · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.546728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.546728Z digest=sha256:a88f18b54a20a96e37d641df476c8d00d0083c9d8b3fac58d054da48b6d9db7b

Observation ddbd4355-7759-44ce-bd94-73a657a3a50a · outbound

This paper cites ChartLlama: A Multimodal LLM for Chart Understanding and Generation.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.550919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.550919Z digest=sha256:a482bf982a1aeacc4e5b092eb3d532616b179b29eb29980cb9690a927b8c0dde

Observation 2036de14-bdd4-4dff-ab7c-0b17df18037c · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.554220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.554220Z digest=sha256:2f208fb1dbbbfb072f900fda97d5b5ef6e596ef1e33deeb9918594afdde8636a

Observation b3716284-a980-40f7-9cfd-662dd62fa70e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Measuring Massive Multitask Language Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.562074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.562074Z digest=sha256:3925b6b964363c1e4e26f766d0c07846feb2292906bea9de771cf1261452d561

Observation 8d5336ba-2e17-4a82-a8ec-3bc5f20d878a · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Measuring Mathematical Problem Solving With the MATH Dataset

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.565972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.565972Z digest=sha256:2ac6dc40d591d4da4d73a896b0a2aa6d392501fbf7b4bfa8bda14792a9ced1c4

Observation 904be90d-30bf-4e36-91d0-1081315e14d6 · outbound

This paper cites Novachart: A large- scale dataset towards chart understanding and generation of multimodal large language models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Novachart: A large- scale dataset towards chart understanding and generation of multimodal large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.569708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.569708Z digest=sha256:43ac30ace592d68cfa982aee8007beeded354b2b7e03297cf8d14ef104b474f8

Observation c3be78fa-f225-494e-99aa-c25eaadca2f7 · outbound

This paper cites Towards Reasoning in Large Language Models: A Survey.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Towards Reasoning in Large Language Models: A Survey

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.573378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.573378Z digest=sha256:6c4e89d85627aa5d4c01ee023c9329f75c3afee10bf04441f86a47ea35b469e7

Observation a0b885fd-fe6b-4bfa-b9fe-70a49c2421cc · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Large Language Models Cannot Self-Correct Reasoning Yet

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.577577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.577577Z digest=sha256:991fa2017ab7ecb24ab5ed0bb4bf5243106e407150d7ee8ca7d0274cbaf22d8c

Observation 52f2157c-e6eb-4379-9ffe-b163cafbc6fe · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.581499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.581499Z digest=sha256:778fdfb2f9449eaa1be66e5ec3872bdfd9a8e6fdee6dc6693b225bd8cac2bdcf

Observation dfe477f8-c437-4249-9e75-feb26edff9df · outbound

This paper cites Cladder: Assessing causal reasoning in language models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Cladder: Assessing causal reasoning in language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.584996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.584996Z digest=sha256:4e44c2dc8ed71e99946a08c53400005cf913b52643ed4e3e95015b5aa9d03111

Observation ea7cef7c-dff0-40a5-8ba3-40574c4ac517 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Gonzalez, Hao Zhang, and Ion Stoica

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.588850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.588850Z digest=sha256:eb7b48da6ccd698c823f50b6abf3e50a7201a3e454779f174d0e9c91f9161d56

Observation e4616ed0-a19d-4ef4-9e42-5903514ce8c5 · outbound

This paper cites SMiCRM: A Benchmark Dataset of Mechanistic Molecular Images.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark SMiCRM: A Benchmark Dataset of Mechanistic Molecular Images

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:19:00.469696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.592774Z digest=sha256:d3254d46e0d282a27e4282713d23ea12ec0bdd1395f26547243158a0bf8704d5

Observation 6b3e57b4-b5e3-4591-917b-d272259502de · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.596503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.596503Z digest=sha256:a36437fec5e8a8be45e38fb2511a57a471450b56a560ec5bf1042fd5a9d6194c

Observation fcd78bbc-f9b5-431e-89a9-a88d5dd2fcbd · outbound

This paper cites Name reactions.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Name reactions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.600489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.600489Z digest=sha256:8eaee70f3ca20bf0580151fc6fca3df517f37eb083a7d986e3d5f6600ab07318

Observation 2172043a-f3c7-4aa5-b2bf-0baf33665701 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.604084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.604084Z digest=sha256:c0e83691c130f68d430244e0ff6010fedf5f4fadea6b58cda9801bd67c2097ec

Observation c46d5a91-088d-4f2a-b0a3-079f476777ea · outbound

This paper cites MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.607732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.607732Z digest=sha256:1778df4e426d5e4b19428c4e8db599770f63109ebeec2f41ed06ed024849ae36

Observation 81907be8-e07d-4f0f-8991-5cd04cc98e82 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.611623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.611623Z digest=sha256:fe7ec31c0fe107b1acc96b32781437e70013b7dcea4fc5091c632f495714f6c1

Observation b91bf0dd-22be-4ac7-a1a6-2f1a1caed467 · outbound

This paper cites Mmsci: A multimodal multi-discipline dataset for phd-level scientific comprehen- sion.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Mmsci: A multimodal multi-discipline dataset for phd-level scientific comprehen- sion

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.615375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.615375Z digest=sha256:e27d9168f026a45228163c845e6da842528b633f410d8a8a607f0e76e28724bc

Observation 418e932c-63b7-4325-bd55-0a3affc2ae66 · outbound

This paper cites Let's Verify Step by Step.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Let's Verify Step by Step

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.620295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.620295Z digest=sha256:24f7f7e6e54b4194d208a8e61eead3caceb5eec7c917e83ab6e3574b5626d1e2

Observation 32a37897-1d1f-465c-90f6-3b860a3363cc · outbound

This paper cites DePlot: One-shot visual language reasoning by plot-to-table translation.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark DePlot: One-shot visual language reasoning by plot-to-table translation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.624870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.624870Z digest=sha256:2af87945af537f01f73d74b2073f05b8efff6f83f9bfe3cca692a003c3d0f159

Observation daa5824b-5ab5-4d97-878e-4b0af8525f67 · outbound

This paper cites Improved baselines with visual instruction tuning.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Improved baselines with visual instruction tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.628663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.628663Z digest=sha256:25535410b60c32481d7d121dc21397951bd873125df0d387b5a2fd8a24a78862

Observation 04b9763d-f4c4-478f-ac6e-c34945a1b47d · outbound

This paper cites Visual instruction tuning.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Visual instruction tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.632519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.632519Z digest=sha256:73319006299da45b470ba88edc2161693d02fae2a0184d73e1a383652d3922f2

Observation ce440d86-6ee9-4f15-97b2-1810d10d93d5 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.636335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.636335Z digest=sha256:2ded20f6386bc82aa29260163077a14feef8f5ecde41db5665697c82411bcc3e

Observation e97a3218-0411-4795-aad3-91cb949c7922 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.640626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.640626Z digest=sha256:e1aecf4711d86a95e17a00015b6dbc0aefd5bb9ae1864448bf81e77992442627

Observation dc3a9eec-b736-468a-85df-6558e857ae3c · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.644283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.644283Z digest=sha256:32f85b5be9c5c02950dee480711c462b367ed3c3b21e851240d48b9e9dbe4481

Observation 4d223824-231b-4af4-bcb2-f15754cbcf7a · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.651901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.651901Z digest=sha256:917cddb9f74e075d5ded2616358b70cbeccfab230790a5de70c3e130c9348cb6

Observation 9efa7098-8324-4f19-8bd5-3dee3d45a696 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.655598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.655598Z digest=sha256:16602c53bc1b0f7af3c4dcec23c36bd1471e115489d9105878fb307448a21ebe

Observation ec75f540-35fc-460b-8cd3-8424254bb626 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.659238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.659238Z digest=sha256:fdc41c99e650738421ff2301224637c234530d5c5403ff09fb6f3d4634f1e26d

Observation bb9c855c-498c-4feb-81e7-5ce5e028ed8f · outbound

This paper cites Plotqa: Reasoning over scientific plots.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Plotqa: Reasoning over scientific plots

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.663104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.663104Z digest=sha256:def1890735da4192a726823a0100b3fff7c1739f9cb9809dc985039998d1cd0f

Observation 20217427-0147-4b64-8e23-c4a60a556ade · outbound

This paper cites Hello gpt-4o.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Hello gpt-4o

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.666668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.666668Z digest=sha256:5292fbbe4e5968f785e6d9b06f92a6e55b4d8b548db4b6d3c84fcb4fc2399066

Observation 5cea753d-68a1-4577-9f8e-c96e07628b5f · outbound

This paper cites Introducing chatgpt pro.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Introducing chatgpt pro

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.670172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.670172Z digest=sha256:d3f94f701908285a6db547887d3c18ddfdb157b12e4b6ca58d4439c72259d821

Observation 406c6e8e-e7bb-4e91-b9d8-cc0ef78359ca · outbound

This paper cites Learning to reason with llms.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Learning to reason with llms

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.673685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.673685Z digest=sha256:661fb3b0614a4010e0f6cb331aa96b4482288b5611679e8fefb87e154489ea5f

Observation 260d2619-6031-4d65-a215-074bc253c68f · outbound

This paper cites Learn ap physics.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Learn ap physics

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.677273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.677273Z digest=sha256:28146fbf1ccd8f382dc45f8d1052bdc2536db0c57366fb2e9fc79d8f463ddff7

Observation 9f10e2cc-25ac-4f8e-82d7-59cb9e5225fd · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Qwq: Reflect deeply on the boundaries of the unknown

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.681278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.681278Z digest=sha256:38c4840ab0acbae4ceebec1fe08ab58c16eff2d2fd0bc3f7440bd23ba55e9d62

Observation 643d4f34-7607-4387-9370-3be0e6c30dc3 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Learning transferable visual models from natural language supervi- sion

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:01.021412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.685072Z digest=sha256:2a6bce688371d6570342aa1be3a8810931e91322e4cbac7b2cdd9175d7a62cad

Observation 903bcecf-2a35-461b-9a52-c0224d9c6cbc · outbound

This paper cites Does Spatial Cognition Emerge in Frontier Models?.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Does Spatial Cognition Emerge in Frontier Models?

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.688926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.688926Z digest=sha256:c8b1674d09244bdecf7e97e2c84fd60b2c6ed739c33117154c4010d5834c7a66

Observation eb48029e-d953-4e6d-b4b0-47d3b7d202d0 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.693064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.693064Z digest=sha256:d79686300fca4ad9c3bdff6c8ce0207f72feabeadbc63b7d97d244d0cad88839

Observation 9eb40ce7-f312-43a6-8f5a-56d6f48d77ac · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.697183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.697183Z digest=sha256:ace93070bf7bc33c2a6bc7b64399f53ca0feee067ddae053771d3d5706629041

Observation b926b98a-5a2b-4ad1-98e8-d7d6a8e3d96d · outbound

This paper cites ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.701191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.701191Z digest=sha256:ec6a351986612219c644d01fd5e37afe31cc3f5ccff461839cdda5d02caed2e9

Observation 04b26edc-cc0b-4a97-b13d-af27afa29223 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.705164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.705164Z digest=sha256:c0c87ca198b1d856e0672504f05b2586ee7dfd378010f13f84be6c75c7c97675

Observation 6b047635-622b-480c-85a1-ddcbde21afd7 · outbound

This paper cites Varco arena: A tourna- ment approach to reference-free benchmarking large lan- guage models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Varco arena: A tourna- ment approach to reference-free benchmarking large lan- guage models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.709933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.709933Z digest=sha256:3713adc348c60dd133919f7e7e89edcfb466b99ea311feadabc445c9a893b7ed

Observation 94cc29a8-17ae-4fe3-aa24-a73ec130fc32 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.713885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.713885Z digest=sha256:bb652b41855cd34a2b2bfd3358702e71acd7b10dae65c0201fc9f473a23650ac

Observation 2c19a6f9-f116-45e4-8275-9764fc426a7e · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.717663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.717663Z digest=sha256:b67894f4719f4c9e3f30146d1f49c971092a891c40a1a94bd377194c0f5b7155

Observation b4744dfb-bc4e-42a2-b3d7-be7e0870efb7 · outbound

This paper cites LXMERT: Learning Cross-Modality Encoder Representations from Transformers.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark LXMERT: Learning Cross-Modality Encoder Representations from Transformers

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.721749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.721749Z digest=sha256:b12c187132b6ab4d6a6f32eea6abc921eda28b85ccfd4e9f48d5cf235cc29651

Observation 51e65c02-b20e-49d0-9a19-59aa7d024572 · outbound

This paper cites Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context (2024).

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context (2024)

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:01.008628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.725362Z digest=sha256:5d264a0aa677709a9fb09d904e1489ce3619d40d3511e2a30ae59b65b1081a30

Observation 0742aa80-6b3c-48fa-9ec1-b196452cf0a0 · outbound

This paper cites Examples.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Examples

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.995760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.729315Z digest=sha256:72ab14af00db7698110c656515a44ce830f4a2f74cecb5b293c7cbd361ff14b6

Observation 7b9733a9-9b99-4fa3-a9df-25512b5dd4e2 · outbound

This paper cites Internvl2: Better than the best— expanding performance boundaries of open-source multi- modal models with the progressive scaling strategy.https: / / internvl.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Internvl2: Better than the best— expanding performance boundaries of open-source multi- modal models with the progressive scaling strategy.https: / / internvl

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.983841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.732995Z digest=sha256:cb8f6de1355f9e48d08dad168084145b276c2df357e7a0d6769ce86d374a5aed

Observation 9f07bf62-9d4d-4e98-af49-6325d5bfaeca · outbound

This paper cites Measuring multimodal mathemat- ical reasoning with math-vision dataset, 2024.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Measuring multimodal mathemat- ical reasoning with math-vision dataset, 2024

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.971755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.736490Z digest=sha256:373be8a6b42796b4b090a81cc307bfe0f70d9f830fe557af0e910eb695622abd

Observation c83163f6-0433-432a-9d37-7562fef9ad19 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.739996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.739996Z digest=sha256:d69ec262186ed43a97740c8d73735e8ee1c69d0f0c11aa905f3d6eea5d8a7a62

Observation 98b439a5-bac1-4166-a0b1-2a79c2f0b81f · outbound

This paper cites CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.744440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.744440Z digest=sha256:54b8cd40a27f40e40defa0ade49202ca86fc194842607a5b8f6b6c61ff3c92d9

Observation 39202520-efa4-4ef1-9b55-5da3841ac2b5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.959274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.749125Z digest=sha256:19d578ec7e9f7b1734a975557799291c682f09849ad5509ca25d29d70a3470da

Observation b025c019-e937-47db-9e0d-03731ffd5513 · outbound

This paper cites an unresolved cited work.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:19:00.945479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.753756Z digest=sha256:5cbd666329775cf1c9077ba7743adb03d53b1dec879cb4f3c5b882a56148c7fd

Observation a7ce70ec-02bc-4ff6-ac9a-b52d084494e3 · outbound

This paper cites Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.757393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.757393Z digest=sha256:d7deae2ad7f7be3013ff63c8060f43f2c8a77794950a003ca8185e40b1516db7

Observation fe070716-1d2f-428c-ab00-7e1f46e5e61c · outbound

This paper cites ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.761214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.761214Z digest=sha256:3f838428e6e15afdef317898f2f40c9baa7b997b54c4b2a48432e807819104a8

Observation 65e9244b-fc31-4104-852f-30695a3c8420 · outbound

This paper cites Qwen2 Technical Report.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Qwen2 Technical Report

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.764946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.764946Z digest=sha256:45b18387adf223711f2718ed5fae84e8fe28e692d3d27ddea992f0c5e1fe7aae

Observation 8a458036-67c4-4727-8596-ce5c897211f8 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.768750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.768750Z digest=sha256:47480e9d4abcbb3d65df77a9a9240d7163f3776a77c0127e13c8442df08aa0f1

Observation d8ed0517-b6c4-4619-9233-68357788667e · outbound

This paper cites Jimenez, Alex L.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Jimenez, Alex L

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.931775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.773299Z digest=sha256:839bce8d7c75a6f78b0cab838f780eeb596db86f310a8ead6d66d02955afb43f

Observation 94ee971f-2195-4d7b-ab91-d332cb320c83 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.777581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.777581Z digest=sha256:096adb03197d5caea111d4cd17cc139c1430a4971e65fe8b1092d22f68b09c03

Observation b32eabd4-57fd-4a1c-b012-c4c8e58938bc · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.781245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.781245Z digest=sha256:c8fc810c55f13795fe2d389f2173f4b1cc56a9aafc569a58e08b732bf258ec3f

Observation 345be9ac-78c1-4c1e-bf27-b86c6d76d8bb · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.784989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.784989Z digest=sha256:e3decdb36b4907d01093b5b9383822736d29b5874a3b226df40e0275f6505659

Observation 21804f5c-f753-458e-b07a-f607a8c784c3 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.919728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.788443Z digest=sha256:0c6be969ea4d903c3867042b727721db869af3b170590bfb9ba7f2be01888621

Observation 5d3f8ba1-f2c9-474f-8a80-b20c75bdf66b · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.792127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.792127Z digest=sha256:43e87f22a9c717649d75bb2d5c076a4a4483010979f2b7fb48bce5fed3c60c8b

Observation 050afd63-1e9e-4872-8d22-36f4c4c66b20 · outbound

This paper cites Raven: A dataset for relational and analogical vi- sual reasoning.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Raven: A dataset for relational and analogical vi- sual reasoning

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.903148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.795675Z digest=sha256:31ddf3c01b4cb792a17ce2596d8d6a86da110426faac424816b9763802aa6aad

Observation d12b1bc9-7eb0-4ade-8d4d-52a7cb3fb672 · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Vinvl: Revisiting visual representations in vision-language models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.799484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.799484Z digest=sha256:a6f5d21d0175159dd3af6cce10bba2c9a71a1714e4a201988ae75b19a53ab7c4

Observation aaae7146-96f1-443b-828f-a01d8c631281 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.802975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.802975Z digest=sha256:597d8456ee0e163e856e2a2ef4cf4f4648e1f5e9099d67ac38baf17f16767a26

Observation d463a017-94cf-4691-95b5-a2027b7c85b5 · outbound

This paper cites Is gpt-4v (ision) all you need for automating academic data vi- sualization? exploring vision-language models’ capability in reproducing academic charts.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Is gpt-4v (ision) all you need for automating academic data vi- sualization? exploring vision-language models’ capability in reproducing academic charts

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.884684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.807191Z digest=sha256:0fc4db082ae1340ced84f678cef9b39b5664009d9941ef883358ea91e79ae9d0

Observation 2711860d-6858-42ec-ad0a-3ce0d9c7f0ce · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Multimodal Chain-of-Thought Reasoning in Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.810896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.810896Z digest=sha256:46f2f9fa2991a28fcc77e0d8671737e593375b4812292e820e642d9b5fac1675

Observation a51f4df2-e185-40cc-9410-f8be7f4d02ca · outbound

This paper cites Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.815179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.815179Z digest=sha256:650e21053e3052f146e3a754b159a2f8ce2c0113adb21d87098caf9601347280

Observation f90a52ac-5fe9-401d-a65e-16d176c36b0c · outbound

This paper cites Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.819521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.819521Z digest=sha256:9ee7808681eadd002c9c30133889e4bf786cb8f88514da521941064300ac6f34

Observation 97664272-fe20-429a-9cd6-eca10c4c8298 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 92

Resolution
malformed identifier
no resolver link, observed 2026-08-10T21:18:59.823334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.823334Z digest=sha256:f3421bec948b1e7c03cca8c78f906996b628f0ecbadb49dd1f82fc6ae775810e

Observation 65fa8981-6ffa-48d0-b0ff-1855e2056abc · outbound

This paper cites an unresolved cited work.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:19:00.872262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.828402Z digest=sha256:7a974ddefbc475a5bacba281e98f5b74ffd3791d0832b790ddb67b59baca2221

Observation 3fc875dd-3c52-4489-a6b8-215e1f8df131 · outbound

This paper cites The vertical distance is 2 units up and 2 units down for each peak.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark The vertical distance is 2 units up and 2 units down for each peak

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.862054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.831987Z digest=sha256:9bc39fe7c32db8bfee6ac4ae3dc90a5ce3558c3aaab6c99b06797ec33e7fce28

Observation 29f8b3f7-ddbd-4776-9343-2d4e6f03fac6 · outbound

This paper cites The vertical distance is 1.5 units up and 1.5 units down for each peak.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark The vertical distance is 1.5 units up and 1.5 units down for each peak

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.851393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.835821Z digest=sha256:827e328189031d41ae21c20d1d009c41c23e6b8a5f6db5e7ebedaade2ea710bc

Observation d7180cf0-e327-42c9-8baa-85222805dbd3 · outbound

This paper cites Therefore, the total distance between the two points is the same for all routes.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Therefore, the total distance between the two points is the same for all routes

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.841284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.839750Z digest=sha256:791aec38dfa5d4f23feb4522797a7b56216231471b4644a7367bf3dcf71f5aa0

Observation 5a8737b9-7df3-4bb5-9c18-f3a409915ee2 · outbound

This paper cites The second cube shows three different faces: a green triangle, a blue circle, a brown arrow.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark The second cube shows three different faces: a green triangle, a blue circle, a brown arrow

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.831711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.843483Z digest=sha256:855ae66b75c493c52d105d3676d80f0e70aaa6e7b672bb653d095d09c9021adf

Observation 522282fa-f8fe-4e66-b9a3-bfb8ed61b6e4 · outbound

This paper cites On the first cube, the green triangle is adjacent to the red square and the yellow star.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark On the first cube, the green triangle is adjacent to the red square and the yellow star

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.821811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.847161Z digest=sha256:6a5fbf498fd51a99baa3cf04d9ffd2b0260c234fd647de99bf86a62fbed32f32

Observation 4fc386dc-f6ed-41b3-9499-db7bcc435702 · outbound

This paper cites Therefore, the face opposite the green triangle must be the kangaroo (which is not shown in <image2> but is given in <image1>).

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Therefore, the face opposite the green triangle must be the kangaroo (which is not shown in <image2> but is given in <image1>)

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.811916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.850900Z digest=sha256:d344750a63b8b5636600820d73b84aca28016329cf77acfb96963d159e5eabe6

Observation 2b062b73-8c22-44b2-bf54-ef27eeab1d6e · outbound

This paper cites The face opposite the kangaroo must be the green triangle.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark The face opposite the kangaroo must be the green triangle

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.801556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.854408Z digest=sha256:8269154ce3415341684d90af06b00288f799896f74f270755e4658329827cbf1

Observation e7ea0d9e-c8ac-4e9f-9377-c2cdd81f6daf · outbound

This paper cites Final Answer: \boxed{B} Figure 17.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Final Answer: \boxed{B} Figure 17

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:19:00.791134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T21:18:59.867228Z digest=sha256:723753da481abfe447eb715b3f72b813bff3c1cc81f7a78587271ad335d0d988

Pith citing papers

Observation 59217165-9cd2-4838-ace1-1f645cccb2b2 · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:32.840359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:08d0b2e772b153a87359835b08e8352edb59e3022a446b60ee04d17effec0152

Observation bea1f8ac-2ae8-4271-97b5-1c28b2ddafb2 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 241

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:40:41.471398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:68eec44be51283dfbe7be63a54dde0ac862b8d8401de23a1d9d28d9beb385323

Observation fc20a570-a8c1-499f-a529-97c46d9fcd85 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 167

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:18:53.278400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:226705278ed897e930b4b4586d28ad54bd618624fd2b819d0de8b573e186bdae

Observation a431e228-57cb-4df6-a848-fcc25fc8e8f5 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T06:59:03.400181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:b09f1f8576b52693734cb509229426d12a69295583f64f382338f1dd22dd66e7

Observation 6de50937-63e5-4a52-8c9a-237f56fc1afa · inbound

Learning to Reason under Off-Policy Guidance cites this paper.

Learning to Reason under Off-Policy Guidance Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:17:02.892270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T23:17:02.701393Z digest=sha256:8e101479a7a65b91b07b2550816cfd427905eab1bb9ab44e781921c8768ce28b

Observation 4e8e5026-c49a-4518-9216-5cdfe46d6385 · inbound

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models cites this paper.

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:47.134304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:33:47.134304Z digest=sha256:4d0d9f27dd922416040b6da83404ff449efae170de024ef96e478e97a29a570b

Observation 8cf6e07d-d06f-4b29-88ea-c79f4e77bcec · inbound

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models cites this paper.

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T05:12:18.221320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:12:18.221320Z digest=sha256:64922120b8635c8f876fe313e1afb03b92f68bb8631c80c65f8e9b35cf52bcf9

Observation df00b970-8667-45ef-b1e6-a5e22a608ddf · inbound

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models cites this paper.

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:35:54.408983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:35:54.408983Z digest=sha256:7480e28769fa9d89facb05037c8544a82f04896911acf656f022e956801ec1d7

Observation ca4d68d7-8604-484e-b562-19be0e03b0cf · inbound

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems cites this paper.

Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:06.836442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:29:06.836442Z digest=sha256:94626597c9b456c17ea683d56ac15d3488448b18ed92793a52d4f102f19fddcd

Observation 8cafd74e-8f58-4fde-a158-02ea21ec09c9 · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:59.566161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:59.566161Z digest=sha256:f2cc96271d2f41c95b6ae70d83bd07330602a8cf228b666b2c0a4f44425db0a6

Observation 3269c397-9cf3-456c-a0c7-69b46502952b · inbound

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? cites this paper.

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:56.253584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:56.253584Z digest=sha256:fd17afdc4112fd4a6acda68ffbbb0bc765478bf46c5dc8ab5a3504f9a399f2ba

Observation 13c6dd49-fa38-4685-86aa-7bee9f02578a · inbound

FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow cites this paper.

FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:16.478809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:16.478809Z digest=sha256:b77fce6d44bc57f80d41803c4d246756c2496150c463fdefd6563f9ebb4cec00

Observation e892892b-c10d-45a6-8065-9432d3feab2f · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:14.468895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:14.468895Z digest=sha256:e54a2e1527a12727a6d52564e27d931dcfef7e2acefc07bb73f8e67b94d892b4

Observation 82ab3d3d-13f9-4382-8cf0-d5cfdb58dcb1 · inbound

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model cites this paper.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.650097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.650097Z digest=sha256:835d9a24b62753309a9b975695655b89116f4bd87eea94c8c32021c9fbb90754

Observation 0af8138d-b50b-4120-96ae-298ddf751c9a · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:03.859705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:03.859705Z digest=sha256:ac48d8ba86b46178d1e4a0b13280dd20007135c4efa38eb1d7c464c5469fa3a2

Observation 060b87ff-030b-457b-bbd4-e859554de971 · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:57.253578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:57.253578Z digest=sha256:280aaa326adb76973fdea4ae81cef205a9f778514d34542e095888669a61a7e7

Observation df31dd7b-4de3-468f-8f15-7c3087c86cca · inbound

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs cites this paper.

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:01.334554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:01.334554Z digest=sha256:22d140bdee20aa9338b93dc68e538db6cc29d3e3a48e69cd040b957d677b7649

Observation a5964cce-ce42-4bca-9940-f8b194a075a7 · inbound

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT cites this paper.

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:36:21.244871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:36:21.244871Z digest=sha256:d7860d786d66c1d514e90dce99cb137cb81e7dfeac345c82f2a0882b184b57ef

Observation 58541b15-d921-4064-b89d-6dee79b1e6b8 · inbound

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models cites this paper.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.367946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.367946Z digest=sha256:9f22eae294ddf32259b8a49d6472cdc0ac3dc4ea0deaf334d62217cf9c5dd7a9

Observation c27708dd-cd6c-4f19-9f5c-898587f63c4b · inbound

Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency cites this paper.

Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:16.540534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:19:16.540534Z digest=sha256:f57afe63d1a31788e13be41b9c3b6d59cacd3dd9f088744d4fb02f987b64ac05

Observation 9024f7a6-09d5-44b8-a2c3-b8d1e33ac7af · inbound

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs cites this paper.

ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:08.211304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:08.211304Z digest=sha256:adfdb0c840821451585898a6de56e6e8d6c9bed07492d7dd11d8a7b14db27b18

Observation 55df135a-9ca6-4fe7-ad75-4bb19f440d8c · inbound

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning cites this paper.

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:34.684819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:34.684819Z digest=sha256:5cadc8821a6637e71994dba4b461e59224a1a2758f04ce78b776d4884adddd30

Observation 6d4fe260-a268-459c-bf38-30fcf3dd1d80 · inbound

CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning cites this paper.

CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:39:18.674276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:39:18.674276Z digest=sha256:21964fac91e858fc72bf592614c11262357480498e84cefa5f6fda7f8d5f126b

Observation d1c9e39c-da3b-451e-b3c2-7c06538b096d · inbound

Skywork-R1V3 Technical Report cites this paper.

Skywork-R1V3 Technical Report Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:03.335779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:03.335779Z digest=sha256:332cfb87e86e77584433f5b4e79fe6d0326a55da9ecf2c712a16bad2f10c5283

Observation b7ae7c65-1c24-4eef-8281-58c46bfdc0d1 · inbound

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning cites this paper.

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:35:11.882839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:35:11.882839Z digest=sha256:54b54bb4d61caa1bcfcdda2a13e8bfa90e73c713d2fb1e3b94a3a067d8d919c0

Observation 9d5b0801-966a-45f1-9e27-1ecb18d0c745 · inbound

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models cites this paper.

SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T16:32:42.934041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:32:42.934041Z digest=sha256:74a914e3ff587c2e87018eecb25984a963e44bdfd0814216b4f184b5225d3d6f

Observation f86f6a0d-31ae-4ba1-9bb2-869119a11ede · inbound

KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts cites this paper.

KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:23:52.822183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:23:52.822183Z digest=sha256:8545cb54774d8fe8bd224990ae12d47150b5f434f77a89d42adaa19dbe3a0c59

Observation aa570e30-2544-4fe1-8f67-24262a221426 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.653975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.653975Z digest=sha256:e7484e6dffe525a6c89a0fa318a389d17877f4e5cba0c02d3b185e692d73aa2c

Observation c66eeefd-99ec-410e-8435-a40e3634a2e5 · inbound

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe cites this paper.

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:07:27.398933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T17:07:27.277040Z digest=sha256:eea4243dda85366c986f9299d877dbeca944896866bbeeb611a6575d8c058cf6

Observation ddc2eb4c-96b9-4766-992b-70958a3ee1c7 · inbound

SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials cites this paper.

SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:56.977544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:56.977544Z digest=sha256:3fa8ed719624d201b6c187ee68e9b3a91aa27d4959ab3682e089d0e54ca91fcd

Observation b5e52865-c1e2-4a43-b5f1-2576227b5ab7 · inbound

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture cites this paper.

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T18:00:26.874973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T17:59:42.921867Z digest=sha256:fb45ec737cf86e074a721909eb899cd038cacbe376d26342213ddc283f769d76

Observation aa028a27-b5f1-47b4-99c1-6f0219a0a6b1 · inbound

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery cites this paper.

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T05:23:33.121778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:23:33.121778Z digest=sha256:eb3af3c07d21e61c7fcfb4327e704816730eb2709c255f9ac956231e2c68cc29

Observation b8dcbd05-29aa-456e-a0fa-05bb7fc6ef92 · inbound

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy cites this paper.

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T20:36:04.457394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:36:04.457394Z digest=sha256:5213781ad495e2bf21967435a6c29735fcdec256c96124d5057d086f9287f29b

Observation 314fcb5f-e6d6-4dd8-9e5f-f82c27d3b6bf · inbound

Seed1.8 Model Card: Towards Generalized Real-World Agency cites this paper.

Seed1.8 Model Card: Towards Generalized Real-World Agency Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T07:45:14.287576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T07:44:02.827006Z digest=sha256:1249999f730174a596fda73190b3d2747ccc05dfd766fd194cb0fa052e179eb3

Observation fabe2c58-eaca-4c41-9005-42d8daed444e · inbound

Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving cites this paper.

Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:46:04.311521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T00:37:55.147350Z digest=sha256:3a4e1637156695823c367c098470b40b3d02c717f6f3d01d8b08ef3bf9526911

Observation 3325c29f-6bc7-4f06-b15b-841b61ccdeb9 · inbound

OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving cites this paper.

OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:21:03.874703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T22:04:19.654714Z digest=sha256:35ac8fd5a77053e4cf9a15f14ea2f581a617b39b84249214ff567a0d4ca69828

Observation 67764f88-0862-4b1c-9f4a-002e83632f6b · inbound

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning cites this paper.

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:41:27.517349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T02:24:12.349405Z digest=sha256:bd0d5e176e441419ca60b812ddfea417668e4d458624169a42d0b4a481198c6e

Observation edaac7b9-5988-4235-82c7-2e1c8513f50e · inbound

Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning cites this paper.

Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal Reasoning Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:15:02.884007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T05:14:28.256032Z digest=sha256:d14fac8b82750631b9facf3be70423aedfea1d5cef6f9c5231fa2253cb6a5a28

Observation eb529714-10e1-4e6b-88ff-28e8a71ad0a3 · inbound

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models cites this paper.

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:48:53.490345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T18:44:54.835575Z digest=sha256:eaa8f78ff04eafe1053362b3554b8791814c17a4f75f331ded15055c8306b99c

Observation 7e1c70c0-7481-4bd7-af97-e107c43eae73 · inbound

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution cites this paper.

AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T00:14:04.635099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T22:56:39.504430Z digest=sha256:71016f3c931ea26a4bbd6673f7a387dfde77eea85e1a9746620910b9ef631a04

Observation 41f675df-f331-4c6b-9d3f-3dd540b15384 · inbound

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs cites this paper.

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:57:06.715624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T23:15:38.962013Z digest=sha256:3cf6e9df7ea5a5644c33f5dc46fd4ec2722f10e8827a4bbfb37f000c5538572e

Observation 01469653-fd64-47f6-bc32-e6f8c0ffde20 · inbound

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP cites this paper.

From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:50.235779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T05:31:17.461916Z digest=sha256:38d0694269eae0a08932ab149d63a4c59629a8a7c2dc74f2d9908b75115a51aa

Observation dcc576b0-35a1-4d70-8ba8-bdb99db859d1 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:07:17.410212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:039b2470b2546e949c95e5cfb8410b250e428ddcad9ca9fd87bc0816367b06fa

Observation 1ad6a5de-5851-4a93-8ca7-e9b1d05f83b4 · inbound

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space cites this paper.

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T06:13:16.894125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:13:16.894125Z digest=sha256:02d58cae4b21cdc9340cd5f2b016befc8cd1256682595c98e20447f2cde4e195

Observation e6992855-794e-4037-b7d3-00ab9c601dcb · inbound

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models cites this paper.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T05:01:28.384980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T05:01:28.384980Z digest=sha256:85214277bbcc37f550aeb3a8a4a57dd8d4f53542ad0ec40e5b67c9898f9fd5e9