Pith. sign in

Paper Citation Record · LEDGER

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

As of 20 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 32 inbound Pith citation observations for arXiv:2502.09621.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09621 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:52:33.342830Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.558808Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved26
  • parse uncertain2
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5268d74f-0072-4c18-9d41-55e57ef0b1e3 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.221914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.221914Z digest=sha256:8d3c8cc5dd7173f1f5a0718e8964e01555bd9988fd6209496fe5d83d051f0938

Observation 60c66316-3e0b-46d4-98f2-843bcc714195 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 2

Resolution
parse uncertain
no resolver link, observed 2026-08-07T20:52:33.225874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.225874Z digest=sha256:cdd7136b39aeb866f33f6397c4aa238ce637765d8c4e6e30326e8169556b1a74

Observation 3d8a2952-ba3d-4120-a0d7-681deac05cc8 · outbound

This paper cites step_index.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency step_index

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.710338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.229574Z digest=sha256:a851cd7dc4b33e78647ba99f449a9a45e5c4f253d083bde9d4fc7680145fe5b7

Observation 749ba094-bf85-43c2-89b7-cb04ac67f73c · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:52:33.546733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.291073Z digest=sha256:9cbf4bac5f96d81141b86480deab10e839f4eedb56bcac7bddad1c1888cf1214

Observation 832d0332-acf3-4264-a511-6511a964f633 · outbound

This paper cites Direct Evaluation Prompt Answer Extraction Prompt You are an AI assistant who will help me to extract an answer of a question.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Direct Evaluation Prompt Answer Extraction Prompt You are an AI assistant who will help me to extract an answer of a question

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.503790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.311478Z digest=sha256:0d4f0ade9d729f9e2b82b4c123dbcb4b0cddf0b6c8caa02e7e6251278a30610f

Observation 1d09d559-de96-461e-9f8a-bcb115ffbfa3 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:52:33.701082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.233364Z digest=sha256:fa93fa12c40da449c36746706f6c8ea559b68596d952b15be10f653c95bddb93

Observation 241ae9be-adf7-4cd6-b396-f9284dd70b26 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:52:33.691491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.236863Z digest=sha256:bab3dd96aa5fcbaa9361150123f8a1ca0b234b3f57a34bb5e38e392ed379f02a

Observation 7508e86b-7e80-43ac-9dac-77cc4b31164d · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.240393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.240393Z digest=sha256:9954fdefc534440699a1897029aff38b728a55e7d1f1cbff829af2eafcd57b12

Observation 7342f0ae-4038-414c-a4c8-f818b9232e99 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:52:33.676389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.244126Z digest=sha256:fbbc5414ebba16b074f2cf718543817a08ef7f8d7d092ac29b74cea26d46e0ef

Observation e591a6b1-a6fc-46be-94cf-34dd91b75fa4 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:52:33.667132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.247459Z digest=sha256:80ab024a2d618846a02be982b58f76c86f42211c17f5485f147b957b6b55d0bb

Observation 746b7039-1864-4aba-8952-4eed97adca3b · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.254203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.254203Z digest=sha256:ad157f4a86510cb69460ed5fee1c6e9a626847510ab0c686b459b2c3fa94eeb7

Observation a29817cf-be52-47a9-bda6-8c4f82248fd9 · outbound

This paper cites step_type.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency step_type

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.644550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.257705Z digest=sha256:690fe97f5cb4e009d3d4f3bd0818bec58ff5fb138e6bffe1aeed78bf8c38eed6

Observation 21d869f2-4f1b-4ad2-a4ed-51e2f37e3a85 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:52:33.633747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.261219Z digest=sha256:171ead9afcb19eacfafeb39aa195d28aad1385e79690c77e73536e285a2ba319

Observation 3e9c6a3a-f02e-4225-b0e6-bf7d80d07348 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.263903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.263903Z digest=sha256:23445a0d2022f390ba0c0a4c29ee15db89d0bd67b6c4819df09ef9586ceef65e

Observation 1032d259-babe-4be4-b0f2-ae96d820a6e9 · outbound

This paper cites IMPORTANT NOTE: Evaluate relevancy independent of correctness.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency IMPORTANT NOTE: Evaluate relevancy independent of correctness

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.616437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.266537Z digest=sha256:5bddb2baa4825495124ba81b94efb923443e535be93294ae4fcfbb683858bc64

Observation 11f8dde7-ef55-4020-aa0d-c3f958bf297b · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:52:33.605578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.269253Z digest=sha256:8a935186200936221b6b6b8ee602880d5e21160ee7e5163dd98c32cab51a35a9

Observation da8972eb-e330-4557-807f-c1afbb1d641b · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:52:33.595247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.271846Z digest=sha256:bf94bb15c6a7e5da377a5f7ab1890c7fde969b257d877fe2fc8b6c3bd6edfe2d

Observation a4dad2aa-bddf-4ab1-bb8c-2ed28d09d63f · outbound

This paper cites step_type.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency step_type

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.586145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.274406Z digest=sha256:28609280b33278956fe6b82efbd59f46ae10eeb8d978bf191b0ff88d166ed2fa

Observation f229736c-cd06-4e7e-a4dd-3cb1bf080ef1 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.277453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.277453Z digest=sha256:530e029a27c79f8e1194737dcf5d43b05238df2fc229a3220795c3881c5d8d38

Observation 11e884f6-1d79-4d8a-9e0b-bb4a451af5dd · outbound

This paper cites Invalid reflections include:.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Invalid reflections include:

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.280116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.280116Z digest=sha256:b824a2d1dbc4734608ce5a206a94fe7ca46297396be81ffd5ecd874e05e84a9e

Observation 43507bb2-4e81-41b0-b8b2-08a2c2fd8f49 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.282828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.282828Z digest=sha256:4adfecf3ac9b75e326cca1bfd6618e7054156bee9fc4eb7a98407a8cd3bdd110

Observation d429e9d3-0ad0-423d-9e8c-2bfa7c08bdae · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.285515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.285515Z digest=sha256:d965a2926c438708aa0fe8935d6fb12537b74f09cd521f6218b325df91920712

Observation 3d04a9a0-b288-4756-af10-3de9cad5a1a4 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.288275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.288275Z digest=sha256:1f16c436e4ffd98eb24c7159dadf6ec9d2408db50d16928ddfc57a0346494db3

Observation 2abfe026-ca1a-44b2-83d3-562527c2b609 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.293949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.293949Z digest=sha256:dbb26b2d92e93c6e0380c8423b0f7faa6293165b42d790f648815f2609be1dd7

Observation 49ea6b87-a217-4af2-92ed-84ad8f2f2634 · outbound

This paper cites conclusion.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency conclusion

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.537854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.296767Z digest=sha256:56b171d160c344ab35d2f61e4186a40e5aaabc7a8576b051dc950b5e1781b1af

Observation 658fb650-a3c3-4139-a9a9-be57ba9058f8 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.299851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.299851Z digest=sha256:32745b6fa56dd43a4058c115fb696a8a6c8d703248926cc0d58ba66bd6c8cae5

Observation efc26d3b-6a51-4a2a-90cd-51fabbc5dad7 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 29

Resolution
parse uncertain
no resolver link, observed 2026-08-07T20:52:33.302709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.302709Z digest=sha256:ad4778c616c3dad8a2b418a883a969dfa0f1d20fa08a074190d32c794d3cd33b

Observation 2eecf3a5-432c-43e4-ad94-76153ff480ea · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.305751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.305751Z digest=sha256:b187bf65649d92d86be063894fd840b9cc4c3ca97bd13b7c32a53143e9f8c745

Observation e986f230-4515-4aff-b65d-5014966819f2 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.308518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.308518Z digest=sha256:d2c1a4ded72d92358a168ea909b13e21e6c7bb8ab6ec2302e8a061ca8cc7b14a

Observation 9b0884fd-23a3-4fe6-bbcc-285e116b8222 · outbound

This paper cites You should directly output the choice letter of the answer.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency You should directly output the choice letter of the answer

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.494822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.314304Z digest=sha256:acf3093576fdb00b1c9bcb2b720568a257bd37270f1684aecc4cd53341426bef

Observation 6bb9ab0a-e049-4ff1-b1fd-3479e22a0c0c · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:52:33.485104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.317252Z digest=sha256:54c2942c832adba3c352aadff8ffbc704fe726adbf639d5012250fe308f86254

Observation 1e297348-8193-4029-9ab6-177516bf3c7f · outbound

This paper cites [Non Multiple choice question].

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency [Non Multiple choice question]

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.475485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.320125Z digest=sha256:41d08bc22f70e7e03d78c7df8d6679ab5dcf27971c22eb2c547f9437d3b6d51a

Observation 8796a0ba-ad41-48ae-89d1-40a8783cbe91 · outbound

This paper cites It could be hidden inside the last step of calculation or inference.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency It could be hidden inside the last step of calculation or inference

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.466549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.323273Z digest=sha256:d53c215f87670b34df88439ddb7c69ab5cd8902a62abaf585f8c78d440bb21a4

Observation fd13c4a4-d569-4979-a6eb-28eafebd606c · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:52:33.456824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.326384Z digest=sha256:949b1b33e5b85bdfff47398048ff0e4e0f5433bb117c5fd2dba5498128f12a33

Observation 48afadfc-f93a-4090-8f79-e0d194c86a7f · outbound

This paper cites Output Format: Directly output the extracted answer of the response.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Output Format: Directly output the extracted answer of the response

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.447385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.329142Z digest=sha256:85da85bb1cf7508aa09fd878a4a30b5774aa1b7797b9accd5996fe2ff070be18

Observation bc756eaa-48ea-4aa0-a715-10c6b348493b · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:52:33.436429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.332115Z digest=sha256:88aaea5d658a104934f3e44b0f4e5f2e468122870a9c5ac6a6b0ae7ba9754228

Observation f3237c9e-de57-449e-b153-4bacada0e294 · outbound

This paper cites [Nan-Multiple-Choice questions].

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency [Nan-Multiple-Choice questions]

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.425389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.334774Z digest=sha256:11eecf24aafd11da6c4da8d6a3700881d3edfcd77048ebb7b60147d49ce0895a

Observation a4529c65-ed23-4002-a967-8c6ca0b8cbf9 · outbound

This paper cites an unresolved cited work.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T20:52:33.414076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.337498Z digest=sha256:35e966463a9d6b3c4d7dada27d4a3ef8276afd644009d2aa144b915ca42622eb

Observation f6f4944a-4182-4228-97af-349f37cbc016 · outbound

This paper cites Output Format: 1.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency Output Format: 1

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.402824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.340188Z digest=sha256:6b539fc5a694551a8c66209520e1a7af339a5d9069f25ae767c2acf0ff450cd6

Observation ee09987c-0794-45a9-96ca-92086d10a819 · outbound

This paper cites {In Context Examples} Question: {question} [Model Answer]: {extract answer} [Standard Answer]: {gt answer} Your output: 35.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency {In Context Examples} Question: {question} [Model Answer]: {extract answer} [Standard Answer]: {gt answer} Your output: 35

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:52:33.392727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T20:52:33.342830Z digest=sha256:8cae42c7413e349b6f71dfbeae7fd970a5a9a5548050dd8a417961b17dcca2a0

Observation 20d4ca02-ffc5-4564-a8f9-efdfbc6e8cab · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T20:52:33.212667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.212667Z digest=sha256:6f42de6aaa97f3880bc1abe8c28ce867a274b9ac6f67443a1de3709dcb8d9751

Observation 7d918680-6fd7-4a71-929a-1afd04f8f45e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency LLaMA: Open and Efficient Foundation Language Models

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-07T20:52:33.216898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:52:33.216898Z digest=sha256:78bf0b8cd01e9d1f2514ddd857eef825409942c1f18d739801795362c14377c5

Pith citing papers

Observation b26f2c23-f8ac-4194-a8a7-e5bcb594f061 · inbound

Chimera: Improving Generalist Model with Domain-Specific Experts cites this paper.

Chimera: Improving Generalist Model with Domain-Specific Experts MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:47.526714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:47.526714Z digest=sha256:22599416031b1e4bb02b8f3bede2bdbc9be3bebf052a0953560f92d54c39ae17

Observation 1da6b807-10dd-4f9b-ad92-86135e062636 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.307174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:6f94f52ad0acbfe28577f072e4ad621f7c15aab7b405cc2a57c7c29754957e03

Observation c05d8c33-2704-4454-b15a-81467a564cc8 · inbound

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark cites this paper.

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:40.558808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:46:40.558808Z digest=sha256:441b8a64a0fb51d793463ac6ee861503e1a9e99e2ae464cb3114c1ba16b5584a

Observation 43395219-3aa2-4dea-8dbb-8e7c711bb315 · inbound

GenCLS++: Pushing the Boundaries of Generative Classification in LLMs Through Comprehensive SFT and RL Studies Across Diverse Datasets cites this paper.

GenCLS++: Pushing the Boundaries of Generative Classification in LLMs Through Comprehensive SFT and RL Studies Across Diverse Datasets MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:55.178439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:43:55.178439Z digest=sha256:df7f59dbf731a17461b599fcf587f9a7ca0f838c08133c22fcf705255ce0a01a

Observation 7f638547-27f0-4112-b319-ff9b6036f745 · inbound

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models cites this paper.

Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:12:18.255202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:12:18.255202Z digest=sha256:c80f09222b3d6024c982fa1bbe0cfdca21040902f327f998acfea24ca49192b8

Observation d553872c-3bdc-45f3-84b6-02daa4b3731c · inbound

T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT cites this paper.

T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:41:47.606617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:41:47.606617Z digest=sha256:18e2808015186de08d42ac8ba9fd46422c2e1b69d44e814bf08153942be538ab

Observation e7c8bc35-93ae-49e7-8dd3-60f36a5cfcb2 · inbound

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning cites this paper.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:25.409128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:25.409128Z digest=sha256:e8ec668e3bdf2d1aa48fb04172c176244fd1c6f3cf16f346833db289deefe0bb

Observation 64fa8169-87f8-49cf-93a9-1ef6df20847e · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:20.174450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:20.174450Z digest=sha256:415c1c15eecf073f6243290153a983aba4a729df8720d79ae900a852c5e0ad60

Observation 62227da7-50b8-4328-beae-82ff7276f698 · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:57.991134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:57.991134Z digest=sha256:d9dec377b06181551a21f7a8a9e27047c3846b669c63dda974251f3597f3d449

Observation e13b0d27-6df3-4901-8c5d-bd1ed2145f42 · inbound

Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation cites this paper.

Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:37.308207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:37.308207Z digest=sha256:1d85c159c24dfad59470826f77679fed7b4a5c2035e3f30191aeb45a3bbe7abc

Observation cd3064e3-fe71-4bab-b0a8-eada1baeb521 · inbound

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT cites this paper.

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:36:19.638964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:36:19.638964Z digest=sha256:11e17551cf73fd708c0c8002e87f25de2bedeaab56c562903d14df559637e6ee

Observation a72d196d-62cc-4209-83be-37aec9055dc7 · inbound

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM cites this paper.

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:33:08.379110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:33:08.379110Z digest=sha256:521b0dc69ddcfb88c2aa4e3cade026a84f47456ecaa89b612c763145fa830a48

Observation 9f385f3f-9ae7-46d2-8098-d70880665cf6 · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:11.542607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:11.542607Z digest=sha256:8f9e228432a83ffa09b4a15c9e018768817e2769ad828b0373567c6cb31ee5a2

Observation 110bcaec-d601-4d7f-807d-775cf6756cf5 · inbound

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark cites this paper.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.357259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.357259Z digest=sha256:ba3772fe5b4b21f8dbcff994226c9af6fe0838b7224354464db3ecec60c4a409

Observation cac6d4c7-64c9-4415-9f84-4ca332bb8d68 · inbound

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models cites this paper.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.371674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.371674Z digest=sha256:4495659d43c919d83a09b51742a27d42658d8ad75db54e24ccfcf6c85d0cf308

Observation 28d5785d-4207-4f33-ae33-38be8ef5b07a · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:12:14.480452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:bf346d3a95e7b5d49be048bdc0bb5b21e0b1d0a733e1de14e72942611d174bf6

Observation 6d2fc3ba-87a4-408e-b6d2-28629dc2b5da · inbound

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI cites this paper.

MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:38.034436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:38.034436Z digest=sha256:88c8267a51324b2c55915c4c57b2d2d17b0f4c9b760ea5221c591955f2a3c2f5

Observation 537b8e27-1d48-45fe-9349-229f5710d9d2 · inbound

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation cites this paper.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.035019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.035019Z digest=sha256:3b9678c3c5153c1b1c91c6c2ce82357aeb735d4804e5d8195869f4d45ce68c38

Observation 6b305162-d8bf-4a8c-b1ec-785991a2c9be · inbound

NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation cites this paper.

NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:40:52.920863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T04:39:58.296388Z digest=sha256:0372bf6e01e934620f8ffed21211e30d95639fed126cc1c8efb42f9552d5f0fb

Observation 3aaa64e7-2f99-4fcc-81b0-00f248706cd0 · inbound

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture cites this paper.

AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:26.936626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T17:59:42.921867Z digest=sha256:4ae0aa66b82d8a365e03c4093b3d2cc48cca7183a450eb6e52b57193c77f5b55

Observation 56a1bbe8-72e1-4ff9-82a5-9707470a5fb8 · inbound

Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning cites this paper.

Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T12:31:13.837612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:31:13.837612Z digest=sha256:2fbda4b1adc3f5aef9ddb9dbac62f287be55de020e138326f43d4266e86bbb6d

Observation f622eb48-a493-4d6c-8d6d-0f1cb57d055e · inbound

Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification? cites this paper.

Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification? MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:53:00.658187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T14:51:26.368439Z digest=sha256:aa0ca809e3f54424a6698cb12f21c5fe6a18be940b99e9d647c7eda9cb9e3e8d

Observation b7c62b84-fba3-43e2-9958-c5d17e716a56 · inbound

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding cites this paper.

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T10:49:58.895997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:49:58.895997Z digest=sha256:02e331febb29f5d81fa81d16a1eccafab8c14d083a48cd5f19dce38dc2467125

Observation 22de22f5-fc81-4bfc-863a-512eface1961 · inbound

LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning cites this paper.

LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:03:52.139078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:03:52.139078Z digest=sha256:dd8146c79021dc577ec0e764788e4698628abc722b384781ba07a15609411d08

Observation 67411b12-62ff-4ea4-aa44-03ef635dfc84 · inbound

Diagnosing Pathological Chain-of-Thought in Reasoning Models cites this paper.

Diagnosing Pathological Chain-of-Thought in Reasoning Models MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T23:26:30.170558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:26:30.170558Z digest=sha256:b5d648275f08a98ec849fa4140ae38b53fc6f49b1d2abe13547f54ec598c6c20

Observation 30637812-e859-4696-a2b8-697ebc23ace7 · inbound

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding cites this paper.

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:51.274219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T18:47:32.778695Z digest=sha256:063e713653cbc818db0dc8b90b7e5047aaae265eed6d225602b03e7f59f6c922

Observation b2fe1677-9a24-4a1f-a230-d30daf5ec570 · inbound

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both cites this paper.

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:14:53.506781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T03:09:58.411261Z digest=sha256:a15c3c034a0b1bba5ef0a150dc458d366e5915771ab78a6addf060208f4bdaec

Observation 06e5d8b0-8dd7-4a95-84c8-b9f2c1406538 · inbound

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning cites this paper.

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:18:05.202239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T06:16:47.650748Z digest=sha256:33542420954d89c9cf9f8fa6fbe6e5a0cc15efcd22c75ce8b99628a3e5f719ee

Observation 11673db8-5d71-4ee3-8f23-1c19f5468d58 · inbound

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs cites this paper.

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-27T23:21:23.361183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T23:15:38.962013Z digest=sha256:d8a15466f3e2e63b11d8501cd926e3469ddc012ab432595bdefd302047b3de6d

Observation 7ed23d51-c75d-4921-97f1-fbf624222c88 · inbound

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning cites this paper.

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:14:19.152287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T06:11:09.693576Z digest=sha256:52c48897711c3bc596c6ada161c85946574b5132660cb499e5f59de1ecf6e4db

Observation 77198966-2992-4550-b956-aed73dcacc83 · inbound

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space cites this paper.

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T06:13:16.894125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:13:16.894125Z digest=sha256:11e7a7333f11993108639c092c1b267058dcba852d77d41e8d3c481a559858a5

Observation 7a8a1784-94f6-4d85-9bb5-47133c928073 · inbound

OpenCoF: Learning to Reason Through Video Generation cites this paper.

OpenCoF: Learning to Reason Through Video Generation MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:46:40.945808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T01:43:37.265723Z digest=sha256:dccee6e3cee0f0675c1ff5ebc8516d1aaf6ff186f4681c8f6325b63df4886a5b