Pith. sign in

Paper Citation Record · LEDGER

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

As of 11 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 4 inbound Pith citation observations for arXiv:2412.12932.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12932 v3

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:38:14.155401Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:18:59.500206Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T15:15:46.368932Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8034891e-c37d-4063-98cf-2afe3b79d316 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.965366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.965366Z digest=sha256:d02d2a6555cd0a676ed566540742adbe4ad42271fa93d16143e8c4733cc81217

Observation fbd2b36c-273f-44ff-aa61-02a06070f876 · outbound

This paper cites write newline.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.969608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.969608Z digest=sha256:44f93d182cd34f9e67a6299689120f7aa4b8d8791b687ab866140f7c24016335

Observation 57a2ac40-063c-4295-8be9-01a130c220f7 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.973903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.973903Z digest=sha256:4b8761ab4757388a622560f7a21fae166762e2ad3ce57b63fa391f74710abccf

Observation dcd8b321-e3f2-449b-8783-839897f77078 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.669899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:13.978345Z digest=sha256:05b70cce82f6c755b63b691b01c8d0e64c737a1e345e1f08cc984ff0b01a3952

Observation 8abbe37d-f615-4282-abd2-c305ace76b5f · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.658652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:13.981961Z digest=sha256:5ca7d9cd24f5aaba105c9a749dd62251a34f71779b3af000b7ebda6016835626

Observation 90671c1b-3e88-46b1-8aa8-74d8d62ec4d3 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.986574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.986574Z digest=sha256:7ef5586db91bed8ff6df7f050b7838b5193655b581236331ee2aff3ff9f7df40

Observation dfd098ce-faa8-454d-bc96-4272b74bfba3 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.990476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.990476Z digest=sha256:6627c50d8ba78b9db7e21cf27bbb0fd9c9b556c3532e6bb9c700159761d66960

Observation 383e9cdc-f93f-4999-8613-e7e20e7c666c · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:13.994300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:13.994300Z digest=sha256:c83d22f654f5ece113a8c29370283a1e46dc8407b36e149c250424662d1c23c6

Observation 68a5ee27-1ea6-4fdb-87c3-c7093dfd2bf8 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.641535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:13.998232Z digest=sha256:5a163e565f7d70de1b92301cdeb5f16b046950b7294aae9e8bcfa77e26c41eed

Observation 69a46e92-6351-43c6-bd9d-ef0b65308156 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.630603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.001637Z digest=sha256:21af64e268be7f8befef2da0199300a042e3a4ccb40d64af65b163c8c7b02cab

Observation 273f5760-434e-4911-b0ce-05243ef18e73 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.620403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.005393Z digest=sha256:360f0bf66ff8d9631391b5d754ea08b5428cb3de4176fe90e0324c40c5c71313

Observation 40d85930-82fa-496e-a1e3-df2c1e391c52 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.610455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.009044Z digest=sha256:ef26b3f5debfbcdd77500020d6f548f3be2da0f47b6eb083e14996ef7014a5b4

Observation c21bbc0c-d03a-441e-8f59-77fc46303491 · outbound

This paper cites P.; Poff, S.; Corredor, M.; Zettlemoyer, L.; Fazel-Zarandi, M.; and Celikyilmaz, A.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models P.; Poff, S.; Corredor, M.; Zettlemoyer, L.; Fazel-Zarandi, M.; and Celikyilmaz, A

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:38:14.599885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.012557Z digest=sha256:14e45a97f47c92cf9efc65c20d04cdd3926b746c5c5829ef38ed52db9bbed515

Observation 4f74b3fe-4c7a-4170-b8ba-e743f1333127 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.587672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.015834Z digest=sha256:94d61c7aa1d2e10195bbfae19239e2a64936bae4a16775734fdd99e6c5c1cde5

Observation ed6278f1-1757-48ac-ada2-5f1a4a9b2a37 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.019407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.019407Z digest=sha256:753a86d94effd1a6c9650dc5768a2fb01e344ca83cb07827f905a1c21f604323

Observation 19065e72-0b5d-4304-be0a-d6fc575e3cf3 · outbound

This paper cites Abstract Visual Reasoning with Tangram Shapes.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Abstract Visual Reasoning with Tangram Shapes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.022709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.022709Z digest=sha256:5a47701d83e370a8b4f06ac3d1d4413c961b281108fe4c1a6aa1bee69a9daa25

Observation 1be8d884-cf09-49e4-a3bc-98a9b2254814 · outbound

This paper cites Y.; Fried, D.; and Salakhutdinov, R.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Y.; Fried, D.; and Salakhutdinov, R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:38:14.571413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.026376Z digest=sha256:5fc90914e28e817dbedc38f23ef4df04c4de078b626b7c84d13958a7df3aa534

Observation c617cfbf-052b-43b6-a360-efd0e7859984 · outbound

This paper cites S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:38:14.559972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.029751Z digest=sha256:84571faa7db1408a938c6dba88fc3633b7be703d8982af68f8573a716a910a61

Observation babae6f3-b5e2-4c33-8fee-c83d9131dda4 · outbound

This paper cites R.; and Koch, G.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models R.; and Koch, G

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:38:14.548776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.032918Z digest=sha256:c5b2cb7c95d85f948d5b91966fbd4cf77f431af282dc39f016480764b9585041

Observation 2019cda5-9f63-43af-b405-3caa3415cba3 · outbound

This paper cites What matters when building vision-language models?.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models What matters when building vision-language models?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.036134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.036134Z digest=sha256:39c3b498d7c51082c0f5ca9ef235320882bf45d6cace24af986c04dcde71e641

Observation d9be3fae-36ec-4b96-906a-d84ef5d741da · outbound

This paper cites Multimodal Reasoning with Multimodal Knowledge Graph.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Multimodal Reasoning with Multimodal Knowledge Graph

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.039890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.039890Z digest=sha256:851c83cb311c9b6ddb488f8a7c7b824b2f7a048885575b6ab4ff0970fc7cd6f3

Observation 6564c27f-cbaa-4ede-97b2-007aef125d7c · outbound

This paper cites D.; Strik, W.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models D.; Strik, W

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:38:14.537450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.043346Z digest=sha256:4b2019f146db410f48254e90b3b8ecf6ea9c1a78f0ea9a7e770ae1e1984cf3b2

Observation 4544a36e-345b-454c-810b-ad01f34d769f · outbound

This paper cites Unified Demonstration Retriever for In-Context Learning.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unified Demonstration Retriever for In-Context Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.046278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.046278Z digest=sha256:4ddd34362207f7158ad1e77bfd7130a3db1a4c1346c0d49c6925b7d8bdef30b3

Observation 810e78d5-86a4-4ec6-8860-1ba54137ca5b · outbound

This paper cites A Comprehensive Evaluation of GPT-4V on Knowledge-Intensive Visual Question Answering.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models A Comprehensive Evaluation of GPT-4V on Knowledge-Intensive Visual Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.049627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.049627Z digest=sha256:9d6aa3e207b06fc0df8ba469487bcd9def14c32da7b1ab943f35cf86292f1d51

Observation 4f169db1-9c86-4ed5-9749-71ed3d63aa53 · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.053159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.053159Z digest=sha256:064e5eb14013c66d8233310cdb699844aec8dc4b132d9351e727ea9a7d4fb9ca

Observation a6208ccc-ec4d-430c-a07c-063dfc4d6b4c · outbound

This paper cites Retrieval-augmented Multi-modal Chain-of-Thoughts Reasoning for Large Language Models.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Retrieval-augmented Multi-modal Chain-of-Thoughts Reasoning for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.056640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.056640Z digest=sha256:78c06fa5ba90ae7f3fad6d1feaf0e796eed4db8d371eda277edcacd246340fda

Observation ec511551-9895-48b9-bd86-8286a527f007 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.525994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.060266Z digest=sha256:b55ccdb4b84d09dc70a3c06e59b1af790472a9f03d319443da4220c7918c4840

Observation 8a417068-b26c-4d4f-925d-5dbbcd1b084b · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.515788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.063317Z digest=sha256:901c8216fbca2f656f658b63a011b664fb5be51b5adb37d49fdac8ea9527e223

Observation 2339f6bd-5f08-4f8e-85f2-f7fdbddac1a9 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.066249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.066249Z digest=sha256:27fae9e39a1b8ec21c62f743b2595a58d0059e17c2015a6585961f71cfd4ad93

Observation b2655f6e-31a8-403b-9d7c-e5ca4003f63c · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.504841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.069654Z digest=sha256:2334c2839574eaaecc006b10eb19b65496f5bbae2d51a9e0c474f0487877c176

Observation 85dd9dc4-7908-409d-9116-467825017e59 · outbound

This paper cites Chain of Images for Intuitively Reasoning.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Chain of Images for Intuitively Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.072618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.072618Z digest=sha256:93447ed95ad363cdf07f33ba053452d06c954c8007d783bd99fb8d26cd30c6d4

Observation c7ca2f4b-34ea-4c76-aab7-aa5be1ba3c1c · outbound

This paper cites KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.076036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.076036Z digest=sha256:738395cfea8eeeb4a4a266ab68f157e0aecef4b62efed6a446cc31c7dba24922

Observation 8807733c-7086-4704-831a-69d459326c88 · outbound

This paper cites What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.079635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.079635Z digest=sha256:28e530e5568683e2338ecb277ea5446eb7f9e262866098cb8b8ad69bd7956adf

Observation e9324f2e-2037-44bc-8182-f6520df2b1f9 · outbound

This paper cites Large Language Models Meet NLP: A Survey.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Large Language Models Meet NLP: A Survey

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.083304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.083304Z digest=sha256:803ff7ec626b6d4ccc8b9a53ba4343e3736eacc2d3fee6d060bdc790e4c50705

Observation f5421ae8-7934-4145-b23e-4498daa03afa · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.494644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.087614Z digest=sha256:30880c268974348a41bf6c684236ac2289ac50c18e60886d2e5142b96e6a193d

Observation 713119a3-79da-4952-b398-a914aba96b10 · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.090582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.090582Z digest=sha256:4c5d35ca6168fced7431cb3f4a9db717744ebbdd5188de34f07de72602233998

Observation 6d6b0eeb-931f-4be8-8f6f-59d562627b0d · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.477738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.093782Z digest=sha256:cbd65319b6232b2f9508e69ae1d873910350bad034db30e1395375aa2092c314

Observation 78611544-0b79-41f8-9451-2bac1a3eb059 · outbound

This paper cites A.; Yasarla, R.; and Patel, V.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models A.; Yasarla, R.; and Patel, V

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:38:14.466729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.096997Z digest=sha256:9c5d2f59e3eb127f3c0c2a9a87dbe2ad3c9320ca36d45c9d999fb508efcdaf2a

Observation d45efeeb-9d63-49fd-8ff4-4b6bf220565a · outbound

This paper cites Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.100180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.100180Z digest=sha256:299373a02b3dceb8e91ede5f430d9c9ccb82edaabfef3c6d893889072f45dd6f

Observation 7911f6a6-a7a4-4fd5-97e7-50dfb40033fd · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.103515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.103515Z digest=sha256:2b4e0527ab660783f6e4d312ec2b9870f0d166effe06511fa21086f4675ee1bb

Observation ef7790aa-6b17-464f-84fc-46772035abfb · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.455530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.106809Z digest=sha256:79a595a88dd704743c8cc2befe2824b87e2ce60c2d38dec68ae7fa7803452af1

Observation 164d0c2f-6277-4f85-9c05-44ffea2f4a6f · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.110031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.110031Z digest=sha256:48f77946b0eaf81b100a2f216529fb90e331c9de65e517b1435907e7632e3217

Observation 4cc2f2db-edce-425b-9e87-24c260e23e27 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.444870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.113646Z digest=sha256:902c1f84965c8a14e6b7a6e8ba8cd023a39d86b20899dc299775be008d2c759c

Observation 3211c793-d67a-4f78-a89c-563cd12de61e · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.434445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.116780Z digest=sha256:8ed83661c653b84e28e07fdf909d8512141028a3ac0d2646a36b4addac3d45a0

Observation 142c6a96-f632-490f-b4e5-d58b7d8eaf93 · outbound

This paper cites Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.120001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.120001Z digest=sha256:1fb691f27ff932f4b161e89f2a1f5e5073a4bac1dc36f29067b6f6cf68f8fc2a

Observation 2a6769f1-860c-4fbf-a0f5-be13c7bc6c1d · outbound

This paper cites The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.123422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.123422Z digest=sha256:56102419dc5bae64bd45cfab79ea887d7cc272ac0b0a414d500668e7749745a1

Observation 47751f60-6b9c-463a-ae0d-459d1c858fb6 · outbound

This paper cites Faithful Logical Reasoning via Symbolic Chain-of-Thought.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Faithful Logical Reasoning via Symbolic Chain-of-Thought

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.126778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.126778Z digest=sha256:e606d8d2bed6b885fff6428f0127ecc1ed318d9089ebe6b602d145dc56688e6c

Observation 1ff5d7f2-ed6e-45dc-bfe8-47372802bdc2 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.130433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.130433Z digest=sha256:8d52e112dee6014117470a37165e0a466397e4c45a9d2e7319256e146fdb6662

Observation 6cb29fa8-a492-4297-887d-5ef4b222f7be · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.424158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.134018Z digest=sha256:2ab50c4ef71c40a5d84df5e6fe1f8e6628e73d207c83a233dbea1d7eb2087bc0

Observation 85b6cef6-180c-469e-a7e4-e54085e4c354 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.137261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.137261Z digest=sha256:0a84d940835157441f93354a8c7492815c2b3c6834462b543d5eaadc16702db7

Observation 5a18b25a-5a27-454c-825a-af4e52c2f3e2 · outbound

This paper cites CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.141106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.141106Z digest=sha256:1112d520f79aa551e2b31140ac557f6600cf535db204cb94f10b46ae379db149

Observation 1cd17fb7-3032-4d16-945a-305b1cbdbe31 · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.145181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.145181Z digest=sha256:e72d84b5a61b3b3a9e02e011f6b88b47426f4afc8308f435cdfa1ba9e69a66d8

Observation 29b5c9b8-f9f7-4945-8f4d-5cbf1548a5f1 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.148756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.148756Z digest=sha256:e3affeaa298397e6f2d3885416e99f9e50cd848a50bd1a5d09bb744adb18a945

Observation 0c7497cd-38ae-4c5e-934b-60964a4b1b1b · outbound

This paper cites an unresolved cited work.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:38:14.407242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T13:38:14.152223Z digest=sha256:c47af05f2e71b272e54978f99672bf6f259113b1c95e947dce374a0f382214c7

Observation bd71a01f-2ac2-4f74-8caf-d03740639725 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T13:38:14.155401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:38:14.155401Z digest=sha256:9e344be51c4e0f217c411f968ddde8071ed979c5c7dcb72aef1c45631d7193d3

Pith citing papers

Observation ee5f873c-6dee-4731-8241-a898d9499189 · inbound

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark cites this paper.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.500206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.500206Z digest=sha256:daf0710aab799586a0b8b48d5dc3f94bea42dfb6ed3c4b13d0f740ee24ab6b2f

Observation 257a7895-bdaa-4beb-a86b-a8c099750e61 · inbound

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL cites this paper.

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:15:46.371485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T15:15:46.255296Z digest=sha256:d1af31ca38ed23f65f7345eb09b508b27b60393726e17d7cdb24cf24713ff5c4

Observation c92d0997-0214-4ac9-a7fc-d42ffd2e7a7c · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

Reference 128

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:42.000176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:04856ce43ad732817577b0f282f45d2154a092bfd6f15526a306107aed8d3907

Observation da70c93b-f4ae-4780-8e09-a7717b5f7a75 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.293580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:1490b0b3588b5546b11417287357d754dd7721645a6e416ddea2ac4cdb5473a4