Pith. sign in

Paper Citation Record · LEDGER

Multimodal Chain-of-Thought Reasoning in Language Models

As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 100 inbound Pith citation observations for arXiv:2302.00923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.00923 v5

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T18:12:27.396607Z

measured 149 of 149 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 100 of 158 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:10.670667Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact19
  • verified fuzzy17
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

98
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7ad35ea1-2a60-458b-b4ef-201ffec9a45b · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Multimodal Chain-of-Thought Reasoning in Language Models Bottom-up and top-down attention for image captioning and visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.663644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:8d7b0913b7dbba8ff46dd94cb91f316e4f00edb152df39aac31576e90b102cee

Observation f8ba0729-c688-4567-a7bc-d91945d5185d · outbound

This paper cites Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al.

Multimodal Chain-of-Thought Reasoning in Language Models Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.452441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:7fd891dac8deaf41870fcf6ab529a0b2c14fa7dadeef61c9b3d0ca01fb1148f5

Observation e4e8e4e3-9895-415e-b56d-399c80cac386 · outbound

This paper cites an unresolved cited work.

Multimodal Chain-of-Thought Reasoning in Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-12T18:12:27.578573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:58e4412f730c625dd871e280c10394328d92fbf77778f50d221fca4c5f2f91e5

Observation bf3623b3-4212-4668-aad6-c56d7d99d596 · outbound

This paper cites End-to-end object detection with transformers.

Multimodal Chain-of-Thought Reasoning in Language Models End-to-end object detection with transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.583463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:8d3e4a39a11a97683b29b9fcb23a544352e3368032d57b0e97d175431f3b8d9b

Observation 821d6c3a-3805-4fef-8cd1-693e7d52fd03 · outbound

This paper cites an unresolved cited work.

Multimodal Chain-of-Thought Reasoning in Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-12T18:12:27.587969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:5b530acbc0a460f6164f7c45dd1d3a643eec60c17adcdd21dab0ef71a3c50585

Observation 1527499e-f4fe-4435-9f89-748255a9274c · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Multimodal Chain-of-Thought Reasoning in Language Models Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:12:27.499329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:50f1097790c8fac9754f85ed1f0f9d7b12545c13da1bec86f0d1c08b03ff9d1b

Observation f804bb78-8fa3-4b9b-918e-eff6e51e1459 · outbound

This paper cites See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning.

Multimodal Chain-of-Thought Reasoning in Language Models See, Think, Confirm: Interactive Prompting Between Vision and Language Models for Knowledge-based Visual Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.534776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:8d409f1ba5b4ee2b967e18e72fcf3b90edf21faa69d135c8a34a55da17f4bb16

Observation 9af814b1-823f-48f3-8120-6ede79c5558f · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Multimodal Chain-of-Thought Reasoning in Language Models PaLM: Scaling Language Modeling with Pathways

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:12:27.549367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:352684b1593fc2121c81bdecf8ea4557c46aa99d7e2e21c1f9bedab2dd966e4e

Observation 2c0d5a16-c27b-4b8b-9306-1d1badee10fd · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Multimodal Chain-of-Thought Reasoning in Language Models Scaling Instruction-Finetuned Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:12:27.494483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:3df6d9592a3c20dba6b097b4dcc2a10fbf9166229c0a9ce0fc423cef98050905

Observation 63675545-d366-4a2b-b33f-9d263b3ea93c · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Multimodal Chain-of-Thought Reasoning in Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.592340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:8934da6df10dcf972149c81caccfa78fba81914835b8448ad117b4879deaec9b

Observation f721fb64-450b-433f-a967-8dbdcee5b4d6 · outbound

This paper cites an unresolved cited work.

Multimodal Chain-of-Thought Reasoning in Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-12T18:12:27.596429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:8fe27a9efd04d69a2f953f75df6f1b994bad0718a20c82ca50ff569ea9516730

Observation 57f5e2e1-7ed5-46d9-bc64-78f66b55bdae · outbound

This paper cites Complexity-Based Prompting for Multi-Step Reasoning.

Multimodal Chain-of-Thought Reasoning in Language Models Complexity-Based Prompting for Multi-Step Reasoning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:12:27.528925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:9dfb891f44922fee31f6f3d5638bcecae31a4962053e2c7e379d22e95f8bffbd

Observation 7c749774-e285-4b05-a420-07d68890ab93 · outbound

This paper cites an unresolved cited work.

Multimodal Chain-of-Thought Reasoning in Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-12T18:12:27.600330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:5117c30162a9a950ac5873525dca04535b497d4053d9804906980f7c6e5ac469

Observation bf5eb678-d7db-4c56-892c-3a45cce3149c · outbound

This paper cites Yaru Hao, Haoyu Song, Li Dong, Shaohan Huang, Zewen Chi, Wenhui Wang, Shuming Ma, and Furu Wei.

Multimodal Chain-of-Thought Reasoning in Language Models Yaru Hao, Haoyu Song, Li Dong, Shaohan Huang, Zewen Chi, Wenhui Wang, Shuming Ma, and Furu Wei

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.437051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:0b70dc835960b5d738143fb90dcd03674c1775380ff8a9ef478c53cca5b83149

Observation b9404238-ea89-4c8a-a7a9-fc0f89a477d4 · outbound

This paper cites Deep residual learning for image recognition.

Multimodal Chain-of-Thought Reasoning in Language Models Deep residual learning for image recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.604691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:48edb0c546fabc93da02b375cfcbe74d43cbd5287f3ac5f4a0512303084f0cee

Observation f6b26b5f-fd45-4736-93ee-998c0b7fd531 · outbound

This paper cites Deep residual learning for image recognition.

Multimodal Chain-of-Thought Reasoning in Language Models Deep residual learning for image recognition

Reference 16

Resolution
metadata mismatch
doi, observed 2026-05-12T18:12:27.471413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:d63fb39d2f8a3c86667ca3a84d075ec228526d1bb3357b82db4ee6193fe19a90

Observation be82e526-8f40-47a9-b560-0438bdaef110 · outbound

This paper cites Towards Reasoning in Large Language Models: A Survey.

Multimodal Chain-of-Thought Reasoning in Language Models Towards Reasoning in Large Language Models: A Survey

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:22:30.236738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:1deb771cea941bddd6432e30cd5b746c9e8b5dcf8bf10099d107cd6b067f37e0

Observation 81b9afc3-05d3-4bed-ad07-0dffd1416a68 · outbound

This paper cites UNIFIEDQA: Crossing format boundaries with a single QA system.

Multimodal Chain-of-Thought Reasoning in Language Models UNIFIEDQA: Crossing format boundaries with a single QA system

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.608549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:304a4fe9f352bf72a5abc96a4a3961fe51828e675beaba5c9482decc2570dbe7

Observation 10537ba7-d59b-4eaa-b875-d2f3a9c9f277 · outbound

This paper cites In Findings of the Association for Com- putational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.).

Multimodal Chain-of-Thought Reasoning in Language Models In Findings of the Association for Com- putational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.)

Reference 19

Resolution
metadata mismatch
doi, observed 2026-05-12T18:12:27.446265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:f259bf37dfec0f15fb9bfe59e56e478451885407aa31f1400bdda7c37ad518c2

Observation b2e7e6eb-0e9c-45d5-a9cb-79dd1201040f · outbound

This paper cites Bilinear attention networks.

Multimodal Chain-of-Thought Reasoning in Language Models Bilinear attention networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.612252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:20303f0ab2c46d7f83f26ca2b8baa8a815240fd95681a0e59e988f7fb8f811b1

Observation af4c2ac4-1ded-41fb-b907-1ea5fa78eee4 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision.

Multimodal Chain-of-Thought Reasoning in Language Models Vilt: Vision-and-language transformer without convolution or region supervision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.616303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:93c27b4cbd79b6a2d7eacc1e0bea4fd5a7357fc5b40b4e55b6b29ae36925a135

Observation 469e3a72-906e-4e29-aeb6-58ae000298d2 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

Multimodal Chain-of-Thought Reasoning in Language Models Large Language Models are Zero-Shot Reasoners

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:04:13.904485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:61349ca0d7514fdc985579c1ad23034bb4c25d5c1d5e4730a23da098839148bb

Observation 9e995384-b5ff-45ab-971a-788f8649b6b8 · outbound

This paper cites On vision features in multimodal machine translation.

Multimodal Chain-of-Thought Reasoning in Language Models On vision features in multimodal machine translation

Reference 23

Resolution
verified exact
doi, observed 2026-05-12T18:12:27.466782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:4c02fad7f9dec3a81a298af49903b0fda10f86e04bd39a76120f80d516994f2a

Observation 0fb9d7f0-8e94-4370-a83c-3715e78ae4a6 · outbound

This paper cites Making Large Language Models Better Reasoners with Step-Aware Verifier.

Multimodal Chain-of-Thought Reasoning in Language Models Making Large Language Models Better Reasoners with Step-Aware Verifier

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:12:27.511336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:d15a8be2cf77448f95f1dea67a4c70dbd669a81d1f38383c0e72e0f9a79fe2ea

Observation 78cad326-9ff8-4673-99e7-855a21908352 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

Multimodal Chain-of-Thought Reasoning in Language Models Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:12:27.523797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:6335acc7cb1c3e19d257bf0a7147102e5845d7c8b6621506c1b1884e9dff94f8

Observation 656bbdf0-841d-4e84-98e4-efd34d95cc44 · outbound

This paper cites Teaching Small Language Models to Reason.

Multimodal Chain-of-Thought Reasoning in Language Models Teaching Small Language Models to Reason

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:12:27.517396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:129ce317ab9406c2ac040d049a91166ea751329a2ca756ffbf58f98b6e65a3d3

Observation a19888a4-dac2-4fbd-baa1-90ca8d1c3b4f · outbound

This paper cites Gpt-4v(ision) system card.

Multimodal Chain-of-Thought Reasoning in Language Models Gpt-4v(ision) system card

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.619807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:5ec792d3dffaba992dca02d0dacd5928ccd1e6b2a2a35dd2ebf3297f21003b58

Observation 4ef36915-1e9e-4c8a-afe9-9b6ffccbdbb3 · outbound

This paper cites an unresolved cited work.

Multimodal Chain-of-Thought Reasoning in Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-12T18:12:27.623055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:682765f272beb170c37549b9871b809fa5b8df122f01d87ed91cd25dd8732e51

Observation ee7709c8-bc65-4e82-a76c-a6167080d958 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Multimodal Chain-of-Thought Reasoning in Language Models Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:12:27.539613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:8dd86c81d2ec9351451bbf2421eec9ac4e8d2d85cb0d8dd228771a0fb6854354

Observation cda73128-84cc-44a0-a79f-d665ea0a59e4 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Multimodal Chain-of-Thought Reasoning in Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:12:27.544680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:8bff1c9d767ea2dc14e16c512b66ef1f00f069e0d717a9b83a8232596b96186a

Observation e0393ae7-7a86-4e00-8d6f-cbad4c222670 · outbound

This paper cites Learning to retrieve prompts for in-context learning.

Multimodal Chain-of-Thought Reasoning in Language Models Learning to retrieve prompts for in-context learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.626501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:a00ba6ce43a079719c87835006ecf9fecfd8c92705067ea24c7c04f61ee6868b

Observation b00368f9-1202-4afe-97c0-d849b4eff5f2 · outbound

This paper cites URL https://aclanthology.org/2022.naacl-main.191.

Multimodal Chain-of-Thought Reasoning in Language Models URL https://aclanthology.org/2022.naacl-main.191

Reference 32

Resolution
verified exact
doi, observed 2026-05-12T18:12:27.457467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:6da7a38e80505aa7ecd13d11c15aeba94f789c34a0ad7fd778c59d8d7d67cdae

Observation aeff0bf7-c854-42ec-9812-2e265e51ffe8 · outbound

This paper cites Alpaca: A strong, replicable instruction-following model.Stanford Center for Research on Foundation Models.

Multimodal Chain-of-Thought Reasoning in Language Models Alpaca: A strong, replicable instruction-following model.Stanford Center for Research on Foundation Models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.630240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:e3cf42e17a18471b8e21dd4f0e0011d6aa5b57e14ee7add5f4abd28b8b5635bd

Observation 273863e7-76eb-4d25-aac4-9bb6dbc8f215 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Multimodal Chain-of-Thought Reasoning in Language Models LaMDA: Language Models for Dialog Applications

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:12:27.555158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:939ef6f25dcab65732e9181625a01f27033ef91d329ecab23c81199187267081

Observation 71bcf958-0f7f-439a-b194-fcb8e5285cb6 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Multimodal Chain-of-Thought Reasoning in Language Models Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.633657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:275b5c17cc19cf7c7b8c59e9e4b381212f8fb68ebfc58224eced5b40cdfbd9c7

Observation af839bf3-d430-4853-9a5c-36811a6a1d97 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Multimodal Chain-of-Thought Reasoning in Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:12:27.560610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:6f394be5ee6d7848777243df7fb53f87baff557558d43f904bbecca502a4e43c

Observation cae2e0b5-7529-438c-9b59-aaaf43daedd2 · outbound

This paper cites doi: 10.18653/v1/2021.acl-long.480.

Multimodal Chain-of-Thought Reasoning in Language Models doi: 10.18653/v1/2021.acl-long.480

Reference 37

Resolution
verified exact
doi, observed 2026-05-12T18:12:27.441860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:4960473d05bf3c77080bfe2b1a60bed45c75e9cf9371ddbb6f12c259816bc043

Observation d7d72da8-8c0d-471a-9b48-539049a53ba9 · outbound

This paper cites Deep modular co-attention networks for visual question answering.

Multimodal Chain-of-Thought Reasoning in Language Models Deep modular co-attention networks for visual question answering

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.637169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:0409fbe21f8eaa65221de6c68b1ee927bd333aaf38b2fc9fc87559089f168fee

Observation 396c4dd5-6a7c-4251-84da-6fe672dfbd37 · outbound

This paper cites an unresolved cited work.

Multimodal Chain-of-Thought Reasoning in Language Models Unresolved cited work

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.463247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:a1e809b0be1504b9450b3c036f826d80a6f578b57eb45dd326f77417098d1163

Observation 80262818-848b-4c79-82ef-eaeef4822ba2 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

Multimodal Chain-of-Thought Reasoning in Language Models LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.485002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:e356a7684361b12b82e5ff5b7629232adad29a7c85a47a9bfd0b8a13f6761db4

Observation 1a39ba30-2706-4a79-bdf0-a3b42487b1f1 · outbound

This paper cites Universal multimodal representation for language understanding.IEEE Transactions on Pattern Analysis and Machine Intelligence, pp.

Multimodal Chain-of-Thought Reasoning in Language Models Universal multimodal representation for language understanding.IEEE Transactions on Pattern Analysis and Machine Intelligence, pp

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.478103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:deb15ea4271e2dbd154f20310a90f3a7d4250959523cfd4a70ca6c0b8a0857ff

Observation 21b84276-9bbd-43bf-a97c-fbbef837aa08 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Multimodal Chain-of-Thought Reasoning in Language Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:12:27.489391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:43e6fd7c30d5cdf12b3cd92ae7201c9b5268051afaccc82e0c10ee9ad860dc7c

Observation 90430afc-f885-4437-9d04-036b3cfa34bb · outbound

This paper cites Here, we present additional examples to illustrate this phenomenon, as depicted in Figure.

Multimodal Chain-of-Thought Reasoning in Language Models Here, we present additional examples to illustrate this phenomenon, as depicted in Figure

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.641184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:2789e08532c5f4b8302282d0a7b88ed223cb51c715080ce71e77cd910f2d24ba

Observation ca4ea2f9-54b9-4cff-a55e-aaccb5841a74 · outbound

This paper cites an unresolved cited work.

Multimodal Chain-of-Thought Reasoning in Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-12T18:12:27.644559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:be2efc72fef664779f270087aa90a07d182ff86092bea9695f2452353fbbfa5f

Observation 7ee6d94f-c59a-42ad-aa4d-28f4b8023501 · outbound

This paper cites •ScienceQA is a large-scale multimodal science question dataset with annotated lectures and explanations.

Multimodal Chain-of-Thought Reasoning in Language Models •ScienceQA is a large-scale multimodal science question dataset with annotated lectures and explanations

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.648925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:d51bc6e0d30c08d9baa82e81111e9eeabf9fc5abfffc8d5fc07b99ccab1efbbf

Observation 2b33669f-bd67-4d6d-9e45-37dc5d96a707 · outbound

This paper cites The vision features are obtained by the frozen ViT-large encoder (Dosovitskiy et al., 2021b).

Multimodal Chain-of-Thought Reasoning in Language Models The vision features are obtained by the frozen ViT-large encoder (Dosovitskiy et al., 2021b)

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.652669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:befb1532394d6b07ed5ade1f14fe47c99026bc36cfee727012ab0779f15a2341

Observation 44f98b01-008f-43c5-8e58-15f468486cd6 · outbound

This paper cites an unresolved cited work.

Multimodal Chain-of-Thought Reasoning in Language Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-12T18:12:27.655991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:a271673ed13ecc942c16779f805c6c4aa3ab470424e3f6a9e42a3c835fd79a48

Observation 469e8305-0116-4b5b-9414-e3f70d359b4c · outbound

This paper cites Then, we use the generated pseudo-rationales as the target rationales for training instead of relying on the human annotation of reasoning chains.

Multimodal Chain-of-Thought Reasoning in Language Models Then, we use the generated pseudo-rationales as the target rationales for training instead of relying on the human annotation of reasoning chains

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.659950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:ff6d1287b0e4b1080fc5a5a9ffbf5e9aec66c71e81e231738f61f5f31ee2d70a

Observation 80dbcc96-326f-47e6-8a24-73ef138993b1 · outbound

This paper cites 80% 14%6% CommonsenseLogicalOthers Figure 11: Categorization analysis.

Multimodal Chain-of-Thought Reasoning in Language Models 80% 14%6% CommonsenseLogicalOthers Figure 11: Categorization analysis

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:12:27.570799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:484a42377e8a5be76aecd377330d7596bc223d63ebd830d8dfea7543a214a942

Pith citing papers

Observation 7c9d7d79-955e-4dd1-b16d-b5ff256b9116 · inbound

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models cites this paper.

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:50:24.184228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T22:50:24.053411Z digest=sha256:884b445c9f970c898bfc22c65c4d899d5a25f45fc0f4011ced4c068795e3f699

Observation d2d319d5-695f-4e9c-8371-74f8fd7000c0 · inbound

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action cites this paper.

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action Multimodal Chain-of-Thought Reasoning in Language Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:17:58.849157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T01:17:58.678036Z digest=sha256:5cd327ba12be8f1c8519825b1c4d789932ee82057a8a85f3190ae61a085c4cce

Observation 1d345092-eea8-491e-9045-de78406b0d1c · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention Multimodal Chain-of-Thought Reasoning in Language Models

Reference 147

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T23:07:42.533156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:34d893ffd094ddf33e81fa5f08e9121c4f9ac83dd760d3492b5273578a5e076f

Observation 19f22747-9d6b-4414-8ef6-69cae187b6a5 · inbound

CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society cites this paper.

CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society Multimodal Chain-of-Thought Reasoning in Language Models

Reference 133

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:40:53.989447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T01:40:53.351795Z digest=sha256:9dc73cb22eca8edbab70d3b354c4c4e513dc15965ba3e9d6cc1523e9318c3244

Observation 5256cbde-1982-4314-bc75-36f929a43ee9 · inbound

Visual Instruction Tuning cites this paper.

Visual Instruction Tuning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T08:22:03.403362Z digest=sha256:83228bdc2b2b1809e0c32c5a9eb4eafc85aa2debd005175e1c07889cd024d4d1

Observation f8f02aaa-3ebb-4513-9d7e-91c6c0f7cc9a · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 188

Resolution
verified exact
local_arxiv, observed 2026-05-16T02:56:42.161866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:f5937906eeeaf4ab20a4b09eea744d7ff14113778f4dcb5392224a98be1cb87f

Observation 083c4f53-0000-4c38-a929-18411a8a6b66 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 287

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:28:39.267291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:31986b7d8c5ae32a67f66586ec1720adde25feb5c6c42ef00cdc659526830b69

Observation a9abddba-d78d-4bfb-850b-d7883baf9c53 · inbound

Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment cites this paper.

Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment Multimodal Chain-of-Thought Reasoning in Language Models

Reference 139

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:30:44.981950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T22:30:44.520703Z digest=sha256:67158bcbe42a7a18972930ff575ef10ae39ef8226ecacde95756f96db3e1f846

Observation dc656a40-a271-46ae-9382-b2922349e537 · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Multimodal Chain-of-Thought Reasoning in Language Models

Reference 151

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:26:06.516043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:d93f5f48f280b146c655518a491a3a0543410dcde054b8d81f3dfea832b24417

Observation 8e663a4c-89fb-439e-b5ba-085ab1d1deba · inbound

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction cites this paper.

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction Multimodal Chain-of-Thought Reasoning in Language Models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:53:33.167426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T22:51:08.753650Z digest=sha256:c5501b8fa8ca0fac1c8bedf131344b682dad74f93e101b2adbc2d53f0acdaaf4

Observation c8340945-5d4a-419d-8176-910feb651aff · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Multimodal Chain-of-Thought Reasoning in Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:18:53.639892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:2764f56ae98e22534a5bc8f9403c931f50f1fb36f477308f983253bb0062cad9

Observation f9b98dac-ae3d-4953-9575-283b2a442322 · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles Multimodal Chain-of-Thought Reasoning in Language Models

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:59:03.343413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:fb71000dafd290d551d6c7325cc1f0ad09a49d88363ac3d4897ce3664263024f

Observation 4969705c-1c80-4067-9738-0eda3b0caf0a · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Multimodal Chain-of-Thought Reasoning in Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:5edc856c7f68221661ab48fdabe8fd217cc2360ae5a0c60fa545a8c6f6f8f8e6

Observation 1637b3ad-0836-4ee4-a4e2-097fff7616fc · inbound

GRIT: Teaching MLLMs to Think with Images cites this paper.

GRIT: Teaching MLLMs to Think with Images Multimodal Chain-of-Thought Reasoning in Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:31:36.128814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:29:47.529564Z digest=sha256:16fb34432a0faea25827d0fda4eed69d9e95a85179d7512ddc2e27c48f179dba

Observation 07e52310-e195-4dca-9353-8f18fd9fbb80 · inbound

Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective cites this paper.

Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective Multimodal Chain-of-Thought Reasoning in Language Models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T13:53:05.648683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T13:52:44.096262Z digest=sha256:72ccd32c7273a45dbc4cbe87c83b600360dfc4cdefc8ffabfcaa8fcbc5731127

Observation a3f638c1-0a06-4065-af59-2a2ccfebf95b · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:05:52.179893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:2081d263de15d6249d94a13f1747388194cb5c844ac8d2600d0a21c108b9083a

Observation 8a1b897e-7d41-43f7-8b74-ee6c45709c95 · inbound

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering cites this paper.

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering Multimodal Chain-of-Thought Reasoning in Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:10.670667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:10.670667Z digest=sha256:be79ec417350a6442577dc8cd4b393c90070e40b6dd0e7c66206a5221e55869a

Observation 5e94eb29-f3b2-481f-9344-cd4227812f63 · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Multimodal Chain-of-Thought Reasoning in Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.744601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.744601Z digest=sha256:7e1316ac377bb1f611e64893035ad08648fcbb9ecaa9bac1e2a8e89638a122ff

Observation 48a225fb-5b1d-4cf3-8168-4a210622fcce · inbound

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark cites this paper.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Multimodal Chain-of-Thought Reasoning in Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.790349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.790349Z digest=sha256:ab185966a306fea3a3e1e8adefad4e663d786de2c555e5978787c00932374424

Observation 5ebbdf54-4206-4f77-9901-8ee38e4e3326 · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems Multimodal Chain-of-Thought Reasoning in Language Models

Reference 227

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:52:16.454085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:d20b7a0c0dba1339726ec0ad5def3c4b8788b8b8a3c3ee9c211e286b225b0407

Observation 36d4ec69-ae2b-41ae-99ab-a8253b6602a5 · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:54.760061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:54.760061Z digest=sha256:7597412fc85e1cfbeb3450068ade39c96c7627440295af49b949bdcb18c865c3

Observation 6c408dae-9419-45a2-9711-7791ca903d3b · inbound

Graph-of-Causal Evolution: Challenging Chain-of-Model for Reasoning cites this paper.

Graph-of-Causal Evolution: Challenging Chain-of-Model for Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:02.608167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:37:02.608167Z digest=sha256:3e2ad2d55727a9204a9a5b20ee103910429f3125ae6de818b8aedddeba785446

Observation 373b2e0a-d748-48a8-967e-69b399de2ec1 · inbound

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning cites this paper.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.050109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.050109Z digest=sha256:1eb3247a8558f70dc699b05c5f01e1af94276d53a9b55ad89dc8e41ee56f12a2

Observation 133a0312-fd14-48e7-a0bb-71019c467bc8 · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism Multimodal Chain-of-Thought Reasoning in Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:23.620651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:23.620651Z digest=sha256:7c1125ef63a772547644369fe8e2edb2a51702dba9b59cba1d979d350fbd930d

Observation 13dd506a-9ce8-4054-af5f-43844a6ab1c3 · inbound

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought cites this paper.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Multimodal Chain-of-Thought Reasoning in Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.403000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.403000Z digest=sha256:0f6e8cb9d593c57cf6db0d1ceaf6e143bce438a51ab9cb8c526fbc2983394252

Observation 52a76e6f-afc4-4d0e-84de-11b6b2635d1d · inbound

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 cites this paper.

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 Multimodal Chain-of-Thought Reasoning in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:13.795355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:13.795355Z digest=sha256:c0bc2b1a7edd941fb0df2193cbad27a2fcd08d0f2db75fb2f476e12c8e5d4d73

Observation 2dca750e-d06e-4844-9e73-de662e458a8d · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond Multimodal Chain-of-Thought Reasoning in Language Models

Reference 253

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:37.239483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:37.239483Z digest=sha256:844185fc3e404d2226c944bd4e2d95fd07f988fb39d3b39dbcad5b55582241f5

Observation e541fac0-dd6d-4160-996c-78d7825e2d88 · inbound

Breaking Thought Patterns: A Multi-Dimensional Reasoning Framework for LLMs cites this paper.

Breaking Thought Patterns: A Multi-Dimensional Reasoning Framework for LLMs Multimodal Chain-of-Thought Reasoning in Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:39:52.632806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:39:52.632806Z digest=sha256:124742fc7ba4e3bcdce9e1021048548a757d86d2c0e248ca62c5a7cfafc26aea

Observation 674066e8-842b-4f03-b0c8-b0e1c546067d · inbound

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? cites this paper.

Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes? Multimodal Chain-of-Thought Reasoning in Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:19:05.501316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:19:05.501316Z digest=sha256:c503f73b729eb1df90551503b13be7f7cd58844c1f0e39c82caa48890c368898

Observation 3dcc9428-0220-4ea6-a573-9c642f348207 · inbound

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning cites this paper.

DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:01.775920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:09:01.775920Z digest=sha256:5ab47a48aaa91d64086f8375fa463795765e98956be75bc070197df4917b1c64

Observation cef7b16d-aaeb-4dea-a7b8-18deed2c501b · inbound

RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought cites this paper.

RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought Multimodal Chain-of-Thought Reasoning in Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:33:02.089330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T08:32:20.566798Z digest=sha256:3d9d445dc01277a550e5e6caba38748f409617d15f42d2bbdae87e5ab6e864b8

Observation df92342c-82be-4b17-ab21-c721e7ac7c18 · inbound

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs cites this paper.

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs Multimodal Chain-of-Thought Reasoning in Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.364377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:27:19.364377Z digest=sha256:c1f2d2b5504a8f9381b7a5fac18d29d782cccd62479713364ebf25e43198fb77

Observation dfcbac5d-2cfb-4fe5-93cb-f41d8aa0b5e6 · inbound

ECCoT: A Framework for Enhancing Effective Cognition via Chain of Thought in Large Language Model cites this paper.

ECCoT: A Framework for Enhancing Effective Cognition via Chain of Thought in Large Language Model Multimodal Chain-of-Thought Reasoning in Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:13.597717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:13.597717Z digest=sha256:50b08ce7f396c32bab161f63256c561c172aef1781b31922d68cb5a389c2b772

Observation 90270684-ff5a-4647-bdab-e290706fab79 · inbound

V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis cites this paper.

V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis Multimodal Chain-of-Thought Reasoning in Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:47.453315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:47.453315Z digest=sha256:62c28b53492cb3b33706faeea27c461bd2a73224d2646e087311651f6c3a289e

Observation 3b8029ff-bebd-4799-878e-2bc272449bc2 · inbound

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling cites this paper.

Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling Multimodal Chain-of-Thought Reasoning in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.620896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.620896Z digest=sha256:b11433ec19392c599c361dc602f958c9f04c3b227e1b964b84c191d727b1ec2b

Observation 91da92d8-23a0-41e2-947d-7e120d2fdbd1 · inbound

Multimodal Mathematical Reasoning with Diverse Solving Perspective cites this paper.

Multimodal Mathematical Reasoning with Diverse Solving Perspective Multimodal Chain-of-Thought Reasoning in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:20.890778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:20.890778Z digest=sha256:98cbff78b9755d762fd85f5ca8e50b5571082bdf2c1555ec37c4bd6c7ddd4a3a

Observation d64192ec-5a48-479b-98c3-4c1ae57b6b90 · inbound

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought cites this paper.

From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought Multimodal Chain-of-Thought Reasoning in Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:35.695869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:35.695869Z digest=sha256:0b1def0d117931fee39a4fc03c5672c419ec52c009a1aad7980a30315eb51155

Observation 8c8a3027-fe41-4a9e-93a3-893590608b0a · inbound

Introspection of Thought Helps AI Agents cites this paper.

Introspection of Thought Helps AI Agents Multimodal Chain-of-Thought Reasoning in Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.893076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.893076Z digest=sha256:79e974140a17e2e8e92d4fcad7c5a02e257be2d4de50457a3da2c25f2be765eb

Observation e766ee2a-fc74-468b-beb6-2c0f3d296603 · inbound

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation cites this paper.

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation Multimodal Chain-of-Thought Reasoning in Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:42.893099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:42.893099Z digest=sha256:4c2561cbb29c22d18a5f3b0bf415b8551eb7fe4d214a1d3bb494184321e9b8f8

Observation d653de49-8eae-4efb-a3c9-30ff8af79890 · inbound

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback cites this paper.

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback Multimodal Chain-of-Thought Reasoning in Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:35.914280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:23:35.914280Z digest=sha256:8fc2e788c143c9847e1077243dbd09030506200cfabcd44c78ee057925f9c9b8

Observation 353e4f03-3de5-4d82-bf48-55b87652f79b · inbound

Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems cites this paper.

Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems Multimodal Chain-of-Thought Reasoning in Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T12:39:42.154856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:39:42.154856Z digest=sha256:a76c5d900a31ad4792b21e6257f387a2aae9363ccf20b34a06527b04d2e38882

Observation d760e15b-049d-4cd4-8727-49c584108094 · inbound

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning cites this paper.

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T23:02:52.657037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T23:02:14.913189Z digest=sha256:d4d2a2ef4c1c9557d97f00c56d80d683add5a3716884bb91e274b71d06182a4a

Observation 7c7ae754-a36f-4a23-8fd3-bcf7d831a020 · inbound

Instant Preference Alignment for Text-to-Image Diffusion Models cites this paper.

Instant Preference Alignment for Text-to-Image Diffusion Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T16:51:04.957340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:51:04.957340Z digest=sha256:0b64e4cd220d874a4ea892ece84c900b1b332dcd61687fb37de85739b75ac8ff

Observation 49e17bd7-3278-408f-ac68-8555d4081583 · inbound

Beyond the Textual: Generating Coherent Visual Options for MCQs cites this paper.

Beyond the Textual: Generating Coherent Visual Options for MCQs Multimodal Chain-of-Thought Reasoning in Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:37.994242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:17:37.994242Z digest=sha256:cd1c0e96c693c46a562c02ccef2bd1b82a5cecf0608368af1857469cfd1db84d

Observation f12eb1c8-495a-418e-b7ac-a83a24d06c18 · inbound

PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation cites this paper.

PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation Multimodal Chain-of-Thought Reasoning in Language Models

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T20:22:50.786017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T20:22:03.697564Z digest=sha256:5258b0503329c1a4362f5a8695390cad7bf3a48417a06cfed4cda90a75bdf11b

Observation bed337f0-7045-4cf1-812a-7404b80d7d3a · inbound

PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation cites this paper.

PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation Multimodal Chain-of-Thought Reasoning in Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:02:50.738468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:02:50.738468Z digest=sha256:fad752bd12341962a36cc71ac63f44a119be4bafaca05289a487a4811b42d325

Observation b3b56c74-004b-42ff-9712-d07230ee392a · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Multimodal Chain-of-Thought Reasoning in Language Models

Reference 221

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:21:48.502694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:c5944fcdd8aff7f09b78655b5ed3a2b483bec8f254a83e523e214dd54441b7fd

Observation 404ae37e-9036-44d2-bbbf-d0cd38a23689 · inbound

TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering cites this paper.

TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering Multimodal Chain-of-Thought Reasoning in Language Models

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-18T19:02:48.699721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:02:01.962726Z digest=sha256:60c9dfdcba466553970acf0ca44a789df3d641f0a1ad38684e52ec12df174fb4

Observation 08d127c0-8ffe-468d-a1e3-8b43d1295ee4 · inbound

VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples cites this paper.

VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples Multimodal Chain-of-Thought Reasoning in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T12:06:04.601125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:06:04.601125Z digest=sha256:4448cf6d064ef57f4815fc89c3e0c521d1d49a9ad5e39f7e833d30935b504cb4

Observation b5d984ed-8311-487e-89d2-a0f09c327f24 · inbound

Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation cites this paper.

Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation Multimodal Chain-of-Thought Reasoning in Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T05:34:40.404905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:34:40.404905Z digest=sha256:78ecd191d77b2329a25ef9123d25c27a29da652824122fbd6ec1979a0a19dbd6

Observation fd290b6e-04a7-40ac-837a-f2d139744a8c · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Multimodal Chain-of-Thought Reasoning in Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:30.357545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:30.357545Z digest=sha256:d9809571d248406d904edc9e67eeee41dfa6d76702dcbdb3aa954953d1b08b0f

Observation b99f0ba0-d221-4aeb-8735-b58887b3660c · inbound

Visual Programmability: A Guide for Code-as-Thought in Chart Understanding cites this paper.

Visual Programmability: A Guide for Code-as-Thought in Chart Understanding Multimodal Chain-of-Thought Reasoning in Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:21.029581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:21.029581Z digest=sha256:b460ef65d86109cd5841efbba1180404ccc4a1972ff5ea0b9b2bde780a0ef71a

Observation 10edc72e-370e-48c3-a010-ed41b3df0c99 · inbound

LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA cites this paper.

LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA Multimodal Chain-of-Thought Reasoning in Language Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-18T17:56:42.215334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:53:08.136677Z digest=sha256:3dc21688bb2746628d5d61163e7b761dbf2db1a2c8bfb17caf0f69b93ffdb5ce

Observation f569ab2b-1434-4ef7-a2e8-4cb778f39a6e · inbound

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework cites this paper.

Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework Multimodal Chain-of-Thought Reasoning in Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:32:36.326769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T12:31:25.257879Z digest=sha256:9478a000bd181ca269129540393355cece0516c5024ceedaa783d3df4babf3ab

Observation acacc8a5-596d-4c7b-9904-daf4fc61de55 · inbound

Latent Visual Reasoning cites this paper.

Latent Visual Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:41:30.484640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:41:30.307521Z digest=sha256:1955a8833b5a80ab7c0ce87926b85fc11755e7b916b2da79907bb58002cb446a

Observation a02e3b88-324c-4d34-a2a8-06494d1fc7f0 · inbound

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning cites this paper.

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:31:24.813779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:30:10.620448Z digest=sha256:9afeab77cd032d80cadde58fab0b6da9ce485758c71a26ba68261c7aac860f47

Observation 966e5e11-9439-4d79-9afa-1a7c4ef1f8f2 · inbound

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning cites this paper.

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:00.973423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:00.973423Z digest=sha256:c3a9e7e37b858e9aa930cf24a15b0061f519aacac729888774d1aed26eb1c169

Observation a1146530-2b00-4072-87af-e1c387abf8e9 · inbound

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials cites this paper.

AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials Multimodal Chain-of-Thought Reasoning in Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T11:26:06.881567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:26:06.881567Z digest=sha256:b05e488190ba11b155b82b0a5cdabbede5eba14051f79ac7cb06f53315e06758

Observation fbb8b3e7-44e2-463f-b63b-96bff2c815aa · inbound

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation cites this paper.

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation Multimodal Chain-of-Thought Reasoning in Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:15:33.896676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T08:13:05.328746Z digest=sha256:600adf8c5594d45cd97e4e5e517281eb4f476ce7f86842f6722328aec00ce644

Observation 7fa3ef87-aebe-42b7-9771-3b3e10d9fd4d · inbound

Logic-Guided Socially-aware Robot Navigation World Model cites this paper.

Logic-Guided Socially-aware Robot Navigation World Model Multimodal Chain-of-Thought Reasoning in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:54:08.862155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:54:08.862155Z digest=sha256:bf7c217e568290acd905f527af9dcc82be73f545190bf98dfed0d2487e92b193

Observation 95dbf2cf-d48c-4ec7-bdbc-ebde4de5fcc8 · inbound

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm cites this paper.

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Multimodal Chain-of-Thought Reasoning in Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:55:35.066474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T00:54:24.641649Z digest=sha256:5005946a043623a875bc4d2dda52e9e2a6a25f6e4e482973cfe2dc274eda5c0f

Observation 896da98d-637d-49c1-963f-0bdd2193e065 · inbound

See, Think, Learn: A Self-Taught Multimodal Reasoner cites this paper.

See, Think, Learn: A Self-Taught Multimodal Reasoner Multimodal Chain-of-Thought Reasoning in Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T19:04:29.050993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:04:29.050993Z digest=sha256:8187ba6341a3e4e88801bab209fd664625f34f14695ddf7d47d092b872997863

Observation f9d79ef7-388b-474a-8279-1dbab18537fe · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Multimodal Chain-of-Thought Reasoning in Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:11:26.653988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:95afb048223352e5bea8127a0147a4b27bc7484546d30ec58a655c324c104d0b

Observation 28a8e559-bec1-457c-8adf-692dc8754528 · inbound

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents cites this paper.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Multimodal Chain-of-Thought Reasoning in Language Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:11:29.850303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T03:09:31.161760Z digest=sha256:7ac59fe111c30b8690581d066f8d79c84cad44362c2032366e43d73c9346a3c5

Observation ad87133e-003b-4ec7-921c-3fec66bddc16 · inbound

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents cites this paper.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Multimodal Chain-of-Thought Reasoning in Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:49.421484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:49.421484Z digest=sha256:a76e81a4975925bd9f09447afa02c6fd4f1566ffeb632af963f4821e06ffac87

Observation 52cad193-c475-4773-909f-2c9223e7d384 · inbound

Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space cites this paper.

Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space Multimodal Chain-of-Thought Reasoning in Language Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:03:38.874678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T23:02:28.588225Z digest=sha256:26d0f2b372cf846daab5b8399b225354d1e6f3c27c771eed8fa332fd6d81990a

Observation 875f1f99-90de-4a37-a7fe-5b2f3508d4f9 · inbound

Thinking with Drafting: Optical Decompression via Logical Reconstruction cites this paper.

Thinking with Drafting: Optical Decompression via Logical Reconstruction Multimodal Chain-of-Thought Reasoning in Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:30:32.904559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T03:28:32.264246Z digest=sha256:03c8e724a1e29499c28c70d9a8ce2b7f7ed3b24622df090e7d30541b7234d573

Observation cf8729e9-749d-462b-b1be-58fc8eed83dd · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets Multimodal Chain-of-Thought Reasoning in Language Models

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:29.065846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:29.065846Z digest=sha256:7ebd38b7cdfa2c773300e01b1f2e001c5ccfa25f0407c128001fbbce3741d859

Observation dcc206f7-9f28-4832-923d-3791a8b20e3d · inbound

LanteRn: Latent Visual Structured Reasoning cites this paper.

LanteRn: Latent Visual Structured Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T17:26:40.248532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:26:40.248532Z digest=sha256:601c304722a85dacfa5f7d85d3cc02e83f089ff41e1167d757d0367f76123cd7

Observation 444c27a7-b7a1-4af7-bbf6-1f567d200795 · inbound

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis cites this paper.

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis Multimodal Chain-of-Thought Reasoning in Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T13:55:53.312853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T13:51:40.334057Z digest=sha256:71597978de208b3ad2244335329b3aaa8bddaed819b1941b4949114e9cbaae27

Observation 8f0119c2-76f2-4f07-8b49-f3d5c788ac6d · inbound

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces cites this paper.

GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces Multimodal Chain-of-Thought Reasoning in Language Models

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T17:33:02.326459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T17:31:08.575993Z digest=sha256:78034d4b96e39bd1d02f47e52689945b7b149d9f904fd5870649c98943963093

Observation 0113667c-c002-43ba-bb8e-7fee2dcf4db3 · inbound

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models cites this paper.

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:06:45.417233Z digest=sha256:b524a31c6b96a607d0de44faa43789edd6b85edd5c18f026cf4f60f93c258d0a

Observation d3cff9b5-13fd-4e3f-a775-4f383422ced1 · inbound

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization cites this paper.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Multimodal Chain-of-Thought Reasoning in Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:264cbccb7bb6d992abd055d4c5663372433ca314a6b37b2247cb5c4b6f74df27

Observation 008179b4-bddc-4f19-8315-2888e22e7dcb · inbound

Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs cites this paper.

Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs Multimodal Chain-of-Thought Reasoning in Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:12:33.231488Z digest=sha256:d8c3e4284aef285045f52db4b449f46cd9ce483794e488f573f05bbc99059901

Observation e53bfd75-6068-4707-8f14-54508e3e4c6d · inbound

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering cites this paper.

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering Multimodal Chain-of-Thought Reasoning in Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:03:14.408704Z digest=sha256:2a32084b0dbf5332231724bdd0c87259e74ac273df078c7ec1ed12091c240520

Observation f911754d-c845-4c89-95e1-b2499f7c68c8 · inbound

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering cites this paper.

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering Multimodal Chain-of-Thought Reasoning in Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T16:33:55.007487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:33:55.007487Z digest=sha256:9c00ae7a09390d69310ad953c0d64eb674a402cbd44ff04148c53a6541a08843

Observation 448c12b8-1bcb-426a-9ce3-913217c88572 · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:03:15.222571Z digest=sha256:725063c41ef09c6caf1df3dfc870342eec81bd7804422a0a0bb8cc882a4bfdef

Observation 6f5664c5-b3a7-497f-b10c-d2c473c6d21e · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T22:48:45.647588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:48:45.647588Z digest=sha256:a6ad8386c83d77b035a3f9e29e3b874678513cf2f75b1b1a5f9a44fe33a702c6

Observation 51135fcb-62cc-4db5-ac2e-48761b49e915 · inbound

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning cites this paper.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:40f01e26f546cacc4318341096a675e47662ac383852d7aec25fdf40f995cceb

Observation 5c0298ad-2e73-42c0-9622-f8d8f9f2c175 · inbound

From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning cites this paper.

From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:35:36.083682Z digest=sha256:128ccb4e750de0644737f0dc34c080628d1ab96636d527b13bebb00b5936ac6f

Observation ea6e9f77-d437-4e91-8bd0-85051ddd1c56 · inbound

Think in Latent Thoughts: A New Paradigm for Gloss-Free Sign Language Translation cites this paper.

Think in Latent Thoughts: A New Paradigm for Gloss-Free Sign Language Translation Multimodal Chain-of-Thought Reasoning in Language Models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T11:44:13.158397Z digest=sha256:1dd38e537d45b3e7b046473fc929df5767e3abbfe9c648ac75727abb72b1b721

Observation 11e719fb-f938-4b88-8698-ce05a887291d · inbound

Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning cites this paper.

Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:28:07.531338Z digest=sha256:7ac4902c9fcfcc53364bbc55c18a471c2d73d25e101655438a63d9d30a9480f0

Observation 632feb4d-d129-48f4-8894-9904f46c762c · inbound

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization cites this paper.

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization Multimodal Chain-of-Thought Reasoning in Language Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:17:33.668451Z digest=sha256:5f5367741498058dfd1004895c947fe87feb071f56bbb8ab580f3bb2259c19ce

Observation 08694fb4-b97b-4ec9-8c88-7477b0e84e8f · inbound

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization cites this paper.

HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization Multimodal Chain-of-Thought Reasoning in Language Models

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T04:30:40.726775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-05T04:29:52.480873Z digest=sha256:4d40b0a32e0597e2f7b24cc30fce56e8007beb8ac9b3d4a4377bb248d2368514

Observation 57a14436-1756-4103-b0de-8b5e7cd5e0a8 · inbound

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization cites this paper.

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization Multimodal Chain-of-Thought Reasoning in Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T23:48:32.613988Z digest=sha256:c7b104939f358d6165bcc1b341be1adfa5a6b50493eb81b289c13ef89f5678e4

Observation 6a9a443f-a542-4e16-a1c1-70402f1e6416 · inbound

AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models cites this paper.

AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:10:32.154619Z digest=sha256:ac7d47081cce93d46b199ee7d38c7f8714901b0b9b96f466e4a2d27de0e0c5b9

Observation 34e1e90d-c6a2-4288-a084-d4d338a30b6b · inbound

Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models cites this paper.

Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:16:34.362000Z digest=sha256:cca41e3200c81e778710d6f6b9c9fd08707f7b6679f32ad59f062f7342305c57

Observation ac698ec7-2f60-4fe5-bcd7-7bfee93b7393 · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Multimodal Chain-of-Thought Reasoning in Language Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:0423080a2c9f00249d6732b44c260682a743d560b665b7347e195ab6d10b3ced

Observation b6533477-c6b4-43e5-8708-eb4dc00684ca · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing Multimodal Chain-of-Thought Reasoning in Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:f3225aa4520c8e36166894f5155f25f65940f145ccabd53321c8fb2b82a7aacc

Observation 7d7aa372-7843-45c9-92b8-7487db071298 · inbound

State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading cites this paper.

State Beyond Appearance: Diagnosing and Improving State Consistency in Dial-Based Measurement Reading Multimodal Chain-of-Thought Reasoning in Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T11:45:57.291112Z digest=sha256:cca03ddb1154a96b3a01bdeaeb1a12566c63edd1569a0fd1d43ab475b3487006

Observation faa228b2-942b-4229-bb22-7be8c2543c37 · inbound

DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA cites this paper.

DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA Multimodal Chain-of-Thought Reasoning in Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T21:11:30.185960Z digest=sha256:d72c403ff0d1b44265f351833dba6257875d994f42a8f1a27ac58ae0d836f8dd

Observation ce9fa366-e0b5-4751-b5c1-a680636d2c3d · inbound

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning cites this paper.

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T19:21:30.235583Z digest=sha256:35f0db04b73df7c1cda27398402ec872aa05af60ae6d28f9e295685b0f5465d7

Observation 411e2a8e-7575-4537-9507-5885224158cd · inbound

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning cites this paper.

MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 131

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T17:19:59.247074Z digest=sha256:3c72e46491895dfc71c5a3e8d3f6161ad31277b7a51d929a54d081b26fb73cea

Observation 151ac1cc-7389-4b8c-9675-c8f91395a498 · inbound

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning cites this paper.

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:12:22.931098Z digest=sha256:fdac3589b43e63efa73499b5a97ead8c5699a64e3cee25f0ca9ac024d658ebdd

Observation 62f447ba-c0f7-4771-b79c-287e0a70530b · inbound

Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning cites this paper.

Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:12:22.228631Z digest=sha256:8b4a5fdf98dfbd130df104a8b154ad8dc59d5aba6e383a1e3af5c5b30350187b

Observation 00ad004f-0338-48b1-ac17-b4b53cd190a7 · inbound

V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning cites this paper.

V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:03:01.612516Z digest=sha256:890f0ac0da87c7d425466136e2e68159b7aa605fafd0405c0ca0e4ef630df991

Observation 45b1477e-25e9-4f63-84a6-3aad6aed15d2 · inbound

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both cites this paper.

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both Multimodal Chain-of-Thought Reasoning in Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:14:53.512879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T03:09:58.411261Z digest=sha256:0f80106308164605dcfc6ec3b2adc436fbef6ec2a43610253ef4c64c4ddf49f4

Observation ca043554-05ea-4d92-bf14-6226200d59f7 · inbound

SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening cites this paper.

SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening Multimodal Chain-of-Thought Reasoning in Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:08:21.157691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T14:03:56.859481Z digest=sha256:e27d2a87fe76d9324f949973851d1fe3500592f2017ecd947fffb0e5db8dbe56

Observation 770e9343-9249-487f-9aac-6508393854e5 · inbound

Pseudocode-Guided Structured Reasoning for Automating Reliable Inference in Vision-Language Models cites this paper.

Pseudocode-Guided Structured Reasoning for Automating Reliable Inference in Vision-Language Models Multimodal Chain-of-Thought Reasoning in Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:38:05.262091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:36:40.361910Z digest=sha256:4d8ca5f3cc1f637dc5171a961b8ca2ddbf681d3971d10c1d30cbe5ad88434ba2

Observation 3c4326ac-86f1-428f-9efd-d911a1e3ca88 · inbound

A Nash Equilibrium Framework For Training-Free Multimodal Step Verification cites this paper.

A Nash Equilibrium Framework For Training-Free Multimodal Step Verification Multimodal Chain-of-Thought Reasoning in Language Models

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:13:05.272603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T06:10:36.290559Z digest=sha256:d54b95907085e2b2c6041071a7540d211554721cd0c750e46f75952be47faedd