Pith. sign in

Paper Citation Record · LEDGER

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?

As of 18 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2505.12766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12766 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:36.673705Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:11:51.138098Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:11:00.663569Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e898e83-0abc-4a92-97a5-49d6a7988c36 · outbound

This paper cites GPT-4 Technical Report.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.449612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.449612Z digest=sha256:853f08de46e0543b4023b154b3fda2ab86c3c9705242f977894029be676e1896

Observation f0310bbd-8149-442c-b476-7ecb8b67f4d5 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.484551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.454637Z digest=sha256:3bd126f9cb228229de856cbcfda5ee25ce501000787406586bfa6dff881caa03

Observation 47eba2d0-4ebc-4b2d-9214-692f035a5915 · outbound

This paper cites Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.470776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.459084Z digest=sha256:65bc1257734aa50ccfc772b85d32e233efa52b67f0a338362a40b7c9655034a7

Observation a3b8a5c6-b2ab-4ff1-9338-5c64f034be4b · outbound

This paper cites Onechart: Purify the chart structural extraction via one auxiliary token.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Onechart: Purify the chart structural extraction via one auxiliary token

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.456499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.464033Z digest=sha256:825ce176ead9ca8a8d5ec5b015d114230cb9573f2c632204e9d118fa0190e6c3

Observation 1edd40f8-57b2-448e-bdb9-ed62261577db · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.468772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.468772Z digest=sha256:0c3fd8be532a50e70655e7b65a14749040d5d2fa2351e1bf890e81f7ed96c235

Observation d7bbec18-b5fd-41fa-97ec-b988064f9b9d · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.443532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.474441Z digest=sha256:58afaddff85d98b5b97d4423bf6c712455c4ae947aae7fc7aec3c8281fbd4d99

Observation e0b87790-3eff-4ca2-93cb-3059c250311e · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.479795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.479795Z digest=sha256:63830c218e6f363e75a991f844302bd0fe7499c16158de779d56780994866014

Observation fa295075-0556-4fdd-9c27-5e5fe8c4f06f · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.486216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.486216Z digest=sha256:a817d7dd1862a9e3995035541f089f6e951f78257c2c215bce04b698212b4a93

Observation 347854c3-2afa-4c36-812e-761718bb78d7 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.491246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.491246Z digest=sha256:fc2ef74ef1d85a40db37c476387c89fead4ac79948704661dce9efb911f0f26f

Observation 76aa7660-cc27-4fa7-8983-518bea5fd6e6 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? LLaVA-OneVision: Easy Visual Task Transfer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.500944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.500944Z digest=sha256:f614c646745761a3d1cb9017d8f25ed06f4b9c7103a620ce5b7300854029b110

Observation 0e6110f5-c63c-467d-9361-79217ce2c16e · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.506036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.506036Z digest=sha256:bdf9c7ca4736fbcdc82c003cbd50249e32fb02f026420861b30658c679c84552

Observation 096892fe-06d3-441d-8eff-bc3abe7fd460 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Monkey: Image resolution and text label are important things for large multi-modal models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.424399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.513878Z digest=sha256:37047d4c71866c77bef8aa8612879432d4e43bee3d065a35113af3a91792a929

Observation 74d3a0f9-77a2-4ba7-b166-6ca7cef8d207 · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.519704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.519704Z digest=sha256:11b993507f52fb0ef950089a25c2152b12edf7b13127ad51617ab8c53eda3b9d

Observation edf02ff3-4d44-46d6-8a90-5b240bf5f2c3 · outbound

This paper cites Mmc: Advancing multimodal chart understanding with large-scale instruction tuning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mmc: Advancing multimodal chart understanding with large-scale instruction tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.378581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.524813Z digest=sha256:871eafa6bbeeef3ce822ef7ec08a765809d12d568153fabd7badc3a733b4b844

Observation 028230fc-9476-4a55-a5db-4e189afedf76 · outbound

This paper cites Improved baselines with visual instruction tuning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Improved baselines with visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.345907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.531578Z digest=sha256:5c1da48aa2fb8157f9bdfe22e9fb695d9887e333bba581c194ea9f253f002ec1

Observation 01a93eae-8f7c-4358-b461-cb0b5dce8b8d · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.537271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.537271Z digest=sha256:add0d2549b75fe8e0d45ca44da3cf8ed660daa4d2caf1a591d4b820356a36175

Observation 041c2c6e-6720-4b88-abc1-4d057104c833 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.260264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.541759Z digest=sha256:c71d973da8af4158c98321b9414deb3cfc452a2f009c825640b62cce684a2713

Observation 38bb003f-b811-44fd-894f-1a7d08ef9f89 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.546422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.546422Z digest=sha256:0f10a8416f44118d0fae0dee78032010d548e01859307a763982401c7bb4da2a

Observation f956029f-b1a5-4fab-b568-ccd13c9d57a9 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.196376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.550750Z digest=sha256:cd466ec9c9489b3d3d51afccd418fd849e4d6eb8fdf3c8c56b59329ff504ae39

Observation 28531fba-c90d-422f-a967-cafcc9bee302 · outbound

This paper cites MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.555576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.555576Z digest=sha256:09e30faa52f73d1982b819154259907dbfc57ef17c77bfb4d2ce8ad423034626

Observation dbae85ca-7c13-41da-bc19-49eadac494a0 · outbound

This paper cites MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.562481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.562481Z digest=sha256:e28c91bc026df8ae0e21808d48f696be3b82ad71f80c2ae0a4bb2c9469c44410

Observation 1540bb76-157e-493b-86b8-84f3145a3d91 · outbound

This paper cites Mmlongbench-doc: Benchmarking long-context document understanding with visualizations.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mmlongbench-doc: Benchmarking long-context document understanding with visualizations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.156732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.573654Z digest=sha256:6ebc0abb411ecb30a66c90ded1bbe50bc2a0a44affa9db0ee65d3ef8a9c01c32

Observation 6dc2d162-4b3d-4398-8e12-eb5de8065b3d · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.143623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.579098Z digest=sha256:37f845ab7b28fd01085d7ef750007b137d40374505dbf1526fc39ae777d5e60f

Observation 786414d7-3b4b-4132-9100-d08dd49d2ef4 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Docvqa: A dataset for vqa on document images

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.130656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.584922Z digest=sha256:00f4682ac54ddd4dae6e0d6d6c7f19ede163a211f78e27a817ec54e6fd732774

Observation 05e07651-8e08-4917-849b-0d9a5d371654 · outbound

This paper cites Gpt-4o system card, 2024.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Gpt-4o system card, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.589136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.589136Z digest=sha256:75c4fb77940f72f71c09419d1f152153b0d9d6b0a981fdb4c684237a2042718e

Observation 154a248a-4893-4a57-a533-c88f53a68975 · outbound

This paper cites Towards vqa models that can read.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Towards vqa models that can read

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.106263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.594518Z digest=sha256:4a09da577ebffe732f44132f97d37c707b4aa4a8589b685d051eecc50267f60a

Observation 784c6e98-c911-49f1-8058-a0620f5ce0b2 · outbound

This paper cites MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.598636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.598636Z digest=sha256:c6552a4981d16c7a3d4dbb4710432ad7511a78153b02a0f8d072e8c7a02daf1a

Observation e82d1835-91d0-4a76-96b7-a9fb8970244e · outbound

This paper cites Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.090992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.603135Z digest=sha256:a246db7254657f0dea75e3262232a553fc118a791d0ab559ae56aee77376940b

Observation 104371b7-565b-4a82-a4f9-d4aaa8b3a135 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.609414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.609414Z digest=sha256:aab62cbea94c446d7f43b6f1341ae2a5b28465c9a3cd07a252524f2c63065f5f

Observation 2381fb54-50f9-4f49-9dc1-d820695f5bb7 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.615071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.615071Z digest=sha256:a31d12075dd86bc177e9cb49d7f7366ab3bb71da142b922cd28e786c40b4b225

Observation 66f9c24f-d165-45a6-ab4e-ab777082572e · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.077729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.619496Z digest=sha256:d589cdde6795017d90e5e3152758a179b8a245d4a185048c7cabf27f28e83017

Observation 6141f1d2-1314-4115-9085-0329bb18c23e · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.623999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.623999Z digest=sha256:8cd54bd3a83ad71f04b6f2acf65ffb673b420b6245579483943770bd2a4fc03b

Observation 5e8e4e41-ade1-42fe-9cb1-8d47ed0b8958 · outbound

This paper cites ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.630121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.630121Z digest=sha256:b54dd41c8e5771922147f2bcceb9c29bcde0878c40a953c774e8467e75125d46

Observation bc6e46c5-4558-4b8e-b9aa-5211e98293bf · outbound

This paper cites ChartBench: A Benchmark for Complex Visual Reasoning in Charts.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? ChartBench: A Benchmark for Complex Visual Reasoning in Charts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.635245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.635245Z digest=sha256:c966a0e94c6c11308f5ed5b574d3b3adcf71095997e08b2abd33450d28d50b4c

Observation 4f33518c-4841-4518-97d9-2af007e2bba8 · outbound

This paper cites If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.640042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.640042Z digest=sha256:4f69b9eb5c997b8cddbc66f764a3c160f070990674e7f50d5862fccdcbce47a9

Observation 1e8e66b1-1f2e-478f-b609-8a51f5b3f20c · outbound

This paper cites CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.645575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.645575Z digest=sha256:feb72f4f8bab161f921674802b5b7898439abe1456780e54acf18742363ff4cc

Observation ac2968f2-de27-417c-af0b-3a2764584527 · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.065030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.650427Z digest=sha256:20d1ce7c0e2985bfb4b0f8e6ca86112a3fe7f5afebafe86f378fbb659cadea98

Observation d59324d4-2907-4833-a643-93d1a6f209fa · outbound

This paper cites Exploring the capabilities of large multimodal models on dense text.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Exploring the capabilities of large multimodal models on dense text

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.051841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.654661Z digest=sha256:a4ac57c2c55d8fda61f3221259f2d99c0da4f7fc895c1f73c75745c18cf0a1e4

Observation c43f4e5d-04f6-473d-acc8-a7322714040c · outbound

This paper cites Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.659746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.659746Z digest=sha256:4bef8f69bf3be899f4ca5a1088a0a7c16672e100fe4f62623ed6c43b2f21f432

Observation f4c93a76-ec00-46f5-b44c-a038b50c953b · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV , pages 169--186, 2025.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV , pages 169--186, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.038092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.664683Z digest=sha256:e3f5a1536314fa779cebac343fe39e770443256ccd04e737c706243bf789de1e

Observation e740d921-a8e9-4ef2-bcff-3851a9807674 · outbound

This paper cites write newline.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? write newline

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.673705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.673705Z digest=sha256:b197a872800aa62b20adaf47414366fa0a85fd489ef0951a8eb99a3944885bc5

Pith citing papers

Observation 5dade1ca-1000-4b86-8708-6d2536cc70fc · inbound

GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts cites this paper.

GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:00.667074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:11:51.138098Z digest=sha256:bca85718ca92b7e73a6abdabbe117e2597e08001b129c9e2b896a829eb4fc416