Pith. sign in

Paper Citation Record · LEDGER

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?

As of 23 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2505.12766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12766 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:31:36.673705Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:11:51.138098Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:11:00.663569Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e898e83-0abc-4a92-97a5-49d6a7988c36 · outbound

This paper cites GPT-4 Technical Report.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.449612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.449612Z digest=sha256:85628a041568310582a57da21e369cee4e87620ed2a6842b8bfad2b49be84a0e

Observation f0310bbd-8149-442c-b476-7ecb8b67f4d5 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.484551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.454637Z digest=sha256:383af74b9f859878a749bde91ef681d66ae9f251b0d5ad8fdd2d645fda550834

Observation 47eba2d0-4ebc-4b2d-9214-692f035a5915 · outbound

This paper cites Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Geoqa: A geometric question answering benchmark towards multimodal numerical reasoning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.470776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.459084Z digest=sha256:a02ec544d33b23fc6f58704da56f0e420dbf4d680273da28fa9914c385fda6f7

Observation a3b8a5c6-b2ab-4ff1-9338-5c64f034be4b · outbound

This paper cites Onechart: Purify the chart structural extraction via one auxiliary token.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Onechart: Purify the chart structural extraction via one auxiliary token

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.456499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.464033Z digest=sha256:c946d6af7af831467f93543d0b485d2b32df3ade63aace7e140b8377f9349b6c

Observation 1edd40f8-57b2-448e-bdb9-ed62261577db · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.468772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.468772Z digest=sha256:22a0c1ed7749ab72f8e7618709421ae9feb5033c493a979cae9e4f2f3db52097

Observation d7bbec18-b5fd-41fa-97ec-b988064f9b9d · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.443532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.474441Z digest=sha256:6c1aca64be9e1e387f2ab2b40284681eef0d6478205867ae4d8650cdb66aab05

Observation e0b87790-3eff-4ca2-93cb-3059c250311e · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.479795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.479795Z digest=sha256:c220ba8399098bda8b08498221e089e724bfb2b5f8a5ab26fee9b88a31f4469e

Observation fa295075-0556-4fdd-9c27-5e5fe8c4f06f · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.486216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.486216Z digest=sha256:356a72e16c3f5101f92f205261e85417fd31866575ab1cb5d13e381542708bd6

Observation 347854c3-2afa-4c36-812e-761718bb78d7 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.491246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.491246Z digest=sha256:d0a7a12794c99e1a958e0b609386cdd0f99a27daba7e76f3fbfd10d106b4e186

Observation 76aa7660-cc27-4fa7-8983-518bea5fd6e6 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? LLaVA-OneVision: Easy Visual Task Transfer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.500944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.500944Z digest=sha256:35169d67352b8675b0c0310a328837c1ea75ea01d685985cbf794fd5807d6292

Observation 0e6110f5-c63c-467d-9361-79217ce2c16e · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.506036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.506036Z digest=sha256:b7e06e0dc1f3431e43d4684e4ed1c2bfead9a58fdaad92edc5fc0c583e1a59c9

Observation 096892fe-06d3-441d-8eff-bc3abe7fd460 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Monkey: Image resolution and text label are important things for large multi-modal models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.424399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.513878Z digest=sha256:42261ef0e137792ee914215915637b8f3d9335b53e8f0d4181c028eacb9b2d0c

Observation 74d3a0f9-77a2-4ba7-b166-6ca7cef8d207 · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.519704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.519704Z digest=sha256:76054c00b093a12907ea6961e55683e3ba278e3079e20aaf06c6a45200269391

Observation edf02ff3-4d44-46d6-8a90-5b240bf5f2c3 · outbound

This paper cites Mmc: Advancing multimodal chart understanding with large-scale instruction tuning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mmc: Advancing multimodal chart understanding with large-scale instruction tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.378581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.524813Z digest=sha256:c23ee046f40459615d73459de2e8b78cbe1051e91e060143a1dbbeb307828698

Observation 028230fc-9476-4a55-a5db-4e189afedf76 · outbound

This paper cites Improved baselines with visual instruction tuning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Improved baselines with visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.345907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.531578Z digest=sha256:f012b9aec794399fbd135e87b574153160fbf71ef24c1307088dc79f47d7023d

Observation 01a93eae-8f7c-4358-b461-cb0b5dce8b8d · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.537271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.537271Z digest=sha256:75023383649172851d1af8d656e64bbc91d3eed58dbd5dab90013b894113e78f

Observation 041c2c6e-6720-4b88-abc1-4d057104c833 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.260264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.541759Z digest=sha256:56cf5a115aaacfa979aeea517f67cf5bd7039b6801946c84073d851f852bff33

Observation 38bb003f-b811-44fd-894f-1a7d08ef9f89 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.546422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.546422Z digest=sha256:1b37993f1659b9779af27c8b604af93f0725b3c3e5dadcc3e3c5b33581f0113f

Observation f956029f-b1a5-4fab-b568-ccd13c9d57a9 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.196376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.550750Z digest=sha256:5acd8f9c1ef87ce559d9f3bafa7255f474e14e8f0c19162ae1d11c6e81946a9b

Observation 28531fba-c90d-422f-a967-cafcc9bee302 · outbound

This paper cites MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.555576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.555576Z digest=sha256:054160bc15d959209b48565a27de72b3939472ab3d1b4c91069349faffdff586

Observation dbae85ca-7c13-41da-bc19-49eadac494a0 · outbound

This paper cites MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.562481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.562481Z digest=sha256:bb17ea2ee60d7f02b4c5c66110755f61f86bb136abe651c66b9e433b96b1da9f

Observation 1540bb76-157e-493b-86b8-84f3145a3d91 · outbound

This paper cites Mmlongbench-doc: Benchmarking long-context document understanding with visualizations.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mmlongbench-doc: Benchmarking long-context document understanding with visualizations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.156732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.573654Z digest=sha256:0cd7ac687ce9a56f5d5e279752e2660f64fe789e2b601251a8331763b6793a1b

Observation 6dc2d162-4b3d-4398-8e12-eb5de8065b3d · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.143623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.579098Z digest=sha256:4e1aeaeadbfa78779dae158648bd616e7b6e4830414a4f7e7900b7fc9f410e63

Observation 786414d7-3b4b-4132-9100-d08dd49d2ef4 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Docvqa: A dataset for vqa on document images

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.130656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.584922Z digest=sha256:7f6ae87b901ccb4ee71dec446cca8fe0fc8017a72c3088820530c41b6fecaf17

Observation 05e07651-8e08-4917-849b-0d9a5d371654 · outbound

This paper cites Gpt-4o system card, 2024.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Gpt-4o system card, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.589136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.589136Z digest=sha256:c7beb70c3c77ae4ffc01628b8179ff97a44a73373b45cbabd8e6970183d536a0

Observation 154a248a-4893-4a57-a533-c88f53a68975 · outbound

This paper cites Towards vqa models that can read.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Towards vqa models that can read

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.106263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.594518Z digest=sha256:3e1efede93d9bebc86b210d9e2999ba695c36822f425bf56d9176e9a8bff0718

Observation 784c6e98-c911-49f1-8058-a0620f5ce0b2 · outbound

This paper cites MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.598636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.598636Z digest=sha256:2b821ed52dcd199b2167748e971ace102e6030ca52549532af35ffdb99f57c58

Observation e82d1835-91d0-4a76-96b7-a9fb8970244e · outbound

This paper cites Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Contextual: Evaluating context-sensitive text-rich visual reasoning in large multimodal models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.090992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.603135Z digest=sha256:7443c471cce55b1edfc5a4d5d59d9d48860c8fb2738ea9acadae6b8e3462de5d

Observation 104371b7-565b-4a82-a4f9-d4aaa8b3a135 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.609414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.609414Z digest=sha256:7f8fbf8f1c473a320523c3a95cdb324a12638ede90758dc89e5ddb4f89adc67f

Observation 2381fb54-50f9-4f49-9dc1-d820695f5bb7 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.615071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.615071Z digest=sha256:f3ccd6af8837e6c334340fcebc3758a8a00b336a627562a5b629c5724b051b88

Observation 66f9c24f-d165-45a6-ab4e-ab777082572e · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.077729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.619496Z digest=sha256:587b9a3a1ea24b1ed985a504324af08bd6b490c79b19b21e2cc8061716419488

Observation 6141f1d2-1314-4115-9085-0329bb18c23e · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.623999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.623999Z digest=sha256:ddeb7e073ef4b279b7b80446e0298c575224f3f195cfc5248bc483831082c82e

Observation 5e8e4e41-ade1-42fe-9cb1-8d47ed0b8958 · outbound

This paper cites ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.630121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.630121Z digest=sha256:9fdda4eb7cb93b8a3211c47fc972e99effd6f588e63d4b55dc8aabc3922ceb52

Observation bc6e46c5-4558-4b8e-b9aa-5211e98293bf · outbound

This paper cites ChartBench: A Benchmark for Complex Visual Reasoning in Charts.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? ChartBench: A Benchmark for Complex Visual Reasoning in Charts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.635245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.635245Z digest=sha256:0fc88cfb6b6b94a443f6bc9805081205b94a71016e89465f36bb6c2fd227ca31

Observation 4f33518c-4841-4518-97d9-2af007e2bba8 · outbound

This paper cites If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.640042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.640042Z digest=sha256:76ca0803b8745f2f4130fef72ab3158eb2f826f2169853f2ef158a52f4d0281b

Observation 1e8e66b1-1f2e-478f-b609-8a51f5b3f20c · outbound

This paper cites CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.645575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.645575Z digest=sha256:8cd1148e892ebd39fa47b5c7c4e163580bc18949cf625efda342900d42588f12

Observation ac2968f2-de27-417c-af0b-3a2764584527 · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.065030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.650427Z digest=sha256:0607c4b35d13fde8b58d31492ac1a0812033ae3915e62fa4766481af4344193e

Observation d59324d4-2907-4833-a643-93d1a6f209fa · outbound

This paper cites Exploring the capabilities of large multimodal models on dense text.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Exploring the capabilities of large multimodal models on dense text

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.051841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.654661Z digest=sha256:08649d09cb5f8c4fc9e285ec6c05dcd569555831165a24a4c5b3912636134535

Observation c43f4e5d-04f6-473d-acc8-a7322714040c · outbound

This paper cites Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.659746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.659746Z digest=sha256:4b0d5e325f6553d2bc39869583c61ee1ca3909a71c12354eae44e225ec20d5af

Observation f4c93a76-ec00-46f5-b44c-a038b50c953b · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV , pages 169--186, 2025.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV , pages 169--186, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:31:37.038092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T20:31:36.664683Z digest=sha256:02ea6cf0a5e3561cce880de158554fb4040cc085155a5c3a7af8062d91c0a001

Observation e740d921-a8e9-4ef2-bcff-3851a9807674 · outbound

This paper cites write newline.

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? write newline

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:36.673705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:31:36.673705Z digest=sha256:c65f2cf7c20584b7217743e8a90c55395bd0ea317021169dc91e13b22b3833b8

Pith citing papers

Observation 5dade1ca-1000-4b86-8708-6d2536cc70fc · inbound

GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts cites this paper.

GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scripts Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:11:00.667074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:11:51.138098Z digest=sha256:70361c57a3c14ab5354c0afae38cf14f4e5d79d02dbfe0b0333c528c1503d70f