Pith. sign in

Paper Citation Record · LEDGER

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

As of 5 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 1 inbound Pith citation observation for arXiv:2410.04509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.04509 v3

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T20:10:59.264484Z

measured 93 of 93 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-23T04:30:38.804702Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-23T04:32:32.712337Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact58
  • verified fuzzy32
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 391c2e3d-8ab0-4bb8-a7e8-0f1680031792 · outbound

This paper cites Complexity in declarative process models: Metrics and multi-modal assessment of cognitive load.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Complexity in declarative process models: Metrics and multi-modal assessment of cognitive load

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.675921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:ed09528c7916af147f92fc1132d470b431ea5ebd5e8debb2b1769ba5c497de6a

Observation 055423cd-2fb1-4a5f-a3ed-bdbab039e4d1 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.713023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:eedea7789f069df956d60d7f7a2b512a7aee1e24baeccc93fa36482823d3a851

Observation 081df14e-80ae-4bb1-ab08-30c62b6666a9 · outbound

This paper cites Scaling laws for generative mixed-modal language models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Scaling laws for generative mixed-modal language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.680024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:8098e958764e92de647de0ae0006e52be19a73f87cfdc01224e445fb2fd679da

Observation 13237bbd-396e-47ce-8b20-4238944258aa · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.615341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:70b718adb0041866d4bb283b5a0e996a1c7ab1e4f57faaf60252a124cc27a2c1

Observation ccb84f09-1457-4698-b7fa-8999f8a525ca · outbound

This paper cites Claude 3, 2024 a.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Claude 3, 2024 a

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.683997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:b9c15bb58248adc88feaac0a36978a75006953e652381d0648a72c33ebf2776c

Observation 5e6cf1ad-5f67-473c-bd17-555a775c03ce · outbound

This paper cites Claude 3.5, 2024 b.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Claude 3.5, 2024 b

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.804632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:2bbbc254b2f1b75a7628c478c303ab96a3a8379642a7388b00b72e7e05e7d5e0

Observation df8fda6c-3f85-440e-8684-a2f46e46d525 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:25.026755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:065569bb178f0affa079089debeeb830c6ef5e238efecae6891dc7e4153c1ebb

Observation 8d8b027e-93cc-4830-a1cc-f388a5eda008 · outbound

This paper cites Turning large language models into cognitive models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Turning large language models into cognitive models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.022104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:807f9b684b7029f0d0f3108e390d79d2c0c64deb4c9405b387b5caaed98653a4

Observation 4ca4f385-91cd-4b24-95d3-a76458269871 · outbound

This paper cites Theoremqa: A theorem-driven question answering dataset.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Theoremqa: A theorem-driven question answering dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.808080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d491d99a942fee5398bdcce6910b7a9b6197b84f3e1594879fa0aad68e7509f7

Observation 3be8ded9-a037-42a7-97eb-24e30023ecc7 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.993510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:4d7c699c020da981a6d6b49b087cd13e4399db420999145d7faee638b59389d6

Observation ce162e8b-2af0-4b96-8ef1-f74d6e85f1d2 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.981889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:94ace674267f0f687e7210d4e9211f61893ac6562a74a4aa0c2d07b26bc75739

Observation 9dc97bf6-0f72-4722-8dcd-ec1e2a400ba1 · outbound

This paper cites A survey on multimodal large language models for autonomous driving.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection A survey on multimodal large language models for autonomous driving

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.793952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:c304ca2a71596a4fdd114e459ee9dd3e13811bdfddabb28d3a1ec63dd99f7d05

Observation 16026aa9-6449-4604-a689-020543105388 · outbound

This paper cites Advancing mathematics by guiding human intuition with ai.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Advancing mathematics by guiding human intuition with ai

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.801048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:b755da18814b29f50147757ec7dfbae4e0b4da65e2ff00ba7e3ba43c4685fc7b

Observation 21518bc5-3b88-484c-822b-15c1805cd1d6 · outbound

This paper cites Visual representations in the human brain are aligned with large language models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Visual representations in the human brain are aligned with large language models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.031867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d1b2aee92aa314027187e65aeb87a0ac6a9c125a8693aaab4fb9fac377e83041

Observation 52922b22-c1d3-45ca-bf92-467dfa39452a · outbound

This paper cites Muffin or chihuahua? challenging multimodal large language models with multipanel vqa.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Muffin or chihuahua? challenging multimodal large language models with multipanel vqa

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.790354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:6a02eacbbabe1f827d76fef4ff9f2577039d4dc4da0a1f2f1eae47862b49ff68

Observation c17fe3b3-ce1f-488e-ab6a-d6b5706b8563 · outbound

This paper cites Trends in Integration of Knowledge and Large Language Models: A Survey and Taxonomy of Methods, Benchmarks, and Applications.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Trends in Integration of Knowledge and Large Language Models: A Survey and Taxonomy of Methods, Benchmarks, and Applications

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.742490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:873a9494b610cda3a8bbb4f5f5c4ccc3f211cbae093850175e8a4ee805ce3c4a

Observation e53d8485-5e44-4821-8c36-a047c854cde0 · outbound

This paper cites IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.891338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:2b55a8230ffda980f233e7d520bac9013dd65ed847306ea88deade07c1afb4ea

Observation 33cedaae-dab5-4308-868d-eabed4c4ea68 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.896437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:2c48dee2c05ef17927cbcc3aaff4bd14e9e7b529f58184753471dbacb4a8f16b

Observation 61b1e7d4-020b-45a2-a222-07c3f5485eaf · outbound

This paper cites UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Socioeconomic Indicator Prediction.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection UrbanVLP: Multi-Granularity Vision-Language Pretraining for Urban Socioeconomic Indicator Prediction

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.871280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:e78a13687c842f2ed70f7a9132df70495c57a60a398fcb2daba969e6ae7ee0d2

Observation 868a8b7a-3584-45a6-9cac-c12faeedf4c6 · outbound

This paper cites PeFoMed: Parameter Efficient Fine-tuning of Multimodal Large Language Models for Medical Imaging.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection PeFoMed: Parameter Efficient Fine-tuning of Multimodal Large Language Models for Medical Imaging

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.825358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:5f181251a9f6b8ee505798f24965d6414905d3867c4bf6b4e39de6d8c0e305a8

Observation afc6ffdf-2f90-4c99-9de8-680fd9e1e13c · outbound

This paper cites CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.003350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:6408131b3450852dce9be7fed38a934f5632fba916b2ed143057f0139b55980f

Observation 61c2cfe1-22c8-449d-b678-ebbc4c2dbc27 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Measuring Mathematical Problem Solving With the MATH Dataset

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.728474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:8b7851701e265c3fea62cc28dea7fd20fa7d8ca56c3e12ba11a6537190205cac

Observation 8900b923-fec1-4134-bf1e-df2e0167bfa0 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.850060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:38acf185c25f9e034d4db45ab3206dd5cd788f85fc08ac548fd85661b1940861

Observation ec757e1c-078f-417a-9525-dc588045b480 · outbound

This paper cites MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.854814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:3814616dd9ff4e14902c84960664f4aeeebe5b12cee41984dbc15980c2f00631

Observation ccc2ac53-320e-476b-a810-444d8c4da6bf · outbound

This paper cites Describe-then-Reason: Improving Multimodal Mathematical Reasoning through Visual Comprehension Training.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Describe-then-Reason: Improving Multimodal Mathematical Reasoning through Visual Comprehension Training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.883048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:6b7dd0e22b826040862a817f91018ce3ca6add523192339c2050db866ec55e6d

Observation 1817a52e-bc03-4163-845e-e23b32334315 · outbound

This paper cites New generation deep learning for video object detection: A survey.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection New generation deep learning for video object detection: A survey

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.797268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:98bf9d94091d1470aeba14ac3e893aa1c2f32f32b2bda90199f3e3db22d24f52

Observation 1e079590-4b40-45db-8bf9-78b8a447502e · outbound

This paper cites Learning instance-level representation for large-scale multi-modal pretraining in e-commerce.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Learning instance-level representation for large-scale multi-modal pretraining in e-commerce

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.823018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:35bd043c0efbaaa0853a709ee4902fb238d442c41ff2cdbf299f98350fc9046f

Observation fd89f66e-3a2a-42cd-9720-86d0ad964e32 · outbound

This paper cites Large language models struggle to learn long-tail knowledge.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large language models struggle to learn long-tail knowledge

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.778712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:98622e69cd3b6ed40beb76b4d8998000710e0c0c60662710de9c2a3eab6f5490

Observation 10420b75-8e73-4209-bad3-e386591f6652 · outbound

This paper cites Scaling Laws for Neural Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Scaling Laws for Neural Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.699095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:b8f32241a296e02b85c0f11bdaf700231039606df29073ed72f45eee3600629c

Observation 73b4a142-6d39-40c3-9099-df106ddfdfb6 · outbound

This paper cites Cognitive load theory: An applied reintroduction for special and general educators.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Cognitive load theory: An applied reintroduction for special and general educators

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.710619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:e9a596717ef23300c5cd3d13edcfbf6b23206144db15586f0078c19e5aa9b191

Observation d3ceb4ea-2a12-403f-92e2-70ad657c6bfc · outbound

This paper cites Large language models are zero-shot reasoners.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large language models are zero-shot reasoners

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.770497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:5ccd398246db17c0b32b7f0bb53deb9426bc1e42865c668852ec06ac00314e55

Observation bc925469-d196-4f80-a92e-8f719e36071c · outbound

This paper cites Solving quantitative reasoning problems with language models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Solving quantitative reasoning problems with language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.766527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:fb734f5fa9bf90d7110eb2d9d6c41a58040a8cd12ec697d12a9b09c9df139880

Observation ca615388-f73a-4a0b-9df3-4df1a512c422 · outbound

This paper cites Bringing Generative AI to Adaptive Learning in Education.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Bringing Generative AI to Adaptive Learning in Education

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.068959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:f295a43f8769ed26e3f199bbed3f26bb7ef9d2b11eb9a0ee1907fa5fe999abe8

Observation 14d40029-233f-43c0-99e7-30d1188c3bdf · outbound

This paper cites Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.840001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:ecd1778ecb3c7a3f0fd87d411d8878a13c260aa2f54534db94c57315d8c97b09

Observation e4a372f3-2979-437e-8640-101eb7831d61 · outbound

This paper cites CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.901744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:5cbc05ad029ab08004ea1d22cf4bc7fbcec1d5f4d431d6b532a4d32911bb8097

Observation 981eb948-3025-48ee-83e2-0d6677026d28 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.758496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:eddce2023efc23233cd0e655efffdcded52c614f18a9f329c07bf8ee1633dec1

Observation a4fb8dfe-28f1-432b-9c51-6b8ead269c3b · outbound

This paper cites MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.804839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:5778b7a895610ab61f77c9d366ab4fd38c54f6a4bf9c6dff2381dbadfe30d222

Observation 213f5ab0-6c1f-454d-809c-372f739c275f · outbound

This paper cites Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.074552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:5aacebdfd58411bee28c307e3e3fbef0689854c1d1ade072afb163416cd150c5

Observation bb2f1cd8-6e24-43f5-8ac0-7e500414bbac · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.784013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d7f7711708b2d3509757cfd69fc203f8909e6a2f9f77709789012849c31ce37f

Observation 4add98e9-d240-4094-9622-d47fb5d22cda · outbound

This paper cites A Survey of Deep Learning for Mathematical Reasoning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection A Survey of Deep Learning for Mathematical Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.737121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:8df42c5d5264e766bb7dffd400a1b22628acf1928519300317c9167c626ab426

Observation 740ff5f3-23ef-4f57-877e-1a072ddbc86d · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.722988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:c4cdab85599b93a36b29b119a82046e14179b93547a6f16d85bf79b5de5e1c61

Observation 5ae4255b-12a5-483b-a923-2ccbaea78250 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Chameleon: Plug-and-play compositional reasoning with large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.762166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:f25d112b84264b4a69ad92db492daa69a43e00364544758265a303170123edb7

Observation a4d2500f-7082-4485-a0e7-69bbf93779de · outbound

This paper cites Large Language Models: A Survey.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large Language Models: A Survey

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.844854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:7ee084f23c624661e42509d67dab303329636155e1f31805d59cadb9b6b00a5c

Observation 1d180948-b687-4322-b825-7e8fb872c1e7 · outbound

This paper cites Scaling data-constrained language models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Scaling data-constrained language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.750446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:82a44c270cd252acfc2032b3add16d4358daa24954bfcc7bebb59c363adbba02

Observation 91aa87ec-9310-4dcc-a70c-71666846f229 · outbound

This paper cites GPT-4 Technical Report.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection GPT-4 Technical Report

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.757361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:8d5c3db19b794a350b8d8af299d63a332ac4f70611ef2c0492c986fa375e1cb6

Observation bc700717-d009-44ea-9e7b-272d59800955 · outbound

This paper cites GPT-4V(ision) system card, 2024 a.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection GPT-4V(ision) system card, 2024 a

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.743969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:4882de03830a672ddabca26001f2e14de9e14734274936d2ea6fed2b8db2a17f

Observation bfbc8825-449e-4ddc-9c43-9c4e997b51e8 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence, 2024 b.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Gpt-4o mini: advancing cost-efficient intelligence, 2024 b

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.738654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:8b22f6e66a1c4f3d9f2eabd7b401dc2c0490160916129155c6374a2ec70833a9

Observation 5d29731a-20f6-4213-8362-89fdc2eb83b0 · outbound

This paper cites Cognitive load theory and instructional design: Recent developments.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Cognitive load theory and instructional design: Recent developments

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.786625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d5f43d5896863d8885e6ef99d4077204edaa5e11fbd0697d3268501767d2dbba

Observation a1e38f9c-de1a-42cd-9fb2-0d325458eb1a · outbound

This paper cites Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.747536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:7cbb4105cbd0e977f280edb6bb566f5c5ec3379e55cfda9423d042e5a8e34f66

Observation 0f08144f-7f53-4bce-967b-e99f9d347ae5 · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.789526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:e787f8af93ec8c1dc637d3cb76af14c2a3014ac31eb8832072ce38e88166f296

Observation b92f1d45-4b97-46a8-b689-3c3579cf5d2b · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.631903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:a8fc5af0916cbe66b6cde4b1bbcf12b073829e9bbb352cab900dda9092e26422

Observation 6109d41c-3620-4bfd-bf38-7ec138d14b65 · outbound

This paper cites Elementary math learning through piaget's cognitive development stages.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Elementary math learning through piaget's cognitive development stages

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.730301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:5ee045432623e5c793a39b945a9f14d6f0cabf44c96b975be5265c11244322b0

Observation c5183aab-f5a4-4d10-ac62-c2ca24b4d04e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.597040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:75cb8b5a2cdcd44c0458634d8ccc7e0ab0fc1948080750b01178218852d42ac4

Observation 1cdd9210-316a-4bb2-a8c4-de08309256d0 · outbound

This paper cites Detecting Pretraining Data from Large Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Detecting Pretraining Data from Large Language Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.831510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:4335ac5fe1d643a1b995b44abf7b22e3c46dffd338d35d7926446cf1484ee734

Observation dd81da68-afc6-487a-b30d-16383dfef233 · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.675026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d323593e95d7d9c0a33655b8aff43f2c135eedd07d357da155ce698231c1bae9

Observation 89ecd939-1a04-4291-8044-b49b18e9cc38 · outbound

This paper cites How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.058416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:f7b27c00ceeb7f428cf81d08e4d24b80d292d70264cc4bf8aceb765be4f1bed1

Observation 81b3a9e7-e85a-4bea-b552-7cf98bfc58c5 · outbound

This paper cites Scieval: A multi-level large language model evaluation benchmark for scientific research.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Scieval: A multi-level large language model evaluation benchmark for scientific research

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.722923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:0b5cef3b59ebb54b59311c49765019c279fec3c5d85657e4c816c731487b4783

Observation 6696d416-a815-4cb5-bea7-38f0a3071df7 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.796726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:0606423e296d5f00c0c17e34d70969786446eefef10f00f87f7fa5b0544d493d

Observation d0ae4849-dc13-4b84-8d16-27965b2d6f13 · outbound

This paper cites Memorization without overfitting: Analyzing the training dynamics of large language models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Memorization without overfitting: Analyzing the training dynamics of large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.734523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:7e37eeae77492c9edbc65fe12fa9183983669fcf28d3d8ced43139c68dfcbc78

Observation d7ddc412-30ba-4a2d-acbc-506790c88a42 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.769190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:674e6cc83243b6b30c22b131a6fdc2a100d10e7be7dbeac7ce9b55e8b7060e1d

Observation 6d1ba5fa-f238-4a5e-aa60-3ad5e5606653 · outbound

This paper cites Large Language Models for Education: A Survey and Outlook.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large Language Models for Education: A Survey and Outlook

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.812679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:ff61a873acd12b698eea52720fef89bca6cd03010433dcd5816e475b3618ab92

Observation dedad807-91d7-4d41-ad1f-0185068474a2 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection CogVLM: Visual Expert for Pretrained Language Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.656773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:8e1d0bb426d3b718bf1d2c5afa14ee84013f29ceb25dbf09d576a2cbc71d09a0

Observation 2841365e-1eb1-47e2-818b-53502139cdcc · outbound

This paper cites Large-scale multi-modal pre-trained models: A comprehensive survey.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large-scale multi-modal pre-trained models: A comprehensive survey

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.706479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:fd75b7d9e5e0c96f43d66648c003f667a76b2bb8aebc2c6d06f516bd98ee591b

Observation e1bbde25-40b2-4a34-b348-a6b4424acfbe · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.665670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:23a300b8638ab19889f2c54c7acd06e35630ca06e8685de53daebd7bceeff8c1

Observation 0b2da07f-e806-4b7e-a552-f13f425ffe47 · outbound

This paper cites Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.651777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:0aa791bb9c7c91cac82aba1f846426f75cdd3fbebdfc6382ba38cda60289ef3e

Observation 2e8d19e2-7d15-4fb1-98f4-a5faad70e5bf · outbound

This paper cites Are deep neural networks adequate behavioral models of human visual perception? Annual Review of Vision Science, 9 0 (1): 0 501--524.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Are deep neural networks adequate behavioral models of human visual perception? Annual Review of Vision Science, 9 0 (1): 0 501--524

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.718670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:84492dae1875c292c40d241744a298892a6bb5d670c7340507fbefa87e5e080f

Observation b7bf2295-098e-429e-bd4b-71e3631f542d · outbound

This paper cites A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.643282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:bdf819b05ef649d12371462c5fae1571e5005b8b24ab06fec12fe860708e9a78

Observation ad01e102-61e5-4ae0-be2e-9cc1545c302f · outbound

This paper cites MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase Understanding.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase Understanding

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.637635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:aafcab2761b0dbfd3c212bd40ec5ef3d34e117ecc7156f2b2eeea9f2862028b8

Observation 752b09f9-3757-4f31-99a5-b587d6533ca1 · outbound

This paper cites SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.609009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:04d7de0f5e9c2e9aff95d0ae512fe81a380102106c8fdc57b7a59dfe68c907f1

Observation 6ed7585e-c8b0-45b2-b324-b47400a2c4e1 · outbound

This paper cites Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.590916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:c1a7d96d732f59abbd19d4eb770afe2cdc7825e8480463521e8cdc6bcdda4924

Observation eeb102f7-ccbe-4efe-9ff6-0a01b33774c4 · outbound

This paper cites Emerging Synergies Between Large Language Models and Machine Learning in Ecommerce Recommendations.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Emerging Synergies Between Large Language Models and Machine Learning in Ecommerce Recommendations

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.049022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:4226a75bed42cd7df53d603c0fa48152241d4d7c82d767f65a0af2f83d5ab6dc

Observation c1a9acb3-51a9-4a54-8acc-a199d0c57ce1 · outbound

This paper cites GeoReasoner: Reasoning On Geospatially Grounded Context For Natural Language Understanding.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection GeoReasoner: Reasoning On Geospatially Grounded Context For Natural Language Understanding

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.693768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:b062e680ccec21d2c402b862a501aeac7723138406cef5e6c16debfca41e9f2a

Observation ae8143ff-b337-457f-ac72-5a7d537990fe · outbound

This paper cites Urbanclip: Learning text-enhanced urban region profiling with contrastive language-image pretraining from the web.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Urbanclip: Learning text-enhanced urban region profiling with contrastive language-image pretraining from the web

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.689996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:f1bfd2dc7192c833147ef84f337e342c3f714859684c4026803602eabd667b95

Observation 0372a539-9a1d-41d3-9673-a0093cb903f5 · outbound

This paper cites Exploring diverse in-context configurations for image captioning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Exploring diverse in-context configurations for image captioning

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.694605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:fd2ae8e4ff6160009f18ce2a7758040bec9c5701a4c56c6824c5a275fe4b84d1

Observation 17c4ad5f-dc70-4c37-ad5e-3797dcbb1393 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.906671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:2e64b02e0ba9becf26425e1e6e9a952f48ec79f34003b164d4194e1d3e4e8e84

Observation 2bab7c3f-75c0-4163-9282-e0ca44687f23 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Yi: Open Foundation Models by 01.AI

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.686872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:ddda5d1a1254698bf81a3a790974fa2c2a60cea639f7ea4e4f5cc726ea8719ed

Observation 7dd62501-1f2d-42ad-929b-c34c289c1fb3 · outbound

This paper cites Large language model as attributed training data generator: A tale of diversity and bias.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large language model as attributed training data generator: A tale of diversity and bias

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.699491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:5dda274a5cdfdb9226cbd7943487cdaafa521732c86b78e378bf660ee6698bcd

Observation fdf4f964-e45d-4c93-ad69-829831adaa9a · outbound

This paper cites MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.818351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:3c506e5330e98170b3e8a2daf6069502216489eb595705a4677991b668b16506

Observation 1fd7739d-d3c1-447c-9fb2-ef9a29d4363c · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.708257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:ec959a161ff46ffaf1b2e330ef44924c177b654f5e5d8439fe6c59a9fca7c630

Observation 60de8a06-b8b0-4a34-a59e-75de3c30507a · outbound

This paper cites A Survey of Large Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection A Survey of Large Language Models

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:13:24.620935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:1c9f31e52653c59761c13d9d054aa3811b326969c12586bfc941e71fbea85dbd

Observation 26e69ea4-d823-4e60-88c0-5da7588b24d4 · outbound

This paper cites Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.681459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:5f1e9d80e12770637128d16522c114073163336c29aab92052f18505ca2d0afc

Observation d7ccc97d-03ac-4148-a9a0-67fb8b57f831 · outbound

This paper cites UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross-Domain Adaptation.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross-Domain Adaptation

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.703873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:a4b0c498e348090c82593981509255c52b6fb45b35b70aacc2d86efafd2718ca

Observation 6a4c28c4-ad13-4c28-8dfe-f0bb8de28cef · outbound

This paper cites Mathscape: Evaluating mllms in multimodal math scenarios through a hierarchical benchmark.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Mathscape: Evaluating mllms in multimodal math scenarios through a hierarchical benchmark

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.864286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:191f0d085194912f464a3fa766d2104f495629e0dddf63b74c95fc28278a43ad

Observation 774c5a7e-bd97-4aff-a820-5a3e8933d4cc · outbound

This paper cites Large Language Model for Participatory Urban Planning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large Language Model for Participatory Urban Planning

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.777188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d258992e2586c3776611051eb71c3701c649ca6cd7a286b4c862b7530253acfc

Observation 6be2948c-6684-4b74-9e22-fbe581ffa754 · outbound

This paper cites Large language models for information retrieval: A survey.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Large language models for information retrieval: A survey

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.037969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:47ea86ee38b812e46c5c028d3b9606f5e7870e42258a63da787227b7e5f5e0b0

Observation 882d7d89-01da-45cb-8219-31bb3569df79 · outbound

This paper cites Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:25.043471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d5ce1ed82156eeeaed57ce169b5ff91510b583124b1c1e665944b7fc16d62375

Observation eb14ca03-8ff0-4d6f-988c-96131b7302d4 · outbound

This paper cites Deep learning for cross-domain data fusion in urban computing: Taxonomy, advances, and outlook.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Deep learning for cross-domain data fusion in urban computing: Taxonomy, advances, and outlook

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.817974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:bc62106d46ceaef9652818b0982bbff3bbea72ff890d51b8a325dbef87eb8540

Observation 0ce46bd0-3a45-4866-b6a3-476d24ef4787 · outbound

This paper cites Object detection in 20 years: A survey.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Object detection in 20 years: A survey

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.829449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:aa709f1a05219ab0276d502920ac2ca5544891090daefe81f236dfa41bc90c39

Observation 6892dd8d-ae52-453f-bd0a-45fa91b59f44 · outbound

This paper cites write newline.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection write newline

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.774675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:98d12d8b00511b7d7cc1b53988af256d2bcce998320f0ac0d814dd390c54f668

Observation 2229d2cb-fbcf-4124-8669-856d25a312ec · outbound

This paper cites @esa (Ref.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection @esa (Ref

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:13:25.754450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:04be08052d322d566ea222b1aa14fe05a418424ca4d98f8e95c57626e17c4c04

Observation b4822c65-1ae0-4888-937b-1572a7a59bd0 · outbound

This paper cites an unresolved cited work.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-05-23T20:13:25.783042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:fe2a53fb06ca8453c48fd6829c678d13f5898632762f130d278f54520f74b84c

Observation 0393e8f1-3007-41c8-a9e9-5de496b1dab9 · outbound

This paper cites an unresolved cited work.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-05-23T20:13:25.814025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:d2d492e8f9ff9551d9cfb56772ca6f684377dd0a3d26a51b764fc3e60e6905c2

Pith citing papers

Observation 0eabe05b-56f9-4448-8a95-946e31910d21 · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection

Reference 224

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:32:32.716530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:e8f0f8002f329a9e0bac7160e2c0ad520d8f23d38408767f7136fbb39afc2cc1