Pith. sign in

Paper Citation Record · LEDGER

Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2403.00231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.00231 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:37.086555Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:48:56.198380Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3a50e86b-4986-4c4a-b5cd-c8852fac54da · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:05:03.916736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:d7da99e458a4687af1c4ace000f712aa1c969647a1f05d6ccc386c32ffc4d0ac

Observation 60f2322e-4ed9-4fe8-ad74-0d9acc528876 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:31.947656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:543b2f5c0bbdca21007134ae365e310b9ac4279f6fbd28ddfa07848866c197e4

Observation d7c5483a-670c-4b63-bdef-68f7bf107a24 · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.783118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:a00caf5fdbc246f72aba6722b2896453825a27d3c46b855a6dedb11fbed735df

Observation 27fa901c-1d2b-47ab-b17f-7e3e7b658d12 · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.148867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:ad28f0e760840561d5416a751e28109640104b63f4080ae5521ae39146f707fb

Observation 763d7f88-6794-437d-8ca5-263415c11c38 · inbound

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices cites this paper.

BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T19:33:00.770072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:33:00.770072Z digest=sha256:5e69b36111f73e2dece1a442253b468fcadd93520af61afad6024a96e17ec150

Observation 601ddb87-e901-45fb-aa69-054dc81ebdb7 · inbound

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? cites this paper.

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation? Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T23:36:24.718590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:36:24.718590Z digest=sha256:5fd09b5f35582129c3292a7e0bfe702f14b2932601c53cb2bd7ba8d977103bf9

Observation d8d921e0-e543-4d53-8b04-811b89fa6e3d · inbound

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios cites this paper.

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:29:59.953617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:29:59.953617Z digest=sha256:fabed7fb75551c0929b8a4005d18a5283427c6917a4e5f47c67818f7796c384c

Observation 53b14f80-db4f-486f-b7ba-a97b6c857c51 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.133266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:578aa8375d2054ed6d957bfba12ea4c78a5bf1156a42d1eaaa945a49d053c8a4

Observation f7fd042b-5362-4d7c-b127-ea1c40380f22 · inbound

Chimera: Improving Generalist Model with Domain-Specific Experts cites this paper.

Chimera: Improving Generalist Model with Domain-Specific Experts Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:13:47.564813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:13:47.564813Z digest=sha256:990072b25597444510f3a2008c4edb420fdda679e53bd5bb8dfb807cdf2cec12

Observation 8c7cbd6e-9714-4531-9133-d38be1457413 · inbound

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM cites this paper.

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:16.804403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:16.804403Z digest=sha256:d4a504206a3fda7831b9931863fe134cc89d74cd913ad7212a727eaa87920f39

Observation f70a4bce-d6c9-4da7-acce-b9facd44ab2f · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:23.926362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:6665db2c0ed24aeac000f27aef49bb0e628d16f1645fd3225a257381c0d95344

Observation 9527b3a2-c3fc-4e2a-8f70-92a477a19b5b · inbound

Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature cites this paper.

Rethinking Comprehensive Benchmark for Chart Understanding: A Perspective from Scientific Literature Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T18:19:03.880670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:19:03.880670Z digest=sha256:fd9048039d4a4d07b8ee286a06399c19da91170209a9c1d845582671892812f1

Observation c93a2ead-e820-4d5d-94bf-0dac3f7f4884 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.282378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:b807357da1394e908a0c5e4306c446ac41cb55cd29997b5bef4ccbaaf971ef79

Observation d046f113-cf1f-414c-9e9d-f21215c262d6 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.374236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.374236Z digest=sha256:f13a3e6e13d76427109610c3e7f7ad9b166c76c55376411444460a50dc4d50f1

Observation d238fc5d-c011-4afe-acad-2d0c4eec1d98 · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:33.338732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:7107e42727ab80fe130e2b9cac4b491291e002d33f1e667ddf4873800f3b3dff

Observation 238d0a45-84fd-4159-8213-2acd3e8df778 · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.925419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:7f8b53f2f74233653b5a60537d2c4a8442472444415db6b27ed08b126bfbcb6d

Observation 26d54d24-5ff5-4297-9863-0ca74008e9f4 · inbound

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding cites this paper.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 134

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:37.086555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:37.086555Z digest=sha256:1fb93de98a5b33127407d7a937fd2d727d8133297b0f9b2219f9fbc1af58eb86

Observation ebfac55c-0657-459d-915c-6b25fd72339e · inbound

MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision cites this paper.

MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:23.409588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:23.409588Z digest=sha256:1bb3cc19b3233fae8b9f7a71e588c0d7bbe2d91b4b969e0b864aab4dfe4be3a0

Observation 24f3f26e-458f-420b-9245-9692e789d51f · inbound

ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering cites this paper.

ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:18.556271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:54:18.556271Z digest=sha256:2c20508a5032f5ea6f085005e2b78aa0eb95918238017f2666590be1a93d31fe

Observation c8a5e407-34d7-4780-8eaf-88fd4a99f7d5 · inbound

Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers cites this paper.

Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:31:05.650761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:31:05.650761Z digest=sha256:bae4a784c3627e13d299e043176ec4d9dc78fcdc07c3a60bf16c21dc738eaf2c

Observation 2181e933-b798-4bb5-957a-695d7eefdb49 · inbound

FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models cites this paper.

FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:52:14.564955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:52:14.564955Z digest=sha256:77f5c12704f455dd913ba5dedccebc3438f0d8e4e0b813faa1131ac423c1ee63

Observation e003bd98-94d9-497e-8c18-a6ab3d9cb463 · inbound

ChartCap: Mitigating Hallucination of Dense Chart Captioning cites this paper.

ChartCap: Mitigating Hallucination of Dense Chart Captioning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:43:36.000015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:43:36.000015Z digest=sha256:72a0546e44d8ffeef9d8118f6a11a78bd45c2bf9f07c58d6661f42972966d89e

Observation 871df512-c395-40d8-bb72-368936f975b7 · inbound

Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant cites this paper.

Adaptive Sparse Softmax: An Effective and Efficient Softmax Variant Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:41:37.990911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:41:37.990911Z digest=sha256:7a81d8378a66ffbacc548ea66ef904e1d4437f9f689a3f96899f973b8379e72c

Observation 5f894236-9082-4dc2-b385-e47f7738b3c1 · inbound

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? cites this paper.

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:35:21.414530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:35:21.414530Z digest=sha256:3bdf236f2bf8179d9ddde3058dec72f7fde2142e2395427493e934b52a69efc5

Observation 7da3b7e4-3912-42b7-8476-7b73c77741b9 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:46.452098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:46.452098Z digest=sha256:22e28b2c5ce89b26ed1ee88b5b1e8f9b339f5d6e0c65caac95ad08f229d236e3

Observation 3eface75-9228-4622-b38e-64ada7d9de03 · inbound

Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models cites this paper.

Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:02:29.362654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:02:29.362654Z digest=sha256:d7bea2422fe9abe243ba951a472abb79907fb189fe9a2fd8afcc49c8959cd3c1

Observation 54c5659c-cbc6-4c45-b550-49f93447698b · inbound

Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization cites this paper.

Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T19:25:50.051755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:25:50.051755Z digest=sha256:2dbeaed749a5fe046505993d2ecaaa67464409388b86c1bb9f893843f61820e6

Observation d6317614-23f2-40cf-9376-adde0a86f484 · inbound

DeepEyesV2: Toward Agentic Multimodal Model cites this paper.

DeepEyesV2: Toward Agentic Multimodal Model Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.342491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T05:32:29.266583Z digest=sha256:a290353e9f47708025c1122702fc71f93accc271c005ecfbc9f3ccad14ac1a1d

Observation b22e4d8c-5e39-4e47-852b-d0e060a42f9c · inbound

Attention Grounded Enhancement for Visual Document Retrieval cites this paper.

Attention Grounded Enhancement for Visual Document Retrieval Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.235200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:53:15.920563Z digest=sha256:04d09fa32f8b1666623b96550dae3297c75f0c86aab3b9d3ac10268a5ad6462d

Observation 4f52da81-0a9f-4f5a-9d0a-ed1ddf0e6d11 · inbound

Spectral Imbalance Causes Forgetting in Low-Rank Continual Adaptation cites this paper.

Spectral Imbalance Causes Forgetting in Low-Rank Continual Adaptation Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T06:04:26.757824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:04:26.757824Z digest=sha256:0023a397f4b8163f37bf08803e1b990a68161bff4e86ed751dcca64c20c658c8

Observation 5ab9584a-fef9-4d1f-ba7b-4957d8e3abe5 · inbound

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models cites this paper.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.177074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:f6ec34039cfa463de5477c1e3c88431ce4af1e1d5d3d0e99a7656f2d6a7796de

Observation a57c095b-5aa7-4e33-b881-eb084d0b059e · inbound

SciFigDetect: A Benchmark for AI-Generated Scientific Figure Detection cites this paper.

SciFigDetect: A Benchmark for AI-Generated Scientific Figure Detection Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:25:57.771036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:39:49.848699Z digest=sha256:52ff8ad805f3a44d38624e9c76704f961efaa8831f90c637631fc09e4a02b4c5

Observation 2cf12307-a70e-41da-8af0-c36a875a6a96 · inbound

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs cites this paper.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:54.443885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:1120158bb6edfe204d7b5d12f5f08d75e37b194261387f7ffde90319187b7746

Observation 1142182e-cf02-4766-a44b-f99b9fedb515 · inbound

Perceptual Flow Network for Visually Grounded Reasoning cites this paper.

Perceptual Flow Network for Visually Grounded Reasoning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.923328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:ac4d88fb0f259601d4937a2c6116f8bee316ce7132b3f038371451398922546b

Observation a6ff8011-8189-494e-a938-9a5a1407433a · inbound

Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models cites this paper.

Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:35:46.686843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T20:56:38.324342Z digest=sha256:01a2c3565335e68d9ec2ca989955c56c38b3fb67b9919387fb35ef2aaa8e200a

Observation 056fba95-f8ea-4e27-99e0-371d26e3b8d1 · inbound

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning cites this paper.

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:18:05.218966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T06:16:47.650748Z digest=sha256:c37a8808794c6121eea8f29b2f5d98894aa4b49a152e690b8aed6dd87d059fcd

Observation c5eaa9fa-2347-4767-9243-568217237c48 · inbound

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models cites this paper.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.507568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:b64205b1ba420ac7dfe7c9b0f5e1dfd9dcd3c21a5c473e909aa3afda6474ad41

Observation 8d972fdb-c888-4348-b7f9-aeb63c740a87 · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:48.431141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:ceafeb421e44af84676830eb2b5a71801ec9e71685cae9bb31ba359ce38c45eb

Observation 28741f6f-23d8-48ae-b58a-05a2601e5872 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.199914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:f92d276ffefe050763996837892eb0369830a7cb10f913e308224586488578fd

Observation 58ba945a-b848-45bf-bee5-0fa8910d1de3 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 168

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:b0d95f87a1cb796de74ae31cf4b1e1adfddf394ddb7b4f359d7a61504a9e27ff

Observation c4c2cf58-85e1-4a9c-9d72-cfc0dda00e0d · inbound

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis cites this paper.

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:26.025156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:13:26.025156Z digest=sha256:57893fce3feeb4b6d2fbc8da3882e1b40b6959ece97fda892aac54fbd3dfd961

Observation 03777384-0678-4999-9811-655003dc4633 · inbound

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis cites this paper.

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T04:31:25.731380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:31:25.731380Z digest=sha256:7159fd986498b82c5037be8f6bb5c374c3d934dcecf8b0551c6092a1977f97d9