Pith. sign in

Paper Citation Record · LEDGER

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling

As of 7 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2507.06183.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06183 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:13:26.851065Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:31:15.525331Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:25:52.598353Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6da5d4b8-7592-4a38-9053-b9b525308c80 · outbound

This paper cites Qwen2.5-VL Technical Report.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.132881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.132881Z digest=sha256:003347a907201472d6ad510e24ebf075edd20b38a776267d462326267386ef6b

Observation 5d27ca47-af39-4d2e-8e17-8919808a8380 · outbound

This paper cites an unresolved cited work.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:13:27.873009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:13:24.213162Z digest=sha256:9a23a3dc8ade589fcddf1e07aee5cf71f0eabb875b960bbc7eeadd20a99716d5

Observation 19b79cb6-d350-4e1a-b55f-207d2b621769 · outbound

This paper cites Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.333743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.333743Z digest=sha256:4485e959c27ade2b385bad1b3fb3e9740adeb6e5253181d9b0334cfdea031787

Observation 651599da-b75b-4bf5-a8a6-0d9e8928bf6b · outbound

This paper cites ChartLlama: A Multimodal LLM for Chart Understanding and Generation.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.426698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.426698Z digest=sha256:b090cf0bcd0a9f516d16ffa83a3db440737561c6973e91da533a1fd351e1772c

Observation e3cac431-4703-4012-a7d0-1aba6291043e · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling LoRA: Low-Rank Adaptation of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.545409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.545409Z digest=sha256:d0af1b486b9aa6d046142e8801644dea63fc532391c5c276dd4069b72cca2dc6

Observation 88c3bd8a-2ae0-4809-8960-838b84349693 · outbound

This paper cites Farhan Ishmam, Md.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Farhan Ishmam, Md

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.670393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.670393Z digest=sha256:b4abf73149773edd9aabaf322c1946bce2a79d7254f3b03b87be640df5f49b2d

Observation 38dfd1a9-4b6d-44c1-b93d-2e265a96ae96 · outbound

This paper cites A Comprehensive Survey on Visual Question Answering Datasets and Algorithms.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling A Comprehensive Survey on Visual Question Answering Datasets and Algorithms

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.765236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.765236Z digest=sha256:3096a59b336f5d84d174d8cf3f20aacf116401df1ee2a14aaa53e16fc898474a

Observation eed876b4-30a5-41f5-aa5c-ebc3e83ca83d · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.886671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.886671Z digest=sha256:e6628a69360a7c4ed8ed7ea5a03df0914d09b35b35b2727a955e0fa32c5d4e61

Observation ff1b53f2-dcc8-4d63-bc0f-25b7544b3dcb · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:24.964823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:24.964823Z digest=sha256:3541f4c0aaabbda76259228c7b5b63ebf9a30c61b5f4da68b396eaaa5c695380

Observation b4289eba-bdfa-4cda-9eb2-9d3971879692 · outbound

This paper cites UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.075905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.075905Z digest=sha256:d8ac3904148bc3c246a795265968ef9e27663dbec0fb881ab28fcbdf17d495d0

Observation 6baaecf6-9a99-4e9e-94b2-f2c15aff8955 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.111434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.111434Z digest=sha256:d34e70f7d840ca0613330f57c5d5158c8304afea21ae5643e4d655059e7f293d

Observation d31e54cf-6ef2-420c-9a96-2ed9844b3045 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.235115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.235115Z digest=sha256:fddd81893833d536c12f9e193c0e9511e30f8350f4d983d38e214f0e985f80a8

Observation 176e75ca-64bb-44ce-bbe3-d5ae04c98d91 · outbound

This paper cites SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.301379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.301379Z digest=sha256:feef04761602982b4ee09a8b5085fd18f57d1d9943a96512f2af987c03c9e058

Observation 10d6402f-023e-44e9-8f6c-cbf7789d6985 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.320183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.320183Z digest=sha256:1370431a8ec6808c049bb3c2a2c976af1d1902040aee40401c3918aabc59a526

Observation a293879e-2f35-481d-8f63-f527c0635758 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.450751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.450751Z digest=sha256:f4a9cfdf2bc91073c3c6718a5dc2ec0cc55e0ee0f1c23ce86e6954043c3b598d

Observation de13ff66-f632-4326-be44-0280c4f4aa6b · outbound

This paper cites Exploring Rewriting Approaches for Different Conversational Tasks.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Exploring Rewriting Approaches for Different Conversational Tasks

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:13:27.260471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:13:25.559058Z digest=sha256:bf9d3f16ac907b0aadabd1e5ac99ea0222d7d1357807d9b11f55c195ceceee50

Observation 007578da-c4f2-4f92-89fe-3b8418d904da · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.672059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.672059Z digest=sha256:fa80730df2b0631ae4d283bd3d82d697404141f80c41664c5d0eefdba2aa895e

Observation 7f3dd424-524f-45f1-b19c-3ed16265bb52 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:25.782023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:25.782023Z digest=sha256:8492fa59f0e7001e0a9039ad7bc7a24aa0508d0f7eb69116729d5e7c205280c1

Observation 2c086010-ccc6-4df4-bad9-24af0fb845d8 · outbound

This paper cites an unresolved cited work.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:13:27.670426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:13:25.915671Z digest=sha256:f877263eaf6c1f08649509ee196ada2ac68bdaa3522622f74a176d02f526e457

Observation c2e00931-b974-435b-9879-eb8708346140 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:26.028106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:26.028106Z digest=sha256:54b30b6a1d44ea17e77f1dc63c1ad5574115ba4507941386b3b074edcfac3782

Observation a23e7867-472e-4edc-8f2a-5adf0d707c99 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:26.141905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:26.141905Z digest=sha256:5a149e6bca6bed2e3deec031616fa1e08724e01cf4da9116706abf55d8080cc2

Observation da7a07fe-dd6e-4d25-8be5-d100c8cd1546 · outbound

This paper cites SPRI: Aligning Large Language Models with Context-Situated Principles.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling SPRI: Aligning Large Language Models with Context-Situated Principles

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:13:27.061026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T19:13:26.256671Z digest=sha256:144a83a8a281862464c2c9228c04131fdfbf57d4e809073b18925999829ccca0

Observation 9766eb6b-7fdc-418d-82cd-7f5121d6651e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:26.374962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:26.374962Z digest=sha256:d092325bb02a6215c3b291ba5007f825119b42ca4c3999ca1c66f7c96983dea3

Observation f79664c0-fd49-4a2b-9e86-14d2de44c813 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:26.528077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:26.528077Z digest=sha256:09d6f4d2ab0b7ae8ed2fd0e64108c04251d13ee35cf4554ce3d9342f1b32dfa0

Observation 4232198f-c475-4031-a872-c21921279274 · outbound

This paper cites URL: " 'urlintro :=.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling URL: " 'urlintro :=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:26.732433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:26.732433Z digest=sha256:121d8687b76699e653806c31b2d9cb57eb83561e99786825eed1799f232379bb

Observation a0195551-95c0-4e78-b3de-2f939d32a89e · outbound

This paper cites write newline.

Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling write newline

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:13:26.851065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:13:26.851065Z digest=sha256:ae739e363d642322da5f3dc9ff88fa1b7aa74546f526b9e4ac1f706c9040ba21

Pith citing papers

Observation c81f5110-650b-476d-88aa-b2734b237dd2 · inbound

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents cites this paper.

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:25:52.617911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:31:15.525331Z digest=sha256:172c6c29c204aa7c190fd0356572c253fa82b40d19462c8a795b10d5e6acd7e7