Pith. sign in

Paper Citation Record · LEDGER

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2507.14544.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14544 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:57:51.244796Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:57:31.114754Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed17e8bd-9b45-4cc5-b20f-4ed2d5ae7b98 · outbound

This paper cites Vqa: Visual question an- swering.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Vqa: Visual question an- swering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.715301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.178667Z digest=sha256:ec899366fc1569d09990c867b958a2a3f0229d8ebe4769843d85f763c2dfab77

Observation 917e16fc-b3fd-496a-a4c1-591ffb87165f · outbound

This paper cites Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:51.182017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:51.182017Z digest=sha256:efec22afa4d5363006c8447ed2625b3799ecfd6f636fe13ad827927f6e61437f

Observation a5c61c48-ebcc-41e9-bced-86d5c1f9d7f1 · outbound

This paper cites Exploratory data analysis.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Exploratory data analysis

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.708624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.185072Z digest=sha256:6c19129a5a426c69fd3df20e3fa5a416da95cf9ac28df306b52b0976f150381b

Observation 19671860-56d1-4ab7-8524-a1604c3096d2 · outbound

This paper cites Mapping medical image-text to a joint space via masked modeling.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Mapping medical image-text to a joint space via masked modeling

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.701691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.187745Z digest=sha256:998ca3f53912c696d901dd50187759752bd867618dc87298331c66a9a275f3b3

Observation 4452c226-e998-4c73-9871-6ea9c39b2f39 · outbound

This paper cites Vision- language transformer and query generation for referring segmentation.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Vision- language transformer and query generation for referring segmentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.694595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.193007Z digest=sha256:21c4aa7205cc3edff38b500f7b37e7770b2c5eb86d83dda5d9748bcc760ed874

Observation 0c7a4cd1-cc89-4c54-b125-36b774af5824 · outbound

This paper cites an unresolved cited work.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:57:51.687329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.195959Z digest=sha256:8aa9c4f06d0a2db266f6e580737a4bb853054ca6826618233ebab94a06171c83

Observation 1722c8a3-45d2-4679-8dc0-42c7e77181c4 · outbound

This paper cites Bridging Multimedia Modalities: Enhanced Multimodal AI Understanding and Intelligent Agents.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Bridging Multimedia Modalities: Enhanced Multimodal AI Understanding and Intelligent Agents

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.679840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.198637Z digest=sha256:6e5c83434b48f83d634a9b20929e5715fb2ea888db395b869cc36336f9d33927

Observation f0beebc1-9999-4b7e-97b3-738c0a256f0c · outbound

This paper cites Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:51.206063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:51.206063Z digest=sha256:3049cd1076672ed6bddd1bf9d19507cb70e99b27c2c6d14ed5dc6e1f40a1b19a

Observation 86233827-70c9-4c20-85fb-799eb5c066fa · outbound

This paper cites Hicks, Vajira Thambawita, P ˚ al Halvorsen, and Michael A.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Hicks, Vajira Thambawita, P ˚ al Halvorsen, and Michael A

Reference 9

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:57:51.376598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.208731Z digest=sha256:ea9c26a769abaa41d02b0949f7dfa8e6ca9f974b624b24c367d4cd8b0a444f4f

Observation e741dce1-0bcc-4023-93c1-4a08f0304e2c · outbound

This paper cites Making the v in vqa matter: Elevating the role of image under- standing in visual question answering.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Making the v in vqa matter: Elevating the role of image under- standing in visual question answering

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.665240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.211245Z digest=sha256:06624e64c89dece5100fbbc393c66eeac45d93d4e58e56e23f469083376aa940

Observation b96e36bc-40d1-4992-98fc-e17246fb0c95 · outbound

This paper cites Unk-vqa: A dataset and a probe into the abstention ability of multi- modal large models.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unk-vqa: A dataset and a probe into the abstention ability of multi- modal large models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.657327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.213584Z digest=sha256:11630b7b46b6c3e4ab7369e5b8ce6c2a376e86da4cae467e2d9d82988e313071

Observation e42d92b9-b616-42be-a79a-447daed339a4 · outbound

This paper cites Unboxing the black box of attention mech- anisms in remote sensing big data using xai.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unboxing the black box of attention mech- anisms in remote sensing big data using xai

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.648487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.215845Z digest=sha256:ecf4e9d0278fdf139ad7a3eeb7df9cb7aea790b64d023a17ee6e7a988592be07

Observation ad9fce89-40e2-408d-a07f-17520ca1cfe4 · outbound

This paper cites an unresolved cited work.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:57:51.639196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.218218Z digest=sha256:e935519a705ce359e3475427b888e7bc921e6a01b6f959521dc5af23679d12a6

Observation 52fce76b-5efc-4398-a560-1fd2926a21f2 · outbound

This paper cites A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:57:51.314658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.220736Z digest=sha256:136682fb7e1f5d67ab3005f2b9b8042c9b228fb943b837fda4b85c32cf6eb162

Observation eec7c4a0-b5c2-4e0c-9742-c6425111f476 · outbound

This paper cites A dataset of clinically generated visual questions and answers about radiology images.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 A dataset of clinically generated visual questions and answers about radiology images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.629812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.223538Z digest=sha256:80592892590fd6fe5f285acd06b15e2aff39a79304c0774255deac7989148017

Observation ad66af8a-1587-48ad-bde3-2ae49004b792 · outbound

This paper cites Blip: Bootstrap- ping language-image pre-training for unified vision-language understanding and generation.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Blip: Bootstrap- ping language-image pre-training for unified vision-language understanding and generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.620877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.225869Z digest=sha256:34ed15a31de15a3e132e4cf1ebf714fd01d131672ff94cecf4b2c9af1e31caa3

Observation 9ceb7cfb-2943-4faf-b52d-6e4488ce4327 · outbound

This paper cites Contrastive pre-training and representation distillation for medical visual question answering based on radiology images.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Contrastive pre-training and representation distillation for medical visual question answering based on radiology images

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.611629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.228230Z digest=sha256:b65a162f8a5b00867bc232d09567b87f9288337823eb726eaa9e2ccfa70c6d0d

Observation 056a0cdb-ad1e-4a66-b930-761fe3d5ccc5 · outbound

This paper cites Slake: A semantically-labeled knowledge-enhanced dataset for medical vi- sual question answering.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Slake: A semantically-labeled knowledge-enhanced dataset for medical vi- sual question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.602511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.230666Z digest=sha256:4bcc158c7512e85861dbfe1cf750a7c9faa17d5d55d00f6725dd0f0435cd8073

Observation 96e95a63-8207-4aca-8982-c9ecab8c1d1f · outbound

This paper cites Improved base- lines with visual instruction tuning.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Improved base- lines with visual instruction tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.592236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.232938Z digest=sha256:e887882581c2ad34dcfdd82a98f5f55fce639fb2e68cee6276a4089e0ed40efc

Observation 07196891-9335-457c-8edd-c9ce5cd95e30 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:51.235217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:51.235217Z digest=sha256:5c8df79a68e9de7ffce8bce6a62be0e7c501c98f7602eeba8aa7a819556643b8

Observation ec70c0d9-9ccf-41fc-883a-365e7fb63635 · outbound

This paper cites A survey of efficient fine-tuning methods for vision-language models—prompt and adapter.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 A survey of efficient fine-tuning methods for vision-language models—prompt and adapter

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.575920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.237728Z digest=sha256:b22151ebf7640d82a5855a58b98f96aa600ec17e2f4dccd02dcc42f004f07476

Observation a61747b3-7918-4215-8f59-8b2ab5cda85c · outbound

This paper cites Multi-modal concept align- ment pre-training for generative medical visual question answering.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Multi-modal concept align- ment pre-training for generative medical visual question answering

Reference 22

Resolution
verified exact
doi, observed 2026-08-06T15:57:51.271628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.239896Z digest=sha256:89921dabc9f822cf2f013448abc8c27e6b5ad39107e7362207bde01f78b1dcdc

Observation 952f75b6-7fa7-4b3e-90e0-63aea0e300d9 · outbound

This paper cites VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:51.242123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:51.242123Z digest=sha256:57189eec9b118fe904caac1d04b658e52687a0e4ef298d30f61b63029ac4d7c6

Observation 2b306bd7-60f4-46a7-abd9-7848f0f9f7fe · outbound

This paper cites Medical visual question answering via conditional reasoning.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Medical visual question answering via conditional reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:57:51.565319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.244796Z digest=sha256:d77b32691a6c80cfc6635d34616fc9dc55e4527578546756914c5cada942353c

Observation ef7d4920-9ac2-4e48-bbe3-717594c224ec · outbound

This paper cites an unresolved cited work.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unresolved cited work

Reference 699

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:57:51.672540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.201059Z digest=sha256:49e682d899552d7583add02f7cb65fcdaef7f29a2c86d045fc6822e2e5301871

Observation 9f4fee1c-b771-45d2-8480-77092804978e · outbound

This paper cites an unresolved cited work.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unresolved cited work

Reference 2023

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:57:51.473870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.203586Z digest=sha256:704fbe2e7b92c1cdaaf3dbda09882e5557058fffff48804aeb6130fb1bcfdaa4

Observation 031d157a-5c45-47c9-a918-6be118d30dd0 · outbound

This paper cites an unresolved cited work.

Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025 Unresolved cited work

Reference 2024

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T15:57:51.556209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:57:51.190555Z digest=sha256:763e16cac0c4c2e5c484103e28deeb1a32de73d0e08d866de69816021be7efc6

Pith citing papers

Observation db7da7da-3c37-4341-aca2-c6eee986dffc · inbound

Measuring and Improving Complex-Atomic Answer Consistency in Endoscopic VQA cites this paper.

Measuring and Improving Complex-Atomic Answer Consistency in Endoscopic VQA Multimodal AI for Gastrointestinal Diagnostics: Tackling VQA in MEDVQA-GI 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T16:57:31.114754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:57:31.114754Z digest=sha256:b3223865c52b554c03cd144f66acf869620623ede6f3a026ddcb95084eb1b813