Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2403.02469.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.02469 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:52:01.605452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T18:22:46.845434Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ca45e851-3e4e-465e-ae15-bbafe474923a · inbound

A Survey of Medical Vision-and-Language Applications and Their Techniques cites this paper.

A Survey of Medical Vision-and-Language Applications and Their Techniques Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T17:52:01.605452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:52:01.605452Z digest=sha256:e7a87041448d2bdb6f821d06b208c31471285ba91b71ca5661f3ef90a096b2f9

Observation 342bdfd6-2dc1-433a-a554-01dee4791942 · inbound

VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge cites this paper.

VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:08:41.740755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:08:41.740755Z digest=sha256:06b64582d8fc195f21893474de9367a658ded7897c671346ddd38fe480808516

Observation 5a290775-cc37-4fae-8a0e-a09f6a5a502a · inbound

Deep Learning-Based Noninvasive Screening of Type 2 Diabetes with Chest X-ray Images and Electronic Health Records cites this paper.

Deep Learning-Based Noninvasive Screening of Type 2 Diabetes with Chest X-ray Images and Electronic Health Records Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:30:50.204633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:30:50.204633Z digest=sha256:9cb929b90cb384a605a9c38cb55d8f846d4c8ac690137d7be86137b4448b1564

Observation 01988932-b5e8-44d9-bded-93c62acbeb64 · inbound

More is Less? A Simulation-Based Approach to Dynamic Interactions between Biases in Multimodal Models cites this paper.

More is Less? A Simulation-Based Approach to Dynamic Interactions between Biases in Multimodal Models Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:59.066469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:29:59.066469Z digest=sha256:fb2d50cceb93c98fc0da66c3ff2463175ed81cf9b79dad59ecd926c18523ca03

Observation 61f1a8b0-5ad0-42c2-8c50-f1bc12f5e64c · inbound

StreamingRAG: Real-time Contextual Retrieval and Generation Framework cites this paper.

StreamingRAG: Real-time Contextual Retrieval and Generation Framework Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T15:24:49.329398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:24:49.329398Z digest=sha256:b54ec8d787412d74630b37d735021cd87aaf2bf1e1b87600db512c036ec5a0bb

Observation 88a74b03-16fe-44ea-9af8-93e70775e2a6 · inbound

Membership Inference Attacks Against Vision-Language Models cites this paper.

Membership Inference Attacks Against Vision-Language Models Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:01:10.860277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:01:10.860277Z digest=sha256:cdb996097109790a5d6edebb15b92ee76f655b0df368c9d916ba8cb68e7804fb

Observation 9cd5f3d5-eaab-4130-8cbc-14b0bebc64d6 · inbound

DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis? cites this paper.

DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis? Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:07.407511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:35:07.407511Z digest=sha256:3c3e314fba9780c2f82afe0a4fe307a7c69b4aea7c55e05ed2d2652de1cf67bf

Observation 6eca3ffc-47e8-4153-aeaf-5e1b8e09c441 · inbound

Analysis of Blood Report Images Using General Purpose Vision-Language Models cites this paper.

Analysis of Blood Report Images Using General Purpose Vision-Language Models Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:22:46.849703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T18:22:10.226073Z digest=sha256:0a3a8b450a6e2156658c49bf1811280179a7f32f30b3054a593f36dab68e6e29

Observation c3f916ae-5c86-4f4a-91d2-83e5e98408f9 · inbound

Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework cites this paper.

Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:45:36.968527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T01:45:29.400398Z digest=sha256:2d4f008a130c8a412ee26b91bb0277f822301904d1ffc63cbcd97b0565a3e20e

Observation 8a39552c-019c-4f81-acd3-883678661fbc · inbound

DentiAsk: A VQA Benchmark for Multimodal Reasoning in Panoramic Dental Radiographs cites this paper.

DentiAsk: A VQA Benchmark for Multimodal Reasoning in Panoramic Dental Radiographs Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T07:15:40.242481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T07:15:40.242481Z digest=sha256:845e270328ef0560d5f789c81abf6dc0b097dc8dd11827c136cd28e5707e401d

Observation 91bb6565-e0dc-45f2-ac94-45ccabee905a · inbound

DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization cites this paper.

DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:40.407120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:40.407120Z digest=sha256:45323733ce86349f2c9d8bd0000fee2b7c002d28056ee60bc20ee4b46ccef78c