Pith. sign in

Paper Citation Record · LEDGER

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models

As of 4 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2604.08815.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.08815 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:56:16.568462Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T00:57:35.205227Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact10
  • verified fuzzy11
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc63a1a0-852a-4fd6-ae75-47a260641b51 · outbound

This paper cites A comprehensive survey on the trustworthiness of large language models in healthcare.arXiv preprint arXiv:2502.15871, 4.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models A comprehensive survey on the trustworthiness of large language models in healthcare.arXiv preprint arXiv:2502.15871, 4

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:50:59.536057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:4354482586aad0d16b41995a5097728a1bd9779b437b9c09f92e72113ccacd28

Observation 1c908649-5975-412b-9170-496d777f81da · outbound

This paper cites Emerging trends in multi- modal artificial intelligence for clinical decision support: A narrative review.Health Informatics Journal, 31(3): 14604582251366141.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Emerging trends in multi- modal artificial intelligence for clinical decision support: A narrative review.Health Informatics Journal, 31(3): 14604582251366141

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:51:47.816806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:eb91b3d311f13fcb8d97b7afb86579ffb546419652cc45c5c692bc7b3bce0969

Observation 5b925679-dc0c-454d-8a62-898cce802656 · outbound

This paper cites arXiv preprint arXiv:2512.16201 (2025).

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models arXiv preprint arXiv:2512.16201 (2025)

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:50:59.527586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:d69fedecf62c148a6990f6de80f1b95fbe8b1e3215880293a06ee3ffd1399db9

Observation b9440404-5816-488d-bf83-33a8dcad6d2d · outbound

This paper cites Medical phrase ground- ing with region-phrase context contrastive alignment.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Medical phrase ground- ing with region-phrase context contrastive alignment

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:51:47.827769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:a40f86c39dcbb7e379a7171e696484edcecd1957398b74ebe7f982e75e1f74f0

Observation eced046a-5ff4-4790-8a8c-6e5e7b1257c2 · outbound

This paper cites Multimodal computing in healthcare: Enhancing clinical decision-making through data fusion.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Multimodal computing in healthcare: Enhancing clinical decision-making through data fusion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:51:47.824923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:38cc8191f971a8639a392056c264a18bf8757f9ac4e3a3e44db49bbfd4b5e645

Observation 173760c1-372f-4ea0-9cb7-64d97a654e12 · outbound

This paper cites Med-glip: Advancing medical language-image pre-training with large-scale grounded dataset.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Med-glip: Advancing medical language-image pre-training with large-scale grounded dataset

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:50:59.375400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:cea9f0d8d4ffc75157e55aaeb83163c75d819ac2b034b17854ef3eea2d154c59

Observation dc371560-6824-4478-a754-098042cf0f08 · outbound

This paper cites an unresolved cited work.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-17T12:51:47.822374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:bcc2b74ff28fd3593cedb9eb4711d184f6c608513a411db679b29318b69c8a39

Observation e92f4b52-4b93-4d15-9da2-98f44ec560e7 · outbound

This paper cites Vision-language mod- els for medical report generation and visual question answer- ing: A review.Frontiers in artificial intelligence, 7:1430984.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Vision-language mod- els for medical report generation and visual question answer- ing: A review.Frontiers in artificial intelligence, 7:1430984

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:51:47.820075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:448a67fda0ec3edc3bc02530eb8b53ba8a065d9151ecbf982f10906318021d1c

Observation edf6cd4c-2d94-41ff-a379-4b7c392eca3d · outbound

This paper cites A survey on hal- lucination in large language models: Principles, taxonomy, challenges, and open questions.ACM Transactions on Infor- mation Systems, 43(2):1–55.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models A survey on hal- lucination in large language models: Principles, taxonomy, challenges, and open questions.ACM Transactions on Infor- mation Systems, 43(2):1–55

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:51:47.836956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:5387cced4c128f7135915d5c66024fb2f48a21f094e26842775d42fa7ff4ab49

Observation 8df619ee-4d27-4366-b392-c727962f524f · outbound

This paper cites Seeing the trees for the forest: rethinking weakly-supervised medical visual grounding.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Seeing the trees for the forest: rethinking weakly-supervised medical visual grounding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:51:47.833884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:f6b8a58e98b4904ecf04cda4894b21b98d522f256fb19bb694d314a62592ece9

Observation 1e8f7d9e-c029-4fcf-b863-72bcd95b5ad5 · outbound

This paper cites Vision Language Models in Medicine.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Vision Language Models in Medicine

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:50:59.513407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:2929070328a17cccb7b7463d73281473b828faee7ee1c3b9147154d7b9f5f6dd

Observation 017301d7-5570-43dc-89e1-3e77e569ed9f · outbound

This paper cites arXiv preprint arXiv:2503.13939 , year=.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models arXiv preprint arXiv:2503.13939 , year=

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:50:59.482224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:404cde155f45e20972d7c59b9894a13d37f2441234ebf1cbce6343aea828a27b

Observation e9d1729e-3ffc-4be9-a435-96bd50d67474 · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Llava-med: Training a large language-and-vision assistant for biomedicine in one day

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:51:47.848542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:95a4cf370a4589305c2dffea18aa965ca85ca4204f44f0421c4c5b6874377369

Observation 2f0c8e08-a9a6-4b9b-a11b-94ccc2b771e7 · outbound

This paper cites AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:50:59.420326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:7647c79e6b5ec7c16c2171ae239e841b5c0911eecb1cdd0faa68e8af69169623

Observation 28a627d3-143e-4e37-a734-ab8b490f6587 · outbound

This paper cites From Classical Machine Learning to Emerging Foundation Models: Review on Multimodal Data Integration for Cancer Research.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models From Classical Machine Learning to Emerging Foundation Models: Review on Multimodal Data Integration for Cancer Research

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:50:59.436766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:c1d130a0c7c0bcac93bf2d477baeaf6fcf454dbbb523564a938f26af8bdc0505

Observation fbab2c31-e2f0-4f72-956c-9b2503d8fe82 · outbound

This paper cites The future of radiology: The path towards mul- timodal ai and superdiagnostics.European Journal of Radi- ology Artificial Intelligence, 2:100014.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models The future of radiology: The path towards mul- timodal ai and superdiagnostics.European Journal of Radi- ology Artificial Intelligence, 2:100014

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:51:47.840189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:edf2a8114ee92fb03757c1b2b34457a25784a63178dd6c5628bf5b5024f77b9a

Observation 5e3cc0c7-f40d-49eb-aef1-62c18e545908 · outbound

This paper cites Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:50:59.555057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:9abc846fd1c9f0a70b8f1784ddf9d476606873d590b5ccfd2d75ba8511dee979

Observation 12c276d5-f2b2-4d71-968c-a26fdd414c91 · outbound

This paper cites Who is re- sponsible? the data, models, users or regulations? a compre- hensive survey on responsible generative ai for a sustainable future.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Who is re- sponsible? the data, models, users or regulations? a compre- hensive survey on responsible generative ai for a sustainable future

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:50:59.547262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:73e1abf1dd89b1396439e319726f61fb73431684ebf5a34b4e51186c7e3a261f

Observation c4ab69ad-30dc-4ba6-a17b-f4aba70745ae · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Process- ing Systems (NeurIPS).

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Process- ing Systems (NeurIPS)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:51:47.830775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:620d30ba411a18a5d8c27e6ed0922b5f6afe3a9c02efce6423954cda4f555922

Observation e3a73f7c-0da0-46d0-807e-31f0fd06ac75 · outbound

This paper cites Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:50:59.365355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:58588b7bfc07a5b5abe0a0f5577c5a9716e0e3563142b5ffa521adb6bfcf1b3d

Observation b0deed74-a62a-4262-88ec-2aa75474bd77 · outbound

This paper cites Multimodal Large Language Models for Medicine: A Comprehensive Survey.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Multimodal Large Language Models for Medicine: A Comprehensive Survey

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:50:59.382029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:62540bb23c8a47b7ac16d97da8bfa884faea7082f9f98f96b6e57ecc9f54c9b6

Observation 4bd6ff36-a363-424b-9b36-9b4d93bd616e · outbound

This paper cites Beyond accuracy: Evaluating visual grounding in multimodal medical reasoning.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Beyond accuracy: Evaluating visual grounding in multimodal medical reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:50:59.497679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:975c4151e6ea734dee787658af137c5ed2a34175b3f4d383afdd73ce221186a0

Observation 6ef0ace3-43a2-44ad-97c5-10eebf60eead · outbound

This paper cites BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:50:59.452214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:11a9aab2cd972b86f99824a0c0e15e2b4327f4b5b1c645f62960b32585da5cbb

Observation 3652f593-4d0d-4fb6-9fbd-ca12ce27aa59 · outbound

This paper cites Can we trust ai doctors? a survey of medical hallucination in large language and vision-language models.Findings of the Asso- ciation for Computational Linguistics (ACL).

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Can we trust ai doctors? a survey of medical hallucination in large language and vision-language models.Findings of the Asso- ciation for Computational Linguistics (ACL)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:51:47.846215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:c7577ef69cd0fcf7ae565ed9d3266692dad7f36caec5afb3df64fbb7c7df32bb

Observation 170e7483-3a73-4a33-870d-ce992c8f523a · outbound

This paper cites Uncertainty-aware medical diagnostic phrase identification and grounding.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models Uncertainty-aware medical diagnostic phrase identification and grounding.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T12:51:47.843457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:56:16.568462Z digest=sha256:86bc70f15ab90abb340ded7786dd83601882c16d517bdc0f722771f505b00dd4

Pith citing papers

Observation 8bfc4113-58e2-4ecc-8dbd-67a041285afb · inbound

Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs cites this paper.

Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs Towards Responsible Multimodal Medical Reasoning via Context-Aligned Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T00:57:35.205227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T00:57:35.205227Z digest=sha256:13c2d6bf78f5657084edecd0eb8077fa7585feb821afd7029fd155ee4634b7c0