Pith. sign in

Paper Citation Record · LEDGER

HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 77 inbound Pith citation observations for arXiv:2406.19280.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.19280 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 77 of 77 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:51.430911Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7253cf51-f030-4c56-8842-6578b2a220e6 · inbound

CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs cites this paper.

CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:04.172031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:04.172031Z digest=sha256:dd8bab0b79a75d2ba60821d32e909b78a21bd498f340a4c7095417fd0d9e1b0d

Observation 00198be9-9255-40dc-a781-d94e84bcdfb4 · inbound

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI cites this paper.

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:16:27.649840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:16:27.649840Z digest=sha256:509d425e5a506d6c70637bdcb2c48dccc7a3fec17abf1f6d6faf51681280aa48

Observation 83739f70-003d-43dd-b32d-3b705eb7fe2a · inbound

On Domain-Adaptive Post-Training for Multimodal Large Language Models cites this paper.

On Domain-Adaptive Post-Training for Multimodal Large Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:56:46.703498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:56:46.703498Z digest=sha256:453f9eae014bc43e2716d11cf4f3db14ca789dc569b2d1d20b8eb4cc18d8c3ab

Observation 6ee5fc8a-b11a-4c47-9a5f-491b582c54de · inbound

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs cites this paper.

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:36:50.113768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T12:36:50.060335Z digest=sha256:e1bda8054f948f923e61fc67644788274ee9358cfc81b0d76f53635418353d15

Observation 895d79ea-704a-479e-8948-f89c959c7064 · inbound

Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging cites this paper.

Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:44.727179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:40:44.727179Z digest=sha256:ea546791865f9a47206d0a66eca00a6fe250a35fc02cbffc1906c81e236563f3

Observation 57ecf7b4-5a5b-4cef-a5ad-e309afef81f9 · inbound

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery cites this paper.

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T18:24:41.782690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:24:41.782690Z digest=sha256:4ba6feae71ba062f059e599f23f2eb8f747959df29299f055d418d50592b6770

Observation d4e072af-8f2f-42b8-90f8-3d6f18fed410 · inbound

From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine cites this paper.

From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-07T22:13:31.837086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:13:31.837086Z digest=sha256:435340a1378e2c8b0da4588c3601a0858e4fff7d18a3459b7ddcfe3cd44505f8

Observation 7815ca95-f077-402d-8f31-5a207dc26b80 · inbound

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation cites this paper.

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:30.364189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:21:30.364189Z digest=sha256:bcb34f1d38ce08df03d3bc4125f09d7d838a58d90d1364968a5b72005c8c067e

Observation 3cc28d30-d7b8-4649-b8f4-24f7a5d66950 · inbound

OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding cites this paper.

OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:51.430911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:46:51.430911Z digest=sha256:1d06e1903cd09b8124f75d795296b4ae20f7d34afdbf7b33c75f8d24d8ecc841

Observation 5c792d77-8aff-41a4-a8b1-48e9d6d2cec6 · inbound

Multimodal Large Language Models for Medicine: A Comprehensive Survey cites this paper.

Multimodal Large Language Models for Medicine: A Comprehensive Survey HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:32:54.016989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:32:54.016989Z digest=sha256:baf4b9bd4135c464e1b219b847ea5e27d7275c6b7441f8a4c607b4de93d8a6f5

Observation c68f42d4-107b-4be0-b3ac-eac6aee055fa · inbound

Keep the General, Inject the Specific: Structured Dialogue Fine-Tuning for Knowledge Injection without Catastrophic Forgetting cites this paper.

Keep the General, Inject the Specific: Structured Dialogue Fine-Tuning for Knowledge Injection without Catastrophic Forgetting HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T06:02:49.092496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:02:49.092496Z digest=sha256:30609c41dc9ceedd68c1a35fc21c8d9eb253bcc94c1abf058be2439936a73062

Observation 0aae5dd9-3380-4448-8881-2bbf1da15a07 · inbound

Ultrasound Report Generation with Multimodal Large Language Models for Standardized Texts cites this paper.

Ultrasound Report Generation with Multimodal Large Language Models for Standardized Texts HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:02:10.384396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:02:10.384396Z digest=sha256:7901c4537775072160e0afcc6e6e254007e6edf82013ff339d44f2cd40af9838

Observation 07975337-5dba-4b77-82b1-1f4658557184 · inbound

MedSG-Bench: A Benchmark for Medical Image Sequences Grounding cites this paper.

MedSG-Bench: A Benchmark for Medical Image Sequences Grounding HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:47.499892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:50:47.499892Z digest=sha256:f499228f1604fd162f0e988e90bc82d470040dfb6b8dc8f18d45d3aa84c64e1c

Observation 3e925468-2865-475f-a3e4-b7789c589afb · inbound

Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning cites this paper.

Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:44.809561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:33:44.809561Z digest=sha256:f13cb508472a4608a38cba2402aabd4d162b47f5ffc41bce20b7d1048d639c98

Observation 4984e00a-9b04-4f94-9a48-9fc79a1f61ca · inbound

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning cites this paper.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.360661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.360661Z digest=sha256:7b72d0d90b878e529e7c80196f0348260c8cbb61727f164b072e512479141a66

Observation 1acf071c-13bb-4a89-b1a7-1a3f015cf456 · inbound

Interpreting Chest X-rays Like a Radiologist: A Benchmark with Clinical Reasoning cites this paper.

Interpreting Chest X-rays Like a Radiologist: A Benchmark with Clinical Reasoning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:57:02.678990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:57:02.678990Z digest=sha256:02c1bf0641ad4e6378699c88e4be577621c9bf43914c2b159742a1e4b820d72c

Observation dd08d31e-90e7-4d0c-8dd7-fb5aee4c4f6b · inbound

DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis? cites this paper.

DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis? HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:06.182529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:35:06.182529Z digest=sha256:d5f98abc49ac30d2c72380c81b9ba2512e696e9ec963fd40b211bbbe2ea0c1ec

Observation c4e31152-36ae-4729-892b-c3b61923bd92 · inbound

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book cites this paper.

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:02.852643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:02.852643Z digest=sha256:1f5ea49bd9231733ede1c698e231c3c553e76c797d1b1b074d3a40845c28528c

Observation 733d3356-3db0-49f4-99cc-eae4d693064b · inbound

Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning cites this paper.

Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:48.640705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:48.640705Z digest=sha256:f61ae0ab6dc2f74a3896a625ed7a63ca195ba46ed4a45555c0145e9491dc99f5

Observation c8d9a6ba-42ad-4486-958d-4bf2f8dcc8ad · inbound

HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding cites this paper.

HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T10:47:15.081327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T10:44:01.880405Z digest=sha256:cb440a4c6eacb9267cdd68ac3166ef8ad4101df58cbbb9fb30661eb7f209c880

Observation a3f3f9be-363b-4006-9834-d829cd011408 · inbound

RARL: Improving Medical VLM Reasoning and Generalization with Reinforcement Learning and LoRA under Data and Hardware Constraints cites this paper.

RARL: Improving Medical VLM Reasoning and Generalization with Reinforcement Learning and LoRA under Data and Hardware Constraints HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:57:22.717492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:57:22.717492Z digest=sha256:8698b5f7723bac2bb8971137a06fcc8d2964bc2481909d5fb1786df2a9ebd79b

Observation 48dcebf1-4bf3-43a9-822f-079ce5e67f00 · inbound

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding cites this paper.

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:47:07.641054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:47:07.641054Z digest=sha256:7390cc64e16d269dfcaf8d1e36e3d9893a7d54618e65411e3acd77a4d79353c8

Observation 8c3e2485-fa62-42fe-9688-a2403bd9e1a3 · inbound

Thought Graph Traversal for Test-time Scaling in Chest X-ray VLLMs cites this paper.

Thought Graph Traversal for Test-time Scaling in Chest X-ray VLLMs HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:17:13.932466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T09:15:46.135084Z digest=sha256:0f438046b188900baaeb83b2799fab847f34c6eb0460285311fbd92555d1791b

Observation 0ec0d310-c99a-4b0b-81d9-446500bb119c · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:17.728304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:17.728304Z digest=sha256:18a29e03b2b65554e1b050d1b8f6cfa8ff1da200eaa4d32a0e72662b215b9616

Observation 592a6bd3-963f-47f0-a361-7739fe31f85f · inbound

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs cites this paper.

Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:08:55.923492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:08:55.923492Z digest=sha256:7b92dd7976ec2aee2e605dd0de5376ffe11199dd467c3bea4de1011f8dfbcc26

Observation 7942a764-7136-4e6a-910c-8649d2c17e1b · inbound

MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration cites this paper.

MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:31:37.237413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:31:37.237413Z digest=sha256:f81a03b786e9c8c4b67f1acd6c83bdb3613b4299f14dc547b9ccf35694a5c4f7

Observation 5829d0ea-3cdc-4841-8729-a2442c5c2caf · inbound

CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale cites this paper.

CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:17.613129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:17.613129Z digest=sha256:29e18deab40e7206d77a72f1dcad6e17a589d4430bf044133c169a1bcf2e92a7

Observation 24cc1d87-4463-494b-9f1d-208e1c2bfc93 · inbound

MIRA: A Novel Framework for Fusing Modalities in Medical RAG cites this paper.

MIRA: A Novel Framework for Fusing Modalities in Medical RAG HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:34:25.360730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:34:25.360730Z digest=sha256:a66013faeb75388fe5f3b473d8e0033b80a537d4779232d3371cc5a430eaff38

Observation ef3b3bd9-d951-4494-8e39-26b57aa1de30 · inbound

Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models cites this paper.

Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:06:23.774162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:06:23.774162Z digest=sha256:9d9a654b099379f8b8641eb27b242895107ae74a06a4b5614e15f869148ba7da

Observation faf3aa55-f2f1-43e7-b1d3-3c55f7592289 · inbound

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study cites this paper.

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:19:46.890585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:19:46.890585Z digest=sha256:dfa55ce5baecbd953fbe8adad1334a0f3746d98b4b65a82951d5f9729b2cf512

Observation 4180b87b-fc99-4342-8b17-1fff9bf367cf · inbound

Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models cites this paper.

Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:32:34.476129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:32:34.476129Z digest=sha256:62b30aa302205916ffa2e1a51efb4f6422c291afcdde782abc1b3721a4fdba21

Observation 4e8e76aa-958b-4d80-92ef-d64a3930fd5f · inbound

CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning cites this paper.

CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T11:00:49.945470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:00:49.945470Z digest=sha256:87eb9212b88d4afaecdc354d34db5fc3e40c84c5b6bfc781df4ff005103eacd5

Observation 199e2a0d-04a8-440f-b669-2c3d141aa345 · inbound

MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models cites this paper.

MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:37:06.417636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:37:06.417636Z digest=sha256:2e00fe3be10ecfb53a61e05318abb3c43d8e68acf7d5a6ba24fd7fcab327b342

Observation 654ea1b8-186c-4fda-b34c-6ae2f4e6c556 · inbound

Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning cites this paper.

Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:43.094255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:21:43.094255Z digest=sha256:c2475e617509807d2475dee6d8b09677f31fbd77eefd3bd6cb6a32ab2d40acd9

Observation 52eb3776-1329-4adb-8b07-ab4f5dc31559 · inbound

Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models cites this paper.

Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T19:48:44.653849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:48:44.653849Z digest=sha256:6f166a1dd2aa49dad744cfb8b6e16c5d79f3d277f31d02dc64ea314b7ad11890

Observation 2dd43cdf-efbc-4911-80f6-fd3a22fc74af · inbound

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis cites this paper.

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:40.611783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:40.611783Z digest=sha256:e8afaf3d7d9c901b0c787978992074089b30f24f0f5cff8ff0b8912093a8c2c9

Observation d8977dd1-3c51-4d54-a3fc-f62988cf15b8 · inbound

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation cites this paper.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:35.849953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:35.849953Z digest=sha256:da29de1905e5f278b6008de7c0dcf8181b1c59a142a6ceba79e29cbf439a23f1

Observation 97cf8811-c31d-467a-8cb8-c73f7373d61d · inbound

CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space cites this paper.

CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T17:52:07.640057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:52:07.640057Z digest=sha256:6309d6008674368a3afb30a4ec39bdc87d5bf21f7fc103620b92d0dcbb9b78b4

Observation 1d3fc894-ee14-4e1f-abd3-91792a3331e9 · inbound

IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation cites this paper.

IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:33:10.132677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T17:31:33.903063Z digest=sha256:83f5ec075ae71c2a61f1ff3aac871831b6bcdcbdc2b8d1c6d912d728e37ce024

Observation bf1e47a4-c6ab-47e7-8e2d-53a55127c2ce · inbound

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding cites this paper.

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T10:49:58.923107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:49:58.923107Z digest=sha256:7a2dfe286140aa470af3557ec856d9a36f288d9c897234f4a44969ecb4d068b1

Observation 3840de88-751e-4afd-ba85-0b28dc408eba · inbound

MEDSYN: Benchmarking Multi-EviDence SYNthesis in Complex Clinical Cases for Multimodal Large Language Models cites this paper.

MEDSYN: Benchmarking Multi-EviDence SYNthesis in Complex Clinical Cases for Multimodal Large Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:36:32.807527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T19:34:42.244909Z digest=sha256:c17d3377c50b522cb3813a92239054982858437cc0d559cee3c5147911905d98

Observation ae4fd3de-ae19-4dd5-8777-e9e4ccb2e5c3 · inbound

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space cites this paper.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:32.110889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:32.110889Z digest=sha256:6abf2851df5a81774d8582228a8f21ae29c1495e174d5ceaa402bcb29ac4593c

Observation 7bb6f50f-f719-47b1-b16e-108788089657 · inbound

Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain cites this paper.

Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T17:47:52.088098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:47:52.088098Z digest=sha256:7aa05d401f4f33bc9276189031d4998cd6edd274efd26e61ce8067e60ffc7cb6

Observation 5ccaa8e1-d88d-4ae9-a63c-0889fc11bedd · inbound

Spatial navigation in preclinical Alzheimer's disease: A review cites this paper.

Spatial navigation in preclinical Alzheimer's disease: A review HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T19:53:53.219607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:53:53.219607Z digest=sha256:5692fc1062b29599bf30c67411000f5464a104a516be6353b24bb92e9a69e79a

Observation d64d33f4-50da-4aa0-9b21-c8c002083005 · inbound

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows cites this paper.

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:08:21.666841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T00:07:22.106337Z digest=sha256:2300089e7ff3fd1729ca7cee5d6ebe5c2d20a88796ea0313b54160bf60c6ff81

Observation 5cd3d405-e309-46e2-9619-58fca984db3e · inbound

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics cites this paper.

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T14:28:05.261852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:28:05.261852Z digest=sha256:7269f5a29e1788ec49da37b72b175aa4224e361a5646db2f7c57cc74d148e8b6

Observation 197074b1-5345-4f8b-ba8f-d687477009ab · inbound

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics cites this paper.

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-15T11:44:19.622453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:44:19.622453Z digest=sha256:0d0d4bc1121a57cbcbec3ae687a305d50c865eef3a816a465fdafbdda542ec7f

Observation 2e54f11e-c003-43d3-b341-021267cad10f · inbound

Scalable and Private Federated Learning Using Distributed Differential Privacy and Secure Aggregation cites this paper.

Scalable and Private Federated Learning Using Distributed Differential Privacy and Secure Aggregation HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T08:39:01.975337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:39:01.975337Z digest=sha256:c971a62a3a90a073da65f657446c4c478bf4c43f82224b8f3fb31f3828e2a7d0

Observation 71bbd91b-1bfc-4079-b77c-f93b36c0136b · inbound

A Utility-preserving De-identification Pipeline for Cross-hospital Radiology Data Sharing cites this paper.

A Utility-preserving De-identification Pipeline for Cross-hospital Radiology Data Sharing HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:59.599810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:05:54.328413Z digest=sha256:f6ccf4d02d40f3481af847d7e32772bc7a34b0ded36fa3fad72dc2b45df408bd

Observation b0a2712d-e26a-41ab-8f96-7c6ee7b040d2 · inbound

MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging cites this paper.

MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:50:25.257886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T12:47:27.551670Z digest=sha256:d5d5ede932dcf9d352e76a1968a6bb9957ad3d6d021ecf35d421a94904bcab61

Observation 7cf6206d-be77-4dd4-904f-efc47039eb89 · inbound

Seeing Through Experts Eyes A Foundational Vision Language Model Trained on Radiologists Gaze and Reasoning cites this paper.

Seeing Through Experts Eyes A Foundational Vision Language Model Trained on Radiologists Gaze and Reasoning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:15:26.144571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T13:14:48.103651Z digest=sha256:6a9c3ff624b7cca0171dd9a19400481af1dae03075a24758260563d3f9a22d6b

Observation 80d0b41f-bebd-4f55-8c46-89536b31d7e9 · inbound

X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis cites this paper.

X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:04:49.687180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T01:04:07.511600Z digest=sha256:8f363b9ea5bb564c002c7de3e54f56e57243b588c8233070cb157885da5ba56d

Observation 2eb0ed3f-43a7-4ef2-a9cb-ba6d34e78ddb · inbound

MedHorizon: Towards Long-context Medical Video Understanding in the Wild cites this paper.

MedHorizon: Towards Long-context Medical Video Understanding in the Wild HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:07.533998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T12:28:35.008604Z digest=sha256:a02e67152761227ff4d8bc9b160655f5be9655098e1cf76f0c7eeb03cd6877a1

Observation 1153f00b-9f48-4dbb-92b8-059a49c3dad1 · inbound

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents cites this paper.

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:41:43.388149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:02:36.159920Z digest=sha256:e92078d8b34b5f98e05d3291ee86040e2ee2a8dcf9f09b9baa59016271155250

Observation ca4e5bde-3444-4c15-9ffc-148920d52e5d · inbound

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks cites this paper.

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:11:21.929441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T05:10:02.941396Z digest=sha256:d4c004fe6a4e0d17ab209a18d61a4add6c20a95e1d7140e70bb01eee251fce2c

Observation afc2741f-a9a7-4cdb-a1c9-2a62f8bc5c6a · inbound

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks cites this paper.

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:25.823602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T05:10:02.941396Z digest=sha256:d793134332cdfacdecd0ffec31c85371a613142431ac874cecdd379d61637b02

Observation e2a5b865-88e4-4acf-8623-8df39f8eb0e1 · inbound

DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making cites this paper.

DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:58:33.487014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T02:56:04.280700Z digest=sha256:36e1f98e2fe1e6fd1781a9eed76fa10fb0ed3d5975e12454637fcd162edc0d8c

Observation 3614521a-79b0-4132-9ed7-dab180e83428 · inbound

RoiMAM: Region-of-Interest Medical Attention Model for Efficient Vision-Language Understanding cites this paper.

RoiMAM: Region-of-Interest Medical Attention Model for Efficient Vision-Language Understanding HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:58:58.612149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T19:58:26.380111Z digest=sha256:d02f86779ad32f5b2efa8ec3d1a3e03d6ee67891559be7ee6c711f56e03cc3e1

Observation 45e836c7-ed99-47e1-8601-d894ef2d6b22 · inbound

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis cites this paper.

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:09:51.696719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-21T08:06:02.176934Z digest=sha256:18383bc68ea15072ad566e7797e8d0c06d7d81d34b8dea4a7599a902033b9573

Observation a9c5687e-b22a-46d2-b86c-2367cd35690e · inbound

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis cites this paper.

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T00:11:26.451809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:11:26.451809Z digest=sha256:4669d6a876011d4ca4c6d9353aa21141c94c5f55c07ae225439c12c2681ba3a5

Observation 98bb59a0-8e3c-4f1c-8873-e1da9ad1562c · inbound

VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs cites this paper.

VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:13:27.291018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T13:10:36.266366Z digest=sha256:65492f2a009adfdf2e5b7fae10fc1dec1807708a303a34ab194fd5c083eb3016

Observation ea2c4426-1b4b-405c-af9a-6ff7799d4ca4 · inbound

ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training cites this paper.

ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 108

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:22:34.684060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T19:16:42.139096Z digest=sha256:d5176ff6e6a8cc5afa76d724fa0c9d1ba9ee631bf84cce640c002c8f323aabce

Observation 300010db-b618-4773-b416-7a2cc7896dc4 · inbound

Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining cites this paper.

Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.274792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T17:26:11.320892Z digest=sha256:deaad2ab8802ca9e2dd666bcaff935d8f58a2ed01544d08a3784f2bd6dc421a9

Observation 8d814340-c964-4fe8-9278-27ee76009c6d · inbound

Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining cites this paper.

Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T02:15:35.558277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T02:15:35.558277Z digest=sha256:e16d072bc943f6f3edeb5aaa448ac1d4e2f2667bf2c6955827af27f9530fd9c1

Observation 7db60428-93a8-4ba5-8b6b-2ad303d57148 · inbound

DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home cites this paper.

DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:05:31.552275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-01T07:40:03.204845Z digest=sha256:adb3f04c96a519f2eeba4041a6782cc3904ad25452de128880a268bd933b91a4

Observation 8e4c68bb-cc2b-4cb5-91d6-c91760874cfc · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.464703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:2c7a4c50762160e0f09e8c273eb7fcc96943999f2b5fa0d9fdc283d23ff1ae44

Observation d11c7951-563c-4a7b-8f3d-432c08c256f9 · inbound

Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection cites this paper.

Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:06.953227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-25T21:06:50.675960Z digest=sha256:f7e78f079d97bf3b1142870c45aafb1f770ddcb0b9c2d76079b9a4dc461d16e8

Observation c560eff4-4e4d-4240-86ae-3f3300c76016 · inbound

Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection cites this paper.

Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:24.287126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T21:30:16.229864Z digest=sha256:16e0eed09f65b060fe4452372e9ffab02c0fbdd1b47b07bfd14dc63bbfb9ac02

Observation b2b49093-3e06-4568-a8bf-e5ed9e92d9d6 · inbound

Aloe-Vision: Robust Vision-Language Models for Healthcare cites this paper.

Aloe-Vision: Robust Vision-Language Models for Healthcare HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:25:57.962003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T02:02:47.472868Z digest=sha256:2866e23beb576a5b44b639b1aa556cf8f0176719fa674e88ecad53dfa46ab113

Observation 37ca057e-28d0-4518-b7ce-500f1bbbedaa · inbound

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning cites this paper.

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.018403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-01T05:36:43.609602Z digest=sha256:fd3217b92bd56574d8e92dcfe4cd27d354801da6b131dc7edd74c1cbbbceb147

Observation 0b9a290a-3e40-4665-9473-1097797549be · inbound

Towards Real-World Ultrasound Understanding: Large Vision-Language Models from Multi-Image Examinations with Long-Form Reports cites this paper.

Towards Real-World Ultrasound Understanding: Large Vision-Language Models from Multi-Image Examinations with Long-Form Reports HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:08:36.894665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T15:59:52.903025Z digest=sha256:cbffd97cdae58ae90ca5b26d3adcc9057bf5986edcc6118c26961c3079f1a885

Observation 2215453b-b7dd-4452-8a2e-6a605180ce9d · inbound

IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation cites this paper.

IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T19:56:14.354465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:56:14.354465Z digest=sha256:11af4f839d4dacf128cfc57a87b3c83b4c7a25a5efa0d9dccebf04219102aecf

Observation 9916d025-d5a6-4d6c-9432-d8edf836a129 · inbound

Evaluating and Understanding Model Editing for Medical Vision Language Models cites this paper.

Evaluating and Understanding Model Editing for Medical Vision Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-07T18:14:03.435995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-07T18:09:33.751019Z digest=sha256:4a8a908abf4a4b3254b015e3d00c302a9c859428f56a9f5f5ae7fab0c8724a60

Observation 2c12473b-74ac-4ee9-8de7-42ac4d543ef4 · inbound

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding cites this paper.

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:29.702013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:29.702013Z digest=sha256:5b623511a20e24f6e771fc10c82b47cf5ebf10ac68c49d8e0c731dbfb656d07e

Observation 2f7a5e43-7b84-4549-a896-6b8611bad252 · inbound

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models cites this paper.

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T20:35:05.706759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:35:05.706759Z digest=sha256:819f696222c68b23389e080db8b65a3f0b812633c170bc17a5308f4cdeecb237

Observation c8d131bf-753d-4450-8c8a-0f1ca67dbe61 · inbound

MIRA: Medical Image Reflection for Agentic Diagnosis cites this paper.

MIRA: Medical Image Reflection for Agentic Diagnosis HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:14.075007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:14.075007Z digest=sha256:32ba1b0ec5c67dbf9ac8817f615ff57d4d3556514b8e37b23f65d3defb5fd1b7

Observation 342cdf9f-2934-4051-8572-375c1816b315 · inbound

HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation cites this paper.

HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:35.751304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:35.751304Z digest=sha256:7ac957218a0407ad96924dd88115fdd394dc09f5d78875a28a52ff20f4e4d83a