Pith. sign in

Paper Citation Record · LEDGER

HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 69 inbound Pith citation observations for arXiv:2406.19280.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.19280 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 69 of 69 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:35:05.706759Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7253cf51-f030-4c56-8842-6578b2a220e6 · inbound

CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs cites this paper.

CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T17:18:04.172031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:18:04.172031Z digest=sha256:1c70c1bdff8bd4a9ab1aecd57f26e14fdfcd3910a0b22a291cfd2fbd16f5651c

Observation 00198be9-9255-40dc-a781-d94e84bcdfb4 · inbound

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI cites this paper.

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:16:27.649840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:16:27.649840Z digest=sha256:61d68700616b53a98a0c825c0ffd2cf3e4f93dc0c23da9f5e2eb96f82ed0a935

Observation 83739f70-003d-43dd-b32d-3b705eb7fe2a · inbound

On Domain-Adaptive Post-Training for Multimodal Large Language Models cites this paper.

On Domain-Adaptive Post-Training for Multimodal Large Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:56:46.703498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:56:46.703498Z digest=sha256:cc1a30d1ac7c8526a6fa58d03f2a6742bb0d4dd2f3e4326111eeed4e5730a42b

Observation 6ee5fc8a-b11a-4c47-9a5f-491b582c54de · inbound

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs cites this paper.

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:36:50.113768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T12:36:50.060335Z digest=sha256:b24c051c171b4da8af7f20f0428986f74a0ee4224ba769107065451bd597be00

Observation 895d79ea-704a-479e-8948-f89c959c7064 · inbound

Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging cites this paper.

Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:44.727179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:40:44.727179Z digest=sha256:216b74bf565044dc492c53f10040d90c7e4daa6f2ae463b59ea5a3fcd369fa61

Observation 57ecf7b4-5a5b-4cef-a5ad-e309afef81f9 · inbound

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery cites this paper.

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T18:24:41.782690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:24:41.782690Z digest=sha256:fa96f489952a0157497d7907793632d70e51b1d13c74d8df20ff27f059c2b6c1

Observation d4e072af-8f2f-42b8-90f8-3d6f18fed410 · inbound

From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine cites this paper.

From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-07T22:13:31.837086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:13:31.837086Z digest=sha256:e28d7d58b0b4f8947426126e9436b789c638f4f204585a772dd708d91ea63dd2

Observation 7815ca95-f077-402d-8f31-5a207dc26b80 · inbound

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation cites this paper.

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:30.364189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:21:30.364189Z digest=sha256:4c09cc10e164e7b33efb3c2530d0835598ff4e3816afa2372df4d2222e451bfd

Observation 3e925468-2865-475f-a3e4-b7789c589afb · inbound

Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning cites this paper.

Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:44.809561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:33:44.809561Z digest=sha256:c62c02b36ece5197200d886fa9cf1bebce1f17ffa77b84a33e2d9dc8a6576aa3

Observation 4984e00a-9b04-4f94-9a48-9fc79a1f61ca · inbound

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning cites this paper.

Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:42.360661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:42.360661Z digest=sha256:e4368e3c5765f81e66956b677827d098d595d23b6135e0240235c12498f9252f

Observation 1acf071c-13bb-4a89-b1a7-1a3f015cf456 · inbound

Interpreting Chest X-rays Like a Radiologist: A Benchmark with Clinical Reasoning cites this paper.

Interpreting Chest X-rays Like a Radiologist: A Benchmark with Clinical Reasoning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:57:02.678990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:57:02.678990Z digest=sha256:86c19f4b5f7bdce2b9cb805836d383de5ab7eca5278adec98a2f8e697bbcb648

Observation dd08d31e-90e7-4d0c-8dd7-fb5aee4c4f6b · inbound

DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis? cites this paper.

DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis? HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:06.182529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:35:06.182529Z digest=sha256:32248c1ce6fe9c9042c14b031044fc88d9cfda64d7975bc83dd6034a860d8e31

Observation c4e31152-36ae-4729-892b-c3b61923bd92 · inbound

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book cites this paper.

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:02.852643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:02.852643Z digest=sha256:ab30cafe8ca0043cef954cefd95b62f39d1ccbb2c17d647054d28380ea1f6978

Observation 733d3356-3db0-49f4-99cc-eae4d693064b · inbound

Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning cites this paper.

Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:48.640705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:48.640705Z digest=sha256:d5852f22e4d7290754b799d36c6c70ab5db0737cb4fc1e9296ef8eaf2d9071c1

Observation c8d9a6ba-42ad-4486-958d-4bf2f8dcc8ad · inbound

HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding cites this paper.

HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T10:47:15.081327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T10:44:01.880405Z digest=sha256:3b3b9462b2befddc0abac33de9c1330c73b07806053c25f4a588214a67ebc661

Observation a3f3f9be-363b-4006-9834-d829cd011408 · inbound

RARL: Improving Medical VLM Reasoning and Generalization with Reinforcement Learning and LoRA under Data and Hardware Constraints cites this paper.

RARL: Improving Medical VLM Reasoning and Generalization with Reinforcement Learning and LoRA under Data and Hardware Constraints HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:57:22.717492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:57:22.717492Z digest=sha256:937223203450f5fd2be5a60364c9c08de099f2aa2d038804335098ef40dca71e

Observation 48dcebf1-4bf3-43a9-822f-079ce5e67f00 · inbound

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding cites this paper.

HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:47:07.641054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:47:07.641054Z digest=sha256:59c85599b555d2cdbbe38363acd844f131ec30dc8438f69b2c38bc1a6bb615a7

Observation 8c3e2485-fa62-42fe-9688-a2403bd9e1a3 · inbound

Thought Graph Traversal for Test-time Scaling in Chest X-ray VLLMs cites this paper.

Thought Graph Traversal for Test-time Scaling in Chest X-ray VLLMs HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:17:13.932466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T09:15:46.135084Z digest=sha256:7aaa1980a15176756fcbd848a3bd4d5080a85582a651e62b1298fafa443f48fb

Observation 0ec0d310-c99a-4b0b-81d9-446500bb119c · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:17.728304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:17.728304Z digest=sha256:1c73b73b616c409f2ffa7cb61c8be9ed64f4e04f6e262d7764ebc35638ec7ca6

Observation 5829d0ea-3cdc-4841-8729-a2442c5c2caf · inbound

CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale cites this paper.

CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual Rationale HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:17.613129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:17.613129Z digest=sha256:f92932c116f767d676ba0b0a473e5df15886a5cde67b3f5cc19665c976edb893

Observation 24cc1d87-4463-494b-9f1d-208e1c2bfc93 · inbound

MIRA: A Novel Framework for Fusing Modalities in Medical RAG cites this paper.

MIRA: A Novel Framework for Fusing Modalities in Medical RAG HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:34:25.360730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:34:25.360730Z digest=sha256:6d4b12ea04013b38ec9b241fdd3e6a6d9cc4b44a5b621cfe5d52c9c24e4e46da

Observation ef3b3bd9-d951-4494-8e39-26b57aa1de30 · inbound

Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models cites this paper.

Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:06:23.774162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:06:23.774162Z digest=sha256:1cb3a533141f67ccf9499a63f29824a3f6c96b9714a2ba5ecede33a62dc5d2a8

Observation faf3aa55-f2f1-43e7-b1d3-3c55f7592289 · inbound

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study cites this paper.

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:19:46.890585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:19:46.890585Z digest=sha256:d601972aa0fc347535dde0f5f8f101d535037497a796dc7bf445bfd2ebb23614

Observation 4180b87b-fc99-4342-8b17-1fff9bf367cf · inbound

Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models cites this paper.

Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:32:34.476129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:32:34.476129Z digest=sha256:4c58d3eab2637571c3b34d7547310eb63bfad6e2bcea7b702f7a7e719ff89802

Observation 4e8e76aa-958b-4d80-92ef-d64a3930fd5f · inbound

CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning cites this paper.

CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T11:00:49.945470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:00:49.945470Z digest=sha256:ca3d9e501b20f8f8d7c1e8f1dfce5ff9ea2f38fc00313e168bdf6479b154be6b

Observation 199e2a0d-04a8-440f-b669-2c3d141aa345 · inbound

MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models cites this paper.

MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:37:06.417636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:37:06.417636Z digest=sha256:63f6655f7b3fbfc157f8c7b3b2d7c4d2a34c3cafb1f4ead5bac3559775d637b3

Observation 654ea1b8-186c-4fda-b34c-6ae2f4e6c556 · inbound

Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning cites this paper.

Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:43.094255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:21:43.094255Z digest=sha256:4ed025863399097e73b561ed0efcf7739ff83b7ba0b8ce4f05e4062615d053ed

Observation 52eb3776-1329-4adb-8b07-ab4f5dc31559 · inbound

Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models cites this paper.

Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T19:48:44.653849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:48:44.653849Z digest=sha256:f8c8adcbec8aba9c02c16d05290c016b1d8acc61e33780300a0fc362c79813dc

Observation 2dd43cdf-efbc-4911-80f6-fd3a22fc74af · inbound

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis cites this paper.

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:40.611783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:40.611783Z digest=sha256:2a8ec3054d1ce84ee280934e6d57802230c4ac0c3fc33992b1dda04e0e36124d

Observation d8977dd1-3c51-4d54-a3fc-f62988cf15b8 · inbound

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation cites this paper.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:35.849953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:35.849953Z digest=sha256:3c05253b1337e602b3be91f54ba70957cc5b177670d9877efcb3d92224512a61

Observation 97cf8811-c31d-467a-8cb8-c73f7373d61d · inbound

CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space cites this paper.

CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T17:52:07.640057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:52:07.640057Z digest=sha256:72ae9be1e66ef9273287b2a626a94c1e4a3833e2c6a76d00504b364b078cdc1a

Observation 1d3fc894-ee14-4e1f-abd3-91792a3331e9 · inbound

IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation cites this paper.

IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:33:10.132677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T17:31:33.903063Z digest=sha256:469a2eb3011505003261a3c5ad6b4db7c699b9b20a036ec9c6798350b89bd8d8

Observation bf1e47a4-c6ab-47e7-8e2d-53a55127c2ce · inbound

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding cites this paper.

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T10:49:58.923107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:49:58.923107Z digest=sha256:84f9f791777c979dcb94ac7814c06236222095b6b2311cb7e8b34701db7ccdde

Observation 3840de88-751e-4afd-ba85-0b28dc408eba · inbound

MEDSYN: Benchmarking Multi-EviDence SYNthesis in Complex Clinical Cases for Multimodal Large Language Models cites this paper.

MEDSYN: Benchmarking Multi-EviDence SYNthesis in Complex Clinical Cases for Multimodal Large Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:36:32.807527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T19:34:42.244909Z digest=sha256:be0f764cd53c35d9cae9e93ab4ea3f361ec31ef8e553279f241852eee90b2947

Observation ae4fd3de-ae19-4dd5-8777-e9e4ccb2e5c3 · inbound

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space cites this paper.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:32.110889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:32.110889Z digest=sha256:12f578a3237eb4dd74fc97d5e722bf5a7c5b0897a64ec8a2f70b63a23faca8f0

Observation 7bb6f50f-f719-47b1-b16e-108788089657 · inbound

Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain cites this paper.

Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T17:47:52.088098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:47:52.088098Z digest=sha256:c4d958febcabad29bc53cfb6116ba1423bd7954db7106dc10c3d261ab1d03256

Observation 5ccaa8e1-d88d-4ae9-a63c-0889fc11bedd · inbound

Spatial navigation in preclinical Alzheimer's disease: A review cites this paper.

Spatial navigation in preclinical Alzheimer's disease: A review HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T19:53:53.219607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:53:53.219607Z digest=sha256:1f9d6ee025d863f8cf733a111f1b564335ced66d96c765d1479d152af512da64

Observation d64d33f4-50da-4aa0-9b21-c8c002083005 · inbound

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows cites this paper.

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:08:21.666841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:07:22.106337Z digest=sha256:9b696c091b6cd871a9265263efc14b37fb51f8aac2950ff9505b62e5924093f5

Observation 5cd3d405-e309-46e2-9619-58fca984db3e · inbound

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics cites this paper.

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T14:28:05.261852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:28:05.261852Z digest=sha256:fa8532915ea9e87955e00d73f1632c944f268b9b8ade0a142129e72178a3be3d

Observation 197074b1-5345-4f8b-ba8f-d687477009ab · inbound

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics cites this paper.

ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-15T11:44:19.622453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:44:19.622453Z digest=sha256:11595b397283d8f1131c44cacf35f4b10352b1430a57ef5f598fe5d3c3b10b7b

Observation 2e54f11e-c003-43d3-b341-021267cad10f · inbound

Scalable and Private Federated Learning Using Distributed Differential Privacy and Secure Aggregation cites this paper.

Scalable and Private Federated Learning Using Distributed Differential Privacy and Secure Aggregation HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T08:39:01.975337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:39:01.975337Z digest=sha256:cf205babc07ebd6aeb3fd053057e1c7d3236d26cfca7e3d9a96bb11821177f89

Observation 71bbd91b-1bfc-4079-b77c-f93b36c0136b · inbound

A Utility-preserving De-identification Pipeline for Cross-hospital Radiology Data Sharing cites this paper.

A Utility-preserving De-identification Pipeline for Cross-hospital Radiology Data Sharing HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:59.599810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:05:54.328413Z digest=sha256:1bca7ac2bfdba15db963d27c72a466bd9a5c9b69ffe781353ae47a58eac24302

Observation b0a2712d-e26a-41ab-8f96-7c6ee7b040d2 · inbound

MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging cites this paper.

MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:50:25.257886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T12:47:27.551670Z digest=sha256:51947878f1f797043802fa9e9412ff8dabd150158ff72e4b5ab0fd68254d6a02

Observation 7cf6206d-be77-4dd4-904f-efc47039eb89 · inbound

Seeing Through Experts Eyes A Foundational Vision Language Model Trained on Radiologists Gaze and Reasoning cites this paper.

Seeing Through Experts Eyes A Foundational Vision Language Model Trained on Radiologists Gaze and Reasoning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:15:26.144571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:14:48.103651Z digest=sha256:5b5b431d6b016076b262b847a511e9ad6a5636f4d727c9180b0b2785c161bcd7

Observation 80d0b41f-bebd-4f55-8c46-89536b31d7e9 · inbound

X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis cites this paper.

X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic Diagnosis HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:04:49.687180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T01:04:07.511600Z digest=sha256:60f7ac00ec2e3296f3df7fe77782d6a835d650458d651438009663a226036dc8

Observation 2eb0ed3f-43a7-4ef2-a9cb-ba6d34e78ddb · inbound

MedHorizon: Towards Long-context Medical Video Understanding in the Wild cites this paper.

MedHorizon: Towards Long-context Medical Video Understanding in the Wild HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:07.533998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-08T12:28:35.008604Z digest=sha256:349b81f1a48eac5e0b0d34fd49721644b010eb70697a3aee52fd7e4fc761d42a

Observation 1153f00b-9f48-4dbb-92b8-059a49c3dad1 · inbound

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents cites this paper.

DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:41:43.388149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T04:02:36.159920Z digest=sha256:d42b8a27b5644cc379228f09a6d080f51633efe2c983a2bcc64e48688f745e18

Observation ca4e5bde-3444-4c15-9ffc-148920d52e5d · inbound

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks cites this paper.

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:11:21.929441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T05:10:02.941396Z digest=sha256:dd9e66cb7e181d140a938649846b6116fc9d64b7c26a009d222b84d175801cc1

Observation afc2741f-a9a7-4cdb-a1c9-2a62f8bc5c6a · inbound

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks cites this paper.

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:31:25.823602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T05:10:02.941396Z digest=sha256:7329d101871f67aefe8d5acc52e983e718ff12b496d8389c907c59def9060386

Observation e2a5b865-88e4-4acf-8623-8df39f8eb0e1 · inbound

DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making cites this paper.

DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:58:33.487014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T02:56:04.280700Z digest=sha256:942f58f8fa8c7d4b09a0221555a1f950fca33a49126061367197c6b6da00a11c

Observation 3614521a-79b0-4132-9ed7-dab180e83428 · inbound

RoiMAM: Region-of-Interest Medical Attention Model for Efficient Vision-Language Understanding cites this paper.

RoiMAM: Region-of-Interest Medical Attention Model for Efficient Vision-Language Understanding HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:58:58.612149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T19:58:26.380111Z digest=sha256:2679a8d49364a657e114f01b775861147b4c139c597dc79f904eeb9f3de91099

Observation 45e836c7-ed99-47e1-8601-d894ef2d6b22 · inbound

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis cites this paper.

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:09:51.696719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-21T08:06:02.176934Z digest=sha256:94f4b6cddae541727a68bdf88bb970f5b99c5c8eb4085ebd086e05e73b8841cd

Observation a9c5687e-b22a-46d2-b86c-2367cd35690e · inbound

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis cites this paper.

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T00:11:26.451809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:11:26.451809Z digest=sha256:f502ff52de6af37dc3f229184eaa588210fa66f492284c7e3d35680d71aac54e

Observation 98bb59a0-8e3c-4f1c-8873-e1da9ad1562c · inbound

VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs cites this paper.

VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:13:27.291018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T13:10:36.266366Z digest=sha256:dc18818e5c1bfac6f9a1b828095f03b4860e4330878683bc545ecf97bdc73885

Observation ea2c4426-1b4b-405c-af9a-6ff7799d4ca4 · inbound

ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training cites this paper.

ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 108

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:22:34.684060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T19:16:42.139096Z digest=sha256:13e3b425bb6788bc0c3769977c260dc31930bcdaa4937e0ee24ed0d847d1e5f8

Observation 300010db-b618-4773-b416-7a2cc7896dc4 · inbound

Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining cites this paper.

Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.274792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T17:26:11.320892Z digest=sha256:f108e2419ff883059d17ad1359d530936f07ae4365812fd25a1a311a43d6279b

Observation 8d814340-c964-4fe8-9278-27ee76009c6d · inbound

Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining cites this paper.

Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T02:15:35.558277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T02:15:35.558277Z digest=sha256:6126b32ad81a95f058b82ccbbdfa9568328776fea3ec08b44ff350ef506b3c87

Observation 7db60428-93a8-4ba5-8b6b-2ad303d57148 · inbound

DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home cites this paper.

DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:05:31.552275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-01T07:40:03.204845Z digest=sha256:fad9fb0b6017b36d61ddafac594806a4431463201f28dceaef5b566719678c05

Observation 8e4c68bb-cc2b-4cb5-91d6-c91760874cfc · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.464703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:63aba1ac36862e7c3ee7a204241f39b30e400f61644bfa556493dd1479b6c797

Observation d11c7951-563c-4a7b-8f3d-432c08c256f9 · inbound

Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection cites this paper.

Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:06.953227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-25T21:06:50.675960Z digest=sha256:b4b821fc3209f39d1b5c0a6fcac3e6075372dac82612e5da57d899f7a4b278b6

Observation c560eff4-4e4d-4240-86ae-3f3300c76016 · inbound

Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection cites this paper.

Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:24.287126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T21:30:16.229864Z digest=sha256:8d5f06844fa2d8a1bf429a3984a1ac99778e4e481ed5e7a98d6f33bae7eedce1

Observation b2b49093-3e06-4568-a8bf-e5ed9e92d9d6 · inbound

Aloe-Vision: Robust Vision-Language Models for Healthcare cites this paper.

Aloe-Vision: Robust Vision-Language Models for Healthcare HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:25:57.962003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T02:02:47.472868Z digest=sha256:f75db983dfba06c6b82e7e5c2a3ec4a6181ec917000b7a6c225d72db15ac0623

Observation 37ca057e-28d0-4518-b7ce-500f1bbbedaa · inbound

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning cites this paper.

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:25:41.018403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-01T05:36:43.609602Z digest=sha256:9b0ed3cdc6abf055d5d8109019fb31517ee5e2727d6f5df284380f79ecb3d0b5

Observation 0b9a290a-3e40-4665-9473-1097797549be · inbound

Towards Real-World Ultrasound Understanding: Large Vision-Language Models from Multi-Image Examinations with Long-Form Reports cites this paper.

Towards Real-World Ultrasound Understanding: Large Vision-Language Models from Multi-Image Examinations with Long-Form Reports HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:08:36.894665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-03T15:59:52.903025Z digest=sha256:ff5ebb959e5073f2ab58bfec3aba706e942738f5577b4269166ae77d19be0d56

Observation 2215453b-b7dd-4452-8a2e-6a605180ce9d · inbound

IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation cites this paper.

IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T19:56:14.354465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:56:14.354465Z digest=sha256:39becc09174f9d38d34445c3ff20f5f5bb7e795aab12ac9762c72c75fbd50e16

Observation 9916d025-d5a6-4d6c-9432-d8edf836a129 · inbound

Evaluating and Understanding Model Editing for Medical Vision Language Models cites this paper.

Evaluating and Understanding Model Editing for Medical Vision Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-07T18:14:03.435995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-07T18:09:33.751019Z digest=sha256:c318bf2583f02aacacf720d15c3802b8f0a3bd3e52dcaa45b6b75bc1c85bfecf

Observation 2c12473b-74ac-4ee9-8de7-42ac4d543ef4 · inbound

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding cites this paper.

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:29.702013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:29.702013Z digest=sha256:fbdf2cd9fbfd2438575b4b440d04be582e1e335e4dad05e0dba0bccbaf54def3

Observation 2f7a5e43-7b84-4549-a896-6b8611bad252 · inbound

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models cites this paper.

MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T20:35:05.706759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:35:05.706759Z digest=sha256:b576a819887d5ec94b81f7fc57d27730e2386501ef1f1d0d6d337f89a4ad96e8

Observation c8d131bf-753d-4450-8c8a-0f1ca67dbe61 · inbound

MIRA: Medical Image Reflection for Agentic Diagnosis cites this paper.

MIRA: Medical Image Reflection for Agentic Diagnosis HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:14.075007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:14.075007Z digest=sha256:3134311d56cb92b8b94fbb023908adeba1d2834e55aabb9ec2c5705f999743be