Pith. sign in

Paper Citation Record · LEDGER

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

As of 11 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2603.13800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.13800 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T18:16:34.491054Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:31:45.670345Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8cdf80f8-7fcc-4a86-9bd2-d10f737ed577 · outbound

This paper cites Radimagenet-vqa: A large-scale ct and mri dataset for radiologic visual question answering.arXiv preprint arXiv:2512.17396,.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space Radimagenet-vqa: A large-scale ct and mri dataset for radiologic visual question answering.arXiv preprint arXiv:2512.17396,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:31.730511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:31.730511Z digest=sha256:d9c6e842d2df688461046b5a3e90cccdba805c2df214fbb14e15000fcf6708b0

Observation affbb95c-25f5-46b9-9509-47599f013b87 · outbound

This paper cites InternLM2 Technical Report.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space InternLM2 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:31.894398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:31.894398Z digest=sha256:19f8e3574f6d860e24b951b890bcd8f3aaca935ef6671082328d91a210ce3ce4

Observation 7a8e7323-7e07-4aa9-9dfe-f5102d46872f · outbound

This paper cites and Weis, S.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space and Weis, S

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:32.074506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:32.074506Z digest=sha256:619b3cffecaee9b8dc6a42440b2f1a808c1439f5b57cb0f23019a1b969b49d30

Observation 10f0d040-f41d-4738-aed2-3cfc1f5e5f1a · outbound

This paper cites The Llama 3 Herd of Models.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:32.225215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:32.225215Z digest=sha256:36ec0f66a3e2bb54bd251931b64e0a2f6b2a9ccfefc06f95c81c5f131075223b

Observation 41aede8a-2506-414e-bf4d-893bad6e9653 · outbound

This paper cites E., Er, S., Almas, F., Simsek, A.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space E., Er, S., Almas, F., Simsek, A

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:32.287343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:32.287343Z digest=sha256:2d6e9ee26bdb02d5ef4416a5036ef3ad5158eacb39f8575f7ce5db68251db1ce

Observation d2890c10-b842-483a-a61d-fbcc0e3ee91c · outbound

This paper cites INSPECT: A Multimodal Dataset for Pulmonary Embolism Diagnosis and Prognosis.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space INSPECT: A Multimodal Dataset for Pulmonary Embolism Diagnosis and Prognosis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:32.462014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:32.462014Z digest=sha256:2963e53cf23173efffdf468c9b1c11bc89389955efb21f02ef3e09a1eb065e8d

Observation 92d04f16-1da0-4392-9ea3-ee45a5c4d21a · outbound

This paper cites Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:32.584875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:32.584875Z digest=sha256:0941507c6153a407e52b7949d1b8cf71f32227e9cd23d3ac873cda7b2ace02a1

Observation e85a6bb2-37af-4a43-addf-454bc49270c9 · outbound

This paper cites Menze, B.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space Menze, B

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:32.767136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:32.767136Z digest=sha256:d9419093ae7463f2eb8b9f21015efab30ce4037f237785f8f321dd8c3869c8ce

Observation b0ee57f9-fc5b-4c1c-b1ab-46d0f6de3f2f · outbound

This paper cites MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:32.946219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:32.946219Z digest=sha256:8d8b8007e84c7920df1aecd65e262a1db7917a8d3bcc6d9e7eac040c4a61e60a

Observation c695bb43-0ca6-465b-b66d-a617325cf295 · outbound

This paper cites Qwen3 Technical Report.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space Qwen3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:33.107574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:33.107574Z digest=sha256:f3a0ec266fd907388ae0b87738b1343b9188486eac0b6bf105a1173a9ab241b1

Observation 45844de7-df57-4c1c-9478-0cdc85e30134 · outbound

This paper cites Thermometer: Towards Universal Calibration for Large Language Models.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space Thermometer: Towards Universal Calibration for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:33.382429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:33.382429Z digest=sha256:1bb0f9031c49cf6ef817b13307cfa5a6ae66df01336205cf59f8acb6404da1d4

Observation 8231a844-0552-495b-ba5b-6254a2e1f2d2 · outbound

This paper cites Med-2e3: A 2d-enhanced 3d medical multimodal large language model.arXiv preprint arXiv:2411.12783,.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space Med-2e3: A 2d-enhanced 3d medical multimodal large language model.arXiv preprint arXiv:2411.12783,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:33.577395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:33.577395Z digest=sha256:dd678173b0c5d9a185d09e66b9c5e41961aa23c869c21143255591280e31af3b

Observation a5c8e9a8-e746-4805-81e5-31572e6effbe · outbound

This paper cites Prs-med: Position reasoning segmentation with vision- language model in medical imaging.arXiv preprint arXiv:2505.11872,.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space Prs-med: Position reasoning segmentation with vision- language model in medical imaging.arXiv preprint arXiv:2505.11872,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:33.764085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:33.764085Z digest=sha256:0486073e0d211d7615af2d5281b0c6f0b62bac47a7f6bb319bba2261d9467f77

Observation abb47cb2-d812-47b1-8d73-addfa802534d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:33.915250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:33.915250Z digest=sha256:d7ba9ba0e2322778e4652f7a93accb0d22a87b923ac07d2d26bf3b1b9e2ba90a

Observation 58c97126-eddf-46c9-9569-f8b15d2f7123 · outbound

This paper cites SA-Med2D-20M Dataset: Segment Anything in 2D Medical Imaging with 20 Million masks.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space SA-Med2D-20M Dataset: Segment Anything in 2D Medical Imaging with 20 Million masks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:34.075552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:34.075552Z digest=sha256:a0d46590a96e27d950696005d32766e488ad9781b98a6541e3256235f557e3d1

Observation 045dcfed-26c4-43a4-8afd-b72a2f621c61 · outbound

This paper cites CLEVRER: CoLlision Events for Video REpresentation and Reasoning.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space CLEVRER: CoLlision Events for Video REpresentation and Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:34.210937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:34.210937Z digest=sha256:5d06d8b1fff8c9acd7a3b37b490c7b84dd1cf7ef15883dd2a69ebbc01479ca29

Observation 40d6e805-7964-4429-bbdd-e9664ac181a4 · outbound

This paper cites All of the embeddings are stored via the index storage of Faiss (Johnson et al., 2019).

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space All of the embeddings are stored via the index storage of Faiss (Johnson et al., 2019)

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:34.491054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:34.491054Z digest=sha256:d7553fddc34a5aed384e15becfcfb80e350e2caa60aa1bfdc676e36a08095fe8

Observation ae4fd3de-ae19-4dd5-8777-e9e4ccb2e5c3 · outbound

This paper cites HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:32.110889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:32.110889Z digest=sha256:336baea2b0997eaae1f43c9f48273db8367edd1110a85755f1e3e570253f66fe

Observation 0a741358-8f30-4963-b68f-e2c7f550051a · outbound

This paper cites MedGemma Technical Report.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space MedGemma Technical Report

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:33.259309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:33.259309Z digest=sha256:131d9e3182783439be2826c37cf657daf57f44a606702a9cb3333aa4252a4f91

Observation a4296c92-2c53-4b3c-afeb-f7c50ed16e66 · outbound

This paper cites Identifying the Best Machine Learning Algorithms for Brain Tumor Segmentation, Progression Assessment, and Overall Survival Prediction in the BRATS Challenge.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space Identifying the Best Machine Learning Algorithms for Brain Tumor Segmentation, Progression Assessment, and Overall Survival Prediction in the BRATS Challenge

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:31.590467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:31.590467Z digest=sha256:292a24bdc952a6e5dc02f5e3ce0cc0fcf091f215940e1893f6900a2047aa0419

Observation 905a0ddd-9284-4ddd-82ac-838b367e6fda · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:34.354163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:34.354163Z digest=sha256:7ce34fee07220ca21d2678138c8a75b2bab027fd066b1904f17cbe669551105c

Observation 4cda2cca-768e-485d-9590-fff01246125c · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:32.339683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:32.339683Z digest=sha256:d1ecf06c10877117a99a74280dc1b291cf916e7d777f80ee19dc577bca31241f

Observation a80220ef-d008-453c-992b-1a3fab8fb7ed · outbound

This paper cites E., Zhang, X., Zhu, M., Alabbad, M.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space E., Zhang, X., Zhu, M., Alabbad, M

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:31.198364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:31.198364Z digest=sha256:246d6bfa334a5abdc8f289cb989292a6f3df46b854292886abdb0b4497202a55

Observation 4f43c40d-00f4-41cb-9f16-410ebc1a8639 · outbound

This paper cites 3d- rad: A comprehensive 3d radiology med-vqa dataset with multi-temporal analysis and diverse diagnostic tasks.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space 3d- rad: A comprehensive 3d radiology med-vqa dataset with multi-temporal analysis and diverse diagnostic tasks

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:32.187614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:32.187614Z digest=sha256:a50dd4ab6875a1a1eb0b89138d1deda4aaa8689e5cb2e269df7889e634f7bae7

Observation 38660bee-18d5-4990-b8d9-8e984c425151 · outbound

This paper cites Qwen2.5-VL Technical Report.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space Qwen2.5-VL Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:31.472901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:31.472901Z digest=sha256:aecc181f242d084dd05cd9dad90b6d87e3eb54d610ca7192d6f96543365dd852

Observation bd8a8576-6537-45f3-ae60-098771088330 · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T18:16:31.288447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:16:31.288447Z digest=sha256:4c5c5bc6ad1532aa436a0b4539f74cfbf69d4f1537f8e490f5ed7134c9360cbf

Pith citing papers

Observation 9915882e-1343-4c71-a4bc-08de70340285 · inbound

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models cites this paper.

Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T03:31:45.670345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:31:45.670345Z digest=sha256:174de893200764522ab1263f005b3c18457d8b798fc5bde9f7cf3e65f1377401

Observation bd428541-2b5b-49af-8af2-ac17ec3a2430 · inbound

Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography cites this paper.

Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:26.906488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:26.906488Z digest=sha256:545c672c207e8c2f9a2795434343f0779e713e11e89f36cf59070b905da39e13