Pith. sign in

Paper Citation Record · LEDGER

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis

As of 3 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2604.10233.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.10233 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:43:02.337806Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact17
  • verified fuzzy31
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71e8a687-9d7f-4d8c-84ea-29cc065c1030 · outbound

This paper cites Qwen2.5 Technical Report.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Qwen2.5 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:58.331205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:0c6e48ddb9dbbb2fd489052315483928c3334fa84c88da3aae321815c6db5078

Observation 885679bd-03eb-437a-8864-b7997f67bcc5 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:43.005263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:0428f0ac91f3075d8178849752657f4d475e3575bf022c2d20f04da8adbc1fbf

Observation 7c24f226-36f0-4a2b-94c2-22c9bdf01b17 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:58.337878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:a59b0add7bf336050faba9615d5d4685f229f8e068ba0866bf5df53267021bbd

Observation 28cc916f-d577-456e-a3b8-9367ed76b9f6 · outbound

This paper cites Sigmoid loss for language image pre-training.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Sigmoid loss for language image pre-training

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:43.009653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:e5e8dccdab4aaf24550d89e57fa6618c56ed6da2d4ce471291da2de53cdd123c

Observation 9e35857a-b8d2-4924-b5df-1874f6dda033 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis DINOv2: Learning Robust Visual Features without Supervision

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:58.346679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:18dd7fd1af34dbe151525f12dc9caba6578404754af5c123eab13032fa36aee6

Observation df48b65a-d775-4173-84f3-50775044ec35 · outbound

This paper cites Spatio-temporal and retrieval-augmented modelling for chest x-ray report generation.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Spatio-temporal and retrieval-augmented modelling for chest x-ray report generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:43.015945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:02583191ffd2e51de6c9317062f2f362b660a7efe3921722561778a4e16308c7

Observation 2972c0e6-4ace-42ae-aaf9-23cb34cffda4 · outbound

This paper cites Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.311872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:0e8cbd4550d8d37508dc0f3e5b460a7438d19ba40cd90d41144180c697a62605

Observation 2acf86df-eca7-49e2-81ed-9137992b1607 · outbound

This paper cites Collaboration between clinicians and vision–language models in radiology report generation.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Collaboration between clinicians and vision–language models in radiology report generation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:43.026368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:33ae82f5fcf60ea3ddc353e60c6124ee85cbb3c1fd76ae0ee6a3347783ecbab9

Observation e442baf4-1c98-47f2-9b54-cf42217ebf8b · outbound

This paper cites Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical Tasks.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.318827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:645c1a9ab2b495aad26b142a6ef64875a0ef5607637729ab1d1a8fb7667b91e5

Observation 1ab44201-8a4c-4ac6-b9d9-5401d5625989 · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Llava-med: Training a large language-and-vision assistant for biomedicine in one day

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:43.035546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:e99d33b98f463e7ca1b73464445a2d8e0571e52b11113295f6216e44c7a08573

Observation cb2eda88-b33c-444e-ba19-e771842c9bfd · outbound

This paper cites M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.323661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:ceef2a01bebee0e7390cfe4eac2e09789590194f2f653b717510449480b7f0e2

Observation 6689e5fc-3971-4ffb-9962-0dbed9760043 · outbound

This paper cites Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:43.032516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:cd029ba3fcf8d73cae814566c1d906c8ff774a1334fd7b3b27f20c1218d2df84

Observation 965cf1b4-ca81-431c-9e77-87e31a1d58ca · outbound

This paper cites Learning transferable visual models from natural language supervision.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Learning transferable visual models from natural language supervision

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:43.019423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:c3affe7e12ea8fc2763b793224a022a2152e362b3114c5ecca98ed4d535b1524

Observation e0e0c55b-773d-436a-bd43-ff8af3433adf · outbound

This paper cites Adaptive mixtures of local experts.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Adaptive mixtures of local experts

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:43.012821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:09b9b166ad3bbd8c36fc5e6868f4192789f81b0429ba1966250edc26d2896a69

Observation 71de56fe-cc81-40ff-a289-89ee9c2648a0 · outbound

This paper cites Learning to prompt for vision- language models.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Learning to prompt for vision- language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:43.023232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:8814b6161ad38dbecbf9c68f203c8bd5b082e1c21d4b20bba649074a1651db24

Observation 3225faba-2b91-446b-98f6-6e90560e9314 · outbound

This paper cites Con- trastive learning of medical visual representations from paired images and text.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Con- trastive learning of medical visual representations from paired images and text

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.996756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:14b284f02c777e3776e6ce55023c980ad1d43c14b4194388eedc63ed9d73bd18

Observation 04e4dda3-4bcc-49ae-aed6-70bb521fefd2 · outbound

This paper cites Procedure-aware surgical video- language pretraining with hierarchical knowledge augmentation.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Procedure-aware surgical video- language pretraining with hierarchical knowledge augmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.936017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:1c2e949c207c3193146c014a40d792a046d3b358d39835d1bd34d5472ebb51bc

Observation af982edf-2232-4d82-8cb3-a40740a428d1 · outbound

This paper cites Merlin: a computed tomography vision–language foundation model and dataset.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Merlin: a computed tomography vision–language foundation model and dataset

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:43.029332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:b8b468b92e65534a849de4c68754ccdf2e9ba81158926d8754a5b4bbf9b11b17

Observation 33b693be-f363-4f52-9241-b30cf95c649b · outbound

This paper cites Triad: Vision foundation model for 3d magnetic resonance imaging.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Triad: Vision foundation model for 3d magnetic resonance imaging

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.939429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:3f0e205fcaf8fb4ae6ceaeaec1930169eda0d2b7bda5e06002b8a0d968a6b42c

Observation f1f07db4-5a90-413c-83b1-1c8fcfcc2ea1 · outbound

This paper cites Learning neuroimaging models from health system-scale data.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Learning neuroimaging models from health system-scale data

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.960277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:dd4bb411221afe6ffc6e6f82ba58ccd3353d87a8ee528ee118b7ae0ff647b976

Observation 4388f767-c5fa-470f-9cb6-3ddaf465984f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis LLaMA: Open and Efficient Foundation Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:58.269293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:f207f5595426ba1bef5ba3411b5d8cb1126803427488dc0d33d6fb1fe26e104c

Observation 0736d530-b06e-4bc9-a8b4-6e097667fb4f · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:58.254258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:bb588ea52f6bd06cc0d6182dcca58bf564b6a42a77da620dc9ad0e9f53557ef7

Observation ef70287c-a201-4846-9dc3-f2773fe76b92 · outbound

This paper cites Towards accurate differential diagnosis with large language models.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Towards accurate differential diagnosis with large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.957568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:4acddf9739fb4ca3b84a53df30a5db42d4a6928edeba43ad49fe9291b8292049

Observation 2d742989-f691-40bd-91ac-9f4e4c02ba5a · outbound

This paper cites Toward expert-level medical question answering with large language models.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Toward expert-level medical question answering with large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.986971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:d09daa749c38f13d5cc47a28e3ebe6be46cbcfe41b75880a92567ed5c5ad92dc

Observation b5421448-b2c4-4a5a-9045-54fee4451c4d · outbound

This paper cites Medla: A logic-driven multi-agent framework for com- plex medical reasoning with large language models.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Medla: A logic-driven multi-agent framework for com- plex medical reasoning with large language models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.276274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:ca1dea886531198290b3d133ffb7cf1753a290d4f741f1c1bf04476472afd96b

Observation b72f02a3-faef-4772-bf46-a5d89c9d0f96 · outbound

This paper cites A generalist vision–language foundation model for diverse biomedical tasks.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis A generalist vision–language foundation model for diverse biomedical tasks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.978035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:bd654551fa74380b484909bb2242edcf381335bb1e5a97abaa0e34204d3e9e8f

Observation 196b1ef7-d7a2-4d0b-801f-1cb8542f6375 · outbound

This paper cites arXiv preprint arXiv:2503.20047 , year=.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis arXiv preprint arXiv:2503.20047 , year=

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:20:58.229416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:a51bd6cf6fcb6afe2d510dedb76e210d00217175add9ed3f2ff219d0dd2a483d

Observation 7127ecd2-47d3-4ba9-bfde-76cdd8bf5981 · outbound

This paper cites Dynamic graph enhanced contrastive learning for chest x-ray report generation.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Dynamic graph enhanced contrastive learning for chest x-ray report generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.946463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:fee4651bb12df8e63b4489611aa7648a01b2ef487ac5a4140f551f2812e050ee

Observation 27716b1b-cddf-4207-abf9-9df74fa96c31 · outbound

This paper cites A medical multimodal large language model for future pandemics.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis A medical multimodal large language model for future pandemics

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.989419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:86b075b55c03da47958b1cb3f782f27de6b655ab8b5199a3eabb67c42785ed28

Observation f34b5b53-04c8-42d3-92d0-3e879f4bb4a7 · outbound

This paper cites Multimodal generative ai for medical image interpre- tation.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Multimodal generative ai for medical image interpre- tation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.969361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:f55006eb11bb2bb11c08e0fa84c1ae6921162d9b7d57127094a3f6bce84a0c42

Observation 73c1a773-8cfb-4ca4-9173-e7e25da55669 · outbound

This paper cites Towards a holistic framework for multimodal llm in 3d brain ct radiology report generation.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Towards a holistic framework for multimodal llm in 3d brain ct radiology report generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.992390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:af18e3150ac85f7812f80f738899b2bbb8ec45a4b6361e92f7a0bd4f69e10744

Observation ec1470e2-6ae5-42a4-bf14-9aeadb954c13 · outbound

This paper cites MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:16:17.195024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:78aacebe9b5deb5630e08f0b305712c142afc3b6d20e580bb718142f787e01a5

Observation 7f52210f-b72d-46ea-9240-3943cd91dd5c · outbound

This paper cites Generating Radiology Reports via Memory-driven Transformer.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Generating Radiology Reports via Memory-driven Transformer

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.284897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:8ccf8717e98dd33bb449e7f5939c0eb50cdd9cd6de478e7b46b683ac20cdb3e4

Observation 156ea7ad-b8e0-4798-94ce-fd58a2803567 · outbound

This paper cites Promptmrg: Diagnosis-driven prompts for medical report generation.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Promptmrg: Diagnosis-driven prompts for medical report generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.950473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:51ae8885c4479fcc50b921348fe7d30aa48d117ed6b88ddcf42ef084088ecd8f

Observation be8c466a-9af3-481e-85be-b5bede69c09e · outbound

This paper cites Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.954280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:30abdb3c728b633ce41a408433a9986eb58e725eb305ed278384509b349eaf57

Observation 10e844f4-7d35-4be5-93c3-562763853f2d · outbound

This paper cites Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.983737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:317942bca3cba79c943b72344d31c4cceb5eebde4663cec1429a6f20b22b790b

Observation 5e236b72-212a-46ea-a51f-a2f8ccda19af · outbound

This paper cites Lmt++: Adaptively collaborating llms with multi- specialized teachers for continual vqa in robotic surgical videos.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Lmt++: Adaptively collaborating llms with multi- specialized teachers for continual vqa in robotic surgical videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.966276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:d6291ed330252941cc1345d0c41033d2d7b0c86ce363cee05c280ccdcd682003

Observation 275fd5fc-8d0a-4034-b588-7936e5fd2672 · outbound

This paper cites Interactive and ex- plainable region-guided radiology report generation.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Interactive and ex- plainable region-guided radiology report generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.980956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:a44090907329718d34f66ca66e9f764e8ee37bb79a470fd5f9fa871c0822e166

Observation 96144497-895a-42ca-8828-9ccef4c52098 · outbound

This paper cites SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.221059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:489ec90679218cc5a8a2e7a9af0d603d7939c0171ee74e005ed3aeb69351ed4a

Observation ac59c447-9289-4d80-beff-b11658f6d429 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:58.261086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:4a69df1978cf13a977e2ea337e7debe411fb6e55cf2a0727b48ef9e3518337d9

Observation 382e72aa-1d4f-44b2-89d4-4a856dc7b8f2 · outbound

This paper cites Roformer: En- hanced transformer with rotary position embedding.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Roformer: En- hanced transformer with rotary position embedding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.943394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:c8c11c4398eeeb1e43a3c81e65b9c25449384fbb2e548c558380b9e25e3e69c9

Observation 85709c67-c6e4-4469-9b7b-63c0de2ee9cf · outbound

This paper cites Do vision transformers see like convolutional neural networks?.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Do vision transformers see like convolutional neural networks?

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.963367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:3a7feb2985199638a76db9494c489758b371ea8c58e7e56edb8c0794aff6a4e6

Observation d72ed521-488d-4648-892b-802c79122c74 · outbound

This paper cites PaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional Experts.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis PaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional Experts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.298006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:88e235fdf3f2219e6b8e35c9a3d9da8f804170ad3f528ff5005f71956590fa27

Observation 6df0f719-58d3-4f66-977d-8a81f876bafd · outbound

This paper cites Scaling vision with sparse mixture of experts.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Scaling vision with sparse mixture of experts

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:43.001977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:494fa990482b58df5d96ee562a56ac6722269372406d6fe59a7a2c561f791ba9

Observation 95961ac8-9d37-4d97-8ab0-2d5c3ffe849c · outbound

This paper cites Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.213409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:a73ba83f5e8f8153428e8dea863b19b67831bcf9a1a5168b9c41422d5449c06b

Observation 90e1b575-01a0-404d-8864-38a4b2c59217 · outbound

This paper cites Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.235552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:e1e600f34194e6584bc1539d08782fa81741a69c2ff5a9a273adbdfebbf9bb69

Observation aefe669f-46d9-440c-8157-bff5acc24a51 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.459169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:3a1d730a23622179d2f75a92bb8a432f79f5420f99b6e1ac3df040bcd38acfef

Observation e759b238-5fae-4070-848e-dcc4c4d264a7 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Lora: Low-rank adaptation of large language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.971861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:9b39600f22c6268751816124d34abbf63f2654f552d82d57274a596e02a4bd24

Observation 18967fb0-cb06-4163-8415-1433c13aa1b3 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis Zero: Memory optimizations toward training trillion parameter models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:59:42.974919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:43:02.337806Z digest=sha256:0b07c99a9268243782eec69d32c7596436b5ba8ec06c8be0649c479d1795abbb

Pith citing papers

No inbound Pith citation observations are available.