Pith. sign in

Paper Citation Record · LEDGER

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

As of 9 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2508.20188.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20188 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:18:46.209624Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c91005ca-6987-4f4c-9641-6d12bc93fcf0 · outbound

This paper cites MedImageInsight: An Open-Source Embedding Model for General Domain Medical Imaging.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study MedImageInsight: An Open-Source Embedding Model for General Domain Medical Imaging

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:44.377415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:44.377415Z digest=sha256:cc3c021b9b841457c3e41b53eb80a5850115d95b69412ba521428af9f56a73b8

Observation b0b28e8d-3d76-440c-8f9e-27eeb4d521ec · outbound

This paper cites Validation of artificial intelligence prediction models for skin cancer diagnosis using dermoscopy images: the 2019 international skin imaging collaboration grand challenge.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Validation of artificial intelligence prediction models for skin cancer diagnosis using dermoscopy images: the 2019 international skin imaging collaboration grand challenge

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:49.080855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:44.439846Z digest=sha256:e8ef456553ee936c2547637e6a8bff4aa07d60b496fa22e017972a0f35bfcdc2

Observation e79d4228-dae4-4244-9455-8673d2e8745c · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study An image is worth 16x16 words: Transformers for image recognition at scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:44.544462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:44.544462Z digest=sha256:dafaf46349ad1e6f4baf49af065b4b446bbf78d193f5b92c227e37fa106c9abb

Observation 5a16eecf-1504-447a-8fa7-1edfd031d9ce · outbound

This paper cites Analysis of trends in geographic distribution of us dermatology workforce density.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Analysis of trends in geographic distribution of us dermatology workforce density

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.831526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:44.610713Z digest=sha256:7b2d90f6edc69f316724798ef70f691971466f6c842f413058a9dadbdbd5b8ba

Observation 933d7108-5a32-45fa-a79d-ddecd706221b · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study LoRA: Low-rank adaptation of large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:44.684834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:44.684834Z digest=sha256:9f32bf0636435c9e8995d40f69067c5d78e53e037af1ead29104de3fd1a867bc

Observation 1983dd57-45d5-400c-aedb-0c3c9b4f10e5 · outbound

This paper cites Transparent medical image ai via an image–text foundation model grounded in medical literature.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Transparent medical image ai via an image–text foundation model grounded in medical literature

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.588042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:44.760942Z digest=sha256:fb0ef415eef2bffc330baf59f870e7bfbc2c0d2faf5cc9426782e1bf84a9fefd

Observation 27d36465-a168-496c-8511-8e2f48e32d3b · outbound

This paper cites Human-ai interaction in skin cancer diagnosis: a systematic review and meta-analysis.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Human-ai interaction in skin cancer diagnosis: a systematic review and meta-analysis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.428291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:44.831224Z digest=sha256:fc7bee8ffa4fb524f28948bbbabbddeaabe0ee172b6eaff97b3b72783f62a95f

Observation a3fee907-6528-46d1-a91f-5d3aac03d17a · outbound

This paper cites The slice-3d dataset: 400,000 skin lesion image crops extracted from 3d tbp for skin cancer detection.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study The slice-3d dataset: 400,000 skin lesion image crops extracted from 3d tbp for skin cancer detection

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.172329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:44.937867Z digest=sha256:ce6885eff1f11147a1a91312795f085a5a2b62627fdd0919253d793efd206f28

Observation c931a68e-0be6-483a-b49d-4896618de5bf · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.014349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.014349Z digest=sha256:7acdea4b2931ded083ceb112940032e592647d44a3ad0f2881a1fe98b1f67745

Observation 8478ded3-3bee-4586-9f57-563c5a983952 · outbound

This paper cites 3d whole-body skin imaging for automated melanoma detection.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study 3d whole-body skin imaging for automated melanoma detection

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.022744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.098968Z digest=sha256:d2141092ff9c07c92af00952cd47292db288b0be1c3eab44b658165d840f455c

Observation fb5e098f-c316-43cf-8753-15d5ed2c38b1 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Learning transferable visual models from natural language supervision

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.894507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.198250Z digest=sha256:0cd2f23aea162b2e566a7b187aa6442ab49e63ca1edbce32525742c376beaea1

Observation e275ef87-4999-4481-954a-b740839a2537 · outbound

This paper cites Steering llama 2 via contrastive activation addition.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Steering llama 2 via contrastive activation addition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.719743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.305119Z digest=sha256:145e4155c1df1f5b5989193e031adb8c54d211d48ace57574ca70b3b38c3841e

Observation 17bb9f56-093f-4dcb-b880-cf22a6bfcc39 · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.416054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.416054Z digest=sha256:4b70240a16efb696df9dfc400e8637f0085b310f5112934cf6404495189631a7

Observation d3e5c552-b236-4167-be21-716c29041771 · outbound

This paper cites Socioeconomic and geographic barriers to dermatology care in urban and rural us populations.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Socioeconomic and geographic barriers to dermatology care in urban and rural us populations

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.508759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.489390Z digest=sha256:79c54c8caa1ec44b8fc197efe34ef0b40aa7a0494043f27ad838ac65e5500c1a

Observation b55f806f-f8ff-45cb-b12e-1bb749c356bc · outbound

This paper cites Composing text and image for image retrieval-an empirical odyssey.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Composing text and image for image retrieval-an empirical odyssey

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.360891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.558953Z digest=sha256:9136a3e721c79088f5664196de7be9aa30cc550b6b93a6d7cb82f62c90fa1247

Observation 83d6baf8-9c44-4186-b411-733ac619071d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.626372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.626372Z digest=sha256:2d96e697ab708f28dc7c12f79b6cffdc26d513585bca9c01090f642fa83eddc2

Observation 49e3d848-5bca-4379-ba46-518088999649 · outbound

This paper cites VisNumBench: Evaluating Number Sense of Multimodal Large Language Models.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study VisNumBench: Evaluating Number Sense of Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.699326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.699326Z digest=sha256:8eb8f6591c0b352b429ac1aa12f228d1ef4c42af5c529d7b3aceb88f09a9a9c3

Observation 1b61aaa9-76cf-41ff-9db0-539acd98ffeb · outbound

This paper cites Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.769309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.769309Z digest=sha256:cfbe57bf547646695558cb6397f02aaf390b6f664eb43a02636c9dbca2c5df28

Observation b776e1cc-ef69-4b29-8767-f9e37f939b3b · outbound

This paper cites A multimodal vision foundation model for clinical dermatology.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study A multimodal vision foundation model for clinical dermatology

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.196661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.818436Z digest=sha256:d3a298fc8ced49085710dcd43d4430152dde539538eb90ff5b32b82d145b041d

Observation 6b2d822d-bf12-435f-bb7f-beb71bf363eb · outbound

This paper cites Qwen2 Technical Report.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Qwen2 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.884802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.884802Z digest=sha256:ee797b59bd41188544d036336d8a319075b4d6bd1b3e6c6872b59497a4a3ce2e

Observation 177cd5a3-b2db-4ad6-8b59-acf93633a55f · outbound

This paper cites Visrag: Vision-based retrieval-augmented generation on multi-modality documents.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Visrag: Vision-based retrieval-augmented generation on multi-modality documents

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:46.973569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.992780Z digest=sha256:e47f35f3f335bd380b52d58667737023e15be8543cd71eba3389334d60a87a09

Observation 49f59181-aba5-4486-86a5-4641af924ae8 · outbound

This paper cites MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:18:46.381344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:46.079905Z digest=sha256:53b2b4c050110d761706c80f522b91100adfe9c90c7eadea05c480283f458ed0

Observation 2a98d987-e749-44ef-a093-65e209631c1c · outbound

This paper cites Revisiting the trustworthiness of saliency methods in radiology ai.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Revisiting the trustworthiness of saliency methods in radiology ai

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:46.821178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:46.140687Z digest=sha256:31e9b6f494d635c83d059156bb2b8a73a306715b21275887027d37e822716e5c

Observation a19ca186-d62a-46ee-b73c-647881d6bf97 · outbound

This paper cites Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:46.651647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:46.209624Z digest=sha256:62d8b54d2d2e5640c1060679c71f22df7eb1ebae715a067b7fa8ebfb77409336

Pith citing papers

No inbound Pith citation observations are available.