Pith. sign in

Paper Citation Record · LEDGER

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

As of 10 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2508.20188.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20188 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:18:46.209624Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c91005ca-6987-4f4c-9641-6d12bc93fcf0 · outbound

This paper cites MedImageInsight: An Open-Source Embedding Model for General Domain Medical Imaging.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study MedImageInsight: An Open-Source Embedding Model for General Domain Medical Imaging

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:44.377415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:44.377415Z digest=sha256:8ca25c5765fd70e8c84bac7e2969d1e0886ca1eb86911441c9714f57efe36824

Observation b0b28e8d-3d76-440c-8f9e-27eeb4d521ec · outbound

This paper cites Validation of artificial intelligence prediction models for skin cancer diagnosis using dermoscopy images: the 2019 international skin imaging collaboration grand challenge.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Validation of artificial intelligence prediction models for skin cancer diagnosis using dermoscopy images: the 2019 international skin imaging collaboration grand challenge

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:49.080855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:44.439846Z digest=sha256:b7d7b8e0701b0f1df0c5cec46e21cdfb9bd599b1530c9d9183a88ed344fc1f06

Observation e79d4228-dae4-4244-9455-8673d2e8745c · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study An image is worth 16x16 words: Transformers for image recognition at scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:44.544462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:44.544462Z digest=sha256:17d51dea2ce5068f711333d9214eb4df9804a85f55d0c7d9718d6ae5166f8d7d

Observation 5a16eecf-1504-447a-8fa7-1edfd031d9ce · outbound

This paper cites Analysis of trends in geographic distribution of us dermatology workforce density.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Analysis of trends in geographic distribution of us dermatology workforce density

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.831526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:44.610713Z digest=sha256:65bc6a08c22ad3498d8e53f3aaf16b9c5fb5dad421307df4fc40481f80c3578c

Observation 933d7108-5a32-45fa-a79d-ddecd706221b · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study LoRA: Low-rank adaptation of large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:44.684834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:44.684834Z digest=sha256:8f41b7d3984ed0dafbb92031a77a628aa2316b36ddfb559da2f4b3cc3f915eaa

Observation 1983dd57-45d5-400c-aedb-0c3c9b4f10e5 · outbound

This paper cites Transparent medical image ai via an image–text foundation model grounded in medical literature.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Transparent medical image ai via an image–text foundation model grounded in medical literature

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.588042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:44.760942Z digest=sha256:eed9ab7cea82dfb7d3202aad4fbe22841e3685b6244a947a78772d3f95b8d96f

Observation 27d36465-a168-496c-8511-8e2f48e32d3b · outbound

This paper cites Human-ai interaction in skin cancer diagnosis: a systematic review and meta-analysis.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Human-ai interaction in skin cancer diagnosis: a systematic review and meta-analysis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.428291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:44.831224Z digest=sha256:08b77aa5009bbe67572dd20e8c52ce397e3d44bae6b94a5e4cf4af2e0c1f6c0a

Observation a3fee907-6528-46d1-a91f-5d3aac03d17a · outbound

This paper cites The slice-3d dataset: 400,000 skin lesion image crops extracted from 3d tbp for skin cancer detection.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study The slice-3d dataset: 400,000 skin lesion image crops extracted from 3d tbp for skin cancer detection

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.172329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:44.937867Z digest=sha256:9bb77132c24def1cb78ea90c6b92f5bde5ece20a3ad960527a2ceac3987907b6

Observation c931a68e-0be6-483a-b49d-4896618de5bf · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.014349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.014349Z digest=sha256:b6294e7fcaeba2fdb96bd364643d62d1e65cfde4f13fb2bdc2b22cafc8fcc269

Observation 8478ded3-3bee-4586-9f57-563c5a983952 · outbound

This paper cites 3d whole-body skin imaging for automated melanoma detection.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study 3d whole-body skin imaging for automated melanoma detection

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.022744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.098968Z digest=sha256:8803ad92f9e4453d082208f9f0daa5c6799b207e4beb001af186ff2d02d17f93

Observation fb5e098f-c316-43cf-8753-15d5ed2c38b1 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Learning transferable visual models from natural language supervision

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.894507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.198250Z digest=sha256:5dfb983d6243d14a82adfd18789a0ec764d372cbeb5d9ea4abef6cd5203c8e58

Observation e275ef87-4999-4481-954a-b740839a2537 · outbound

This paper cites Steering llama 2 via contrastive activation addition.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Steering llama 2 via contrastive activation addition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.719743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.305119Z digest=sha256:d00a2774b92b3f8ef84d6cd8ea66474d38435cb0ecb6299d94bafbbc0f1e723c

Observation 17bb9f56-093f-4dcb-b880-cf22a6bfcc39 · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.416054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.416054Z digest=sha256:03b4fdd410e886c65c92e79297cbbd5551bf951f7956d477b4b7f455a16edc33

Observation d3e5c552-b236-4167-be21-716c29041771 · outbound

This paper cites Socioeconomic and geographic barriers to dermatology care in urban and rural us populations.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Socioeconomic and geographic barriers to dermatology care in urban and rural us populations

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.508759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.489390Z digest=sha256:a9f666b9a23efdaf9f7533c25cbbe1ba54cf18d95e49c491a67c7702cacb5e69

Observation b55f806f-f8ff-45cb-b12e-1bb749c356bc · outbound

This paper cites Composing text and image for image retrieval-an empirical odyssey.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Composing text and image for image retrieval-an empirical odyssey

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.360891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.558953Z digest=sha256:078a40812425e01340eb6f61f8bade0b07ad21b696b77cd26501fa05d58d3997

Observation 83d6baf8-9c44-4186-b411-733ac619071d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.626372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.626372Z digest=sha256:9ff82c0342f326f8db2aace0419b81cbb3ec618bf25c9c2f449813d1e1ac5028

Observation 49e3d848-5bca-4379-ba46-518088999649 · outbound

This paper cites VisNumBench: Evaluating Number Sense of Multimodal Large Language Models.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study VisNumBench: Evaluating Number Sense of Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.699326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.699326Z digest=sha256:b298ba8e9517c9f01f8e938711059bb1d8688caacc3194099c73a06fc5e628c5

Observation 1b61aaa9-76cf-41ff-9db0-539acd98ffeb · outbound

This paper cites Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.769309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.769309Z digest=sha256:7c7ea0946a5c40691f82d1acb18f8630693c0320d24318a5c3d7d24cef020c50

Observation b776e1cc-ef69-4b29-8767-f9e37f939b3b · outbound

This paper cites A multimodal vision foundation model for clinical dermatology.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study A multimodal vision foundation model for clinical dermatology

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.196661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.818436Z digest=sha256:f0676d25d94857275884d8299b40df93eaeb9e8a24b2b4b8aedc362a1d838240

Observation 6b2d822d-bf12-435f-bb7f-beb71bf363eb · outbound

This paper cites Qwen2 Technical Report.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Qwen2 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.884802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.884802Z digest=sha256:4e64d6f0bc8119c3f5e4ca57f3b3c34817378f0887e33e30996b64854491ed89

Observation 177cd5a3-b2db-4ad6-8b59-acf93633a55f · outbound

This paper cites Visrag: Vision-based retrieval-augmented generation on multi-modality documents.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Visrag: Vision-based retrieval-augmented generation on multi-modality documents

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:46.973569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:45.992780Z digest=sha256:e57cf80d01959b93f908028f1e59a7e9b253618c6048d8758514af7ce104cf59

Observation 49f59181-aba5-4486-86a5-4641af924ae8 · outbound

This paper cites MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:18:46.381344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:46.079905Z digest=sha256:890358d6fba2742d97d5b41bc74126ae35243fe2ee8c0d517c1406e513d6a326

Observation 2a98d987-e749-44ef-a093-65e209631c1c · outbound

This paper cites Revisiting the trustworthiness of saliency methods in radiology ai.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Revisiting the trustworthiness of saliency methods in radiology ai

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:46.821178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:46.140687Z digest=sha256:0a198f330e3644d50d40153fb6c5a3e3db90c8a79ea50f11067c93de93ef104e

Observation a19ca186-d62a-46ee-b73c-647881d6bf97 · outbound

This paper cites Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:46.651647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:18:46.209624Z digest=sha256:51f8f993e76394d95c250abdebea6173c669f70f0a77cdf897de22f316737603

Pith citing papers

No inbound Pith citation observations are available.