Pith. sign in

Paper Citation Record · LEDGER

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning

As of 18 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2506.17645.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17645 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:32:35.305415Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8cc23c65-a2ca-44da-9458-d6c17bb96daa · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Bottom-up and top-down attention for image captioning and visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.580028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.224862Z digest=sha256:d7f5208328f84c56ee40b41cea045097e1c8680d02c4144fbbd7043c70b12360

Observation 92a41822-b70e-43f3-9ee0-3f738d7de9b3 · outbound

This paper cites Improving diagnostic accuracy through feedback: The diagnosis learning cycle.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Improving diagnostic accuracy through feedback: The diagnosis learning cycle

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.565101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.231800Z digest=sha256:f5d53c75a354e11a48d1238d0cd35721581218242672d615aace5018fde522e4

Observation 36faa479-fbb2-4991-a684-b397d6576d12 · outbound

This paper cites an unresolved cited work.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:32:35.550041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.238497Z digest=sha256:6e398ca621d82080569d633112d696556d4952ccb1b8b82213874e0fdd9ab03b

Observation 4ef30ea5-f1dc-4879-8658-fe0b25ccc5b6 · outbound

This paper cites Wsicaption: Multiple instance generation of pathology reports for gigapixel whole-slide images.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Wsicaption: Multiple instance generation of pathology reports for gigapixel whole-slide images

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.534492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.243951Z digest=sha256:63ea11cdb36c75a7e42860be6a76c3fbc7ade0f002269be09607e21dbae4b975

Observation 5a33e654-c73a-47e6-bd2e-6f1040da5214 · outbound

This paper cites Generating radiology reports via memory-driven transformer.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Generating radiology reports via memory-driven transformer

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.518570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.249315Z digest=sha256:951e860f0332845556d34e2b481c174d22bc055ace6123efa033412f4e96bbb8

Observation 8d710a77-3c1d-42d0-84df-e4241bc0daef · outbound

This paper cites Cross-modal memory networks for radiology report generation.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Cross-modal memory networks for radiology report generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.503695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.254714Z digest=sha256:751c4237b2f5aacbf86f86f558576ad9e4b8057efac174484e4ec43a9aa93cc8

Observation e256619f-cfea-4e5f-9dc6-919addd5f14d · outbound

This paper cites Meshed-memory transformer for image captioning.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Meshed-memory transformer for image captioning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.487915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.259937Z digest=sha256:90d96db78bf54374d9a70a082c38f945cdfe138174c901a00f33bb8a6278a2f2

Observation 06a2b681-89ad-4a71-8eee-380cc1e5160f · outbound

This paper cites Anatomic pathology quality assurance through peer review.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Anatomic pathology quality assurance through peer review

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.472092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.264991Z digest=sha256:9b59e7649a0d7a51a457ef318a49538dda2b7cd571e3b4f88ad14e405defe3b1

Observation ead99c06-f324-4be9-8ef3-d22252b89824 · outbound

This paper cites Histgen: A local-global encoding framework for pathology report generation.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Histgen: A local-global encoding framework for pathology report generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.457182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.270208Z digest=sha256:0a3fa2361fa497be10a2dae0b5273d6992b83f58e3b049e4c8171668d74b2b31

Observation 079348f2-6b39-4a8d-bc02-2f20f3a05237 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Lora: Low-rank adaptation of large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.441528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.275250Z digest=sha256:05c1cb6af6e0ebaa7ec9ef39d1ac232e4ee7eb16fa83f62495e47e1265339b36

Observation 22c56e16-71e5-47be-ad87-3b61b9a2f6df · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.424255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.280246Z digest=sha256:d82596de6aebcbc9397d059c69fc6d05863c274cbb84ce169f1d71df791ea297

Observation b68315e8-bcbd-494f-af58-e4f1f0f60c03 · outbound

This paper cites Llava-med: Training a large language-and-vision assistant for biomedicine in one day.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Llava-med: Training a large language-and-vision assistant for biomedicine in one day

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.408117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.284848Z digest=sha256:6a60d971164af789ee713d1711b18aa672463b3504ba3ef04ed71e29fcf08872

Observation 761e1ef5-6bd5-4d2f-85dc-1d8f8116607c · outbound

This paper cites Improving factual completeness and consistency of image-to-text radiology report generation.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Improving factual completeness and consistency of image-to-text radiology report generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.393332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.290054Z digest=sha256:74eed37ba5cb55fc2931756f4268e51ea653f139a116d26556b7c62176cd48ea

Observation de8d5563-737d-4acf-b3df-d12cecef4cdd · outbound

This paper cites Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:32:35.294994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:32:35.294994Z digest=sha256:95d2744495b88018fac29c5f6ead8de9c46946ce77a78325a7cb820775dbebc1

Observation 61be891b-0c9a-4df0-afc1-15bf991767a0 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.376829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.300180Z digest=sha256:31390dc51e670e9caa4dec894d371174f8fcf319c4f0fd23aa196227b54b1e2c

Observation 6c652973-3b90-4eae-b8c8-6cfe09c24fe4 · outbound

This paper cites Show and tell: A neural image caption generator.

Histopathology Image Report Generation by Vision Language Model with Multimodal In-Context Learning Show and tell: A neural image caption generator

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:32:35.360454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:32:35.305415Z digest=sha256:ab2ba290346f8bcaa7fdd3a02060027ecd32ed39b993d3d07a07c129230cb89a

Pith citing papers

No inbound Pith citation observations are available.