Pith. sign in

Paper Citation Record · LEDGER

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2507.07006.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07006 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:55:42.909470Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy12
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9932535-b6c4-47bc-a3f4-c420b54bb7be · outbound

This paper cites Inference of captions from histopatho- logical patches,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Inference of captions from histopatho- logical patches,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.712752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.712752Z digest=sha256:a455c1512d874863d8e0a881b88091404d2c4efac87dbe2f946c22464a3f4a8e

Observation 739cf3c5-fac6-455f-a33e-610924a0e4a3 · outbound

This paper cites Artificial intelligence for digital and compu- tational pathology,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Artificial intelligence for digital and compu- tational pathology,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.438872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.719966Z digest=sha256:0b73679d68c319e6b230b2dc07a59226d680ed22d8ac63739c244e4abe5c15bd

Observation 2139d102-de3e-4115-acf9-04a56fdd76db · outbound

This paper cites PathAlign: A vision-language model for whole slide images in histopathology.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning PathAlign: A vision-language model for whole slide images in histopathology

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.731503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.731503Z digest=sha256:0580c9e5618572a6fd1680cad97c27e030a436a568349cdfeca5d465ed16b11a

Observation 7befc8a2-bffe-4772-b5eb-76a5584d584c · outbound

This paper cites Attention-based deep multiple instance learning,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Attention-based deep multiple instance learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.422584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.741430Z digest=sha256:2224a3eba01462b753db32775d8f85664cae773ecefbb57caaf02c93a0c7207c

Observation 728e977a-1ac4-40dd-ad78-842a2b4fe611 · outbound

This paper cites Transmil: Transformer based correlated multiple instance learning for whole slide image classification,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Transmil: Transformer based correlated multiple instance learning for whole slide image classification,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.403249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.749496Z digest=sha256:75896d3eaee131e04978f51a99afd78886f245c6be5b31c0422589512928b07c

Observation 6f2a7822-8f9b-4d26-8dac-0a14827ebf07 · outbound

This paper cites A survey on graph-based deep learning for computational histopathology,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning A survey on graph-based deep learning for computational histopathology,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.384796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.757346Z digest=sha256:edeee73d25de5ef3ec3183c41dddf69e53c4e28f26764a7ffc6952d9e37fa72a

Observation cf583242-523d-4b31-a16a-a31f1ad19426 · outbound

This paper cites Enhanced descriptive captioning model for histopathological patches,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Enhanced descriptive captioning model for histopathological patches,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.368148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.765499Z digest=sha256:9113e1b569ede27b4c2bcd8182169f90afc2443b42cfdfe502ec641578eaeb2b

Observation 7b837ee8-7723-49bf-8f44-306d52212c12 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning LLaMA: Open and Efficient Foundation Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.773499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.773499Z digest=sha256:9838f1b96dd12b257dd59de62664ed9731356a04de1acd3ffb79eb90f8154d9e

Observation 02665d5e-0043-4810-b3c6-686424d71821 · outbound

This paper cites Clinicalt5: A generative language model for clinical text,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Clinicalt5: A generative language model for clinical text,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.780443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.780443Z digest=sha256:3e59605bb48df02e24e4d515f031886bf55ccc674926eefa73813e914af59f3d

Observation b8c755d5-f623-452f-9812-c7db0bba8579 · outbound

This paper cites Biogpt: generative pre-trained transformer for biomedical text genera- tion and mining,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Biogpt: generative pre-trained transformer for biomedical text genera- tion and mining,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.787842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.787842Z digest=sha256:09c5f2c104b87b290f95d4067bb93a03b265bd5d1baa3177f5116d6922fee932

Observation 6efab8ef-5069-41e8-84b1-25c65161e0c8 · outbound

This paper cites A generalist vision–language foundation model for diverse biomedical tasks,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning A generalist vision–language foundation model for diverse biomedical tasks,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.326042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.793771Z digest=sha256:8a3be4292053818c0c846727bfc3739ba7909d459b36cce3e366582e5a7d42a3

Observation 65261c9b-15de-480b-a638-794e88d896c9 · outbound

This paper cites Imagenet large scale visual recognition challenge,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Imagenet large scale visual recognition challenge,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.799479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.799479Z digest=sha256:881b77a2f0ee1d651e53d9ddc4aee9ba066f8125fa300a3b96f42b480cb25f92

Observation e266c0f8-8819-4563-a81e-561eb8c19ea5 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Learning transferable visual models from natural language supervision,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.807114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.807114Z digest=sha256:1abcfbe10ca4e23cb9f77f32698b894a9828a7df9b43fd931c1b6f907de1dc47

Observation 48256ba8-a2cf-45c3-8012-278cf66882fb · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.812086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.812086Z digest=sha256:030777499928849fb92e9458fbb617b7c4e53f889bd48df8a405ece228d65d6c

Observation 495defc2-7de4-4268-b6b0-fb972fd55273 · outbound

This paper cites PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:55:43.039262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.817005Z digest=sha256:a624ac353195d37e93bf520ea10b7e33782c715bdac5f4f0724d5fb4ec453715

Observation f7dc63f0-fc8b-47cf-8f76-817d84553ec2 · outbound

This paper cites Dual-stream multiple instance learn- ing network for whole slide image classification with self-supervised contrastive learning,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Dual-stream multiple instance learn- ing network for whole slide image classification with self-supervised contrastive learning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.827274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.827274Z digest=sha256:8a1c7a2b673f6c67e11ef0202d5c6b7cd02078b07322b0d26264ba5c38e5b2ff

Observation b4a3ca2f-297d-4129-bd08-f09cd1b4717c · outbound

This paper cites Dtfd-mil: Double-tier feature distillation multiple instance learning for histopathology whole slide image classification,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Dtfd-mil: Double-tier feature distillation multiple instance learning for histopathology whole slide image classification,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.833688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.833688Z digest=sha256:a945ad36fbfb6bbbe4a39bd40c2878a30c2819b73cf61cca8bcf5fb2f424bc54

Observation 97216124-38e4-4580-af9d-2cfcef796500 · outbound

This paper cites Visual language pretrained multiple instance zero-shot transfer for histopathology images,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Visual language pretrained multiple instance zero-shot transfer for histopathology images,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.244937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.839603Z digest=sha256:0c76efb8196d64c5c69dc2cf28a5bf032c87502f34d72b847a35911a4c3a5090

Observation 5109a3e4-46d6-4c05-8712-f734fc9d1172 · outbound

This paper cites CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.848632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.848632Z digest=sha256:496e22176fa57691a3b5e79c65a6d63a4cd6bdf7efbb7fa4a4c9f925546313ce

Observation 2658958b-a776-418b-bdd7-870f4b9e42d3 · outbound

This paper cites Large language models in healthcare and medical domain: A review,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Large language models in healthcare and medical domain: A review,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.221309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.856878Z digest=sha256:3a72b6f8d023ea7a867fbbf3829548d4d68d6303cee5b9c9debb35986b202a84

Observation 02e84621-ff0b-464c-b8bd-93d432a6204c · outbound

This paper cites What a Whole Slide Image Can Tell? Subtype-guided Masked Transformer for Pathological Image Captioning.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning What a Whole Slide Image Can Tell? Subtype-guided Masked Transformer for Pathological Image Captioning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:55:42.992013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.863470Z digest=sha256:3ec797a40c9fa78dfacb09d55c1671c0a8b122b297e8339ba77abcbf96f54090

Observation 1de277b8-78c8-49df-8c52-4116cf10e648 · outbound

This paper cites Patch-based convolutional neural network for whole slide tissue image classification,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Patch-based convolutional neural network for whole slide tissue image classification,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.197120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.869493Z digest=sha256:a8df1ffae36274b999bf58097dfbbbdbced685063fb816ee50613b4055b707e8

Observation 68dadbd5-0921-47af-b719-846e4ebc45d4 · outbound

This paper cites Unsupervised deep embedding for clustering analysis,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Unsupervised deep embedding for clustering analysis,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.173062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.874812Z digest=sha256:4bc2848bcc65c86b33b492fa9674294c6f75a4d703ffa6b02df95f5cce473e7d

Observation 424ea3b9-5158-4c0b-b301-59e52794ab7d · outbound

This paper cites Attention is all you need,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Attention is all you need,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.879763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.879763Z digest=sha256:6e742df361876df47cf5388a93723ce6d26d212658d885a865e6dd7752c33e6b

Observation 3b64879b-1ced-4796-877f-520141799fea · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Categorical Reparameterization with Gumbel-Softmax

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.885997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.885997Z digest=sha256:f68d1eb5f700cf69a5836a20360750e4a84cf13c5ed7deee8ba012e22405b2cb

Observation 7e0bb6d2-6446-4a3e-9a87-5c3cbc466655 · outbound

This paper cites Graph attention networks,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Graph attention networks,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.893311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.893311Z digest=sha256:6d05bdcaab9d443b5af407a2cb9d85c430d11acc6cb8e5baabe6c7ceed6734d6

Observation 18ce9874-2ea8-4468-a38a-e9a3918433be · outbound

This paper cites Rectifier nonlinearities improve neural network acoustic models,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning Rectifier nonlinearities improve neural network acoustic models,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:42.898607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:42.898607Z digest=sha256:9efcc1ad604375aeef57b6ae59665880e4c9f6f6be2518b65b16a860952e6381

Observation 54c78102-d8bc-4517-8bc0-9e5faaa68d60 · outbound

This paper cites A dataset for breast cancer histopathological image classification,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning A dataset for breast cancer histopathological image classification,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.122386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.904383Z digest=sha256:cb91f84f984b83f7ad422e554675f05d9e7593aa65dbae9b7e72f750762c3ab3

Observation be848d36-dd34-4587-b09c-cf81813c94c5 · outbound

This paper cites A survey of evaluation metrics used for nlg systems,.

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning A survey of evaluation metrics used for nlg systems,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:55:43.099400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:55:42.909470Z digest=sha256:0ed8c3dcecea3233d177d931479f5bb553b2d7545ded7d7dce285708c4951ac8

Pith citing papers

No inbound Pith citation observations are available.