REVIEW 3 cited by
When an Image is Worth 1,024 x 1,024 Words: A Case Study in Computational Pathology
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This technical report presents LongViT, a vision Transformer that can process gigapixel images in an end-to-end manner. Specifically, we split the gigapixel image into a sequence of millions of patches and project them linearly into embeddings. LongNet is then employed to model the extremely long sequence, generating representations that capture both short-range and long-range dependencies. The linear computation complexity of LongNet, along with its distributed algorithm, enables us to overcome the constraints of both computation and memory. We apply LongViT in the field of computational pathology, aiming for cancer diagnosis and prognosis within gigapixel whole-slide images. Experimental results demonstrate that LongViT effectively encodes gigapixel images and outperforms previous state-of-the-art methods on cancer subtyping and survival prediction. Code and models will be available at https://aka.ms/LongViT.
Forward citations
Cited by 3 Pith papers
-
From Pixels to Gigapixels: Bridging Local Inductive Bias and Long-Range Dependencies with Pixel-Mamba
Pixel-Mamba, an end-to-end Mamba-based architecture with progressive token expansion, reports tumor staging and survival scores on three TCGA datasets that match or exceed several pathology foundation models without p...
-
Diagnostic Text-guided Representation Learning in Hierarchical Classification for Pathological Whole Slide Image
PathTree improves whole slide image classification by organizing disease categories as a binary tree and using expert-written pathological text to guide feature aggregation, achieving state-of-the-art results on three...
-
Cross-Modality Learning for Predicting IHC Biomarkers from H&E-Stained Whole-Slide Images
HistoStainAlign uses contrastive alignment between paired H&E and IHC whole-slide embeddings to predict P53, PD-L1, and Ki-67 status from H&E slides alone, achieving moderate F1 scores on small internal datasets.
Discussion (0). Continue with ORCID to comment.