Pith. sign in

REVIEW 5 cited by

ST-Align: A Multimodal Foundation Model for Image-Gene Alignment in Spatial Transcriptomics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.16793 v1 pith:RPOSQLU6 submitted 2024-11-25 cs.CV q-bio.GN

ST-Align: A Multimodal Foundation Model for Image-Gene Alignment in Spatial Transcriptomics

classification cs.CV q-bio.GN
keywords st-alignalignmentinsightsmultimodalspatialacrosseffectivelyencoders
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Spatial transcriptomics (ST) provides high-resolution pathological images and whole-transcriptomic expression profiles at individual spots across whole-slide scales. This setting makes it an ideal data source to develop multimodal foundation models. Although recent studies attempted to fine-tune visual encoders with trainable gene encoders based on spot-level, the absence of a wider slide perspective and spatial intrinsic relationships limits their ability to capture ST-specific insights effectively. Here, we introduce ST-Align, the first foundation model designed for ST that deeply aligns image-gene pairs by incorporating spatial context, effectively bridging pathological imaging with genomic features. We design a novel pretraining framework with a three-target alignment strategy for ST-Align, enabling (1) multi-scale alignment across image-gene pairs, capturing both spot- and niche-level contexts for a comprehensive perspective, and (2) cross-level alignment of multimodal insights, connecting localized cellular characteristics and broader tissue architecture. Additionally, ST-Align employs specialized encoders tailored to distinct ST contexts, followed by an Attention-Based Fusion Network (ABFN) for enhanced multimodal fusion, effectively merging domain-shared knowledge with ST-specific insights from both pathological and genomic data. We pre-trained ST-Align on 1.3 million spot-niche pairs and evaluated its performance through two downstream tasks across six datasets, demonstrating superior zero-shot and few-shot capabilities. ST-Align highlights the potential for reducing the cost of ST and providing valuable insights into the distinction of critical compositions within human tissue.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

    cs.CV 2026-05 unverdicted novelty 8.0

    VISTA is the first large-scale interaction-aware benchmark that decomposes videos into entities, actions, and relations to diagnose spatio-temporal biases in vision-language models.

  2. Gene Ontology-Guided Hierarchical Spatial Gene Expression Prediction from Histopathology Images

    cs.AI 2026-08 conditional novelty 6.0

    MSGR's Gene Ontology-guided hierarchical decoder improves spatial gene expression prediction from histology images, with the biological structure adding a +0.027 gain over an equivalent random hierarchy.

  3. HierarchicalDAEW: Domain-Aware Edge-Weighted Graph Convolution with Evidential Uncertainty for Multi-Section Spatial Gene Expression Prediction from H&E Histology

    cs.LG 2026-07 reject novelty 6.0

    HierarchicalDAEW predicts spatial gene expression with expression-derived domain edge typing and calibrated uncertainty, beating 13 baselines on breast Visium sections.

  4. VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

    cs.CV 2026-05 unverdicted novelty 6.0

    VISTA is a new ~12K-pair benchmark and taxonomy for open-set multi-entity spatio-temporal understanding in VLMs that decomposes videos into entities, actions, and relational dynamics for multi-axis diagnostics.

  5. Spatial Transcriptomics as Images for Large-Scale Pretraining

    cs.CV 2026-03 conditional novelty 5.0

    Cropping ST slides into fixed multi-channel gene patches preserves local spatial context, multiplies training samples, and beats spot- and slice-based pretraining on domain detection.