Pith. sign in

REVIEW 2 cited by

ViTs for SITS: Vision Transformers for Satellite Image Time Series

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.04944 v3 pith:TOZDIO2F submitted 2023-01-12 cs.CV cs.LG

classification cs.CVcs.LG
keywords sitsmodeltimevisionavailableimagenovelprocessing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper we introduce the Temporo-Spatial Vision Transformer (TSViT), a fully-attentional model for general Satellite Image Time Series (SITS) processing based on the Vision Transformer (ViT). TSViT splits a SITS record into non-overlapping patches in space and time which are tokenized and subsequently processed by a factorized temporo-spatial encoder. We argue, that in contrast to natural images, a temporal-then-spatial factorization is more intuitive for SITS processing and present experimental evidence for this claim. Additionally, we enhance the model's discriminative power by introducing two novel mechanisms for acquisition-time-specific temporal positional encodings and multiple learnable class tokens. The effect of all novel design choices is evaluated through an extensive ablation study. Our proposed architecture achieves state-of-the-art performance, surpassing previous approaches by a significant margin in three publicly available SITS semantic segmentation and classification datasets. All model, training and evaluation codes are made publicly available to facilitate further research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Retrieval of Surface Solar Radiation through Implicit Albedo Recovery from Temporal Context

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An attention-based model retrieves surface solar radiation from satellite image sequences and matches albedo-informed models when given about 40 hours of temporal context.

  2. Smartflow: Enabling Scalable Spatiotemporal Geospatial Research

    cs.CV 2025-06 reject novelty 4.0 of 10

    A cloud framework for scalable geospatial analysis is described together with a qualitative demonstration of a construction-detection model.

Pith tools