Pith. sign in

REVIEW 2 cited by

Revisiting the Encoding of Satellite Image Time Series

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.02086 v2 pith:DJQZDC7G submitted 2023-05-03 cs.CV

classification cs.CV
keywords sitsimagelearningrepresentationsatellitesegmentationtemporaladvances
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Satellite Image Time Series (SITS) representation learning is complex due to high spatiotemporal resolutions, irregular acquisition times, and intricate spatiotemporal interactions. These challenges result in specialized neural network architectures tailored for SITS analysis. The field has witnessed promising results achieved by pioneering researchers, but transferring the latest advances or established paradigms from Computer Vision (CV) to SITS is still highly challenging due to the existing suboptimal representation learning framework. In this paper, we develop a novel perspective of SITS processing as a direct set prediction problem, inspired by the recent trend in adopting query-based transformer decoders to streamline the object detection or image segmentation pipeline. We further propose to decompose the representation learning process of SITS into three explicit steps: collect-update-distribute, which is computationally efficient and suits for irregularly-sampled and asynchronous temporal satellite observations. Facilitated by the unique reformulation, our proposed temporal learning backbone of SITS, initially pre-trained on the resource efficient pixel-set format and then fine-tuned on the downstream dense prediction tasks, has attained new state-of-the-art (SOTA) results on the PASTIS benchmark dataset. Specifically, the clear separation between temporal and spatial components in the semantic/panoptic segmentation pipeline of SITS makes us leverage the latest advances in CV, such as the universal image segmentation architecture, resulting in a noticeable 2.5 points increase in mIoU and 8.8 points increase in PQ, respectively, compared to the best scores reported so far.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. $T^{3}S$: Think in Thermal Time for Generalizable Crop Mapping from Satellite Image Time Series

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Sampling satellite time series by cumulative growing degree days instead of calendar time improves cross-year crop classification accuracy and uncertainty calibration.

  2. A Joint Learning Framework with Feature Reconstruction and Prediction for Incomplete Satellite Image Time Series in Agricultural Semantic Segmentation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A joint teacher-student training framework with feature reconstruction and prediction improves cropland extraction and crop classification accuracy on incomplete satellite image time series.

Pith tools