Pith. sign in

REVIEW 2 cited by

How Effective is Pre-training of Large Masked Autoencoders for Downstream Earth Observation Tasks?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.18536 v1 pith:DV72O6EA submitted 2024-09-27 cs.CV

classification cs.CV
keywords tasksdownstreampre-trainingclassificationeffectivemodelsreconstructiontask
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised pre-training has proven highly effective for many computer vision tasks, particularly when labelled data are scarce. In the context of Earth Observation (EO), foundation models and various other Vision Transformer (ViT)-based approaches have been successfully applied for transfer learning to downstream tasks. However, it remains unclear under which conditions pre-trained models offer significant advantages over training from scratch. In this study, we investigate the effectiveness of pre-training ViT-based Masked Autoencoders (MAE) for downstream EO tasks, focusing on reconstruction, segmentation, and classification. We consider two large ViT-based MAE pre-trained models: a foundation model (Prithvi) and SatMAE. We evaluate Prithvi on reconstruction and segmentation-based downstream tasks, and for SatMAE we assess its performance on a classification downstream task. Our findings suggest that pre-training is particularly beneficial when the fine-tuning task closely resembles the pre-training task, e.g. reconstruction. In contrast, for tasks such as segmentation or classification, training from scratch with specific hyperparameter adjustments proved to be equally or more effective.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Masked Autoencoders for Ultrasound Signals: Robust Representation Learning for Downstream Applications

    cs.LG 2025-08 conditional novelty 6.0 of 10

    MAE pre-training on synthetic ultrasound signals transfers to real measured signals and beats from-scratch and CNN baselines on time-of-flight classification, with the biggest gains in low-label regimes.

  2. MultiMAE Meets Earth Observation: Pre-training Multi-modal Multi-task Masked Autoencoders for Earth Observation Tasks

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A ViT-based MultiMAE pre-trained on MMEarth with split Sentinel-2 bands, elevation, and segmentation labels transfers to several EO classification and segmentation datasets.

Pith tools