Pith. sign in

REVIEW 3 cited by

A Billion-scale Foundation Model for Remote Sensing Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.05215 v4 pith:TK5BLWSM submitted 2023-04-11 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords foundationmodelsparameterstasksdownstreammodelperformancepretraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As the potential of foundation models in visual tasks has garnered significant attention, pretraining these models before downstream tasks has become a crucial step. The three key factors in pretraining foundation models are the pretraining method, the size of the pretraining dataset, and the number of model parameters. Recently, research in the remote sensing field has focused primarily on the pretraining method and the size of the dataset, with limited emphasis on the number of model parameters. This paper addresses this gap by examining the effect of increasing the number of model parameters on the performance of foundation models in downstream tasks such as rotated object detection and semantic segmentation. We pretrained foundation models with varying numbers of parameters, including 86M, 605.26M, 1.3B, and 2.4B, to determine whether performance in downstream tasks improved with an increase in parameters. To the best of our knowledge, this is the first billion-scale foundation model in the remote sensing field. Furthermore, we propose an effective method for scaling up and fine-tuning a vision transformer in the remote sensing field. To evaluate general performance in downstream tasks, we employed the DOTA v2.0 and DIOR-R benchmark datasets for rotated object detection, and the Potsdam and LoveDA datasets for semantic segmentation. Experimental results demonstrated that, across all benchmark datasets and downstream tasks, the performance of the foundation models and data efficiency improved as the number of parameters increased. Moreover, our models achieve the state-of-the-art performance on several datasets including DIOR-R, Postdam, and LoveDA.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CGEarthEye:A High-Resolution Remote Sensing Vision Foundation Model Based on the Jilin-1 Satellite Constellation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 15-million-image sub-meter remote sensing dataset built from Jilin-1 imagery and a multi-scale self-supervised ViT framework achieve benchmark results comparable to or better than prior remote sensing foundation models.

  2. Distributed Cross-Channel Hierarchical Aggregation for Foundation Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) reduces memory and boosts throughput for multi-channel vision foundation models by spreading tokenization and channel fusion across GPUs with only a small qu...

  3. SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A unified multi-modal remote sensing foundation model with adaptive patch merging, modality prompt tokens, mixture of experts, and query-based semantic aggregation contrastive learning outperforms SkySense by 1.8 poin...

Pith tools