Pith. sign in

REVIEW 2 cited by

One for All: Toward Unified Foundation Models for Earth Vision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.07527 v2 pith:AIMTKV3J submitted 2024-01-15 cs.CV

classification cs.CV
keywords foundationbackbonedownstreammodelsdataremotesensingsingle
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Foundation models characterized by extensive parameters and trained on large-scale datasets have demonstrated remarkable efficacy across various downstream tasks for remote sensing data. Current remote sensing foundation models typically specialize in a single modality or a specific spatial resolution range, limiting their versatility for downstream datasets. While there have been attempts to develop multi-modal remote sensing foundation models, they typically employ separate vision encoders for each modality or spatial resolution, necessitating a switch in backbones contingent upon the input data. To address this issue, we introduce a simple yet effective method, termed OFA-Net (One-For-All Network): employing a single, shared Transformer backbone for multiple data modalities with different spatial resolutions. Using the masked image modeling mechanism, we pre-train a single Transformer backbone on a curated multi-modal dataset with this simple design. Then the backbone model can be used in different downstream tasks, thus forging a path towards a unified foundation backbone model in Earth vision. The proposed method is evaluated on 12 distinct downstream tasks and demonstrates promising performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distributed Cross-Channel Hierarchical Aggregation for Foundation Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) reduces memory and boosts throughput for multi-channel vision foundation models by spreading tokenization and channel fusion across GPUs with only a small qu...

  2. Deploying Geospatial Foundation Models in the Real World: Lessons from WorldCereal

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A structured protocol for deploying geospatial foundation models is introduced and validated in WorldCereal, where fine-tuned Presto outperforms a fully-supervised CatBoost baseline in crop mapping.

Pith tools