Pith. sign in

REVIEW 8 cited by

reBEN: Refined BigEarthNet Dataset for Remote Sensing Image Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.03653 v5 pith:SU25YHR5 submitted 2024-07-04 cs.CV eess.IV

classification cs.CVeess.IV
keywords bigearthnetrebendatasetimagepatchesremotesensingsentinel-2
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents refined BigEarthNet (reBEN) that is a large-scale, multi-modal remote sensing dataset constructed to support deep learning (DL) studies for remote sensing image analysis. The reBEN dataset consists of 549,488 pairs of Sentinel-1 and Sentinel-2 image patches. To construct reBEN, we initially consider the Sentinel-1 and Sentinel-2 tiles used to construct the BigEarthNet dataset and then divide them into patches of size 1200 m x 1200 m. We apply atmospheric correction to the Sentinel-2 patches using the latest version of the sen2cor tool, resulting in higher-quality patches compared to those present in BigEarthNet. Each patch is then associated with a pixel-level reference map and scene-level multi-labels. This makes reBEN suitable for pixel- and scene-based learning tasks. The labels are derived from the most recent CORINE Land Cover (CLC) map of 2018 by utilizing the 19-class nomenclature as in BigEarthNet. The use of the most recent CLC map results in overcoming the label noise present in BigEarthNet. Furthermore, we introduce a new geographical-based split assignment algorithm that significantly reduces the spatial correlation among the train, validation, and test sets with respect to those present in BigEarthNet. This increases the reliability of the evaluation of DL models. To minimize the DL model training time, we introduce software tools that convert the reBEN dataset into a DL-optimized data format. In our experiments, we show the potential of reBEN for multi-modal multi-label image classification problems by considering several state-of-the-art DL models. The pre-trained model weights, associated code, and complete dataset are available at https://bigearth.net.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EarthShift: a benchmark for measuring robustness to real-world distribution shifts in Earth observation

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    EarthShift is a new benchmark using paired datasets to measure robustness of geospatial foundation models to realistic distribution shifts, finding consistent 15-20% performance drops out-of-distribution across 8 mode...

  2. OSMGraphCLIP: Learning Global Location Representations from OpenStreetMap Graphs

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    OSMGraphCLIP learns global location embeddings from OSM graphs via multi-scale graph encoding and contrastive alignment that match or exceed satellite baselines on many socioeconomic, health, and environmental tasks.

  3. Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Sentinel2Cap provides human-annotated captions for multimodal Sentinel satellite images, with zero-shot tests showing RGB outperforming SAR and prompts helping performance.

  4. Adaptive Gradient Calibration for Single-Positive Multi-Label Learning in Remote Sensing Image Scene Classification

    cs.CV 2025-10 unverdicted novelty 6.0 of 10

    AdaGC adaptively applies gradient calibration with dual EMA in SPML for RS imagery to recover full labels from single-positive annotations and reports SOTA results on two benchmarks.

  5. Using Multiple Input Modalities Can Improve Data-Efficiency and O.O.D. Generalization for ML with Satellite Imagery

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Adding auxiliary geographic data layers to satellite imagery improves label efficiency and out-of-sample generalization across four SatML tasks, with frozen or hand-coded fusion beating fine-tuned variants.

  6. DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    DiffuSAM fuses diffusion-based localization cues with SAM models to deliver over 14% higher Acc@0.5 in zero-shot object grounding for remote sensing imagery compared to prior methods.

  7. Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    Introduces SGR and CGR refinement pipelines plus majority-voting ensemble to improve visual grounding accuracy in remote sensing by combining RemoteSAM and SAM3.

  8. MoSAiC: Multi-Modal Multi-Label Supervision-Aware Contrastive Learning for Remote Sensing

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Jointly training intra-modal SimCLR, inter-modal Sentinel-1/2 contrast, and MulSupCon with BCE yields better low-label multi-label land-cover classification than the tested baselines.

Pith tools