{"id":"c6edb245-e59d-4064-adc7-b4993bf0f9b0","arxiv_id":"2412.14123","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A single JEPA-based model with scale-adaptive encoders is pre-trained on five heterogeneous Earth observation datasets and reaches state-of-the-art results across nine downstream tasks.","lead":"AnySat is a single artificial intelligence model trained on satellite and aerial imagery from 11 different sensors with very different resolutions and scales. After self-supervised pre-training on five heterogeneous datasets, it matches or beats specialized models on nine Earth observation tasks, sometimes with just a linear probe.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The weakest point is not geographic coverage—external transfer is demonstrated—but protocol fairness: SICKLE is scored on the validation split, PASTIS-HD test imagery is in the SSL pretraining set, and competing models had modified input layers.","rationale":"The paper's central claim has two components: a method (AnySat's scale-adaptive JEPA pretraining) and an empirical claim (SOTA on nine downstream tasks). The reader's weakest assumption focuses on GeoPlex's geographic coverage; that concern is substantially answered by the paper's own external dataset results. If GeoPlex were too Europe/US-centric to transfer, we would expect the largest failures on SICKLE, BraDD, and TimeSen2Crop; instead these are among the largest gains. The more fragile premise is that the reported SOTA numbers are measured under the same protocol as the baselines. The manuscript itself flags three protocol asymmetries: SICKLE is evaluated on the validation split because the test split is unreleased; competing foundation models are adapted by modifying their input layers; and PASTIS-HD is included in the SSL pretraining set in its entirety, so the PASTIS-HD test folds were seen without labels. Any one of these could explain a reported margin, and PASTIS-HD's margins are small enough (+0.2 mIoU on segmentation) to be within run-to-run noise, as the authors themselves note for FLAIR. This does not invalidate the architectural contribution or the large-margin external gains; it does mean the abstract's blanket \"state-of-the-art on the test sets of GeoPlex\" overstates what the controlled comparisons show. A conditional verdict remains appropriate, with conditions on official-test-set confirmation and leakage-free PASTIS-HD evaluation.","tokens_in":22181,"tokens_out":13567,"duration_ms":118319,"concrete_test":"Obtain the official SICKLE test set from the dataset authors (or confirm in writing that the published UTAE/Unet3d numbers were also computed on the validation split), and evaluate the released AnySat checkpoint against the same baselines on the identical split with identical preprocessing. If AnySat's +3.6 mIoU edge over UTAE collapses or reverses on the official test set, the external SOTA claim for SICKLE is not supported. As a complementary within-paper check, rerun PASTIS-HD evaluation with GeoPlex SSL pretraining that excludes all PASTIS-HD test-fold tiles; if the reported +2.8/+0.2 gains disappear, the GeoPlex-test-set SOTA claim is leakage-inflated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that AnySat, after self-supervised pretraining on GeoPlex, reaches state-of-the-art results on the test sets of GeoPlex and six external datasets. The reader's weakest assumption—geographic representativeness of GeoPlex—is not the most fragile point: the external results on SICKLE (India), BraDD (Amazon), TimeSen2Crop (Slovenia), Sen1Floods11 (global), and So2Sat (global cities) are direct evidence that the learned representations transfer across regions. The load-bearing fragility is instead evaluation protocol. Three asymmetries weaken the SOTA comparisons: (1) Section 4.1 states that SICKLE's test set has not been released and that the authors \"use the validation set\"; baseline numbers in Table B are taken from prior papers, and if those papers reported test-set results, the +3.6 mIoU margin is not an apples-to-apples SOTA. (2) Section 4.1 reports that competing foundation models had their input layers modified to match target band counts; for models whose pretraining was tuned to their native sensors, this adaptation can systematically reduce performance relative to AnySat's flexible encoder. (3) Section 4.1 states that PASTIS-HD is included in GeoPlex \"in its entirety (without labels)\"; because PASTIS-HD is normally evaluated by 5-fold cross-validation, the test-fold images were seen during SSL pretraining. AnySat's PASTIS-HD gains (+2.8 F1 classification, +0.2 mIoU segmentation) may therefore be transductive rather than generalizable. If these protocol asymmetries explain the reported margins, the \"SOTA\" part of the central claim is overstated, even though the underlying architecture may still be a strong contribution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"AnySat proposes a single JEPA-based multimodal Earth observation model whose scale-adaptive patch encoder accepts heterogeneous resolutions, patch sizes, and modalities. The model is pretrained self-supervised on GeoPlex, a newly compiled collection of five datasets and eleven sensors, and then fine-tuned or linearly probed on nine downstream tasks, including six external datasets. The paper reports state-of-the-art or near-state-of-the-art results across these tasks and provides ablations of patch dropping, the contrastive loss, and subpatch-based segmentation.","tokens_in":22489,"tokens_out":6861,"duration_ms":61588,"significance":"If the evaluation were strictly apples-to-apples, the result would be a meaningful step: one architecture simultaneously trained on highly heterogeneous EO data and transferred to unseen sensor configurations without per-dataset retraining. The paper is unusually transparent about several limitations, including the use of the SICKLE validation set, the inclusion of all PASTIS-HD tiles in self-supervised pretraining, and the modification of competitor input layers. The code and models are released, and the ablation study provides useful evidence for the main design choices. The architecture itself is simple and GeoPlex is a valuable resource. However, the protocol asymmetries listed below currently prevent the central 'state-of-the-art' claim from being fully established; I do not consider geographic coverage of GeoPlex to be the most fragile point, since the external transfer results themselves are evidence of generalization.","major_comments":[{"comment":"The paper states that SICKLE's test set has not been released and that the authors use the validation set, while the baseline numbers in Table B are taken from earlier papers. If those baselines were reported on the official test split, the claimed +3.6 mIoU improvement is not an apples-to-apples comparison. Please report AnySat on the published SICKLE test split if it is now available, or rerun the comparison methods on the same validation split and report all numbers under an identical protocol.","section":"Section 4.1 and Appendix D (SICKLE)"},{"comment":"PASTIS-HD is included in GeoPlex 'in its entirety (without labels),' and PASTIS-HD is normally evaluated by 5-fold cross-validation. The SSL pretraining therefore sees the imagery of the test folds, and the reported +2.8 macro-F1 / +0.2 mIoU gains may be transductive rather than evidence of generalizable representation learning. The ablation section describes pretraining on folds 1 through 4 and evaluating on a held-out fold; this protocol should be used for the main PASTIS-HD results as well, and both versions should be compared explicitly.","section":"Section 4.1 and 4.2 (PASTIS-HD)"},{"comment":"For foundation models that cannot consume the target modalities, the paper modifies their input layers to match the target number of spectral bands. This is a known source of disadvantage for pretrained models, especially when the pretraining was tied to native sensors, such as CROMA with Sentinel-2 and ScaleMAE/SatMAE with RGB input. The paper should justify each modification, report the exact adaptation procedure, or rely on published baselines with their original protocols; otherwise the comparisons in Tables A and B are not controlled.","section":"Section 4.1 (Competing Methods)"},{"comment":"The reported margins include several very small differences, notably PASTIS-HD segmentation (+0.2 mIoU), FLAIR (within roughly 0.5-1.8 points of the best baseline), and So2Sat (59.1 vs. DOFA's 59.3). The paper gives no error bars, confidence intervals, or multi-seed runs, so these differences are within plausible run-to-run noise. Quantitative state-of-the-art claims need variance estimates or at least a clear statement that such small differences are not being used as evidence.","section":"Sections 4.2 and Appendix A (Tables A-B)"}],"minor_comments":[{"comment":"The text says the ablation is performed on Fold 5 of PASTIS-HD, while the Table 2 caption says Fold 1; please reconcile this inconsistency.","section":"Section 4.2 and Table 2"},{"comment":"The introduction claims that 'over 75% of the learnable parameters' are shared, while Appendix C reports 73.6% modality-agnostic parameters; the numbers should be consistent.","section":"Introduction and Appendix C"},{"comment":"The sentence 'We set he contrastive temperatureγ to 0.1 to n Eq. X' contains a typo and should read 'We set the contrastive temperature gamma to 0.1 in Eq. (2)'.","section":"Appendix C (Optimization Parameters)"},{"comment":"The column header layout in Table 1 is very hard to follow; please reformat it so that each benchmark and metric is clearly aligned with its column.","section":"Table 1"},{"comment":"The text notes that only 1.3 of the 2.3 million PLANTED images used in the original paper are publicly available; please clarify explicitly whether GeoPlex uses only the public 1.3 million images for pretraining.","section":"Section 4.1 and Table C (PLANTED)"}],"recommendation":"major_revision","confidential_remarks":"This is a solid empirical systems paper, and the authors are candid about known limitations, but the state-of-the-art claims rest on split-hygiene and baseline-adaptation issues that should be fixed before publication. A revision focusing on PASTIS-HD train-fold-only pretraining, an identical SICKLE protocol for all methods, and explicit reporting of baseline adaptation details would substantially increase confidence in the main claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth reading, but don't take the SOTA headline at face value. AnySat's real contribution is architectural: a single self-supervised model, based on JEPA with scale-adaptive patch encoders and a shared combiner, pretrained on five datasets with 11 sensors and then fine-tuned or probed on nine tasks. That is new in the EO world, and the design is sensible. JEPA's feature-space prediction avoids per-modality decoders, the sub-patch encoders handle varying resolutions and channel counts, and 73.6% of parameters are shared across modalities and resolutions. The ablations are informative, the disclosures are unusually candid, and the code and models are released.\n\nWhat is actually strong is the external transfer. BraDD-S1TS (+10.2 mIoU) and TimeSen2Crop (+11.0 OA) over prior SOTA are large margins, and neither dataset is in GeoPlex. That demonstrates the learned representations transfer across regions and sensor configurations rather than just memorizing the pretraining set. The paper also tells you its own limitations: PASTIS-HD is in GeoPlex 'in its entirety (without labels)', and SICKLE's test set is unreleased so the validation set is used. That transparency earns credit.\n\nThe problem is that those same disclosures sink the clean SOTA story. For PASTIS-HD, the test-fold images were seen during SSL pretraining, so the +2.8 F1 over OmniSat is transductive, not generalizable. For SICKLE, the +3.6 mIoU margin over Unet3d is only meaningful if the baseline numbers were also computed on the same validation split, which is doubtful since prior papers report test-set results. Add the modified input layers for competing foundation models, and the comparison is systematically tilted in AnySat's favor. The paper acknowledges this but never quantifies the tilt. There are also no error bars or multiple seeds; the 0.5 FLAIR deficit is waved off as noise.\n\nNone of this kills the architectural claim. The clean external results, especially BraDD and TimeSen2Crop, show that one flexible model can transfer across sensors and regions. GeoPlex itself is a useful artifact. But the paper overstates what it proves: 'SOTA on 9 tasks' should be softened to 'competitive on some, clearly better on others.'\n\nRecommendation: send it to peer review. The evaluation weaknesses are fixable—matched baselines, multiple seeds, and a PASTIS-HD protocol that pretrains only on folds 1–4 for the main number. The self-citation to OmniSat is legitimate; it is the direct antecedent for the combiner and contrastive loss. I would bring this to a reading group because it is a great case study in how SSL contamination and validation-set evaluation can inflate SOTA claims.","headline":"Solid architectural contribution, but the SOTA claims run ahead of the evidence; the protocol asymmetries are disclosed but they still undercut the headline numbers.","tokens_in":23115,"tokens_out":3835,"would_cite":true,"duration_ms":31873,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single self-supervised model pretrained on eleven heterogeneous sensors — 0.2 m to 250 m resolution — beats task-specific specialists on nine downstream Earth-observation benchmarks.","keywords":["Earth observation","foundation model","joint embedding predictive architecture","self-supervised learning","multimodal remote sensing","multi-resolution","scale-adaptive encoder","GeoPlex"],"falsifier":"A decisive experiment: pretrain AnySat on GeoPlex with the European and US datasets removed (or on an equal-size corpus drawn only from the tropics and southern hemisphere), then fine-tune on the same external benchmarks — SICKLE (India), BraDD-S1TS (Amazon), Sen1Floods11 (global). If the fine-tuned gains over task-specific baselines largely survive, the representations are genuinely resolution- and modality-driven; if they collapse, the reported external gains owe more to geographic overlap or dataset-specific shortcuts. A cheaper second check: measure linear-probe performance on a geographically held-out continent that shares no biome with GeoPlex.","tokens_in":21928,"feed_emoji":"🛰️","tokens_out":12017,"duration_ms":95811,"temperature":0.7,"pith_summary":"Earth observation data arrive in wildly different forms — 20 cm aerial photos, 10 m radar and optical time series, 250 m climate bands — and until now each dataset or task has needed its own model with a fixed input format. AnySat is a single architecture pretrained in a self-supervised way on all of these at once: five datasets, eleven sensors, resolutions from 0.2 to 250 m, without resizing or per-dataset retraining. After pretraining, the same model fine-tunes or linearly probes to state-of-the-art results on nine monitoring tasks, including six external datasets whose sensor combinations never appeared in pretraining. If the claim holds, one Earth-observation foundation model can replace the current zoo of per-dataset models and adapt to a new sensor with little more than a new input projector.","feed_headline":"One model masters Earth observation from 0.2 m to 250 m resolution","feed_subtitle":"Self-supervised on five datasets, one model then beats task specialists on nine benchmarks.","key_machinery":"The load-bearing mechanism is the scale-adaptive patch encoder coupled with a multimodal JEPA training objective. A tile is cut into patches of a fixed physical size $P \\times P$ meters, shared by all modalities; each patch is split into fixed-size sub-patches of $\\delta_m$ pixels, embedded by a modality-specific projector (with a lightweight temporal attention module for time series), and aggregated by a shared spatial transformer whose token count — but never its embedding size — changes with resolution and patch size. This turns 'many resolutions and modalities' into a token-count problem solvable with a single weight set. The training signal is then the semantics shared across sensors: a contrastive loss aligns same-patch embeddings from different modalities, and a JEPA prediction loss forces a masked student to reproduce the unmasked teacher's embeddings for dropped patches.","core_discovery":"AnySat's central claim is that spatial alignment alone can carry multimodal Earth-observation learning: when different sensors observe the same ground patch, their representations should agree and should be predictable from one another in feature space, with no need to reconstruct pixels. The paper realizes this with a multimodal Joint Embedding Predictive Architecture (JEPA): a student network processes heavily masked and modality-dropped patches and must reproduce the embeddings produced by an unmasked teacher whose weights are an exponential moving average of the student's, trained by an L2 prediction loss plus a contrastive loss that pulls same-patch, different-modality embeddings together. Because prediction happens in latent space, no modality-specific decoders are required, and the scale-adaptive patch encoder means one shared weight set handles resolutions from 0.2 m to 250 m. Pretrained self-supervised on GeoPlex (five datasets, eleven sensors, 249K km²), AnySat is then fine-tuned or linearly probed to state-of-the-art results on nine downstream tasks — land cover, crop type, tree species, flood, burn scar, deforestation, and climate zone mapping — including six external datasets whose sensor configurations never appeared during pretraining.","pith_inferences":["The paper leaves its own geographic stress test implicit: GeoPlex skews heavily toward Europe and the United States, so the decisive experiment is whether the external-benchmark wins in India, the Amazon, and global floods survive when pretraining is rebuilt from tropical and southern-hemisphere data alone; that test would separate genuine resolution/modality generalization from geography-driven s","The recipe — align modalities through shared ground location, predict in embedding space rather than pixel space — is not obviously satellite-specific; applying the identical architecture to multimodal medical imaging, where CT, MRI, and ultrasound view the same anatomy, would be a direct transfer test the authors do not run.","The ablation in which removing the contrastive loss costs 4.3 F1 points on classification while removing masking costs little suggests a leaner recipe — strong cross-modal contrast with minimal masking — might reach similar representations at lower pretraining cost; this follows from the paper's own numbers but is not an experiment it performs."],"forward_implications":["Adaptation cost collapses: a linear probe with about 6,000 trainable parameters beats or matches specialized models and a 630M-parameter foundation model on flood segmentation and Amazon deforestation detection, so new monitoring tasks become cheap to deploy.","Unseen sensors and configurations work after minimal surgery: single-date Sentinel pairs, single-pixel time series, and a wholly unseen sensor (HLS) all reach state-of-the-art results, the last via a newly initialized projector.","Pretraining pays off most where labels are scarce — the largest gains are on the smallest labeled datasets (TreeSatAI-TS, PASTIS classification) — so label-poor environmental monitoring stands to benefit most.","GeoPlex is extensible: because each training step samples a random dataset and patch size, new datasets can be folded into the same pretraining run, letting every added region or sensor sharpen the shared backbone."],"supporting_citations":[{"why":"Supplies the joint-embedding predictive architecture and the student/EMA-teacher training paradigm that AnySat adapts to multimodal Earth observation.","marker":"[6]"},{"why":"Supplies the cross-attention modality combiner and the same-patch cross-modal contrastive loss, and contributes two GeoPlex datasets (TreeSatAI-TS, PASTIS-HD).","marker":"[7]"},{"why":"Supplies the ground-sampling-distance positional encoding and the scale-aware training idea that the scale-adaptive patch encoder builds on.","marker":"[56]"},{"why":"Supplies the lightweight temporal attention encoder used to collapse time-series channels in the patch encoder.","marker":"[26]"},{"why":"PLANTED is one of the five GeoPlex pretraining datasets; its five-sensor time series, including ALOS-2 and MODIS, supply the coarsest resolutions in the mix.","marker":"[54]"},{"why":"FLAIR is one of the five GeoPlex datasets; it supplies 0.2 m aerial imagery with elevation and dense land-cover labels at the 102 m tile scale.","marker":"[25]"},{"why":"SatlasPretrain is the source of S2NAIP-URBAN, the largest GeoPlex component, adding NAIP very-high-resolution imagery and Landsat 8/9 to the pretraining mix.","marker":"[11]"},{"why":"Provides the momentum/EMA teacher update mechanism that stabilizes the JEPA teacher network during self-supervised training.","marker":"[33]"},{"why":"Provides the InfoNCE objective that AnySat modifies into the cross-modal contrastive loss enforcing spatial consistency across sensors.","marker":"[52]"},{"why":"BraDD-S1TS is the external Amazon deforestation-change benchmark where AnySat reports its largest gain (+10.2 mIoU), demonstrating adaptation to a sensor configuration absent from GeoPlex.","marker":"[39]"}],"fun_headline_variants":["AnySat: One model, 11 sensors, beats specialists on 9 tasks","AnySat: One model learns from 11 sensors, any resolution","AnySat: One model for every Earth sensor resolution and task","Unified Earth observation: AnySat learns across resolutions and modalities"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that GeoPlex — 249K km² concentrated in Europe and the United States — is representative enough of global land cover and imaging conditions that representations trained on it transfer to India, the Amazon, and worldwide flood scenes, rather than the external gains being geographic or dataset-specific shortcuts.","fun_headline_variants_meta":{"raw":{"variants":["AnySat: One model, 11 sensors, beats specialists on 9 tasks","AnySat: One model learns from 11 sensors, any resolution","AnySat: One model for every Earth sensor resolution and task","Unified Earth observation: AnySat learns across resolutions and modalities"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000992,"raw_usage":{"total_tokens":4215,"prompt_tokens":966,"completion_tokens":3249,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":3173}},"tokens_in":582,"tokens_out":3249,"duration_ms":19939,"temperature":1.0,"reasoning_tokens":3173,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:27:54.122356+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive experiment: pretrain AnySat on GeoPlex with the European and US datasets removed (or on an equal-size corpus drawn only from the tropics and southern hemisphere), then fine-tune on the same external benchmarks — SICKLE (India), BraDD-S1TS (Amazon), Sen1Floods11 (global). If the fine-tuned gains over task-specific baselines largely survive, the representations are genuinely resolution- and modality-driven; if they collapse, the reported external gains owe more to geographic overlap or dataset-specific shortcuts. A cheaper second check: measure linear-probe performance on a geographically held-out continent that shares no biome with GeoPlex.","supporting_citations":[{"cited_title":"Self-supervised learning from images with a joint-embedding predictive architecture","cited_arxiv_id":null,"evidence_quote":"Supplies the joint-embedding predictive architecture and the student/EMA-teacher training paradigm that AnySat adapts to multimodal Earth observation."},{"cited_title":"Omnisat: Self-supervised modality fusion for earth observation","cited_arxiv_id":null,"evidence_quote":"Supplies the cross-attention modality combiner and the same-patch cross-modal contrastive loss, and contributes two GeoPlex datasets (TreeSatAI-TS, PASTIS-HD)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ground-sampling-distance positional encoding and the scale-aware training idea that the scale-adaptive patch encoder builds on."},{"cited_title":"Lightweight temporal self-attention for classifying satellite images time series","cited_arxiv_id":null,"evidence_quote":"Supplies the lightweight temporal attention encoder used to collapse time-series channels in the patch encoder."},{"cited_title":"Planted: A dataset for planted forest identification from multi-satellite time series","cited_arxiv_id":null,"evidence_quote":"PLANTED is one of the five GeoPlex pretraining datasets; its five-sensor time series, including ALOS-2 and MODIS, supply the coarsest resolutions in the mix."},{"cited_title":"FLAIR: A country-scale land cover semantic segmentation dataset from multi-source optical imagery","cited_arxiv_id":null,"evidence_quote":"FLAIR is one of the five GeoPlex datasets; it supplies 0.2 m aerial imagery with elevation and dense land-cover labels at the 102 m tile scale."},{"cited_title":"SatlasPretrain: A large-scale dataset for remote sensing image under- standing","cited_arxiv_id":null,"evidence_quote":"SatlasPretrain is the source of S2NAIP-URBAN, the largest GeoPlex component, adding NAIP very-high-resolution imagery and Landsat 8/9 to the pretraining mix."},{"cited_title":"Momentum contrast for unsupervised visual representation learning","cited_arxiv_id":null,"evidence_quote":"Provides the momentum/EMA teacher update mechanism that stabilizes the JEPA teacher network during self-supervised training."},{"cited_title":"Deforestation detection in the Amazon with Sentinel-1 SAR image time series","cited_arxiv_id":null,"evidence_quote":"BraDD-S1TS is the external Amazon deforestation-change benchmark where AnySat reports its largest gain (+10.2 mIoU), demonstrating adaptation to a sensor configuration absent from GeoPlex."}],"review_version":1}