{"id":"297ddd38-244c-42a9-91f4-b9565e9edc46","arxiv_id":"2607.19766","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"An independent multi-grid benchmark finds GBA and TEMPO most accurate global building-area products overall, with accuracy dropping sharply in Africa, Asia, and high-density urban areas.","lead":"This study measures how accurate four global maps of building footprint area are, using a large manually labeled reference dataset. It finds that no single product wins on every metric—two of them, GBA and TEMPO are generally best—and that all products are weaker in Africa, Asia, and dense urban areas.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Temporal mismatch between ORBITaL-Net (2017) and TEMPO (2023) may drive the GBA-over-TEMPO MAE/WMAPE ranking; no temporally matched check is provided.","rationale":"The reader's weakest assumption—temporal mismatch—is indeed the most load-bearing concern. The paper explicitly acknowledges this bias, and the reported GBA-over-TEMPO margins on the TEMPO grid are so small that even a modest temporal trend could flip them. Since TEMPO provides a time series, a direct falsification test is feasible. The reader's conditional verdict is appropriate: the benchmark is useful but the headline ranking should not be treated as settled until this confound is addressed. My analysis does not change the verdict, so UNCHANGED.","tokens_in":14937,"tokens_out":5102,"duration_ms":48739,"concrete_test":"Re-run the comparison on the TEMPO reference grid using only ORBITaL-Net chips with acquisition dates in 2018–2020, and compare GBA (2019) against TEMPO 2019Q4 (or 2020Q4) rather than 2023Q4. If the MAE/WMAPE ordering flips or the margin becomes indistinguishable (e.g., |ΔMAE| < 1 m²), the original ranking is a temporal artifact. Alternatively, compute TEMPO's own predicted area on the same cells for 2018Q1 and 2023Q4; if the mean absolute change exceeds the GBA–TEMPO MAE gap, temporal mismatch alone can account for the reported advantage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that GBA outperforms TEMPO on MAE/WMAPE rests on a comparison against ORBITaL-Net ground truth whose median image date is 2017, while GBA is 2019 and TEMPO is 2023Q4. The authors concede in Sec. 5.4 that this creates a systematic bias in favor of GBA. The magnitude of the suspected bias is large relative to the reported differences: on the TEMPO reference grid, GBA's MAE advantage over TEMPO is only 0.55 m² (50.89 vs 51.44) and its WMAPE advantage is 0.25 percentage points (34.09 vs 34.34). Over a 6-year period, building area can change by more than this in active urban areas. Because TEMPO is the only product with a temporal series, it is possible to quantify the temporal effect directly, but the paper does not do so. Without a temporally matched subset or an explicit sensitivity analysis, the 'GBA always best' ranking is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates four global building-area products (GHSL, TEMPO, GBA, Overture) against the ORBITaL-Net manual building-mask dataset. The authors compare products on three reference grids (two GHSL grids and the TEMPO grid), using five metrics: total area, aggregate difference Δ(m²), Δ(%), MAE, WMAPE, and R². They report that GBA has the lowest MAE and WMAPE on all three grids, while TEMPO has the lowest aggregate bias and highest R²; GHSL generally performs worst and is prone to overestimation. They further stratify results by continent, population density, and income group, finding that all products are less accurate in Africa and Asia and that most degrade in high-density urban areas. The central conclusion is that no single product is uniformly best and that metric choice and geographic context matter. The authors acknowledge a temporal mismatch between ORBITaL-Net (median 2017), GBA (2019), and TEMPO (2023Q4), and note this creates a systematic bias favoring GBA.","tokens_in":15187,"tokens_out":2747,"duration_ms":31829,"significance":"If the conclusions are valid, this would be a valuable independent benchmark for the growing field of global building-area products. The study has notable strengths: it uses a large, globally distributed, manually labeled ground-truth dataset; it makes its comparison protocol explicit through Algorithms 1 and 2; it evaluates on multiple reference grids to mitigate grid-choice bias; it includes a baseline estimator; and it reports stratified results by region, density, and income. These features are genuinely useful and go beyond most prior assessments. The claim that no single product is uniformly best is plausible and practically important for downstream users. However, the headline GBA-versus-TEMPO ranking is not yet established because of the temporal confounding acknowledged in Sec. 5.4, and the reported differences are small relative to the suspected bias. The paper therefore needs a targeted sensitivity analysis or matched-subset check before its central ranking claim can be accepted.","major_comments":[{"comment":"The central claim that GBA always performs best by MAE and WMAPE is confounded by the temporal mismatch that the authors themselves acknowledge. ORBITaL-Net imagery has median year 2017, GBA is 2019, and TEMPO is 2023Q4. On the TEMPO reference grid, GBA's advantage over TEMPO is only 0.55 m² in MAE (50.89 vs 51.44) and 0.25 percentage points in WMAPE (34.09 vs 34.34). Over a five-to-six-year window, differential building growth alone could plausibly account for differences of this size, especially since TEMPO is the only product with a temporal series. The paper does not provide a temporally matched comparison (e.g., using TEMPO's ~2019 epoch or restricting the evaluation to cells with no detectable change over time), nor a sensitivity analysis bounding the effect of the mismatch. Without such an analysis, the 'GBA always best on MAE/WMAPE' headline is not established.","section":"§5.4 and Table 4"},{"comment":"MAE, WMAPE, and Δ are reported as point estimates with no uncertainty intervals. The only metric with reported uncertainty is R² (bootstrap). Given that the key GBA-versus-TEMPO differences are small (sub-meter MAE differences and sub-percentage-point WMAPE differences), confidence intervals, paired tests, or bootstrap resampling over cells are needed to determine whether these differences are statistically distinguishable from noise. This is especially important because the valid cell sets differ across the three reference grids, so cross-grid comparisons are not repeated measurements on the same sample.","section":"§4 and Table 3"},{"comment":"The three reference-grid experiments use different sets of valid cells, because the 99% containment filter depends on the reference grid. The paper notes that the ground-truth total area changes from ~67 million m² on the GHSL 100 m grid to ~142 million m² on the TEMPO 77 m grid. This means that the metrics on different grids are computed over different populations of cells. The claim that 'the relative behavior of all products remains stable' is not supported by any statistical test or by an explicit comparison of the overlapping cell sets. If the goal is to show that product rankings are robust to the reference grid, the authors should either restrict to cells valid under all three grids or report how much of the observed metric variation is due to sample composition.","section":"§5.1 and Algorithm 1"},{"comment":"The ORBITaL-Net labeling protocol has well-documented limitations that are acknowledged but not quantified: off-nadir imagery includes facades rather than rooftops, and cloud-occluded buildings cannot be labeled, causing potential underestimation of ground-truth building area. Since the products under evaluation have different definitions of 'building area' (e.g., footprint polygons versus density rasters), these label artifacts could affect both absolute accuracy and relative ranking. The authors should provide a sensitivity analysis or at least a quantitative bound on how large these label effects would need to be to change the GBA-versus-TEMPO ranking. As it stands, the reported accuracy numbers inherit an unknown label bias.","section":"§5.4 and Fig. 4"}],"minor_comments":[{"comment":"The abstract says ORBITaL-Net imagery has approximately 0.47 m resolution, while §2 says 0.45 m. Please reconcile these numbers.","section":"Abstract / §2"},{"comment":"Please report the number of valid cells n used for each reference grid, together with the sample statistics. This would help the reader understand why total ground-truth area varies so much across grids and would make the cross-grid comparison more interpretable.","section":"Table 4"},{"comment":"There is a typo: 'facaces' should be 'facades'.","section":"§5.4"},{"comment":"The finding that 'GHSL always achieves the highest error rates' is stated before the evidence in Table 4; since Table 4 is the primary evidence, it would be clearer to cite it immediately after the claim and to specify which cells are included in each comparison.","section":"§5.1"},{"comment":"TEMPO's temporal coverage is listed as 2018–2025 in Table 1, but the text in §3.3 says Q1 2018 through Q2 2025. Please ensure consistency and clarify why the 2023Q4 epoch was chosen rather than a more recent one.","section":"§3.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful contribution and the protocol is carefully designed, but the central GBA-versus-TEMPO ranking is not yet supported by the evidence because of the acknowledged temporal mismatch. The differences in MAE/WMAPE are small enough that a factor the authors themselves identify as biasing in favor of GBA could be decisive. I would be willing to see a revision that includes a temporally matched analysis (e.g., using TEMPO's 2019 epoch) or an explicit sensitivity analysis, along with uncertainty intervals for the primary metrics. The label-bias issues (facades, cloud occlusion) should also be quantified or bounded. I would not reject the paper on these grounds because they are addressable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a useful benchmark paper, but the central GBA-over-TEMPO ranking on MAE/WMAPE is not as solid as the abstract implies. The authors themselves concede in Sec. 5.4 that the temporal mismatch between ORBITaL-Net (median 2017), GBA (2019), and TEMPO (2023Q4) creates a systematic bias in favor of GBA. The margins are tiny — on the TEMPO grid, GBA's MAE advantage is 0.55 m² (50.89 vs 51.44) and its WMAPE advantage is 0.25 percentage points (34.09 vs 34.34). Six years of urban change can easily be larger than that. Since TEMPO has a quarterly time series, they could have checked how much a temporally matched epoch changes the numbers, but they didn't. So the \"GBA always best\" claim should be treated as a hypothesis, not a settled result.\n\nWhat's new and good: it's the first independent joint comparison of GHSL, TEMPO, GBA, and Overture against ORBITaL-Net, on three reference grids, with explicit reprojection algorithms and a sensible multi-metric setup. The stratification by continent, population density, and income is valuable and mostly more robust than the headline ranking. The finding that all products degrade in Africa and Asia, and that most degrade in high-density urban areas, is worth taking seriously. The paper is also honest about its limitations — the temporal bias, the facade/occlusion labeling issues, and the absence of European samples are all acknowledged, not hidden.\n\nSoft spots beyond the temporal issue: no uncertainty intervals on MAE and WMAPE (only R² gets bootstrap); the 99% cell-coverage threshold is a free parameter but not obviously consequential; and the ground truth itself has known label biases that could interact with product differences. The Europe gap is a coverage hole, though the authors' reasoning that North America and high-income results are similar is plausible.\n\nWho this is for: anyone choosing a global building-area product for urban, energy, or emissions work, and anyone doing accuracy assessment of geospatial products. It deserves a serious referee — the protocol is a good template and the stratified results are informative. But the paper needs revision before publication: re-run the TEMPO comparison on a temporally matched epoch (or at least do a sensitivity analysis) and report uncertainty on the primary metrics. If that doesn't change the ranking, fine; if it does, the conclusions need redrawing.","headline":"A genuinely useful first benchmark, but the GBA-over-TEMPO ranking on MAE/WMAPE is thinner than it looks because of an acknowledged temporal bias favoring GBA, and the paper doesn't quantify it.","tokens_in":15649,"tokens_out":2526,"would_cite":true,"duration_ms":23066,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"No single global building-area product is uniformly most accurate: GBA wins on per-cell error, TEMPO on bias and R².","keywords":["building area estimation","global building footprint products","accuracy assessment","remote sensing benchmark","satellite imagery","reference grid","semantic segmentation","ORBITaL-Net"],"falsifier":"Re-run the evaluation using only ORBITaL-Net chips whose imagery acquisition date matches each product's epoch (2019 for GBA, 2023Q4 for TEMPO), then check whether GBA still beats TEMPO on MAE and WMAPE; if TEMPO wins the temporally matched comparison, the headline ranking is an artifact of the reference data. A second check: recompute ORBITaL-Net areas from rooftop-only labels (excluding facades) and see whether GBA's per-cell error advantage persists.","tokens_in":14811,"feed_emoji":"🗺️","tokens_out":5811,"duration_ms":52271,"temperature":0.7,"pith_summary":"This paper tries to establish, for the first time, an independent and fair comparison of four global building-area products—GHSL, TEMPO, the Global Building Atlas (GBA), and Overture—using a globally distributed set of manually labeled building footprints as ground truth. Its central finding is that no product is uniformly the most accurate: GBA achieves the lowest mean absolute error and weighted absolute percentage error on all three reference grids, while TEMPO achieves the smallest aggregate bias and highest R² on those same grids. GHSL consistently performs worst and systematically overestimates building area. The paper also claims that all four products are substantially less accurate in Africa and Asia, and that most of them lose accuracy in high-density urban areas. A sympathetic reader would take away that product choice should be driven by the error metric of interest and by geographic context, rather than a single global leaderboard.","feed_headline":"No global building map wins every accuracy metric","feed_subtitle":"Four products tested: GBA leads on per-cell error, TEMPO on bias; all lag in Africa and Asia.","key_machinery":"The load-bearing mechanism is the multi-reference-grid accuracy protocol: ground-truth building area from ORBITaL-Net is rasterized onto each of three reference grids (GHSL WGS84 3 arcsec, GHSL Mollweide 100 m, TEMPO 77 m), keeping only cells at least 99% covered by an annotated image chip; vector products (GBA, Overture) are reprojected by polygon intersection, raster products (GHSL, TEMPO) by area-weighted density aggregation; and every product is scored on the same matched cells with MAE, WMAPE, Δ(m²), Δ(%), and R². Repeating on three grids controls for the systematic advantage a product gains when compared on its own native grid, and the 99% coverage filter removes partial-boundary cells","core_discovery":"The paper claims that, when four major global building-area products are evaluated against ORBITaL-Net manually labeled building footprints on multiple reference grids and five metrics, the accuracy ranking is metric-dependent. GBA always has the lowest MAE and WMAPE; TEMPO always has the smallest total-area bias (Δ m², Δ %) and highest R²; GHSL has the largest errors and overestimates; Overture sits in between. The paper further claims that accuracy degrades markedly in Africa and Asia for all products, and in high-density urban areas for most products, and that GHSL's severe overestimation is concentrated in Africa and lower-middle-income regions. These results are presented as the first g","pith_inferences":["Because the paper concedes the reference imagery is ~2 years older than GBA's and ~6 years older than TEMPO's, the GBA-over-TEMPO margin on MAE/WMAPE is itself a testable artifact: re-running the benchmark on temporally matched ORBITaL-Net chips (2019-only chips for GBA, 2023-only for TEMPO) could flip or shrink that margin.","The multi-grid protocol implies that raster products are being compared under an extra handicap when evaluated off their native grid; a fairer 'best achievable accuracy' comparison would report each product only on its own grid, which favors TEMPO more than the headline numbers show.","The stratifications suggest an ensemble or region-weighted blend (e.g., TEMPO in urban areas, GBA in rural areas) could beat any single product; the paper's tables provide the calibration data to test that.","Since ORBITaL-Net has no European samples, the headline ranking is really a ranking of Africa, Asia, and the Americas; adding European labels—where building geometry differs—could change the aggregate ordering, despite the authors' conjecture that it would not."],"forward_implications":["Applications that need accurate per-cell building area (e.g., local energy demand, urban planning) should prefer GBA over TEMPO by these results, since GBA wins MAE and WMAPE on every grid.","Applications whose downstream error tolerances depend on total bias or on capturing spatial variation (e.g., emissions inventories, exposure proxies) should prefer TEMPO, which wins Δ and R² on every grid.","Any use of these products in Africa or Asia should carry substantially inflated uncertainty; the paper shows all four products lose accuracy there, with GHSL's overestimation in Africa exceeding +130%.","All products except GHSL tend to underestimate building area, and most underestimate more severely in high-density urban areas, so corrections for systematic undercount should be applied in dense cities.","GHSL's consistent overestimation, especially in lower-middle-income regions, makes it a risky choice for building-area-based population or emissions studies in those regions."],"fun_headline_variants":["Accuracy of global building maps depends on metric and location","No global building dataset tops every accuracy test","Building-area maps: best varies by metric, worst in Africa and Asia","GBA leads on error, TEMPO on bias, but all miss Africa and Asia","Global building products: accuracy gap widens in Africa and Asia"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire ranking rests on treating ORBITaL-Net's 2017-era manual labels—which include facades in off-nadir imagery and cannot mark buildings hidden by clouds—as the true building area for products made from 2019 and 2023 imagery; the authors concede this creates a systematic bias favoring GBA.","fun_headline_variants_meta":{"raw":{"variants":["Accuracy of global building maps depends on metric and location","No global building dataset tops every accuracy test","Building-area maps: best varies by metric, worst in Africa and Asia","GBA leads on error, TEMPO on bias, but all miss Africa and Asia","Global building products: accuracy gap widens in Africa and Asia"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000549,"raw_usage":{"total_tokens":2451,"prompt_tokens":731,"completion_tokens":1720,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":475,"completion_tokens_details":{"reasoning_tokens":1633}},"tokens_in":475,"tokens_out":1720,"duration_ms":10883,"temperature":1.0,"reasoning_tokens":1633,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T11:45:01.709152+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the evaluation using only ORBITaL-Net chips whose imagery acquisition date matches each product's epoch (2019 for GBA, 2023Q4 for TEMPO), then check whether GBA still beats TEMPO on MAE and WMAPE; if TEMPO wins the temporally matched comparison, the headline ranking is an artifact of the reference data. A second check: recompute ORBITaL-Net areas from rooftop-only labels (excluding facades) and see whether GBA's per-cell error advantage persists.","supporting_citations":[],"review_version":1}