Pith. sign in

REVIEW 4 major objections 5 minor 33 references

Global Building Area Estimation Products: How Accurate Are They?

T0 review · 4 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read No single global building-area product is uniformly most accurate: GBA wins on per-cell error, TEMPO on bias and R².

desk verdict A genuinely useful first benchmark, but the GBA-over-TEMPO ranking on MAE/WMAPE is thinner than it looks because of an acknowledged temporal bias favoring GBA, and the paper doesn't quantify it. read the letter →

arxiv 2607.19766 v1 pith:HEUKHNXP submitted 2026-07-22 cs.CV

classification cs.CV
keywords buildingareaestimationglobalfootprintproductsaccuracyassessmentremotesensingbenchmarksatelliteimageryreferencegridsemanticsegmentationORBITaL-Net
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish, for the first time, an independent and fair comparison of four global building-area products—GHSL, TEMPO, the Global Building Atlas (GBA), and Overture—using a globally distributed set of manually labeled building footprints as ground truth. Its central finding is that no product is uniformly the most accurate: GBA achieves the lowest mean absolute error and weighted absolute percentage error on all three reference grids, while TEMPO achieves the smallest aggregate bias and highest R² on those same grids. GHSL consistently performs worst and systematically overestimates building area. The paper also claims that all four products are substantially less accurate in Africa and Asia, and that most of them lose accuracy in high-density urban areas. A sympathetic reader would take away that product choice should be driven by the error metric of interest and by geographic context, rather than a single global leaderboard.

What carries the argument

The load-bearing mechanism is the multi-reference-grid accuracy protocol: ground-truth building area from ORBITaL-Net is rasterized onto each of three reference grids (GHSL WGS84 3 arcsec, GHSL Mollweide 100 m, TEMPO 77 m), keeping only cells at least 99% covered by an annotated image chip; vector products (GBA, Overture) are reprojected by polygon intersection, raster products (GHSL, TEMPO) by area-weighted density aggregation; and every product is scored on the same matched cells with MAE, WMAPE, Δ(m²), Δ(%), and R². Repeating on three grids controls for the systematic advantage a product gains when compared on its own native grid, and the 99% coverage filter removes partial-boundary cells

What would settle it

Re-run the evaluation using only ORBITaL-Net chips whose imagery acquisition date matches each product's epoch (2019 for GBA, 2023Q4 for TEMPO), then check whether GBA still beats TEMPO on MAE and WMAPE; if TEMPO wins the temporally matched comparison, the headline ranking is an artifact of the reference data. A second check: recompute ORBITaL-Net areas from rooftop-only labels (excluding facades) and see whether GBA's per-cell error advantage persists.

Watch

Extended reading notes

Core claim

The paper claims that, when four major global building-area products are evaluated against ORBITaL-Net manually labeled building footprints on multiple reference grids and five metrics, the accuracy ranking is metric-dependent. GBA always has the lowest MAE and WMAPE; TEMPO always has the smallest total-area bias (Δ m², Δ %) and highest R²; GHSL has the largest errors and overestimates; Overture sits in between. The paper further claims that accuracy degrades markedly in Africa and Asia for all products, and in high-density urban areas for most products, and that GHSL's severe overestimation is concentrated in Africa and lower-middle-income regions. These results are presented as the first g

Load-bearing premise

The entire ranking rests on treating ORBITaL-Net's 2017-era manual labels—which include facades in off-nadir imagery and cannot mark buildings hidden by clouds—as the true building area for products made from 2019 and 2023 imagery; the authors concede this creates a systematic bias favoring GBA.

Editorial extensions

If this is right

  • Applications that need accurate per-cell building area (e.g., local energy demand, urban planning) should prefer GBA over TEMPO by these results, since GBA wins MAE and WMAPE on every grid.
  • Applications whose downstream error tolerances depend on total bias or on capturing spatial variation (e.g., emissions inventories, exposure proxies) should prefer TEMPO, which wins Δ and R² on every grid.
  • Any use of these products in Africa or Asia should carry substantially inflated uncertainty; the paper shows all four products lose accuracy there, with GHSL's overestimation in Africa exceeding +130%.
  • All products except GHSL tend to underestimate building area, and most underestimate more severely in high-density urban areas, so corrections for systematic undercount should be applied in dense cities.
  • GHSL's consistent overestimation, especially in lower-middle-income regions, makes it a risky choice for building-area-based population or emissions studies in those regions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the paper concedes the reference imagery is ~2 years older than GBA's and ~6 years older than TEMPO's, the GBA-over-TEMPO margin on MAE/WMAPE is itself a testable artifact: re-running the benchmark on temporally matched ORBITaL-Net chips (2019-only chips for GBA, 2023-only for TEMPO) could flip or shrink that margin.
  • The multi-grid protocol implies that raster products are being compared under an extra handicap when evaluated off their native grid; a fairer 'best achievable accuracy' comparison would report each product only on its own grid, which favors TEMPO more than the headline numbers show.
  • The stratifications suggest an ensemble or region-weighted blend (e.g., TEMPO in urban areas, GBA in rural areas) could beat any single product; the paper's tables provide the calibration data to test that.
  • Since ORBITaL-Net has no European samples, the headline ranking is really a ranking of Africa, Asia, and the Americas; adding European labels—where building geometry differs—could change the aggregate ordering, despite the authors' conjecture that it would not.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper evaluates four global building-area products (GHSL, TEMPO, GBA, Overture) against the ORBITaL-Net manual building-mask dataset. The authors compare products on three reference grids (two GHSL grids and the TEMPO grid), using five metrics: total area, aggregate difference Δ(m²), Δ(%), MAE, WMAPE, and R². They report that GBA has the lowest MAE and WMAPE on all three grids, while TEMPO has the lowest aggregate bias and highest R²; GHSL generally performs worst and is prone to overestimation. They further stratify results by continent, population density, and income group, finding that all products are less accurate in Africa and Asia and that most degrade in high-density urban areas. The central conclusion is that no single product is uniformly best and that metric choice and geographic context matter. The authors acknowledge a temporal mismatch between ORBITaL-Net (median 2017), GBA (2019), and TEMPO (2023Q4), and note this creates a systematic bias favoring GBA.

Significance. If the conclusions are valid, this would be a valuable independent benchmark for the growing field of global building-area products. The study has notable strengths: it uses a large, globally distributed, manually labeled ground-truth dataset; it makes its comparison protocol explicit through Algorithms 1 and 2; it evaluates on multiple reference grids to mitigate grid-choice bias; it includes a baseline estimator; and it reports stratified results by region, density, and income. These features are genuinely useful and go beyond most prior assessments. The claim that no single product is uniformly best is plausible and practically important for downstream users. However, the headline GBA-versus-TEMPO ranking is not yet established because of the temporal confounding acknowledged in Sec. 5.4, and the reported differences are small relative to the suspected bias. The paper therefore needs a targeted sensitivity analysis or matched-subset check before its central ranking claim can be accepted.

major comments (4)
  1. [§5.4 and Table 4] The central claim that GBA always performs best by MAE and WMAPE is confounded by the temporal mismatch that the authors themselves acknowledge. ORBITaL-Net imagery has median year 2017, GBA is 2019, and TEMPO is 2023Q4. On the TEMPO reference grid, GBA's advantage over TEMPO is only 0.55 m² in MAE (50.89 vs 51.44) and 0.25 percentage points in WMAPE (34.09 vs 34.34). Over a five-to-six-year window, differential building growth alone could plausibly account for differences of this size, especially since TEMPO is the only product with a temporal series. The paper does not provide a temporally matched comparison (e.g., using TEMPO's ~2019 epoch or restricting the evaluation to cells with no detectable change over time), nor a sensitivity analysis bounding the effect of the mismatch. Without such an analysis, the 'GBA always best on MAE/WMAPE' headline is not established.
  2. [§4 and Table 3] MAE, WMAPE, and Δ are reported as point estimates with no uncertainty intervals. The only metric with reported uncertainty is R² (bootstrap). Given that the key GBA-versus-TEMPO differences are small (sub-meter MAE differences and sub-percentage-point WMAPE differences), confidence intervals, paired tests, or bootstrap resampling over cells are needed to determine whether these differences are statistically distinguishable from noise. This is especially important because the valid cell sets differ across the three reference grids, so cross-grid comparisons are not repeated measurements on the same sample.
  3. [§5.1 and Algorithm 1] The three reference-grid experiments use different sets of valid cells, because the 99% containment filter depends on the reference grid. The paper notes that the ground-truth total area changes from ~67 million m² on the GHSL 100 m grid to ~142 million m² on the TEMPO 77 m grid. This means that the metrics on different grids are computed over different populations of cells. The claim that 'the relative behavior of all products remains stable' is not supported by any statistical test or by an explicit comparison of the overlapping cell sets. If the goal is to show that product rankings are robust to the reference grid, the authors should either restrict to cells valid under all three grids or report how much of the observed metric variation is due to sample composition.
  4. [§5.4 and Fig. 4] The ORBITaL-Net labeling protocol has well-documented limitations that are acknowledged but not quantified: off-nadir imagery includes facades rather than rooftops, and cloud-occluded buildings cannot be labeled, causing potential underestimation of ground-truth building area. Since the products under evaluation have different definitions of 'building area' (e.g., footprint polygons versus density rasters), these label artifacts could affect both absolute accuracy and relative ranking. The authors should provide a sensitivity analysis or at least a quantitative bound on how large these label effects would need to be to change the GBA-versus-TEMPO ranking. As it stands, the reported accuracy numbers inherit an unknown label bias.
minor comments (5)
  1. [Abstract / §2] The abstract says ORBITaL-Net imagery has approximately 0.47 m resolution, while §2 says 0.45 m. Please reconcile these numbers.
  2. [Table 4] Please report the number of valid cells n used for each reference grid, together with the sample statistics. This would help the reader understand why total ground-truth area varies so much across grids and would make the cross-grid comparison more interpretable.
  3. [§5.4] There is a typo: 'facaces' should be 'facades'.
  4. [§5.1] The finding that 'GHSL always achieves the highest error rates' is stated before the evidence in Table 4; since Table 4 is the primary evidence, it would be clearer to cite it immediately after the claim and to specify which cells are included in each comparison.
  5. [§3.3] TEMPO's temporal coverage is listed as 2018–2025 in Table 1, but the text in §3.3 says Q1 2018 through Q2 2025. Please ensure consistency and clarify why the 2023Q4 epoch was chosen rather than a more recent one.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the benchmark compares independent external products against an external manually-labeled ground-truth dataset; acknowledged limitations are validity threats, not circular reasoning.

full rationale

The paper's accuracy estimates are derived by reprojecting independent building products onto reference grids and comparing them to ORBITaL-Net, an externally produced, manually labeled ground-truth dataset. None of the evaluated products were trained on ORBITaL-Net: GBA was trained using OSM labels and Chinese-city data, TEMPO used weak supervision from Google Open Buildings and Overture, GHSL used a composite of existing built-up products, and Overture is a conflation of external footprint sources. Algorithm 1 and Algorithm 2 perform only geometric reprojection, intersection, and area-weighted aggregation; they contain no fitted parameters and no quantity defined in terms of the target ranking. The baseline estimator, which uses the mean ground-truth area, is explicitly a sanity-check reference and is not presented as a predictive product. The paper's own limitations (Sec. 5.4) — temporal mismatch between ORBITaL-Net's 2017 median imagery and TEMPO's 2023Q4 epoch, facade/occlusion labeling issues, and absence of European samples — are threats to the validity or generalizability of the comparison, not circularity. The self-citations to Lancellotti et al. 2025 and Markakis et al. 2026 appear only as application-motivation references in the introduction and do not support the benchmark's central claims. Thus the derivation chain is self-contained against external data and exhibits no load-bearing circular step.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central claim rests on treating ORBITaL-Net as unbiased temporal-matched ground truth and on the choice of reference grids, completeness threshold, and stratification bins. No new physical or algorithmic entities are introduced.

free parameters (1)
  • Reference-cell completeness threshold = 0.99
    Algorithm 1 retains only reference-grid cells with at least 99% overlap with ORBITaL-Net imagery. This hand-chosen threshold determines the evaluated cell set and changes the ground-truth sample across grids (67M vs 142M m² total GT area).
assumptions (3)
  • domain assumption ORBITaL-Net's manual labels provide an unbiased measure of true building footprint area.
    Central premise for treating ORBITaL-Net as ground truth. Sec. 5.4 undermines it by noting off-nadir facades are included and occluded buildings are unlabeled.
  • ad hoc to paper Temporal mismatch between ground truth (median 2017) and product epochs (2019-2026) does not change the accuracy ranking.
    The paper states the opposite for GBA versus TEMPO in Sec. 5.4 ('systematic bias in favor of GBA estimates'), yet retains the ranking conclusions in Sec. 6.
  • domain assumption Conclusions from ORBITaL-Net chips generalize globally despite no European samples and only ~6,480 km² coverage.
    Used to frame the comparison as 'global' (Sec. 6). The authors hypothesize Europe would not change conclusions but provide no test.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Global Building Area Estimation Products: How Accurate Are They?." pith.science (2026). https://pith.science/paper/HEUKHNXP

@misc{pith2026260719766,
  author       = {Pith},
  title        = {Pith review of: Global Building Area Estimation Products: How Accurate Are They?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HEUKHNXP}},
  note         = {Machine review of arXiv:2607.19766}
}
read the original abstract

Geo-spatial rasters of building footprint area are useful for a variety of tasks, such as monitoring urbanization, improving energy efficiency, and tracking greenhouse gas emissions. There are now multiple global building raster datasets, however there lacks an independent, comprehensive, and fair assessment of their accuracy. In this work, we evaluate the accuracy of four major global building products: Global Human Settlement Layer (GHSL), Microsoft's TEMPO (TEMPO), The Global Building Atlas (GBA), and Overture. As ground truth for assessing their accuracy, we use ORBITaL-Net, a globally diverse dataset of manually labeled building footprints. To ensure fairness, we evaluate products on grids of multiple spatial resolutions, and several conventional performance metrics. Our results indicate that either GBA or TEMPO generally achieves the highest overall accuracy, depending upon the particular evaluation criteria. We also stratify the accuracy of each product by several factors: geographic location, population density, and income groups. The results reveal that product accuracy can sometimes vary significantly with respect to these factors. Notably, all products are significantly less accurate in Africa and Asia. Most products also suffer significant accuracy reduction in high-density urban areas.

Figures

Figures reproduced from arXiv: 2607.19766 by the authors.

Figure 1
Figure 1. Three examples (a-c) of satellite imagery (first column) and corresponding [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Left: A sample ORBITaL-Net optical image, overlaid by reference grid cells [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗
Figure 3
Figure 3. Qualitative samples showing the pipeline used to compare products using the [PITH_FULL_IMAGE:figures/full_fig_p021_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Examples of: (a) partially occluded imagery due to clouds, (b) fully occluded [PITH_FULL_IMAGE:figures/full_fig_p025_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

33 extracted references · 2 linked inside Pith

  1. [1]

    Spatial Statistics , volume=

    Accuracy of areal interpolation methods for count data , author=. Spatial Statistics , volume=. 2015 , publisher=

  2. [2]

    Earth System Science Data , volume=

    3D-GloBFP: The first global three-dimensional building footprint dataset , author=. Earth System Science Data , volume=. 2024 , publisher=

  3. [3]

    Publications Office of the European Union , year=

    Applying the degree of urbanisation: A methodological manual to define cities, towns and rural areas for international comparisons: 2021 edition , author=. Publications Office of the European Union , year=

  4. [4]

    IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , volume=

    Auditing geospatial datasets for biases: Using global building datasets for disaster risk management , author=. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , volume=. 2024 , publisher=

  5. [5]

    2025 , author =

    GADM: Global Administrative Areas , howpublished =. 2025 , author =

  6. [6]

    2026 , note =

    World Bank Country and Lending Groups , author =. 2026 , note =

  7. [7]

    Scientific data , volume=

    WorldPop, open data for spatial demography , author=. Scientific data , volume=. 2017 , publisher=

  8. [8]

    Earthquake Spectra , volume=

    Global building exposure model for earthquake risk assessment , author=. Earthquake Spectra , volume=. 2023 , publisher=

Show all 33 references
  1. [9]

    3.2 Global Atlas of the three major greenhouse gas emissions for the period 1970--2012 , author=

    EDGAR v4. 3.2 Global Atlas of the three major greenhouse gas emissions for the period 1970--2012 , author=. Earth System Science Data , volume=. 2019 , publisher=

  2. [10]

    Proceedings of the National Academy of Sciences , volume=

    Global scenarios of urban density and its impacts on building energy use through 2050 , author=. Proceedings of the National Academy of Sciences , volume=. 2017 , publisher=

  3. [11]

    PloS one , volume=

    A meta-analysis of global urban land expansion , author=. PloS one , volume=. 2011 , publisher=

  4. [12]

    arXiv preprint arXiv:2511.19277 , year=

    Closing Gaps in Emissions Monitoring with Climate TRACE , author=. arXiv preprint arXiv:2511.19277 , year=

  5. [13]

    Estimating Global, High-resolution Onsite Building Emissions , author=

  6. [14]

    GIScience & Remote Sensing , volume=

    Iterative self-organizing SCEne-LEvel sampling (ISOSCELES) for large-scale building extraction , author=. GIScience & Remote Sensing , volume=. 2022 , publisher=

  7. [15]

    Journal of Remote Sensing , volume=

    The last puzzle of global building footprints—Mapping 280 million buildings in East Asia based on VHR images , author=. Journal of Remote Sensing , volume=. 2024 , publisher=

  8. [16]

    2025 , note =

    Global ML Building Footprints , howpublished =. 2025 , note =

  9. [17]

    arXiv preprint arXiv:2107.12283 , year=

    Continental-scale building detection from high resolution satellite imagery , author=. arXiv preprint arXiv:2107.12283 , year=

  10. [18]

    2025 , note =

    OpenStreetMap , howpublished =. 2025 , note =

  11. [19]

    doi:10.5281/zenodo.5571936 , url =

    Zanaga, Daniele and Van De Kerchove, Ruben and De Keersmaecker, Wanda and Souverijns, Niels and Brockmann, Carsten and Quast, Ralf and Wevers, Jan and Grosu, Alex and Paccini, Audrey and Vergnaud, Sylvain and Cartus, Oliver and Santoro, Maurizio and Fritz, Steffen and Georgiev...

  12. [20]

    Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

    A convnet for the 2020s , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

  13. [21]

    Proceedings of the European conference on computer vision (ECCV) , pages=

    Unified perceptual parsing for scene understanding , author=. Proceedings of the European conference on computer vision (ECCV) , pages=

  14. [22]

    Remote Sensing of Environment , volume=

    A deep learning method for building height estimation using high-resolution multi-view imagery over urban areas: A case study of 42 Chinese cities , author=. Remote Sensing of Environment , volume=. 2021 , publisher=

  15. [23]

    arXiv preprint arXiv:2511.12104 , year=

    TEMPO: Global Temporal Building Density and Height Estimation from Satellite Imagery , author=. arXiv preprint arXiv:2511.12104 , year=

  16. [24]

    Earth System Science Data Discussions , volume=

    GlobalBuildingAtlas: an open global and complete dataset of building polygons, heights and LoD1 3D models , author=. Earth System Science Data Discussions , volume=. 2025 , publisher=

  17. [25]

    Scientific Data , volume=

    Outlining where humans live, the World Settlement Footprint 2015 , author=. Scientific Data , volume=. 2020 , publisher=

  18. [26]

    ISPRS Journal of Photogrammetry and Remote Sensing , volume=

    Breaking new ground in mapping human settlements from space--The Global Urban Footprint , author=. ISPRS Journal of Photogrammetry and Remote Sensing , volume=. 2017 , publisher=

  19. [27]

    Scientific data , volume=

    A crowdsourced global data set for validating built-up surface layers , author=. Scientific data , volume=. 2022 , publisher=

  20. [28]

    International Journal of Applied Earth Observation and Geoinformation , volume=

    Spatially explicit accuracy assessment of deep learning-based, fine-resolution built-up land data in the United States , author=. International Journal of Applied Earth Observation and Geoinformation , volume=. 2023 , publisher=

  21. [29]

    Remote sensing of environment , volume=

    Assessing the accuracy of multi-temporal built-up land layers across rural-urban trajectories in the United States , author=. Remote sensing of environment , volume=. 2018 , publisher=

  22. [30]

    2021 IEEE international geoscience and remote sensing symposium IGARSS , pages=

    Global land use/land cover with Sentinel 2 and deep learning , author=. 2021 IEEE international geoscience and remote sensing symposium IGARSS , pages=. 2021 , organization=

  23. [31]

    International Journal of Digital Earth , volume=

    Advances on the Global Human Settlement Layer by joint assessment of Earth Observation and population survey data , author=. International Journal of Digital Earth , volume=. 2024 , publisher=

  24. [32]

    Plos one , volume=

    Accuracy assessment of Global Human Settlement Layer (GHSL) built-up products over China , author=. Plos one , volume=. 2020 , publisher=

  25. [33]

    Scientific Data , volume=

    ORBITaL-Net: A labeled training library for large-scale building feature extraction , author=. Scientific Data , volume=. 2025 , publisher=

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.