Pith. sign in

REVIEW 3 major objections 5 minor 9 references

mwBTFreddy: A Dataset for Flash Flood Damage Assessment in Urban Malawi

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read The mwBTFreddy dataset pairs pre- and post-Cyclone Freddy satellite images of Blantyre with building-level damage labels to give machine learning a local benchmark.

desk verdict A genuinely useful new dataset for a data-scarce region, but the damage labels are unvalidated and the pre/post imagery dates the authors admit are not a faithful reflection of the flood event—worth a serious referee, not a desk reject. read the letter →

arxiv 2505.01242 v1 pith:JUHUOYUP submitted 2025-05-02 cs.LG

classification cs.LG
keywords mwBTFreddydatasetsatelliteimageryflooddamageassessmentbuildingclassificationCycloneFreddyMalawixBDscalegeoreferencedannotations
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces mwBTFreddy, a dataset of 696 georeferenced satellite image tiles—348 pre- and 348 post-Cyclone Freddy—covering three flood-hit urban areas of Blantyre, Malawi. Each image comes with a JSON file of building polygons and a damage level (no, minor, major, or destroyed) assigned by visual inspection following the xBD damage scale. The authors' claim is that this localized resource fills a gap: existing disaster datasets are built on imagery and building styles from other regions, so models trained on them may not transfer to Malawian urban settings. If the dataset is fit for purpose, it gives researchers a benchmark for building detection and damage classification in an African urban context and supports spatial analysis for relocation, drainage, and emergency-response planning.

What carries the argument

The load-bearing mechanism is the paired-image grid cell: each tile is captured before and after the cyclone, georeferenced from four corner coordinates, so the same building can be compared across time. Building polygons are drawn manually in GIS software and assigned one of four xBD damage levels (0 no damage, 1 minor, 2 major, 3 destroyed); the JSON annotation ties each polygon to pixel and geographic coordinates. The pre/post pairing is what lets a model separate flood damage from ordinary building differences, and the xBD scale is what makes the labels comparable to existing disaster datasets.

What would settle it

Select a random sample of annotated buildings and have independent annotators re-label them from the same images, or survey the buildings on the ground; if agreement falls well short of the level needed for reliable training labels, or if field checks contradict the JSON damage levels, the dataset's central value is not established.

Watch

Extended reading notes

Core claim

The central discovery is that a compact, locally sourced dataset can capture building-level flood damage in a data-scarce African city. The dataset consists of 696 annotated instances drawn from 1,026 image tiles over Chilobwe, Ndirande, and Chirimba, with each instance pairing a pre-disaster (August 2022) and post-disaster (May 2023) image of the same grid cell; buildings too blurred, too obscured by trees, or split across tiles were excluded. Damage labels follow the four-level xBD scale and are attached to manually drawn building polygons with both pixel and geographic coordinates. The paper argues this resource enables building detection, damage classification, and flood-damage visualization specifically tuned to region-specific construction styles, and notes the dataset was already used in a disaster-damage challenge.

Load-bearing premise

The load-bearing premise is that the manually assigned damage labels are accurate enough to be trusted; the annotations were produced by visual assessment by the authors with no reported inter-annotator agreement, independent ground-truth validation, or error analysis, so if those labels are unreliable, models trained or evaluated on the dataset would not measure true flood damage.

Editorial extensions

If this is right

  • Machine-learning models for building detection and damage classification can be trained or evaluated on imagery and building styles from a Malawian urban setting rather than only on datasets from other countries.
  • Because pre- and post-disaster images are paired by grid cell, models can learn change signals that isolate flood effects from background variation.
  • The JSON annotations support spatial analysis and visualization, so planners can map damage concentrations around the affected mountains and target relocation or drainage decisions.
  • The 330 unannotated tiles retained separately allow users to reconstruct the broader landscape or study why some areas were excluded.
  • The dataset's prior use in a public disaster-damage challenge provides an external benchmark for comparing detection and classification approaches on this data.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A dataset of 696 annotated images from three neighborhoods is small for training deep models from scratch, so its realistic role is likely as a benchmark or fine-tuning set rather than a standalone training corpus.
  • The labels encode the annotators' visual judgments following the xBD rubric; models trained on it will inherit those judgments, so adding inter-annotator agreement statistics or field-validated subsets would materially strengthen downstream conclusions.
  • Comparing mwBTFreddy against xBD or other flood datasets could quantify domain shift between Malawian urban building styles and the regions represented in existing benchmarks, which is a direct way to test the paper's motivation.
  • The dataset could be extended to other Malawian cities or to rural and peri-urban areas the cyclone also hit, where building types and damage patterns likely differ.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces mwBTFreddy, a dataset of paired pre- and post-disaster satellite images of urban areas in Blantyre, Malawi, affected by Cyclone Freddy in March 2023. The dataset contains 696 annotated images (348 pre-disaster and 348 post-disaster) with JSON annotation files providing building polygons, geographic coordinates, and damage labels following the xBD four-class scale (no damage, minor, major, destroyed). The images were exported from Google Earth Pro, georeferenced in QGIS, and manually annotated by members of the Kuyesera AI Lab. The paper is structured as a datasheet following the standard template, with sections on motivation, composition, collection process, preprocessing, uses, distribution, and maintenance. The authors claim the dataset fills a gap in localized disaster-response resources for African urban contexts and is suitable for building detection and damage classification tasks. The dataset is publicly available on Zenodo and has been used in a Zindi competition.

Significance. If the dataset's labels are reliable, mwBTFreddy addresses a genuine gap: there are few publicly available, localized satellite datasets of flood damage in African informal urban settlements, and the paper openly documents many sources of noise and limitations. The detailed datasheet format, the explicit discussion of exclusion criteria, and the public Zenodo release are strengths, as is the fact that the dataset has already been exercised in a competitive machine-learning setting. However, the central claim that the dataset can support trustworthy damage classification rests on the accuracy and reproducibility of the manually assigned labels, and that premise is not demonstrated. The authors themselves note that the pre/post imagery may not faithfully reflect the damage, and no inter-annotator agreement, ground-truth comparison, or baseline experiment is provided. The significance of the contribution is therefore conditional on label validation, which is the main load-bearing issue in the manuscript.

major comments (3)
  1. [Composition, 'Is any information missing from individual instances?' and Collection Process] The central claim that mwBTFreddy is suitable for training and evaluating machine-learning models for damage classification is not yet supported because the target labels are unvalidated. The Collection Process section describes manual visual assessment by the authors following the xBD scale, but no inter-annotator agreement, independent expert review, or ground-truth comparison is reported. This concern is amplified by the datasheet's own admission in Composition that 'the post and pre images may not be a faithfull reflection of the damage,' and by the temporal gaps: pre-disaster images are from August 2022, the flood event occurred in March 2023, and post-disaster images are from May 2023. Changes in that window could reflect construction, demolition, vegetation change, or post-cyclone cleanup rather than flood damage. Please add an inter-annotator agreement study or external validation (e.g., against official damage assessments or field reports), and quantify how the temporal window may affect label reliability.
  2. [Uses and Distribution; datasheet answer to 'Are there recommended data splits?'] The paper asserts that the dataset supports building detection and damage classification, but it provides no dataset statistics, no class distribution, no recommended train/validation/test splits, and no baseline experiments. The datasheet explicitly answers 'No' to the question of recommended splits, and no evaluation is reported anywhere in the manuscript. Without a baseline model run or at least a label-distribution analysis, a prospective user cannot assess whether the dataset is usable for the claimed tasks. Please report per-class building counts, geographic distribution of labels, and a simple baseline experiment (e.g., a standard object-detection or classification model) to demonstrate practical utility.
  3. [Table I, Table III, and Table V; Composition section] The image-count accounting is internally inconsistent and will confuse users. Table I reports 1,026 images for the three areas (Chilobwe 1,000, Ndirande 20, Chirimba 6), Table V reports 696 annotated images and 330 unannotated images, and Table III reports 348 pre-disaster images and 348 post-disaster images as the total annotated instances. If 696 is the total number of annotated images, then the 348/348 split means 348 paired locations, not 696 independent tiles; if the original 1,026 images include both pre and post tiles, the relationship between tiles, annotated images, and pre/post pairs should be stated explicitly. Please clarify the exact counting scheme and reconcile the tables.
minor comments (5)
  1. [Throughout] There are numerous typos and formatting errors, including 'anotated' for 'annotated', 'faithfull' for 'faithful', 'od March' for 'of March', and the consistently broken 'T able' table headers. A careful proofreading pass is needed.
  2. [Composition, 'Does the dataset contain all possible instances?'] The answer 'The dataset contains all possible instances' is contradicted by the later statement that 330 unannotated images were retained in a separate folder and by the exclusion criteria for buildings. Please reword to distinguish the annotated core dataset from the full set of captured tiles.
  3. [Preprocessing/cleaning/labeling] The text says 'No pre-processing of images was done' but then states that JPEG images downloaded from Google Earth Pro were converted to TIFF after geo-tagging and annotation. Conversion, cropping, and resizing described in the Collection Process are preprocessing steps; please make the description internally consistent.
  4. [Distribution] The license information is insufficient: 'Refer to Zenodo's terms of use' does not specify which license applies to the dataset. The authors should state an explicit license (e.g., CC-BY 4.0) and confirm that the Google Earth Pro imagery can be redistributed under that license.
  5. [Collection Process / Maintenance] The annotation and preprocessing scripts are only 'available upon reasonable request,' which weakens reproducibility. Since the dataset is archived on Zenodo, the scripts and a description of the annotation environment should be included in the same archive.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a dataset description with no derivation chain; its utility claims rest on externally hosted data and manual annotation, not on self-referential reasoning.

full rationale

This paper contains no mathematical derivation, fitted parameters, or predicted quantities, so there is no chain of reasoning that could reduce to its own inputs. The central claim is that mwBTFreddy is a useful localized dataset of paired pre/post satellite images with building-level damage labels. The labels are produced by manual visual assessment following the xBD scale, which is a process external to the paper's own claims rather than a definition of the claim itself. The only self-references are to the Zenodo dataset record, which is the artifact being described rather than a load-bearing justification for its content. The datasheet's explicit admission that 'the post and pre images may not be a faithfull reflection of the damage' is a genuine data-quality and label-validity limitation, but it concerns whether the dataset measures flood damage accurately, not whether the paper's argument is circular. Similarly, the absence of inter-annotator agreement or ground-truth validation is a correctness or reliability risk, not a circularity risk. Under the stated criteria, no circular step can be exhibited: no equation equals another by construction, and no fitted value is renamed as a prediction. The honest finding is therefore no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The dataset relies on the assumption that satellite imagery and manual visual labels are accurate and representative. No validation is provided, and the paper itself flags that the pre/post timing and image quality may not faithfully capture flood damage.

assumptions (4)
  • domain assumption Satellite imagery from Google Earth Pro provides sufficient resolution to identify buildings and classify damage.
    The entire annotation pipeline depends on visual interpretation of these images; the paper notes some buildings could not be identified due to resolution.
  • domain assumption Manual visual assessment following the xBD Joint Damage Scale yields accurate damage labels.
    No inter-annotator agreement or ground-truth comparison is reported; the paper admits georeferencing and image quality introduce noise.
  • domain assumption Pre-disaster images from August 2022 and post-disaster images from May 2023 bracket the event well enough to represent flood damage.
    The paper itself states that the chosen images may not be a faithful reflection of the damage because of what Google Earth Pro had available.
  • domain assumption The exclusions of buildings (occluded, unclear, appearing in multiple tiles) do not bias the remaining damage distribution.
    The paper excludes certain buildings and images but does not analyze whether these exclusions alter class balance or geographic coverage.

how reviews work

0 comments
Cite this review

Pith. "Pith review of mwBTFreddy: A Dataset for Flash Flood Damage Assessment in Urban Malawi." pith.science (2026). https://pith.science/paper/JUHUOYUP

@misc{pith2026250501242,
  author       = {Pith},
  title        = {Pith review of: mwBTFreddy: A Dataset for Flash Flood Damage Assessment in Urban Malawi},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JUHUOYUP}},
  note         = {Machine review of arXiv:2505.01242}
}
read the original abstract

This paper describes the mwBTFreddy dataset, a resource developed to support flash flood damage assessment in urban Malawi, specifically focusing on the impacts of Cyclone Freddy in 2023. The dataset comprises paired pre- and post-disaster satellite images sourced from Google Earth Pro, accompanied by JSON files containing labelled building annotations with geographic coordinates and damage levels (no damage, minor, major, or destroyed). Developed by the Kuyesera AI Lab at the Malawi University of Business and Applied Sciences, this dataset is intended to facilitate the development of machine learning models tailored to building detection and damage classification in African urban contexts. It also supports flood damage visualisation and spatial analysis to inform decisions on relocation, infrastructure planning, and emergency response in climate-vulnerable regions.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

9 extracted references · 8 canonical work pages

  1. [1]

    T., Kowe, P., Maviza, A., Magidi, J., Chikwiramakomo, L., Mavaringana, M

    Gumindoga, W., Liwonde, C., Rwasoka, D. T., Kowe, P., Maviza, A., Magidi, J., Chikwiramakomo, L., Mavaringana, M. d. J. P., and Tshitende, E. (2024). Urban flash floods modeling in M zuzu C ity, M alawi based on S entinel and MODIS data. Frontiers in Climate , Volume 6 - 2024

  2. [2]

    T., Choset, H., and Gaston, M

    Gupta, R., Hosfelt, R., Sajeev, S., Patel, N., Goodman, B., Doshi, J., Heim, E. T., Choset, H., and Gaston, M. E. (2019). xbd: A dataset for assessing building damage from satellite imagery. CoRR , abs/1911.09296

  3. [3]

    Kadzuwa, H. (2023). P ost C yclone F reddy D isaster M apping in M alawi. Edinburgh; July 2023. Accessed: 2024-11-09

  4. [4]

    mw BTF reddy (1.0) [data set]

    KuyeseraAI (2024). mw BTF reddy (1.0) [data set]. https://zenodo.org/records/14190390. Accessed: 2025-04-28

  5. [5]

    Lubanga, A., Uwishema, O., and Nazir, A. (2023). The devastating effect of C yclone F reddy amidst the deadliest cholera outbreak in M alawi: A double burden for an already weak healthcare system. Published in Annals of Medicine and Surgery, 85(7), pp. 3761--3763

  6. [6]

    Ma, Y., Zhou, F., Wen, G., Gen, H., Huang, R., Liu, G., and Pei, L. (2022). Assessment of buildings and electrical facilities damaged by flood and earthquake from satellite imagery. The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences , XLVI-3/W1-2022:133--140

  7. [7]

    Mawenda, J., Watanabe, T., and Avtar, R. (2020). An analysis of urban land use/land cover changes in B lantyre C ity, S outhern M alawi (1994–2018). Published in Sustainability, 12(6), Article 2377. DOI: 10.3390/su12062377

  8. [8]

    and Nambazo, O

    Nazombe, K. and Nambazo, O. (2023). Monitoring and assessment of urban green space loss and fragmentation using remote sensing data in the four cities of M alawi from 1986 to 2021. Published in Scientific African, 20, Article e01639

Show all 9 references
  1. [9]

    K uyeseraai disaster damage and displacement challenge

    Zindi (2025). K uyeseraai disaster damage and displacement challenge. https://zindi.africa/competitions/ . [Accessed 30-04-2025]

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.