{"id":"43cf4d08-7d26-4ae5-94ec-170a25132ed9","arxiv_id":"2505.08202","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A narrative review of generative AI for disaster damage assessment that claims to be the first comprehensive survey but suffers from citation mismatches and overlaps with prior reviews.","lead":"This paper surveys how AI and generative AI are used to assess damage from earthquakes, wildfires, and cyclones. It is a narrative review, not a new experiment, and its usefulness depends on whether its cited sources actually support the claims.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's core value depends on accurate literature attribution, and the manuscript contains numerous citation-referent mismatches (e.g., [32], [35], [36], [39]) that break that foundation. The reader's REJECT verdict stands.","rationale":"The decisive issue is not disagreement with any external consensus but an internal evidentiary failure: the paper's genre is a narrative survey, so its only evidence for claims is the references it attaches to them. Those attachments are wrong in multiple places that matter—e.g., [32] does not discuss attention/spatial pyramid pooling, [35] does not discuss video damage classification, [39] does not discuss distress-call analysis, and [40] does not discuss sound-event detection. Because these mismatches are spread across sections and modalities, they indicate a systematic verification gap rather than a simple typo. The central novelty claim also rests on shaky ground: cited prior reviews ([4], [9], [31]) already cover adjacent ground, and the manuscript does not specify an inclusion protocol or a comparison that would establish first-ness. I therefore agree with the reader's weakest-assumption diagnosis and with the REJECT verdict. A full citation-claim audit would settle whether the errors are pervasive; if they are, revision and re-review are required, not minor edits.","tokens_in":11450,"tokens_out":5260,"duration_ms":48066,"concrete_test":"Construct a complete citation-claim matrix: for each of the 77 references, extract the sentence or clause it is attached to in the full text and compare it with the title and abstract of the cited paper. Classify each as supporting, partially supporting, or mismatched. Then recompute the paper's assertions with mismatches removed. The test is decisive because the current manuscript already shows gross mismatches at [32], [35], [36], [37], [39], [40], [41], [46], [47], and [48]; if a full audit confirms a mismatch rate above, say, 10%, the survey's attribution layer is unreliable and the verdict should remain REJECT. If the mismatches turn out to be isolated typos (e.g., one wrong reference repeated), a revised version could be acceptable after correction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is to be the first comprehensive survey of GenAI in disaster assessment and response. For a narrative survey, the evidence for every survey claim is its bibliography, so the accuracy of each citation-to-claim mapping is load-bearing. That mapping fails systematically. In Section III.B, 'attention mechanisms and spatial pyramid pooling ... [32]' cites Yang (2020), a conditional GAN study for aero-engine vibration, not an image-recognition method. In Section III.C, real-time damage classification in video streams is attributed to [35], which is Dong et al.'s image super-resolution paper; detection of building collapses from video frames is attributed to [36], SenseGen's synthetic sensor data generator; and SfM/MVS 3D reconstruction is attributed to [37], Viola & Jones's 2001 face detector. In Section III.D, distress-call analysis is attributed to [39], a photogrammetry paper on historical buildings, and sound-event detection to [40], a book on data anonymization. Similar mismatches appear at [41], [46], [47], and [48]. These are not peripheral; they are the citations that supposedly support the core technical descriptions of how GenAI is applied across modalities. With this many wrong referents, a reader cannot trust the survey's synthesis without independently re-deriving every claim from the primary literature, which defeats the purpose of a survey. The 'first comprehensive' novelty claim is also not established: the manuscript itself cites earlier reviews such as [4] (Generative Deep Learning for Natural Hazard Analysis), [9] (AI/ML earthquake damage assessment), and [31] (natural disaster detection survey) without explaining what makes the present scope novel. Thus the central contribution—a trustworthy organizing reference—is not supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents a narrative survey of AI and generative AI applied to disaster damage assessment, with domain-specific reviews for earthquakes, wildfires, and cyclones, and a modality-based organization covering text, image, video, and audio data. It also discusses misinformation, adversarial attacks, privacy, explainability, and future research directions. The paper's stated central contribution is that it is the first comprehensive survey of GenAI techniques for disaster assessment and response. The survey is entirely literature-based; it contains no experiments, datasets, or machine-checked artifacts, so its evidentiary value rests wholly on the accuracy of its citations.","tokens_in":11722,"tokens_out":6837,"duration_ms":56884,"significance":"If its citation base were accurate, this survey would be a useful organizing reference for researchers and practitioners at the intersection of generative AI and disaster management. It correctly identifies a real trend: the use of GANs, VAEs, and large multimodal models for damage classification, synthetic data augmentation, and multimodal fusion. However, the significance is not realized in the current version. The survey's only evidence is its bibliography, and that bibliography is systematically unreliable. Multiple core technical claims across Sections III.B, III.C, III.D, and V are attached to references that do not address the claimed topic. A reader cannot trust the survey's synthesis without independently re-verifying every citation against the primary literature, which defeats the purpose of a survey. For these reasons, the claimed contribution cannot be accepted as it stands.","major_comments":[{"comment":"The image-data section cites [32] (Yang 2020, a conditional GAN for aero-engine vibration analysis) to support the claim that attention mechanisms and spatial pyramid pooling extract multi-scale contextual information for damage extraction; [33] (Xing et al., flood vulnerability from remote sensing and street-view imagery) to support image inpainting; and [34] (Ghimire et al., text-based generative AI in construction) to support image super-resolution. None of these references addresses the techniques described in the accompanying sentences. Because image analysis is one of the core modalities of the survey, these mismatches break the evidentiary chain for a central section.","section":"§III.B, refs [32]–[34]"},{"comment":"The video-data section attributes real-time damage classification in video streams to [35] (Dong et al., image super-resolution), building-collapse detection from standalone frames to [36] (Alzantot et al., a synthetic sensor data generator), and SfM/MVS 3D reconstruction to [37] (Viola & Jones, a face detector). These references do not support the specific claims with which they are paired, leaving the entire video-modality discussion without credible evidentiary support.","section":"§III.C, refs [35]–[37]"},{"comment":"The audio-data section attributes distress-call analysis to [39] (Galantucci & Fatiguso, photogrammetry of historical buildings), sound-event detection to [40] (Arbuckle & El Emam, data anonymization), and cross-modal correlation to [41] (Smadi et al., speech recognition). The cited references do not address these claims; in particular, [39] and [40] are topically unrelated to audio-based disaster analysis. The audio section therefore lacks valid literature support for its main technical assertions.","section":"§III.D, refs [39]–[41]"},{"comment":"The privacy section cites [47] (Zanardelli et al., an image forgery detection survey) as the source for k-anonymity and l-diversity, and [48] (Szegedy et al., intriguing properties of neural networks) as the source for the Laplace mechanism of differential privacy. Neither reference discusses privacy or differential privacy, so the technical discussion of privacy protections is unsupported.","section":"§V, refs [47]–[48]"},{"comment":"In the text-modality section, the manuscript credits 'Tarasconi, Francesco, et al. (2017)' with work on relationship extraction and event correlation, but the citation at that point is [23] (Li et al., Data-Driven Techniques in Disaster Information Management). The actual Tarasconi et al. paper appears as [29]. This is another citation-referent mismatch in a section that otherwise discusses core text-analysis claims.","section":"§III.A, ref [23] vs [29]"},{"comment":"The paper's central claim of being 'the first comprehensive survey of GenAI techniques used for disaster assessment and response' is not established. The manuscript itself cites earlier surveys with overlapping scope, notably [4] (Ma et al., generative deep learning in natural hazard analysis) and [9] (Bhadauria, AI/ML in earthquake engineering), but it does not compare its coverage against these works or justify why its scope is distinct. A survey's novelty claim requires a systematic comparison with prior surveys, which is absent.","section":"Abstract and §VI, novelty claim"}],"minor_comments":[{"comment":"References [28] and [30] are identical (both are Havas et al., E2mC). The duplicate [30] is cited in §III.B to support the importance of image data, but the E2mC paper concerns social media and crowdsourcing, not image data.","section":"Reference list, [28] and [30]"},{"comment":"References [25], [26], and [27] appear in the reference list but are never cited in the body of the manuscript.","section":"Reference list, [25]–[27]"},{"comment":"The sentence 'Tarasconi, Francesco, et al(2017) published research suggesting that GenAI also supports relationship extraction' is anachronistic, since GenAI as a term and the underlying large generative models were not the subject of 2017 research; moreover, the cited reference number is wrong, as noted in Major Comment 5.","section":"§III.A, 'Tarasconi et al. (2017)'"},{"comment":"Figure 1 is never referenced in the text; the caption 'Taxonomy of Natural Disaster surveys and methods in this survey' does not explain how the figure relates to the surrounding discussion.","section":"Figure 1 caption"},{"comment":"The abstract contains a subject-verb disagreement: 'AI and Generative AI ... presents a breakthrough solution'; the sentence should read 'present.' Additionally, the reference to 'Gen-AI' with a hyphen in the final sentence of the abstract is inconsistent with the 'GenAI' spelling used elsewhere.","section":"Abstract and title"}],"recommendation":"reject","confidential_remarks":"The citation-referent mismatches are extensive and systematic, spanning nearly every modality section and several sections beyond. This is not a case of a few inattentive citations; the bibliography appears to have been assembled without checking whether each reference supports the specific claim it is attached to. For a survey, whose sole evidentiary basis is its references, this is a foundational failure. The novelty claim is also asserted without comparison to existing surveys. The paper would require a complete re-verification and likely rewriting of the bibliography and several sections, which is beyond the scope of a normal revision. I therefore support rejection. The authors might be encouraged to resubmit a thoroughly re-referenced version, ideally with a documented search and selection methodology."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know up front. First, the survey's architecture is fine: it organizes GenAI-for-disaster work by disaster type and by modality (text, image, video, audio), and it covers privacy, security, and misinformation. That's a sensible skeleton. Second, the citation-to-claim mapping is broken in every modality section. I checked a sample: [32] is a CGAN paper about aero-engine vibration, cited for attention mechanisms and spatial pyramid pooling; [35] is the classic super-resolution paper, cited for video-stream damage classification; [36] is SenseGen synthetic sensor data, cited for detecting building collapses in video; [37] is Viola & Jones face detection, cited for SfM/MVS 3D reconstruction; [39] is photogrammetry of historical buildings, cited for distress call analysis; [40] is an anonymization book, cited for sound event detection. I could go on. For a narrative survey, citations are the evidence, and this evidence is unreliable.\n\nWhat the paper does do well: the recent references (2024–2025) on earthquake damage assessment with GPT-4o and Gemini, wildfire segmentation with a quantum-compatible VQ-VAE, and the hurricane damage ML studies are real and relevant. The discussion of CycleGAN-based data augmentation and the ethics section are reasonable. The taxonomy figure is a good idea.\n\nWhere it falls down: the mismatches are not typos. They are the support for the technical descriptions. A reader cannot trust the synthesis without going back to primary sources, which defeats the purpose of a survey. The 'first comprehensive survey' claim also doesn't stand, since the paper itself cites earlier reviews on generative deep learning for natural hazards [4], AI/ML earthquake damage assessment [9], and natural disaster detection surveys [31]. The paper never explains what new scope or organization it adds beyond those. Minor issues include [30] duplicating [28] and a few listed references never being cited in the text.\n\nFor a survey, this is a load-bearing flaw. The topic is timely and the authors' framing is sensible, but the execution is not reliable enough for readers. If every citation were rebuilt and the novelty claim recalibrated, this could become a useful organizing resource. In current form, I would not publish it or cite it.\n\nMy recommendation: if this lands on your desk, do not desk-reject without comment. Send it back for major revision with a request to verify every citation and to reposition the contribution relative to the prior reviews it cites. That is referee time well spent, because the topic matters and the fix is concrete.","headline":"A well-intentioned survey whose citation-to-claim errors in every modality section make it unusable as a reference; the 'first comprehensive' claim also collapses on contact with its own bibliography.","tokens_in":12323,"tokens_out":3089,"would_cite":false,"duration_ms":31086,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims to be the first comprehensive review of generative AI techniques for disaster assessment and response.","keywords":["Generative AI","Disaster response","Damage assessment","Multimodal data","Earthquake","Wildfire","Cyclone","Misinformation"],"falsifier":"Read the abstracts of the cited papers directly and compare them with the survey's in-text claims, starting with [32], [35], and [39]. If a material fraction of the survey's attributions fail this check, its reliability as a map of the field is not established.","tokens_in":11272,"feed_emoji":"🛰️","tokens_out":6365,"duration_ms":58826,"temperature":0.7,"pith_summary":"This paper is a survey of how artificial intelligence, including generative AI, is being applied to damage assessment after natural disasters. Its central claim is that this is the first comprehensive review devoted specifically to GenAI in disaster assessment and response, and that the field now has enough results to map. The authors argue that generative models can combine text, image, video, and audio data to make damage assessment faster and more complete across earthquakes, wildfires, and cyclones. They also catalog the risks, including fake media that could corrupt assessments, privacy concerns, and model bias, and they call for secure and ethical deployment. A sympathetic reader would take this as an organizing reference: a structured inventory of what techniques exist, what evidence supports them, and where research gaps lie.","feed_headline":"First survey maps generative AI for disaster damage assessment","feed_subtitle":"Earthquake, wildfire, and cyclone damage tools across text, image, video, and audio.","key_machinery":"The organizing device is a taxonomy that cross-cuts three disaster types — earthquakes, wildfires, and cyclones — with four data modalities: text, image, video, and audio. The technical core is a family of generative models: Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs) such as CycleGAN, Vector-Quantized VAEs (VQ-VAEs), and transformer-based large language models. These models do three main jobs in the survey's account: they generate synthetic training data to compensate for scarce disaster imagery, they translate pre-disaster scenes into simulated post-disaster scenes, and they fuse multimodal inputs for situational awareness. The taxonomy is what carries the survey's argument, because it lets the authors claim coverage of a whole field rather than isolated applications.","core_discovery":"The paper's central claim, stated in its own words, is that it represents the first comprehensive survey of GenAI techniques used for disaster assessment and response. On the paper's own terms, the discovery is that generative AI has moved from being an abstract possibility to a documented set of applications: VAEs and GANs synthesize training data and simulate damaged scenes, CycleGANs convert pre-fire images to post-fire imagery for wildfire detection, and fine-tuned multimodal models classify building damage from post-earthquake photos. The survey organizes these results by disaster type and by data modality, and it reports guarded evidence, such as GPT-4o achieving 28.6 to 75.0 percent accuracy on EMS-98 damage classification, which it reads as promising but not yet operational. It also argues that the same generative power creates a threat surface, since manipulated images and videos can poison assessment pipelines.","pith_inferences":["An implication the authors leave implicit is that the survey's practical value depends on its attributions being accurate, so a reader who plans to act on a specific technique should check the original source before relying on the survey's framing.","A testable extension of the survey's thesis is to run the same fine-tuned multimodal model used for earthquake damage classification on cyclone and wildfire imagery, to see whether the reported accuracy range generalizes across disaster types.","The survey's modality taxonomy suggests a natural benchmark design: fuse text posts, drone video, and audio calls from a single event and measure whether multimodal fusion beats any single modality for damage severity scoring.","One consequence the authors gesture at but do not develop is that the same generative models used for assessment could be used for privacy protection, for example by generating sanitized or de-identified versions of sensitive disaster media before analysis."],"forward_implications":["If the survey's picture is right, damage maps could be produced in near real time by feeding social media text and images, drone video, and audio calls into generative pipelines.","Synthetic data from CycleGANs and related models could directly address the chronic shortage of labeled disaster imagery that currently limits supervised damage detection.","Fine-tuned multimodal large language models could eventually triage building damage, but the reported accuracy range indicates they are not yet reliable enough for operational use without human review.","Because generative media can be weaponized, any deployed assessment system would need provenance checks such as watermarking and perceptual hashing as part of the pipeline.","A shared benchmark built from unified multimodal disaster datasets would be needed to measure progress, since the survey identifies the lack of such benchmarks as a gap."],"supporting_citations":[{"why":"Introduces variational autoencoders, one of the two core generative model families the survey builds its assessment story on.","marker":"[2]"},{"why":"Introduces generative adversarial networks, the other core model family used for synthetic damage imagery.","marker":"[3]"},{"why":"Supplies the general case that generative deep learning is applicable to natural hazard data generation.","marker":"[4]"},{"why":"Provides the lead earthquake example of multimodal social media data estimating shaking intensity.","marker":"[7]"},{"why":"Provides the post-earthquake damage classification accuracy figures that anchor the survey's cautious assessment of current GenAI performance.","marker":"[8]"},{"why":"Supplies the wildfire segmentation method using a modified vector-quantized VAE.","marker":"[14]"},{"why":"Supplies the CycleGAN data-augmentation result that improved wildfire detection accuracy.","marker":"[15]"},{"why":"Provides the cyclone-domain machine learning baseline for hurricane flood damage risk estimation.","marker":"[19]"},{"why":"Supplies the transfer-learning results for hurricane damage classification from images.","marker":"[20]"},{"why":"Supports the claim that generative models can predict infrastructure vulnerability by simulating post-storm states.","marker":"[21]"}],"fun_headline_variants":["GenAI enters disaster response: first survey of damage tools","Survey: generative AI aids disaster damage assessment","AI survey charts damage tools for quakes, fires, cyclones","First look at GenAI for disaster damage assessment","Generative AI for disaster damage: survey of tools and threats"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every citation in the survey actually supports the claim it is attached to, for example that [32] really covers attention mechanisms, [35] really covers video streams, and [39] really covers distress-call analysis.","fun_headline_variants_meta":{"raw":{"variants":["GenAI enters disaster response: first survey of damage tools","Survey: generative AI aids disaster damage assessment","AI survey charts damage tools for quakes, fires, cyclones","First look at GenAI for disaster damage assessment","Generative AI for disaster damage: survey of tools and threats"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1285,"prompt_tokens":925,"completion_tokens":360,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":281}},"tokens_in":541,"tokens_out":360,"duration_ms":3544,"temperature":1.0,"reasoning_tokens":281,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:59:33.580203+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Read the abstracts of the cited papers directly and compare them with the survey's in-text claims, starting with [32], [35], and [39]. If a material fraction of the survey's attributions fail this check, its reliability as a map of the field is not established.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the wildfire segmentation method using a modified vector-quantized VAE."}],"review_version":1}