{"id":"279f8653-88cf-418a-af68-c540ae232cde","arxiv_id":"2608.00012","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new benchmark evaluates MLLMs on raw satellite sounding streams across disaster lifecycle phases; all tested models score below 0.30, exposing large capability gaps.","lead":"Obshazard-bench is a new test that feeds AI models raw satellite temperature/humidity soundings and station data to answer disaster questions before, during, and after events. On 127 past disasters, the best model still scores below 0.30, showing current AI is far from ready for autonomous disaster response.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Event metadata in each VQA sample (country/location/date/category) allows answers from memorized disaster records; low scores don't isolate raw-observation reasoning, so the central claim is not yet established.","rationale":"The reader's weakest assumption—that event context dominates the raw sounding signal—is exactly the load-bearing concern here. The central claim requires that low scores reflect an inability to reason from raw AMSU-A/HIRS/MHS observations. However, the benchmark supplies country, location, date, and disaster category as text context, and many questions (e.g., total deaths, magnitude) have answers that can be recalled from historical disaster records. The authors even acknowledge that some performance benefits come from world knowledge, not from the soundings. A visual-input ablation would settle whether the raw tensor contributes anything. The reader's CONDITIONAL verdict already accounts for this gap, so my stress test does not change the verdict; it reinforces it with a concrete check.","tokens_in":14680,"tokens_out":4548,"duration_ms":119751,"concrete_test":"Run a visual-input ablation: for every VQA sample, replace the sounding tensor with zero-filled noise of identical shape while keeping the same prompt and attached event metadata. If model scores remain statistically unchanged relative to the full-input condition, then the raw observations are not used, and the central claim collapses. Companion check: run a metadata-ablation that removes country/location/date; a large score drop would indicate context-driven answering. The visual-input ablation is the single decisive check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that MLLMs fail to transform raw sounding streams into decision-relevant disaster reasoning—requires that the evaluation isolates physical reasoning from the raw observations. The benchmark attaches event metadata (country, location, event date, disaster category) to every sample (Section III-C, Fig. 1). Many target variables, especially Humanitarian Burden and Magnitude, are historically recorded in EM-DAT; a model with broad world knowledge can answer these from memory of the event, without using the AMSU-A/HIRS/MHS tensor at all. The paper's own Section IV-C attributes GPT-5.5's strong earthquake/wildfire performance to 'broader world knowledge and semantic reasoning about disaster mechanisms,' confirming that context contributes. Without a control condition that removes or randomizes the observation input, or ablates the metadata, the reported scores (all <0.30) cannot distinguish failure at raw physical inference from failure at recalling contextual facts or from the fact that some questions (e.g., total deaths) are underdetermined by atmospheric soundings alone. The absence of negative controls and the daily temporal compression (Section III-B) further weaken the link between the claimed 'real-time high-frequency' paradigm and the actual evaluation. Thus the benchmark's central finding of 'substantial limitations in transforming raw multi-channel physical observations' is not uniquely supported.","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-04T01:45:19.458450+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}