{"id":"4ce6ed61-080a-4008-be6e-fb7f44b3aa30","arxiv_id":"2505.14043","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A Mamba-based multispectral detector with three new modules reports state-of-the-art accuracy on VEDAI at real-time speed and with 17 MB size, though it trails some methods on larger objects.","lead":"This paper builds a small-target detection model on the Mamba architecture, combining visible and infrared imagery with three added modules to improve local detail and suppress background. It reports state-of-the-art accuracy on one benchmark and real-time speeds on another, making the result potentially useful for drone and edge-device deployments.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1's baseline mAP values are inconsistent with their own per-class AP columns, so the VEDAI state-of-the-art claim is not currently reproducible.","rationale":"The paper's central claim is that S4^6-MSTD achieves state-of-the-art small-target detection accuracy, specifically the 81.2% mAP on VEDAI. That claim lives or dies with the comparability of Table 1. The table provides per-class AP values for every method, so the reported mAP should follow from those values under a single, stated averaging rule. It does not: the Ours row exactly equals the unweighted mean of its per-class APs, while at least two baseline rows deviate by more than two points from the same computation. This is not a matter of missing error bars; it is an internal inconsistency in the evidence table that the SOTA number is built on. If the baseline numbers were generated with a different metric, or the per-class numbers were taken from different runs, then the headline 'state-of-the-art' comparison is not currently valid. The reader's weakest_assumption correctly flagged cross-method comparability, and this concern is a concrete instance of that failure. I do not conclude the method is bad: the ablations are internally consistent, the efficiency numbers are plausible, and the architecture story is coherent. But the central accuracy claim needs to be re-established under one protocol before the paper can be accepted. For that reason I keep the reader's CONDITIONAL verdict unchanged, while sharpening the condition: the authors should release the evaluation code and verify Table 1's mAP arithmetic, or explain the averaging rule that reconciles the rows.","tokens_in":15138,"tokens_out":11288,"duration_ms":116659,"concrete_test":"Write a short script that takes the per-class AP columns of Table 1, computes both the unweighted class mean and the instance-weighted mean (using VEDAI's per-class object counts), and compares each row's reported mAP. Then run the official VEDAI evaluation script on the authors' released checkpoint. If YOLOFusion's row remains 78.6 while its class mean is 75.9, or DMM's remains 75.0 while its class mean is 72.6, the baseline numbers come from a different protocol and the SOTA claim must be re-derived under one consistent protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 1's reported mAP values are not internally consistent with the per-class APs shown in the same table, and the inconsistency sits exactly on the SOTA claim. Averaging the eight per-class APs for Ours gives (91.8+69.3+82.3+78.5+84.6+85.8+75.6+81.6)/8 = 81.19, matching the printed 81.2. Applying the same unweighted class-mean formula to the YOLOFusion row gives 75.9, not the printed 78.6, and to DMM gives 72.6, not 75.0. If VEDAI mAP is the usual unweighted class mean, Table 1 mixes at least two averaging rules and the claimed state of the art is not established; if it is instance-weighted, then Ours's exact equality with the unweighted mean is unexplained. Either way, the central '81.2% mAP, state-of-the-art' assertion depends on a comparison that is not currently reproducible from the table. This is a concrete comparability failure, not just missing error bars.","agreement_with_reader":"partial"},"referee_report":null,"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-07T15:41:21.792764+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}