{"id":"96df01f2-e77c-4a01-a13f-27b0e72d5827","arxiv_id":"2507.22685","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A multi-modal dataset and benchmark for leaf wetness detection using mmWave SAR and RGB images, with baseline results for the Hydra model.","lead":"This paper releases a benchmark dataset of synchronized mmWave radar, synthetic aperture radar (SAR), and RGB images for detecting whether leaves are wet or dry. It reports accuracy benchmarks using the authors' Hydra model across indoor and outdoor settings, and provides the dataset and SAR imaging code for community use.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The benchmark claim is undermined by contradictory accuracy figures and the absence of any train/test split; as written, the reported results cannot be reproduced or compared against future methods.","rationale":"The reader's rejection is appropriate. The main load-bearing weakness is not primarily label noise or sample size, though those matter; it is that the benchmark protocol is undefined and the numbers as printed are internally inconsistent. Section 4 contains two overlapping reports of scan-distance accuracy with different values for identical distances. Because no split is defined, a user cannot tell which numbers are held-out. This makes the headline contribution—a benchmark—unsupported. A revised paper that supplies the missing train/test protocol, resolves the contradictory numbers, and provides confusion matrices could make the dataset usable, but the current version does not support the benchmark claim. I therefore agree with the rejection verdict, with the emphasis shifted from label quality to evaluation protocol.","tokens_in":7307,"tokens_out":5452,"duration_ms":63049,"concrete_test":"Download the public dataset and inspect the released code and file metadata. First, check whether any partition or manifest file exists; the naming convention in Section 3.5 appears not to include split information. Second, using only the 200 mm, 150 mm, and 100 mm filenames, recompute Hydra's accuracy under one explicit train/test split with the released code. If the result does not uniquely reproduce one of the two reported triples (95.52/91.32/82.1 or 95.43/93.38/84.10), the benchmark's evaluation is unreproducible as written.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that Hydra-Bench is a benchmark for multi-modal fusion and SAR algorithm optimization. For that claim to hold, a third party must be able to reproduce the reported numbers under a defined protocol. Section 1 says the dataset is used for training and testing Hydra, and Section 3.6 says the evaluation centers on Hydra, but neither Section 3.4 (file structure) nor Section 3.5 (naming convention) records a train/test split, and Section 4 never states how the 292 samples were partitioned. The filename convention encodes only group, date, distance, and index. Without split metadata, the reported accuracies cannot be distinguished from in-sample fits, and future methods cannot be compared fairly. The internal inconsistencies support this reading: the first scan-distance analysis reports 95.52% ± 1.46% at 200 mm, 91.32% ± 2.65% at 150 mm, and 82.1% ± 4.13% at 100 mm, while the paragraph immediately after reports 95.43% ± 1.47% at 200 mm, 93.38% ± 2.56% at 150 mm, and 84.10% ± 4.53% at 100 mm. The sentence saying accuracy 'improves to 85.41%' immediately after reporting 96% is similarly hard to reconcile. The downloadable SAR code and image data are real artifacts, but the evaluation layer that would convert them into a benchmark is missing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Hydra-Bench, a multi-modal dataset for leaf wetness detection that contains synchronized mmWave raw data, SAR images, and RGB images collected over six months from five plant species, with 268 indoor and 24 outdoor sample pairs. The authors report benchmark results using their own Hydra model, including indoor accuracy of 96%, outdoor accuracy of approximately 90%, single-modality baselines (camera-only 86.13%, SAR-only 83.43%), and comparisons across fusion strategies and scan distances. The stated contribution is that the dataset can serve as a benchmark for future multi-modal fusion and SAR imaging algorithm optimization.","tokens_in":7623,"tokens_out":3538,"duration_ms":40802,"significance":"If the dataset artifact is made available as described, it is a potentially useful resource: synchronized raw mmWave data, SAR images, and RGB images with an accompanying SAR imaging codebase are not common in agricultural sensing, and the six-month, five-species collection adds diversity. However, the evaluation layer that would turn this artifact into a benchmark is currently unreliable. The reported accuracy numbers are internally contradictory, the train/test protocol is absent, and the evaluation is exclusively performed with a model developed by the same authors on the same dataset. These issues do not invalidate the dataset itself, but they do invalidate the quantitative claims as written and prevent third-party comparison against the reported baselines.","major_comments":[{"comment":"The two consecutive scan-distance analyses report different accuracies for the same distances: the first reports 95.52% ± 1.46% at 200 mm, 91.32% ± 2.65% at 150 mm, and 82.1% ± 4.13% at 100 mm, while the next paragraph reports 95.43% ± 1.47% at 200 mm, 93.38% ± 2.56% at 150 mm, and 84.10% ± 4.53% at 100 mm. These contradictory numbers mean the benchmark results cannot be reproduced or compared against future methods, and the paper does not explain which set of numbers is authoritative.","section":"Section 4, scan-distance paragraphs"},{"comment":"No train/test split is defined anywhere in the manuscript. The file structure (Section 3.4) and naming convention (Section 3.5) encode only group, date, distance, and index, and Section 4 never states how the 292 samples are partitioned into training, validation, and test sets. Since Section 1 says the dataset was used for training and testing Hydra, the reported accuracies cannot be distinguished from in-sample fits, and future researchers have no protocol under which to reproduce or compare against the numbers.","section":"Sections 3.4, 3.5, and 4"},{"comment":"The text states that Hydra achieves 96% ± 2.14% accuracy and then says this 'improves to 85.41% ± 2.24% with our data enhancement techniques.' A number lower than the reported baseline cannot be an improvement, and the sentence is irreconcilable with the reported accuracy values. This is not a mere wording issue because the fusion-strategy comparison is one of the paper's central benchmark claims.","section":"Section 4, fusion results"},{"comment":"The outdoor accuracy is reported only as 'approximately 90%' with no sample size, confidence interval, or per-condition breakdown. Given that the entire outdoor subset contains only 24 samples across dynamic conditions (rainy, dawn, and low-light night), the stated outdoor robustness claim is not quantitatively evaluable, and the paper should either provide the full protocol and error bars or explicitly mark the outdoor numbers as preliminary.","section":"Section 4, outdoor results"}],"minor_comments":[{"comment":"The frequency slope is written as 'k = 70.295M Hz/µ'; the units are malformed and should be MHz/µs.","section":"Section 3.1"},{"comment":"The example filename '0 0119 200 1 200.jpg' contains spaces where the naming convention specifies underscores; the example should match the declared format exactly.","section":"Section 3.5"},{"comment":"The dataset URL appears only in the Introduction; a dedicated data-availability statement with access conditions, license, and expected citation format would be more appropriate for a dataset paper.","section":"Section 1"},{"comment":"The paper repeatedly says the dataset is 'precisely calibrated' (Sections 2.3 and 3.1) but does not report any calibration error or drift check across the six-month collection period; a sentence describing the calibration validation would strengthen the resource description.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The dataset artifact itself may be salvageable, but the benchmark layer is under-specified and internally inconsistent. I would suggest asking the authors to provide the exact train/test split, corrected accuracy tables, and a clear statement of which numbers are held-out test results before considering the paper for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: the dataset is real and worth having. Synchronized mmWave raw data, SAR images, and RGB images for 292 samples across five species, six months, indoor and outdoor, with binary wet/dry labels — that's a useful public resource for an under-served sensing problem. The SAR imaging code is also a concrete artifact. Credit where due: the collection effort and the file naming convention are careful enough that a third party could probably reconstruct the samples.\n\nBut the evaluation layer doesn't hold. The same scan distances get different numbers in adjacent paragraphs: 200 mm is both 95.52% ± 1.46% and 95.43% ± 1.47%; 100 mm is both 82.1% ± 4.13% and 84.10% ± 4.53%. Then \"improves to 85.41% ± 2.24%\" right after reporting 96% — that reads as a typo at best, but as published it's contradictory. There's no train/test split anywhere. Section 3.5 gives the filename convention; it encodes group, date, distance, index, and cross-section, but no split indicator. Section 4 never says how the 292 samples were partitioned. So the reported accuracies could be in-sample fits, and future methods have no defined protocol to compare against.\n\nThe circularity is real but moderate: evaluating Hydra on the dataset described as the one used to train Hydra is fine for a \"here's how our model does on our data\" paper, but it's not a benchmark. The outdoor result (\"approximately 90%\") is too vague to be a baseline. The reader's concern about label noise and representativeness of 24 outdoor samples is secondary; the inconsistent numbers are the load-bearing flaw.\n\nThe paper is salvageable. Split the data, report on a fixed split, keep one consistent set of numbers, state the protocol, and make clear that the Hydra results are a reference implementation rather than an independent test. Then the dataset has value.\n\nWho this is for: people building mmWave or multi-modal agricultural sensors. It should get a serious referee — not because the current numbers are trustworthy, but because the artifact is real and the evaluation issues are fixable. I would not cite the benchmark numbers as they stand, but I'd keep an eye on the dataset.\n\nRecommendation: send it to peer review with a request for major revision, and ask the authors to add split metadata and correct the internal inconsistencies.","headline":"A genuinely new dataset for leaf wetness sensing, but the reported benchmark numbers are internally inconsistent and the evaluation protocol is missing, so the paper as written can't serve as a benchmark.","tokens_in":8106,"tokens_out":2681,"would_cite":false,"duration_ms":26494,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Hydra-Bench offers a public multi-modal dataset for leaf wetness detection, with depth-aware radar-plus-camera fusion reaching 96% indoor accuracy.","keywords":["leaf wetness detection","multi-modal dataset","mmWave radar","synthetic aperture radar","RGB imaging","sensor fusion","agricultural IoT","benchmark"],"falsifier":"Take 30 published samples, re-measure each leaf's wetness with an independent method such as pre- and post-wetting weight or a commercial leaf-wetness sensor, and compare against the filename labels; if more than a small handful disagree, the reported accuracy rests on mislabeled ground truth rather than sensor capability.","tokens_in":7024,"feed_emoji":"🌿","tokens_out":8261,"duration_ms":83901,"temperature":0.7,"pith_summary":"Leaf wetness—water sitting on a leaf surface—is a key trigger for plant disease, but existing detectors rely on synthetic leaves or single modalities that degrade in real fields. This paper presents Hydra-Bench, a public dataset of 292 synchronized samples gathered over six months from five plant species, each pairing raw mmWave radar data, a Synthetic Aperture Radar image, and an RGB image under both indoor and outdoor conditions. Using the Hydra model as a benchmark, the authors report 96% accuracy indoors and roughly 90% on a small farm subset, with camera-only and SAR-only baselines around 86% and 83%. The central offer is a shared testbed: because raw radar signals are included, researchers can improve SAR imaging algorithms and fusion strategies without rebuilding the hardware.","feed_headline":"292 radar-camera samples benchmark leaf wetness detection","feed_subtitle":"Five plant species, six months, indoor and outdoor scenes; depth-aware fusion hits 96 percent indoors.","key_machinery":"The load-bearing object is the dataset sample: a synchronized trio of raw mmWave chirp data from a 77–81 GHz FMCW radar, a SAR image formed by a range-migration algorithm with phase compensation and a plane-wave decomposition theorem, and an RGB frame from a camera calibrated to the radar's field of view. A filename convention encodes the dry/wet label, collection date, sensor distance, sample index, and SAR cross-section depth, so every capture is self-describing. The benchmark model is Hydra's two-stage fusion pipeline: a CNN extracts features from depth-aligned SAR slices and RGB imagery, then a transformer with depth-aware positional encoding models relationships across SAR depths for the final wet/dry decision.","core_discovery":"The paper's central claim is that the missing piece for reliable leaf wetness detection is a calibrated, multi-modal public dataset, and that Hydra-Bench supplies it. Each sample is a three-way synchronized capture of raw mmWave reflections, a reconstructed SAR image, and an RGB image, with the two imaging modalities aligned to a common field of view. On this data, the depth-aware fusion model reaches 96% ± 2.14% indoor accuracy and about 90% in outdoor farm settings, outperforming camera-only (86.13% ± 2.13%) and SAR-only (83.43% ± 3.05%) baselines as well as early and late fusion variants. The scan-distance results show accuracy staying above 90% when the SAR scan window is 150 mm or wider and degrading at shorter distances, and the release of raw data turns the dataset into a benchmark for SAR imaging algorithm optimization rather than only a classification benchmark.","pith_inferences":["Beyond the paper: the 24 outdoor samples are too few to certify real-world performance, so the roughly 90% farm accuracy should be read as a pilot result until a larger outdoor split exists.","Beyond the paper: the binary dry/wet labels ignore partially wet leaves and wetness duration, which are what disease models actually track; the naming convention could be extended to continuous or multi-class wetness labels.","Beyond the paper: the raw data release makes it possible to learn SAR image formation end-to-end instead of using the fixed range-migration pipeline, an optimization direction the paper names but does not explore.","Beyond the paper: an obvious stress test is low-light RGB plus SAR fusion, since the paper motivates multi-modality partly by RGB's lighting sensitivity; isolating that subset would show how much SAR contributes when vision fails."],"forward_implications":["Researchers can train and compare leaf wetness classifiers on a common resource without needing their own radar-camera rig.","The reported fusion advantage—96% versus 86% camera-only and 83% SAR-only—gives a concrete target for future multi-modal fusion methods.","The raw mmWave files let imaging researchers benchmark SAR reconstruction algorithms directly, not just downstream classification.","The scan-distance curves quantify the resolution and coverage trade-off of SAR scanning, so future systems can choose scan windows against a known degradation curve.","The small outdoor subset documents what happens when lighting and weather vary, giving a baseline for field-deployment claims."],"supporting_citations":[{"why":"Supplies the Hydra model and the calibration protocol whose dataset this paper benchmarks.","marker":"[1]"},{"why":"Presents mmLeaf, the single-modality mmWave baseline that Hydra is compared against.","marker":"[11]"},{"why":"Specifies the 77–81 GHz radar module that generates the raw mmWave data.","marker":"[16]"},{"why":"Describes the data capture board that records and stores the raw radar signals.","marker":"[17]"},{"why":"Identifies the RGB camera used as the visual modality in each sample.","marker":"[18]"},{"why":"Provides the FMCW SAR signal processing foundation for the imaging pipeline.","marker":"[19]"},{"why":"Supplies the near-field MIMO-SAR imaging formulation used for reconstruction.","marker":"[20]"},{"why":"Provides the wave-representation theorem used to approximate spherical waves during image formation.","marker":"[22]"}],"fun_headline_variants":["Radar, SAR, RGB fused for leaf wetness at 96% indoor","Depth-aware fusion tops leaf wetness sensing benchmarks","Multi-modal leaf wetness dataset beats single-sensor baselines","Synchronized radar-SAR-RGB raises leaf wetness accuracy","Leaf wetness benchmark: radar+Camera fusion hits 96%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's utility rests on the assumption that the 292 dry/wet labels are correct and that radar-to-camera calibration stayed consistent across six months and five plant species; if label noise crept in or calibration drifted, the reported accuracies and the dataset's value as a benchmark both erode.","fun_headline_variants_meta":{"raw":{"variants":["Radar, SAR, RGB fused for leaf wetness at 96% indoor","Depth-aware fusion tops leaf wetness sensing benchmarks","Multi-modal leaf wetness dataset beats single-sensor baselines","Synchronized radar-SAR-RGB raises leaf wetness accuracy","Leaf wetness benchmark: radar+Camera fusion hits 96%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000156,"raw_usage":{"total_tokens":1186,"prompt_tokens":879,"completion_tokens":307,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":218}},"tokens_in":495,"tokens_out":307,"duration_ms":4566,"temperature":1.0,"reasoning_tokens":218,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T11:22:54.031473+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take 30 published samples, re-measure each leaf's wetness with an independent method such as pre- and post-wetting weight or a commercial leaf-wetness sensor, and compare against the filename labels; if more than a small handful disagree, the reported accuracy rests on mislabeled ground truth rather than sensor capability.","supporting_citations":[{"cited_title":"Hydra-Bench: A Benchmark for Multi-Modal Leaf Wetness Sensing","cited_arxiv_id":"2507.22685","evidence_quote":"Supplies the Hydra model and the calibration protocol whose dataset this paper benchmarks."},{"cited_title":"Effects of leaf wetness duration, temperature, and host phenologi- cal stage on infection of walnut by xanthomonas arbori- cola pv. juglandis,","cited_arxiv_id":null,"evidence_quote":"Presents mmLeaf, the single-modality mmWave baseline that Hydra is compared against."},{"cited_title":"Validation of weather and leaf wetness forecasts for a lettuce downy mildew warning system,","cited_arxiv_id":null,"evidence_quote":"Specifies the 77–81 GHz radar module that generates the raw mmWave data."},{"cited_title":"An optical method of dew estimation,","cited_arxiv_id":null,"evidence_quote":"Describes the data capture board that records and stores the raw radar signals."},{"cited_title":"Poster: mmleaf: Versatile leaf wetness detection via mmwave sensing,","cited_arxiv_id":null,"evidence_quote":"Identifies the RGB camera used as the visual modality in each sample."},{"cited_title":"PHYTOS 31 Manual Web,","cited_arxiv_id":null,"evidence_quote":"Provides the FMCW SAR signal processing foundation for the imaging pipeline."},{"cited_title":"A bio-mimetic leaf wetness sensor from replica molding of leaves,","cited_arxiv_id":null,"evidence_quote":"Supplies the near-field MIMO-SAR imaging formulation used for reconstruction."},{"cited_title":"Proteus: : Enhanced mmwave leaf wetness detection with cross-modality knowledge transfer,","cited_arxiv_id":null,"evidence_quote":"Provides the wave-representation theorem used to approximate spherical waves during image formation."}],"review_version":1}