{"id":"9a30b563-fe5f-4777-9b6f-fbb682a7bc67","arxiv_id":"2607.25310","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"On one UAV hyperspectral scene, the ACE detector with an operator-in-the-loop signature bootstrap found all seven PFM-1 targets in nine inspections; SAM variants needed 2,897–4,558.","lead":"This study tests four classical algorithms for spotting PFM-1 landmines in drone hyperspectral imagery, comparing ground-measured, in-scene, and simulated operator-updated target signatures. The most efficient algorithm found all seven mine locations after nine candidate inspections, while two spectral-angle variants required thousands.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Bootstrap's signature update (Eqs. 5–6) uses ground-truth pixel masks, not human feedback; 'reaching in-scene performance' is a label-in-the-loop artifact, and Table 2 counts depend on this unrealistic oracle.","rationale":"The paper is a clean retrospective comparison with a clear protocol and code, and I do not see a fatal internal inconsistency. The reader's CONDITIONAL verdict is appropriate. My concern deepens the reader's weakest assumption: it is not just that v(q) may be optimistic, but that the subsequent signature update uses ground-truth masks rather than operator output, so the 'human-in-the-loop' name overstates what is demonstrated. This is load-bearing because the central quantitative results—especially the small review counts for ACE—depend on a perfect mask oracle. The proposed test distinguishes whether the protocol works with operator-level information or only with labels. Until that test is run (or the claims are explicitly scoped to a label-assisted retrospective), the paper should remain conditional, with the conclusions tempered and the by-construction equivalence flagged.","tokens_in":6296,"tokens_out":7250,"duration_ms":79494,"concrete_test":"Re-run the bootstrap with P(T) in Eq. (5) replaced by the union of the confirmed inspection neighborhoods I_eta(q) (or the top-scoring pixels within those neighborhoods), keeping v(q) unchanged. Compare the final maps and Table 2 milestones (ACE 9, CEM 22, MF 38, SAM 2897/4558). If these counts and the final AP values are essentially unchanged, the exact-mask oracle is not load-bearing; if they diverge, the operational conclusions are tied to the unrealistic label-in-the-loop assumption. A secondary variant with v(q) corrupted by 5–10% false positives/negatives would test sensitivity to the human model, but the mask-oracle replacement is the decisive first check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The weakest load-bearing point is not simply the verifier v(q); it is the signature update. Eq. (5) sets the next signature to the mean of P(T_a^(k)), and Eq. (6) constructs P(T) from the ground-truth positive mask R_r for each confirmed region, taking the central 50% of those labeled pixels. So as soon as v(q) returns r, the algorithm is given the exact per-pixel mine mask and averages those pixels. A human operator in the field provides a categorical confirmation at a candidate location, not a segmentation of the mine. Thus the loop is not merely 'human confirmation' as Sec. 4 states; it is a label-in-the-loop oracle for the target spectrum. The abstract's statement that full-review bootstrapping reaches the fully informed in-scene case after all seven regions are verified is true by construction (once T contains all regions, d is the mean of the same core pixels as the core in-scene signature), so it cannot be read as empirical support for the method. More importantly, Table 2's discovery counts (ACE 9, CEM 22, MF 38, SAM 2897/4558) are generated under this oracle: the detector that first places a candidate near a target is rewarded with clean target pixels. If a real system must infer which pixels inside the inspection neighborhood are mine pixels—e.g., by taking the neighborhood or high-scoring subregions—the adaptation trajectory and the final maps may differ. The single scene and manual merging of nine mask fragments into seven regions (Sec. 2.3) make this concern harder to dismiss, since the oracle is especially favorable when target regions are small and well separated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a single-scene, retrospective case study of PFM-1 landmine detection from a UAV VNIR hyperspectral image. It compares four classical detectors (SAM, MF, ACE, CEM) under three target signatures: an external SVC ground spectrum, a fully informed in-scene core-pixel spectrum, and a simulated 'human-in-the-loop' bootstrap that updates the target spectrum after operator confirmation of detector-proposed candidates. Performance is evaluated with ROC-AUC, AP, cumulative target-discovery curves, and spatial candidate-review counts. The headline results are that the full-review bootstrap attains pixel-level metrics identical to the in-scene core signature, and that ACE requires only 9 candidate reviews to confirm all seven target regions, while the SAM variants require thousands.","tokens_in":6705,"tokens_out":6629,"duration_ms":61451,"significance":"The paper's contribution is the operator-facing evaluation methodology (target-discovery curves and candidate-review counts) and an open implementation. These metrics are indeed more relevant to demining than pixel-level AUC alone, and the authors deserve credit for making the code available and for clearly stating the retrospective, oracle-assisted nature of their setup. However, the equivalence between bootstrap and in-scene signatures is an algebraic identity, not an empirical result, and the confirmation oracle supplies the exact target spectrum. The operational conclusions are therefore not yet established. With a reframing of the bootstrap as an oracle upper bound and a sensitivity analysis, the case study could be a useful benchmark for the community.","major_comments":[{"comment":"The equality between the full-review bootstrap and the fully informed in-scene case is true by construction. The core in-scene signature is the mean of central target pixels of all seven regions (Sec. 2.4). Once T_a^(k) contains all seven regions, Eq. (5) sets d to the mean of P(T_a^(k)), and Eq. (6) defines P(T) as the central 50% of the ground-truth masks of those regions — exactly the core in-scene signature. Therefore the identical ROC-AUC/AP values in Table 1 are forced and cannot be cited as evidence that bootstrapping recovers in-scene performance. The abstract's sentence 'Full-review bootstrapping reaches the fully informed in-scene signature case...' is a statement about the protocol, not a result. Please rephrase as a consistency check and move the emphasis to the discovery counts.","section":"Sec. 2.4, Eqs. (5)–(6); Table 1; Abstract"},{"comment":"The simulated human-in-the-loop uses a much stronger oracle than a human operator. v(q) in Eq. (2) returns a target-region label based on the ground-truth mask, and the signature update in Eqs. (5)–(6) averages the central 50% of the ground-truth positive pixels of each confirmed region. An operator who verifies a candidate location provides a binary decision, not a pixel-level segmentation of the mine. Thus the bootstrap is effectively a ground-truth label-in-the-loop procedure that hands the algorithm the target spectrum after each confirmation. The candidate-review counts in Table 2 are therefore optimistic and may not transfer to field use. The authors should either use an update rule based only on the operator-available information (e.g., mean of the inspection neighborhood) or explicitly characterize the results as an oracle bound.","section":"Sec. 2.4, Eqs. (2), (6), Sec. 4"},{"comment":"The quantitative conclusions rest on a single scene with seven target regions and on several fixed parameters (N=6, rho=25, eta=12, f=0.5). No sensitivity analysis is reported. The confirmation radius eta=12 px corresponds to approximately 0.15 m at the stated GSD, which is larger than the 0.12 m length of a PFM-1; this lenient criterion likely affects the early discovery of ACE and CEM. Because the discovery counts are the paper's main operational output, the authors should report how Table 2 changes under plausible variations of eta, N, and f, or explicitly limit the claims to this scene and parameter setting.","section":"Sec. 2.3–2.4, Table 2"}],"minor_comments":[{"comment":"The title and abstract contain an extraneous space: 'SIGNA TURE' should be 'SIGNATURE'.","section":"Title/Abstract"},{"comment":"The SAM centered row is garbled: '1: 3+5; 20: 2; 21: 1; 23: 7; 450: 4; 760: 613 23 136 760 4558' — '613' is likely '6', and the milestone columns are misaligned. Please reformat.","section":"Table 2"},{"comment":"The combinatorial arg-min definition of the core set is impractical and does not reflect how such a set would be computed. State the equivalent implementation (sort pixels by distance to centroid and take the closest ceil(f|R_r|) pixels).","section":"Eq. (6)"},{"comment":"The mean-centered SAM variant is mentioned but never defined. Provide its formula or a precise reference so the reader can reproduce it.","section":"Sec. 2.2"},{"comment":"The x-axis label 'False alarms before target-region discovery + 1' and the tick positions (0, 9, 99, ...) are confusing; clarify whether the values are pixel-based false alarms or candidate reviews.","section":"Fig. 3"},{"comment":"Because the full-review bootstrap signature equals the core in-scene signature once all regions are confirmed, the bootstrap maps in Figs. 2–4 are identical to the core in-scene maps. This should be stated explicitly to avoid misleading the reader.","section":"Figs. 2–4"}],"recommendation":"major_revision","confidential_remarks":"The paper is an honest, well-structured case study, but the headline equivalence between bootstrap and in-scene signatures is circular. The strongest contribution is the operator-facing evaluation protocol. If the authors reframe the bootstrap as an oracle upper bound and add a sensitivity analysis, the paper could be acceptable; in its current form, the central claim overstates the evidence. Please also verify that the dataset reference [4] and the code repository are publicly accessible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the headline equivalence is true by construction, not by experiment, but the actual deliverable—per-detector discovery counts on a real PFM-1 scene—is solid and useful.\n\nThe paper does its scope-setting well: it claims no new detector or theory, and instead reports target-discovery curves and candidate-review counts for SAM, MF, ACE, and CEM under three signature sources. ACE confirming all seven regions in nine candidate inspections, versus 2,897 and 4,558 for the SAM variants, is the kind of operator-facing number that is rare in this literature. The protocol in Eqs. (1)–(6) is concrete enough to reproduce, the code is promised, and using AP rather than ROC-AUC for a 0.0042% target fraction is the right call.\n\nThe soft spots are real, but not fatal. The abstract's statement that full-review bootstrapping reaches the fully informed in-scene case is an identity: Eq. (5) sets the bootstrap signature to the mean of the same core pixels that define the in-scene reference, once all regions are confirmed. That claim should be explicitly flagged as a consistency check, not empirical support. The larger issue, which the stress-test note gets right, is that the simulated operator in Eq. (2) supplies the exact per-pixel mine mask for each confirmed region, and Eq. (6) then averages the central half of those labeled pixels. A human operator in the field says 'yes, mine' at a candidate location; they do not provide a segmentation. So Table 2's counts are generated under a label-in-the-loop oracle, which is a stronger assumption than 'human confirmation' as the text sometimes suggests. The single scene, the manual merging of nine mask fragments into seven regions, and the heuristic radii all narrow how much carries over to a real deployment. To its credit, the paper does call the experiment retrospective and simulated, but the conclusions lean on the numbers more than that caveat allows.\n\nThe fix is mostly framing: state plainly that the bootstrap/in-scene equality is an identity, temper the operational claim, and consider adding a no-update baseline so the value of verification is actually measured. For what this paper is—a controlled case study on one scene with a transparent oracle—the discovery-curve comparison is a genuinely useful template for demining HSI evaluation. It deserves a serious referee, not a desk reject.","headline":"The headline equivalence is true by construction, but the per-detector discovery counts are new, reproducible, and worth publishing as a case study.","tokens_in":7187,"tokens_out":2919,"would_cite":true,"duration_ms":29291,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A human-in-the-loop signature bootstrap, starting from a ground-measured spectrum and refining it with operator-verified in-scene target pixels, lets classical hyperspectral detectors match fully informed in-scene performance, but the inspe","keywords":["hyperspectral imaging","PFM-1 landmine detection","UAV VNIR","human-in-the-loop","target signature bootstrap","target discovery curve","adaptive coherence estimator","candidate review count"],"falsifier":"Run the same seven-region experiment with a real human operator reviewing the candidate lists (or an emulator with realistic false-positive and false-negative rates) and compare the confirmed target sets and review counts to those produced by the ground-truth-based v(q); if the operator misses targets or confirms false alarms, the review counts and the bootstrap's convergence to in-scene performance would not transfer.","tokens_in":6213,"feed_emoji":"🚁","tokens_out":5892,"duration_ms":50995,"temperature":0.7,"pith_summary":"The paper establishes that in hyperspectral PFM-1 mine screening, the target signature need not be known a priori: a human-in-the-loop bootstrap that starts from an external spectroradiometer signature and is updated with operator-verified in-scene target pixels converges to the fully informed in-scene core-pixel signature, matching its detection performance. The decisive new result is the inspection effort required, which varies by more than two orders of magnitude across standard detectors—ACE confirms all seven target regions in two rounds and nine candidate reviews, CEM in 22, MF in 38, and the SAM variants in thousands. Because real demining is bottlenecked by the false alarms a human must inspect, the paper argues that target-discovery curves and candidate-review counts are more operationally informative than ROC-AUC, which stays high even when many false alarms precede target hits. The finding matters because it quantifies, for one realistic scene, how much operator effort each classical detector actually demands.","feed_headline":"ACE finds all PFM-1 mine targets in 9 reviews; SAM needs 4,558","feed_subtitle":"Why inspection effort, not ROC-AUC, decides whether hyperspectral mine detection works in the field.","key_machinery":"The central mechanism is the bootstrap signature-update loop defined in Eqs. (1)–(6). Each round, non-maximum suppression proposes the top N equal to six spatially distinct unreviewed candidate locations; a simulated operator v(q) confirms a candidate if its 12-pixel inspection radius intersects a ground-truth target region; the reviewed regions are blocked; and if at least one new target is confirmed, the next-round signature is the mean of the central 50% of pixels (selected by distance to each region's centroid) from all confirmed regions. This loop turns the external SVC spectrum gradually into an in-scene core-pixel signature, with the detector's ranking behavior determining how many re","core_discovery":"On the paper's own terms, the central claim is that a bootstrap signature updated from verified target pixels recovers the advantages of an in-scene signature, and the question is how quickly. Using Eqs. (1)–(6), the protocol proposes six spatially distinct candidates per round, simulates operator confirmation with ground-truth overlap, blocks reviewed neighborhoods, and recomputes the target signature as the mean of the central 50% of pixels from all confirmed target regions. Across five detectors—SAM (centered and uncentered), MF, ACE, and CEM—the bootstrap converges to the fully informed core in-scene case once all seven regions are verified, meaning the final score maps and pixel-level m","pith_inferences":["The paper leaves implicit that the bootstrap's convergence to the in-scene reference is guaranteed by construction once all regions are confirmed, since Eq. (5) defines the bootstrap signature as the mean of the same core pixels; the genuinely informative quantity is the convergence rate, which the review counts capture.","A natural stress test would replace the perfect ground-truth emulator v(q) with a noisy operator model that occasionally confirms false alarms or misses targets; this could change the relative ranking of detectors, especially those that verify early and therefore shape the signature from a small number of regions.","The inspection radius eta = 12 px, set to about the PFM-1 footprint, implies a trade-off worth testing: a smaller radius would reduce blocked area and allow more candidates per round, but could fragment target verification; a larger radius would suppress false alarms but might merge nearby false positives with true targets.","The seven target regions were obtained by manually merging nine connected-components; a different merging choice would alter the discovery curves and review counts, so the exact numbers in Table 2 are scene- and preprocessing-specific."],"forward_implications":["Demining-oriented HSI studies should report target-discovery curves, candidate-review counts, and false-alarm-before-detection values alongside pixel-level metrics, because inspection burden is often the operational bottleneck.","A detector like ACE, which confirms all seven targets in two rounds and nine reviews, could be deployed with a very small review budget; SAM variants cannot, because their last target arrives only after thousands of reviews.","Bootstrapping that stops after a small candidate batch may reject detectors whose first confirmed target appears deeper in the spatial ranking, so practical systems should set a patience rule or inspection budget rather than an early stopping rule.","Because the final bootstrap signature equals the core in-scene signature as soon as all seven regions are verified, the paper's protocol is a concrete way to obtain an in-scene signature without any a priori knowledge of target locations—at a detector-dependent cost."],"fun_headline_variants":["ACE finds PFM-1 mines in 9 reviews, SAM needs 4,558","Bootstrap signature matches full info after 7 verifications","For mine detection, inspection count beats ROC-AUC","Hyperspectral screening: ACE trims false alarms to 9","SAM vs ACE on landmines: 4,558 reviews vs 9"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the simulated operator—who confirms any candidate whose 12-pixel radius touches a ground-truth target region and never confirms a false alarm—matches how a real human would behave, and that the central half of each confirmed target region is a clean PFM-1 spectrum.","fun_headline_variants_meta":{"raw":{"variants":["ACE finds PFM-1 mines in 9 reviews, SAM needs 4,558","Bootstrap signature matches full info after 7 verifications","For mine detection, inspection count beats ROC-AUC","Hyperspectral screening: ACE trims false alarms to 9","SAM vs ACE on landmines: 4,558 reviews vs 9"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000159,"raw_usage":{"total_tokens":1067,"prompt_tokens":745,"completion_tokens":322,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":228}},"tokens_in":489,"tokens_out":322,"duration_ms":3957,"temperature":1.0,"reasoning_tokens":228,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T02:47:28.488059+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same seven-region experiment with a real human operator reviewing the candidate lists (or an emulator with realistic false-positive and false-negative rates) and compare the confirmed target sets and review counts to those produced by the ground-truth-based v(q); if the operator misses targets or confirms false alarms, the review counts and the bootstrap's convergence to in-scene performance would not transfer.","supporting_citations":[],"review_version":1}