REVIEW 3 major objections 6 minor 1 cited by
Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection
T0 review · 3 major / 6 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read A human-in-the-loop signature bootstrap, starting from a ground-measured spectrum and refining it with operator-verified in-scene target pixels, lets classical hyperspectral detectors match fully informed in-scene performance, but the inspe
desk verdict The headline equivalence is true by construction, but the per-detector discovery counts are new, reproducible, and worth publishing as a case study. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the bootstrap signature-update loop defined in Eqs. (1)–(6). Each round, non-maximum suppression proposes the top N equal to six spatially distinct unreviewed candidate locations; a simulated operator v(q) confirms a candidate if its 12-pixel inspection radius intersects a ground-truth target region; the reviewed regions are blocked; and if at least one new target is confirmed, the next-round signature is the mean of the central 50% of pixels (selected by distance to each region's centroid) from all confirmed regions. This loop turns the external SVC spectrum gradually into an in-scene core-pixel signature, with the detector's ranking behavior determining how many re
What would settle it
Run the same seven-region experiment with a real human operator reviewing the candidate lists (or an emulator with realistic false-positive and false-negative rates) and compare the confirmed target sets and review counts to those produced by the ground-truth-based v(q); if the operator misses targets or confirms false alarms, the review counts and the bootstrap's convergence to in-scene performance would not transfer.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that a bootstrap signature updated from verified target pixels recovers the advantages of an in-scene signature, and the question is how quickly. Using Eqs. (1)–(6), the protocol proposes six spatially distinct candidates per round, simulates operator confirmation with ground-truth overlap, blocks reviewed neighborhoods, and recomputes the target signature as the mean of the central 50% of pixels from all confirmed target regions. Across five detectors—SAM (centered and uncentered), MF, ACE, and CEM—the bootstrap converges to the fully informed core in-scene case once all seven regions are verified, meaning the final score maps and pixel-level m
Load-bearing premise
The load-bearing premise is that the simulated operator—who confirms any candidate whose 12-pixel radius touches a ground-truth target region and never confirms a false alarm—matches how a real human would behave, and that the central half of each confirmed target region is a clean PFM-1 spectrum.
Editorial extensions
If this is right
- Demining-oriented HSI studies should report target-discovery curves, candidate-review counts, and false-alarm-before-detection values alongside pixel-level metrics, because inspection burden is often the operational bottleneck.
- A detector like ACE, which confirms all seven targets in two rounds and nine reviews, could be deployed with a very small review budget; SAM variants cannot, because their last target arrives only after thousands of reviews.
- Bootstrapping that stops after a small candidate batch may reject detectors whose first confirmed target appears deeper in the spatial ranking, so practical systems should set a patience rule or inspection budget rather than an early stopping rule.
- Because the final bootstrap signature equals the core in-scene signature as soon as all seven regions are verified, the paper's protocol is a concrete way to obtain an in-scene signature without any a priori knowledge of target locations—at a detector-dependent cost.
Reading between the lines
- The paper leaves implicit that the bootstrap's convergence to the in-scene reference is guaranteed by construction once all regions are confirmed, since Eq. (5) defines the bootstrap signature as the mean of the same core pixels; the genuinely informative quantity is the convergence rate, which the review counts capture.
- A natural stress test would replace the perfect ground-truth emulator v(q) with a noisy operator model that occasionally confirms false alarms or misses targets; this could change the relative ranking of detectors, especially those that verify early and therefore shape the signature from a small number of regions.
- The inspection radius eta = 12 px, set to about the PFM-1 footprint, implies a trade-off worth testing: a smaller radius would reduce blocked area and allow more candidates per round, but could fragment target verification; a larger radius would suppress false alarms but might merge nearby false positives with true targets.
- The seven target regions were obtained by manually merging nine connected-components; a different merging choice would alter the discovery curves and review counts, so the exact numbers in Table 2 are scene- and preprocessing-specific.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a single-scene, retrospective case study of PFM-1 landmine detection from a UAV VNIR hyperspectral image. It compares four classical detectors (SAM, MF, ACE, CEM) under three target signatures: an external SVC ground spectrum, a fully informed in-scene core-pixel spectrum, and a simulated 'human-in-the-loop' bootstrap that updates the target spectrum after operator confirmation of detector-proposed candidates. Performance is evaluated with ROC-AUC, AP, cumulative target-discovery curves, and spatial candidate-review counts. The headline results are that the full-review bootstrap attains pixel-level metrics identical to the in-scene core signature, and that ACE requires only 9 candidate reviews to confirm all seven target regions, while the SAM variants require thousands.
Significance. The paper's contribution is the operator-facing evaluation methodology (target-discovery curves and candidate-review counts) and an open implementation. These metrics are indeed more relevant to demining than pixel-level AUC alone, and the authors deserve credit for making the code available and for clearly stating the retrospective, oracle-assisted nature of their setup. However, the equivalence between bootstrap and in-scene signatures is an algebraic identity, not an empirical result, and the confirmation oracle supplies the exact target spectrum. The operational conclusions are therefore not yet established. With a reframing of the bootstrap as an oracle upper bound and a sensitivity analysis, the case study could be a useful benchmark for the community.
major comments (3)
- [Sec. 2.4, Eqs. (5)–(6); Table 1; Abstract] The equality between the full-review bootstrap and the fully informed in-scene case is true by construction. The core in-scene signature is the mean of central target pixels of all seven regions (Sec. 2.4). Once T_a^(k) contains all seven regions, Eq. (5) sets d to the mean of P(T_a^(k)), and Eq. (6) defines P(T) as the central 50% of the ground-truth masks of those regions — exactly the core in-scene signature. Therefore the identical ROC-AUC/AP values in Table 1 are forced and cannot be cited as evidence that bootstrapping recovers in-scene performance. The abstract's sentence 'Full-review bootstrapping reaches the fully informed in-scene signature case...' is a statement about the protocol, not a result. Please rephrase as a consistency check and move the emphasis to the discovery counts.
- [Sec. 2.4, Eqs. (2), (6), Sec. 4] The simulated human-in-the-loop uses a much stronger oracle than a human operator. v(q) in Eq. (2) returns a target-region label based on the ground-truth mask, and the signature update in Eqs. (5)–(6) averages the central 50% of the ground-truth positive pixels of each confirmed region. An operator who verifies a candidate location provides a binary decision, not a pixel-level segmentation of the mine. Thus the bootstrap is effectively a ground-truth label-in-the-loop procedure that hands the algorithm the target spectrum after each confirmation. The candidate-review counts in Table 2 are therefore optimistic and may not transfer to field use. The authors should either use an update rule based only on the operator-available information (e.g., mean of the inspection neighborhood) or explicitly characterize the results as an oracle bound.
- [Sec. 2.3–2.4, Table 2] The quantitative conclusions rest on a single scene with seven target regions and on several fixed parameters (N=6, rho=25, eta=12, f=0.5). No sensitivity analysis is reported. The confirmation radius eta=12 px corresponds to approximately 0.15 m at the stated GSD, which is larger than the 0.12 m length of a PFM-1; this lenient criterion likely affects the early discovery of ACE and CEM. Because the discovery counts are the paper's main operational output, the authors should report how Table 2 changes under plausible variations of eta, N, and f, or explicitly limit the claims to this scene and parameter setting.
minor comments (6)
- [Title/Abstract] The title and abstract contain an extraneous space: 'SIGNA TURE' should be 'SIGNATURE'.
- [Table 2] The SAM centered row is garbled: '1: 3+5; 20: 2; 21: 1; 23: 7; 450: 4; 760: 613 23 136 760 4558' — '613' is likely '6', and the milestone columns are misaligned. Please reformat.
- [Eq. (6)] The combinatorial arg-min definition of the core set is impractical and does not reflect how such a set would be computed. State the equivalent implementation (sort pixels by distance to centroid and take the closest ceil(f|R_r|) pixels).
- [Sec. 2.2] The mean-centered SAM variant is mentioned but never defined. Provide its formula or a precise reference so the reader can reproduce it.
- [Fig. 3] The x-axis label 'False alarms before target-region discovery + 1' and the tick positions (0, 9, 99, ...) are confusing; clarify whether the values are pixel-based false alarms or candidate reviews.
- [Figs. 2–4] Because the full-review bootstrap signature equals the core in-scene signature once all regions are confirmed, the bootstrap maps in Figs. 2–4 are identical to the core in-scene maps. This should be stated explicitly to avoid misleading the reader.
Circularity Check
Full-review bootstrap's convergence to in-scene performance is guaranteed by the definition of its update rule, not demonstrated.
-
self definitional
[Sec. 2.4, Eqs. (5)-(6); Sec. 2.4 core in-scene definition; Sec. 3 Table 1]
"If at least one target region has been confirmed, the next signature is the mean of the core pixels from all confirmed regions: d(k+1)a = ... (5). ... For each target region r, let Rr be its positive mask pixels ... P(T) = ... with f = 0.5. Thus, the central 50% of each confirmed target region is used for signature averaging. ... The core in-scene case uses the mean spectrum of central target pixels from all seven target regions."
After all seven regions are in T, Eq. (6) selects the central 50% of the ground-truth positive pixels of every region and Eq. (5) averages those pixels, which is exactly the definition of the core in-scene signature. Therefore the final bootstrap signature is identical to the core in-scene signature before any detector score map is computed. The abstract's claim that 'Full-review bootstrapping reaches the fully informed in-scene signature case' and Table 1's equal ROC-AUC/AP values are identities, not empirical evidence. Additionally, the update is fed the ground-truth mask Rr, not the categorical human confirmation v(q), so the equality is not a test of operator feedback.
-
self definitional
[Sec. 3, Results (first paragraph)]
"The full-review bootstrap matches the core in-scene case because its final signature is built after all seven regions have been verified."
This sentence gives the causal explanation, and the cause is the construction in Eqs. (5)-(6): the final signature is, by definition, the mean of the central ground-truth pixels of all seven regions. Framing this definitional identity as a 'match' presents a tautology as an empirical result.
full rationale
The central circular step is the headline convergence: Eq. (5) sets the bootstrap signature to the mean of the central 50% of each confirmed region's ground-truth positive pixels, while the core in-scene signature is defined as the mean of the same central pixels from all seven regions. Once all seven regions are confirmed, the two signatures are the same mathematical object, so Table 1's identical metrics and the abstract's 'reaches the fully informed in-scene signature case' are true by construction. The paper itself acknowledges this in Sec. 3. The detector-specific review counts (ACE 9, CEM 22, MF 38, SAM ~3k-4.5k) are empirical properties of the score maps and are not circular, but they are produced under a ground-truth oracle v(q) that grants the exact mine mask Rr for the next signature; this is an optimism/validity concern rather than an equation-level circularity. No load-bearing self-citation or imported-uniqueness pattern is present. Overall, the central 'reaches in-scene' claim reduces to a definition, while the comparative detector ranking retains independent content; hence partial circularity.
Assumptions & free parameters
free parameters (5)
- Candidate batch size N =
6
- NMS suppression radius rho =
25 px (~0.32 m)
- Inspection radius eta =
12 px (~0.16 m)
- Core fraction f =
0.5
- Review round cap =
2001
assumptions (4)
- standard math Detector formulations (SAM, MF, ACE, CEM) as in refs [6,7] are accepted without derivation.
- domain assumption The nine connected components are merged into seven target regions by the authors' manual judgment.
- domain assumption Human verification is perfectly emulated by ground truth via v(q) in Eq. 2 (intersection with target region within radius eta).
- domain assumption A single UAV scene with 248 labeled target pixels is representative enough to compare detectors operationally.
Cite this review
Pith. "Pith review of Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection." pith.science (2026). https://pith.science/paper/UQO3ZVUQ
@misc{pith2026260725310,
author = {Pith},
title = {Pith review of: Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/UQO3ZVUQ}},
note = {Machine review of arXiv:2607.25310}
}
read the original abstract
Hyperspectral imaging (HSI) is useful for material discrimination, but operational mine screening also depends on how many false alarms must be inspected before targets are found. This paper studies PFM-1 landmine detection in unmanned aerial vehicle (UAV) visible and near-infrared (VNIR) HSI using spectral angle mapper (SAM), matched filter (MF), adaptive coherence estimator (ACE), and constrained energy minimization (CEM). We compare a ground-measured SVC signature, a fully informed in-scene core-pixel signature, and a simulated human-in-the-loop signature bootstrap. Besides receiver operating characteristic area under the curve and average precision, we report target-discovery curves and spatial candidate-review counts. Full-review bootstrapping reaches the fully informed in-scene signature case after all seven target regions are verified, but the required inspection effort varies strongly: ACE confirms all regions in two rounds and nine candidate inspections, whereas the SAM variants need thousands of candidate reviews for their final target locations. Code is available at https://github.com/SagarLekhak/IEEE_WHISPERS_2026_UAV_HSI_PFM1.
Forward citations
Cited by 1 Pith paper
-
SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift
SULAND v2, a manually re-annotated landmine dataset, raises YOLOv8 IID mAP@50 by 14.6–19.6 points and OOD mAP@50 by ~25 points after fixing an inverted class-label convention.
Reference graph
Works this paper leans on
-
[1]
UA Vs make this sensing mode more practical for close-range surveys while reducing direct personnel exposure in contaminated areas [2, 3]
INTRODUCTION HSI has been investigated for stand-off landmine detection because calibrated wavelength-resolved spectra can help sep- arate surface materials [1]. UA Vs make this sensing mode more practical for close-range surveys while reducing direct personnel exposure in contaminated areas [2, 3]. Recent UA V VNIR datasets now support controlled studies...
-
[2]
DA TA AND METHODS 2.1. UA V hyperspectral scene The experiment uses a low-altitude PFM-1 subset from a UA V VNIR benchmark dataset [4, 5]. The analyzed image contains 1705×3461pixels and 272 retained bands spanning approx- imately 400–1000 nm. A binary pixel mask identifies visible PFM-1 target support. The mask contains 248 positive pix- els; all remaini...
arXiv 2026
-
[3]
round: newly confirmed region IDs
RESULTS Table 1 summarizes the pixel-level results. ROC-AUC is high for MF, ACE, and CEM in all signature cases, but AP gives a 0.0 0.2 0.4 0.6 0.8 1.0 Recall 0.0 0.2 0.4 0.6 0.8 1.0 Precision SAM centered AP=0.080 SAM uncentered AP=0.010 MF AP=0.155 ACE AP=0.314 CEM AP=0.255 Fig. 2. Precision-recall curves for the full-review bootstrap score maps. 100 10...
-
[4]
ROC-AUC can remain strong even when high-ranking false positives precede some target regions
DISCUSSION The main implication is that HSI mine screening should be evaluated with operator-facing metrics, not only pixel- wise separability. ROC-AUC can remain strong even when high-ranking false positives precede some target regions. Target-discovery curves and candidate-review counts directly address how much non-target area must be inspected before ...
-
[5]
Classical detectors remain useful base- lines, but their operational value depends on signature quality and early false-alarm behavior
CONCLUSION This case study evaluated SAM, MF, ACE, and CEM on UA V hyperspectral imagery of PFM-1 targets under external, fully informed in-scene, and simulated human-in-the-loop signature settings. Classical detectors remain useful base- lines, but their operational value depends on signature quality and early false-alarm behavior. ACE was most efficient...
-
[6]
A survey of landmine detection using hyper- spectral imaging,
I. Makki, R. Younes, C. Francis, T. Bianchi, and M. Zuc- chetti, “A survey of landmine detection using hyper- spectral imaging,”ISPRS Journal of Photogrammetry and Remote Sensing, vol. 124, pp. 40–53, 2017
2017
-
[7]
Developing a hy- perspectral non-technical survey for minefields via uav and helicopter,
M. Bajic, T. Ivelja, and A. Brook, “Developing a hy- perspectral non-technical survey for minefields via uav and helicopter,”The Journal of Conventional Weapons Destruction, vol. 21, no. 1, pp. 49–63, 2017
2017
-
[8]
How to implement drones and machine learning to reduce time, costs, and dangers associated with landmine detection,
J. Baur, G. Steinberg, A. Nikulin, K. Chiu, and T. de Smet, “How to implement drones and machine learning to reduce time, costs, and dangers associated with landmine detection,”The Journal of Conventional Weapons Destruction, vol. 25, no. 1, 2021, Article 29
2021
Show all 19 references
-
[9]
A uav-based vnir hyperspectral benchmark dataset for landmine and uxo detection,
S. Lekhak, E. J. Ientilucci, J. Baur, and S. Ghosh, “A uav-based vnir hyperspectral benchmark dataset for landmine and uxo detection,”arXiv preprint arXiv:2510.02700, 2025
2025
-
[10]
Benchmarking deep learning and statistical target detection methods for pfm-1 landmine detec- tion in uav hyperspectral imagery,
S. Lekhak, P. R. Pulakurthi, R. Bhatta, and E. J. Ien- tilucci, “Benchmarking deep learning and statistical target detection methods for pfm-1 landmine detec- tion in uav hyperspectral imagery,”arXiv preprint arXiv:2602.10434, 2026
2026 arXiv
-
[11]
Detection algorithms for hyperspectral imaging applications,
D. Manolakis and G. Shaw, “Detection algorithms for hyperspectral imaging applications,”IEEE Signal Pro- cessing Magazine, vol. 19, no. 1, pp. 29–43, 2002
2002
-
[12]
Hyperspectral target detection: An overview of current and future challenges,
N. M. Nasrabadi, “Hyperspectral target detection: An overview of current and future challenges,”IEEE Signal Processing Magazine, vol. 31, no. 1, pp. 34–44, 2014
2014
-
[13]
Target detection in hyperspectral imagery using forward modeling and in-scene information,
M. Axelsson, O. Friman, T. V . Haavardsholm, and I. Renhorn, “Target detection in hyperspectral imagery using forward modeling and in-scene information,”IS- PRS Journal of Photogrammetry and Remote Sensing, vol. 119, pp. 124–134, 2016
2016
-
[14]
An automatic robust iteratively reweighted unstructured detector for hyper- spectral imagery,
T. Wang, B. Du, and L. Zhang, “An automatic robust iteratively reweighted unstructured detector for hyper- spectral imagery,”IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 7, no. 6, pp. 2367–2382, 2014
2014
-
[15]
Towards robust hyperspectral target detection via test-time spectrum adaptation,
R. Gerster and P. St ¨utz, “Towards robust hyperspectral target detection via test-time spectrum adaptation,”Re- mote Sensing, vol. 17, no. 16, pp. 2756, 2025
2025
-
[16]
Advanced hyperspectral detection based on elliptically contoured distribution models and opera- tor feedback,
A. Schaum, “Advanced hyperspectral detection based on elliptically contoured distribution models and opera- tor feedback,” inProceedings of the IEEE Applied Im- agery Pattern Recognition Workshop, 2009, pp. 1–5
2009
-
[17]
The relationship between precision-recall and roc curves,
J. Davis and M. Goadrich, “The relationship between precision-recall and roc curves,” inProceedings of the 23rd International Conference on Machine Learning, 2006, pp. 233–240
2006
-
[18]
The precision-recall plot is more informative than the roc plot when evaluating binary classifiers on imbalanced datasets,
T. Saito and M. Rehmsmeier, “The precision-recall plot is more informative than the roc plot when evaluating binary classifiers on imbalanced datasets,”PLoS ONE, vol. 10, no. 3, pp. e0118432, 2015
2015
-
[19]
Explosive ordnance guide for ukraine, third edi- tion,
Geneva International Centre for Humanitarian Demi- ning, “Explosive ordnance guide for ukraine, third edi- tion,” 2025, Geneva, Switzerland: GICHD
2025
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.