{"id":"b506641b-6116-41ed-aa11-65c337ee8540","arxiv_id":"2607.10553","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A local sliding map plus history-pose observation inference and incremental viewpoint clustering lets aerial robots search large unknown scenes with less memory and lower decision latency than dense-map baselines.","lead":"SLIDER is a UAV target-search system that keeps only a sliding local map plus sparse historical poses, instead of a dense global map. It claims lower memory, faster decisions, and better search completion than prior methods in large outdoor scenes.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"History-aware labeling is the real soft spot, but the paper already flags it; the bigger untested risk is whether completeness claims survive realistic occlusion without a dense observation map.","rationale":"The Reader correctly isolates the history-aware frontier module (Sec. IV-B, Alg. 1) as the load-bearing modeling choice. Empirical gains in Table II and real flights are real under the tested conditions, so the paper is not broken; the concern is whether those gains generalize when the optimistic labeling assumption is stressed. No derivation error or circular construction exists, and the incremental clustering / sliding-map contributions stand independently. Because the authors already flag the occlusion limit and the sim/real setups do not stress it, CONDITIONAL remains the right call; the concrete occlusion-injection test would decide whether the verdict should stay or move toward ACCEPT. I see no stronger internal inconsistency.","tokens_in":12944,"tokens_out":517,"duration_ms":7121,"concrete_test":"In the Campus scene, inject synthetic occluders (thin vertical walls or dense foliage) that hide 20–30% of target surfaces from historical poses while remaining visible from later viewpoints; re-run SLIDER vs. EPIC and measure (a) fraction of targets still detected and (b) final frontier residual volume. If SLIDER’s completeness drops >10% relative to EPIC while memory stays low, the history-aware claim fails under realistic occlusion.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central efficiency claim (Table II: 30–50% lower Exp. Tm., all targets found) rests on Algorithm 1 lines 3–8 correctly labeling voxels as sufficiently observed via sparse historical poses + sensor model, without a dense global observation map. The paper’s own Conclusions admit theoretical limits under extreme occlusion. In the geometry-based sim model (Sec. V-A), targets are declared detected solely by FoV + r_good with no image noise or partial occlusion; real flights (Sec. V-D) use AprilTags that are high-contrast and easy to re-observe. If the backward-inference step systematically under-marks occluded surfaces as ‘good’, frontiers are pruned too early, completeness is overstated, and the memory/latency gains become an artifact of incomplete coverage rather than superior planning. The ablation (Table IV) only shows runtime, not labeling accuracy or missed-target rate under controlled occlusion.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"SLIDER is a lightweight aerial target-search framework that replaces dense global occupancy/observation maps with a robot-centric sliding local map (ROG-MAP + local point cloud and observation-quality maps) plus sparse historical poses (M_hist, T_hist). Observation quality of voxels is inferred online from current and historical poses together with the known sensor model (FoV, raycast visibility, r_good), enabling frontier detection without a global observation map (Algorithm 1, Sec. IV-B). Frontiers are clustered with a normal-aware metric; viewpoints are generated and maintained via incremental viewpoint clustering (VCs) that only resets/updates locally affected clusters; a sparse topological graph built from refined viewpoints supplies long-horizon guidance to the SUPER local planner. Simulations in three large MARSIM scenes (forest, garage, campus; 5 runs each) and two real flights report 30–50% lower exploration time, higher average speed, lower memory, and full target completeness versus adapted SSearcher, FALCON, and EPIC baselines (Table II, III, IV; Sec. V).","tokens_in":13278,"tokens_out":1308,"duration_ms":15527,"significance":"If the history-aware labeling is reliable, the work offers a practical, memory-efficient alternative to grid- or global-point-cloud-based exploration for large-scale UAV target search. The combination of sliding local maps, sparse pose history, and truly incremental viewpoint clustering is a clear engineering advance over repeated global re-clustering (SSearcher) and growing global ikd-trees (EPIC). Strengths include multi-scene simulation with ablations that isolate memory, frontier runtime, and incremental clustering; real-world flights with AprilTags; and an open project page. The contribution is primarily systems-level rather than theoretical, but it is timely for resource-constrained aerial search-and-rescue and security applications.","major_comments":[{"comment":"Sec. IV-B / Algorithm 1 lines 3–8 and Conclusions: the central completeness claim rests on the axiom that a voxel is ‘sufficiently observed’ if it lies in FoV, is raycast-visible, and is within r_good of at least one historical or current pose. The paper itself notes theoretical limits under extreme occlusion. Table IV only reports frontier-detection runtime (F_local vs F_global); there is no controlled measurement of labeling accuracy, false-negative rate on occluded surfaces, or missed-target rate when history is sparse. Without such evidence the 100% Completeness numbers in Table II (and the associated time/memory gains) could partly reflect premature frontier pruning rather than superior planning. A quantitative occlusion stress test or an explicit free-space/visibility residual map would make the claim load-bearing.","section":null},{"comment":"Sec. V-A and V-D: simulation uses a purely geometric detection model (target inside camera FoV and r_good; no image noise, partial occlusion, or lighting). Real flights use high-contrast AprilTags that are easy to re-observe. Both settings under-stress the history-aware inference relative to realistic visual search. The performance ranking versus baselines is therefore only partially transferable; either a more realistic perception model in simulation or a quantitative discussion of failure modes under partial visibility is needed to support the ‘search efficiency’ claim.","section":null},{"comment":"Table II / Sec. V-B: baselines were adapted (SSearcher FoV alignment, FALCON LiDAR + 30 m partitions, EPIC angular constraint replaced by camera visibility). While the intent is fairness, the adaptations are not fully specified (exact parameter values, whether original authors’ recommended settings were retained). Small differences in observation model or map resolution can change completeness and timing. A short appendix or supplementary table listing every modified parameter would strengthen the comparative claim.","section":null}],"minor_comments":[{"comment":"Acronym inconsistency: abstract and title use SLIDER; Sec. I expands it as ‘Sparse gLobal Information-DrivenEfficient target seaRch’ (missing space, awkward capitalization). Align the expansion with the title.","section":null},{"comment":"Fig. 2 caption and body: color legend (red/white/yellow/blue/purple) is dense; a small legend inset would improve readability.","section":null},{"comment":"Table I lists M_obs_rc with q_i ∈ Q but never defines the discrete set Q; a one-sentence definition would help.","section":null},{"comment":"Sec. IV-C.2: the virtual viewpoint insertion at p_c is described narratively; a short pseudocode line or reference back to Algorithm 1 would clarify the order of operations.","section":null},{"comment":"References: several recent large-scale exploration works (e.g., EDEN, HPHS) are cited; ensure the comparison discussion in Sec. II-B explicitly positions SLIDER against their hierarchical schemes rather than only against the three experimental baselines.","section":null},{"comment":"Typographical: ‘V oxels’ (space after V) appears multiple times in Sec. IV-B; ‘gLobal’ and ‘seaRch’ in the expansion; ‘Compl. (%)’ column header is fine but the ✗ symbols in Table II could be replaced by explicit ‘fail’ for accessibility.","section":null}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid systems paper for RA-L. The history-aware labeling soft spot is real but already partially acknowledged by the authors; requiring a controlled occlusion ablation (or at least a quantitative discussion of residual free-space risk) is the main gate for acceptance. Novelty relative to SUPER + EPIC + SSearcher is incremental but the engineering integration and real-world results are valuable. No citation or dual-submission concerns noted."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a clean engineering paper that actually ships a usable trade-off. The core move is simple and effective: keep only a robot-centric sliding map (ROG-MAP style) plus sparse historical poses and orientations, then use the known sensor model to back-infer whether a voxel was already well observed. That lets them skip a dense global observation map. They pair it with truly incremental viewpoint clustering that only resets and re-clusters the local changes instead of the full set every step, plus a sparse topo map built from the refined viewpoints for long-horizon guidance. That combination is the real contribution; the individual pieces exist, but the integrated system is new and the numbers show it.\n\nWhat it does well is the evaluation. Three large sim scenes (forest, garage, campus), five runs each, against adapted public baselines (SSearcher, FALCON, EPIC). Table II is clear: 30–50 % lower exploration time, higher average speed, all targets found, and candidate-generation times that stay low while the others blow up. Memory table and the IVC ablation isolate the gains. Two real flights with AprilTags and a real platform close the loop. Algorithms are readable, parameters are the usual domain knobs, citations hit the right priors without padding. No circular math or fitted constants pretending to be theory.\n\nThe soft spot is exactly the one the stress-test and the authors themselves name: the history-pose labeling (Alg. 1 lines 3–8) can under-mark occluded surfaces as “good,” pruning frontiers too early. Sim detection is pure geometry + r_good; real tags are high-contrast. They admit the theoretical limit in the conclusions and list future free-space work. It is a real caveat for completeness claims under heavy clutter, but it does not collapse the efficiency results that were measured. Everything else is solid systems work.\n\nThis is for people who actually fly large-scale search or care about onboard memory/latency. If that is your lane, read it and look at the github. It deserves serious referee time and is already the kind of paper that moves the practical baseline. Engage.","headline":"Practical systems win for large-scale UAV target search: sliding local maps + sparse pose history + incremental clustering deliver real memory and latency gains, with the main risk (occlusion mislabeling) already flagged by the authors.","tokens_in":13867,"tokens_out":542,"would_cite":true,"duration_ms":14949,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Aerial robots can search large unknown spaces efficiently by keeping only a sliding local map plus sparse pose history, not dense global maps.","keywords":["aerial robotics","target search","frontier detection","sliding local map","viewpoint clustering","sparse topological map","memory-efficient exploration"],"falsifier":"In a large cluttered environment containing deep U-shaped traps or heavy multi-layer occlusion, run SLIDER and a dense-global-map baseline side-by-side; if SLIDER systematically leaves target surfaces unmarked as frontiers while the dense baseline finds them, or if SLIDER’s completeness drops below 100 percent while the baseline succeeds, the history-inference claim fails.","tokens_in":13855,"feed_emoji":"🚁","tokens_out":601,"duration_ms":7773,"temperature":0.7,"pith_summary":"SLIDER shows that large-scale aerial target search does not need a dense global occupancy or observation map. A robot-centered sliding local map, together with a sparse record of past poses and a known sensor model, is enough to decide which surfaces have already been well observed and which still need attention. Frontiers are found by replaying history against the sensor model rather than storing per-voxel global quality; viewpoints are clustered incrementally so only the local change is reprocessed; and a sparse topological graph of refined viewpoints supplies long-horizon guidance. In forest, garage and campus simulations and in two outdoor real flights, the method finishes faster, flies at higher average speed, uses far less memory, and still finds every target, while three strong baselines either slow down or fail to finish. The practical claim is that real-time, memory-bounded search remains possible even as the environment grows to thousands of square meters.","feed_headline":"Drones search large spaces without dense global maps","feed_subtitle":"Sliding local maps plus pose history cut memory and decision time while still finding every target","key_machinery":"History-aware frontier detection (Algorithm 1): for each new point, first check the current pose; if still under-observed, query nearby historical poses and the sensor FoV/visibility model to decide whether the voxel was already sufficiently seen, then cluster only the updated frontier voxels.","core_discovery":"A local sliding map plus sparse historical poses and the sensor model can replace dense global observation maps for aerial target search, while incremental viewpoint clustering and a sparse topological map keep planning real-time, yielding lower memory, lower decision latency and higher search efficiency than current map-heavy methods.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Sliding local maps replace dense globals for drone target search","Sparse pose history lets aerial robots search without heavy maps","Local sliding maps cut memory and decision time in target search","Incremental clustering and sparse topology speed aerial searches","Drones find targets efficiently via local maps plus pose history"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Sparse past robot poses plus the known sensor model are assumed sufficient to correctly mark surfaces as already well observed, even when occlusions or limited field of view hide parts of the scene.","fun_headline_variants_meta":{"raw":{"variants":["Sliding local maps replace dense globals for drone target search","Sparse pose history lets aerial robots search without heavy maps","Local sliding maps cut memory and decision time in target search","Incremental clustering and sparse topology speed aerial searches","Drones find targets efficiently via local maps plus pose history"]},"model":"grok-4.5","effort":"low","cost_usd":0.00505,"raw_usage":{"total_tokens":1327,"prompt_tokens":680,"num_sources_used":0,"completion_tokens":59,"cost_in_usd_ticks":50500000,"prompt_tokens_details":{"text_tokens":680,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":588,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":680,"tokens_out":59,"duration_ms":5921,"temperature":1.0,"reasoning_tokens":588,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T10:51:50.554925+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"In a large cluttered environment containing deep U-shaped traps or heavy multi-layer occlusion, run SLIDER and a dense-global-map baseline side-by-side; if SLIDER systematically leaves target surfaces unmarked as frontiers while the dense baseline finds them, or if SLIDER’s completeness drops below 100 percent while the baseline succeeds, the history-inference claim fails.","supporting_citations":[],"review_version":1}