{"id":"94770f0b-f649-4ed8-81e8-8a21030fde8e","arxiv_id":"2508.00288","paper_version":5,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"UAV-ON is a new benchmark of 14 Unreal Engine environments with 1270 annotated objects that tests whether aerial agents can navigate to goals described by semantic instance-level instructions.","lead":"The paper introduces UAV-ON, a benchmark of 14 simulated open-world environments for drone navigation, with 1270 target objects described by category, size, and appearance. It reports that all tested navigation agents fail to reliably find these objects, which marks a gap in current aerial semantic navigation technology.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Benchmark validity rests on baseline strength and environment fidelity; abstract evidence is insufficient to support the claimed difficulty gap.","rationale":"The reader identified external validity (environment fidelity) as the weakest assumption. I agree that is a concern, but the more immediately load-bearing issue is baseline representativeness: without evidence that the baselines are strong, the central empirical claim that 'all baselines struggle' could be trivially true. This directly affects whether the benchmark measures the intended difficulty gap. Since this is an abstract-only review, the appropriate verdict remains UNVERDICTED; my concern does not move it to a different category, hence UNCHANGED. The concrete test would resolve the concern by checking whether a strong baseline succeeds and whether the aerial modality itself is the cause of difficulty.","tokens_in":736,"tokens_out":1879,"duration_ms":21284,"concrete_test":"Run a state-of-the-art ObjectNav policy (e.g., a semantic-goal variant of a strong embodied navigation agent) on UAV-ON with a comparable compute budget, and report success rate, SPL, and goal-conditioned distance curves. In addition, ablate the aerial dimension by running the same baselines with a simulated ground agent in the same environments. If the strong agent achieves materially higher success (e.g., >50%) or the ground agent performs comparably to the aerial agents, the 'all baselines struggle' claim is an artifact of weak baselines or environment design rather than a property of aerial semantic ObjectNav.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that all baselines struggle, demonstrating the compounded difficulty of aerial ObjectNav. This claim is load-bearing because the benchmark's utility as a discriminator depends on it. Yet the abstract reports no baseline details: no model capacities, no training budgets, no comparisons to established ObjectNav or VLN agents. If the baselines are deliberately simple or undertuned, their failures reflect implementation weakness rather than task difficulty. Similarly, the 14 Unreal Engine environments are described as 'high-fidelity,' but no fidelity metrics, domain-randomization protocols, or sim-to-real transfer evidence are provided. The benchmark's external validity for real UAV operations therefore rests on an unverified assertion. This is not an internal inconsistency but a missing security argument for the central result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces UAV-ON, a benchmark for open-world Object Goal Navigation (ObjectNav) by aerial agents, built on 14 Unreal Engine environments covering urban, natural, and mixed-use settings. The benchmark provides 1270 annotated target objects, each with an instance-level semantic instruction encoding category, physical footprint, and visual descriptors, intended to create realistic ambiguity and grounding challenges. The authors implement several baselines, including Aerial ObjectNav Agent (AOA), a modular policy, and report in the abstract that all baselines struggle in this setting. The paper argues that this difficulty gap demonstrates the compounded challenges of aerial navigation and semantic goal grounding.","tokens_in":978,"tokens_out":2573,"duration_ms":27819,"significance":"If the benchmark is valid and its difficulty gap is genuinely attributable to task features rather than baseline weakness, UAV-ON would fill a meaningful gap in aerial embodied navigation research, moving beyond language-conditioned VLN paradigms. The concrete contributions listed in the abstract—14 high-fidelity environments, 1270 annotated objects, and instance-level semantic instructions—are potentially reusable resources for the community. The work does not appear to derive new theory, but benchmark construction and baseline evaluation are empirically valuable. However, the abstract alone does not support the central empirical claim, and the external validity of the simulated environments to real UAV operations is asserted rather than demonstrated. The paper's utility will depend critically on the full text providing quantitative results, reproducible protocols, and fidelity evidence.","major_comments":[{"comment":"The central empirical claim—'all baselines struggle in this setting'—is stated without any quantitative support. No success rates, SPL, distance-to-goal metrics, error bars, or statistical comparisons are reported. Because the benchmark's discriminating power rests on this claim, the abstract is insufficient to establish that the difficulty is due to task complexity rather than undertuned baselines. The full text must report per-baseline performance, model capacities, training budgets, and hyperparameters.","section":"Abstract"},{"comment":"The 14 Unreal Engine environments are described as 'high-fidelity,' but no fidelity metrics, domain-randomization protocols, or sim-to-real transfer evidence are provided. The claim that UAV-ON reflects large-scale open-world aerial conditions therefore rests on an unverified assertion. The paper should include at least basic quantitative characterizations of the environments (e.g., semantic label accuracy, visual diversity measures) or a discussion of known limitations for real-world transfer.","section":"Abstract"},{"comment":"The evaluation protocol is not specified. The abstract does not state the number of episodes per object/environment, the success criteria, episode length limits, or random-seed handling. Without this information, the benchmark's reproducibility and comparability cannot be assessed. The full text must define a concrete evaluation protocol and report aggregated results across multiple runs if stochasticity is involved.","section":"Abstract"}],"minor_comments":[{"comment":"The term 'ObjectNav' is used without definition, though it is common in the field; a brief expansion would improve accessibility for readers outside embodied AI.","section":"Abstract"},{"comment":"The acronym 'UAV-ON' is not expanded at first use; consider stating 'Unmanned Aerial Vehicle Object Navigation' or similar to clarify its meaning.","section":"Abstract"},{"comment":"The phrase 'diverse semantic regions and complex spatial layouts' is qualitative; the paper would benefit from concrete statistics, such as the number of rooms, typical floorspace, or distribution of object categories across environments.","section":"Abstract"},{"comment":"The paper should state the intended availability of the benchmark, the Unreal Engine assets, and the baseline implementations to support community adoption and verification.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This review is based on the abstract only, as the full text was not provided. The central empirical claim and the benchmark's external validity cannot be assessed from the available material. If the full paper contains careful baseline tuning, quantitative results, and fidelity analysis, it could be a significant contribution; conversely, if the baselines are trivial and the environments are unvalidated, the claims would weaken substantially. I would recommend requesting the full manuscript before reaching a verdict."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take on UAV-ON: this is a legitimate benchmark contribution—first open-world ObjectNav benchmark for aerial agents, with instance-level instructions that encode category, footprint, and visual descriptors, 14 Unreal Engine environments, and 1270 annotated objects. That combination is new, and the modular baseline AOA is a reasonable starting point. The authors also correctly distinguish their task from VLN, which is a real gap.\n\nWhat I like: the task definition is clean, the scale is moderate but useful, and the instance-level instructions add realistic ambiguity that ground-agent ObjectNav benchmarks usually avoid. This should attract the embodied navigation crowd.\n\nWhere I'm more cautious: the abstract's central empirical claim, that all baselines struggle, is load-bearing but has no numbers behind it. No metrics, error bars, baseline capacities, or training budgets. If the baselines are deliberately simple, the 'compounded challenges' conclusion could be an artifact of implementation weakness. The stress-test note is right about this. Also, 'high-fidelity' environments are asserted, not demonstrated—no fidelity metrics or sim-to-real evidence. These are missing arguments, not internal contradictions, so they may be resolved in the full paper. But as of the abstract, the difficulty gap is unverified.\n\nI agree with the reader's cautious take; the low soundness score reflects the abstract-only constraint, not a known flaw. If the full paper ships baseline details, benchmark statistics, and reproducibility code, this deserves a serious referee.\n\nRecommendation: send to peer review. Ask reviewers to demand full per-baseline results and environment fidelity analysis. If those hold up, this is a citable asset.","headline":"New aerial ObjectNav benchmark with a plausible design, but the abstract doesn't yet support the claim that the task is hard for all baselines.","tokens_in":1348,"tokens_out":1916,"would_cite":true,"duration_ms":21446,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces UAV-ON, a benchmark that asks aerial agents to find target objects from high-level semantic goals rather than step-by-step instructions, and reports that all tested baselines struggle at the task.","keywords":["aerial object goal navigation","UAV benchmark","open-world navigation","semantic goal grounding","Unreal Engine simulation","embodied AI","long-horizon exploration"],"falsifier":"The central claim would be called into question if a baseline trained on UAV-ON succeeded at a comparable rate on real outdoor drone flights searching for objects described only by category, footprint, and visual descriptors; for example, finding ten real-world objects across several urban and natural areas.","tokens_in":582,"feed_emoji":"🛸","tokens_out":2403,"duration_ms":26315,"temperature":0.7,"pith_summary":"This paper introduces UAV-ON, a benchmark for object-goal navigation by aerial agents in large open-world environments. Its aim is to measure whether autonomous drones can locate a specified target object when given only a high-level semantic goal, without the sequential linguistic commands used in vision-and-language navigation. The benchmark comprises 14 high-fidelity virtual environments spanning urban, natural, and mixed-use settings, with 1270 annotated target objects, each described by category, physical footprint, and visual descriptors. The authors implement a modular baseline, the Aerial ObjectNav Agent, and report that all evaluated methods perform poorly, arguing that the setting compounds the difficulties of aerial navigation with those of semantic goal grounding.","feed_headline":"Drone benchmark: autonomous aerial object search defeats all baselines","feed_subtitle":"New benchmark gives aerial robots semantic goals instead of step-by-step instructions; none of the tested agents solves them.","key_machinery":"The load-bearing object is the UAV-ON benchmark itself, and within it the instance-level goal instruction: each target object is encoded with its category, physical footprint, and visual descriptors, turning the navigation task into a grounded reasoning problem rather than a following-instructions problem. The Aerial ObjectNav Agent serves as the evaluation instrument, a modular policy intended to integrate instruction semantics with egocentric observations for long-horizon exploration.","core_discovery":"The paper's central claim is that UAV-ON captures a real capability gap in embodied intelligence: aerial agents operating in large-scale open-world environments cannot yet reliably complete object-goal navigation when the goal is a semantic description rather than a route instruction. The benchmark's 14 Unreal Engine environments and 1270 instance-level goal annotations are designed to force agents to reason about what an object looks like, where it is likely to be, and how to navigate over long horizons. The authors' own baseline, Aerial ObjectNav Agent, combines instruction semantics with egocentric observations, and its failure alongside the other baselines is presented as evidence that the compounded challenges of aerial navigation and semantic goal grounding remain unsolved.","pith_inferences":["The instance-level goal annotations could be reused as a testbed for zero-shot sim-to-real transfer, since they encode visual descriptors that a real deployment would also need to ground.","An implicit consequence is that the same benchmark design could be applied to ground robots, allowing researchers to isolate the aerial-specific difficulty by comparing against existing ground-based object navigation results.","Because every tested baseline struggles, the benchmark currently offers a ceiling measurement rather than a meaningful ranking; adding easier sub-tasks or progressive goal complexity would improve its diagnostic value."],"forward_implications":["Aerial object-goal navigation can be studied as a standalone problem, separate from the sequential-instruction paradigm of vision-and-language navigation.","Semantic goal descriptions of the kind UAV-ON defines could become a practical interface for tasking autonomous drones in real operations.","Future methods can be compared against the reported baseline failures, giving the community a concrete measure of progress in long-horizon aerial exploration.","The difficulty of the current baselines points research toward combining language grounding with spatial exploration and planning in a single policy."],"supporting_citations":[],"fun_headline_variants":["Aerial robots fail on new open-world object search benchmark","New drone benchmark: semantic goals stump all tested agents","Open-world aerial Nav benchmark leaves every baseline behind","Drone ObjectNav: no agent passes semantic goal benchmark","UAV benchmark exposes gap in long-horizon object search"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 14 Unreal Engine environments with their annotated objects faithfully represent large-scale open-world aerial conditions, so that the measured difficulty transfers to real UAV operations.","fun_headline_variants_meta":{"raw":{"variants":["Aerial robots fail on new open-world object search benchmark","New drone benchmark: semantic goals stump all tested agents","Open-world aerial Nav benchmark leaves every baseline behind","Drone ObjectNav: no agent passes semantic goal benchmark","UAV benchmark exposes gap in long-horizon object search"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000441,"raw_usage":{"total_tokens":2233,"prompt_tokens":943,"completion_tokens":1290,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":1225}},"tokens_in":559,"tokens_out":1290,"duration_ms":9425,"temperature":1.0,"reasoning_tokens":1225,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:13:14.204768+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The central claim would be called into question if a baseline trained on UAV-ON succeeded at a comparable rate on real outdoor drone flights searching for objects described only by category, footprint, and visual descriptors; for example, finding ten real-world objects across several urban and natural areas.","supporting_citations":[],"review_version":1}