{"id":"ca1f90e8-77d6-41f0-8c46-0863f39bf7f2","arxiv_id":"2508.05855","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":3.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey paper that maps attack strategies, defense mechanisms, and evaluation methods for safety in embodied navigation; no new experimental result is claimed.","lead":"This paper is a survey of safety in embodied navigation, where AI agents physically move through environments to reach targets. It organizes known attacks, defenses, and evaluation methods into one map, aiming to guide safer robot design.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Supplied full text does not match the claimed paper: header reads arXiv:2508.05854v3 [quant-ph] 11 Jun 2026, not 2508.05855 cs.AI, and the text is unreadable; the survey's central claim cannot be audited.","rationale":"The reader's weakest assumption was precisely that the supplied full text is the actual paper and that the abstract accurately summarizes it. My review independently confirms the same blocking issue: the visible header references a different arXiv paper (2508.05854v3, quant-ph) and the body is unreadable mojibake. This is not a manufactured concern about consensus or internal consistency; it is a direct obstacle to assessing the paper's central claim. The only way the abstract's 'comprehensive analysis' claim can be checked is by having the actual survey text. Since the provided text is not it, the verdict should remain UNVERDICTED, with low confidence. I am not asking for rejection because there is no evidence the underlying survey is wrong or fabricated; there is simply insufficient verifiable content. The concrete test is a straightforward retrieval-and-extraction check that will settle whether this concern is a real property of the submission or an artifact of the review pipeline. If the test passes, the survey claim can then be reviewed in the normal way. If it fails, the submission as provided is not the claimed survey and cannot be assessed. Agreement with the reader is full: the weakest assumption is the same, and the recommended disposition is unchanged.","tokens_in":20725,"tokens_out":2216,"duration_ms":22561,"concrete_test":"Use the arXiv API to retrieve the canonical metadata for the identifier 2508.05855: query export.arxiv.org/api/query?id_list=2508.05855 and confirm that the returned title matches 'Safety of Embodied Navigation: A Survey', that the category is cs.AI, and that the abstract is identical to the one supplied. Then download the PDF from arXiv (e.g., https://arxiv.org/pdf/2508.05855) and run a text extraction (pdftotext) on it. Verify that the extracted text is readable, matches the supplied abstract in relevant opening passages, contains sections on attack strategies, defense mechanisms, and evaluation methodologies, and that no occurrence of '2508.05854' or 'quant-ph' appears in the header. If the canonical paper is readable and matches the abstract, the concern is resolved and the survey claim becomes assessable. If the canonical paper differs, or if the body is still unreadable/quant-ph, t","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that the paper provides a comprehensive survey of safety in embodied navigation, covering attacks, defenses, and evaluation methods. For a survey, that claim is meaningful only if the body actually contains the cited literature, taxonomy, and synthesis. The supplied full text is unreadable mojibake, and its visible header line, 'arXiv:2508.05854v3 [quant-ph] 11 Jun 2026', refers to a different arXiv identifier and category than this submission (arXiv:2508.05855, cs.AI). Under the reviewing rule that every manuscript passage is in-scope evidence, this header and corrupted body are not dismissed as pipeline artifacts. They constitute direct evidence that the document provided for review is not the claimed survey, or at least that its substantive content is inaccessible. Therefore the load-bearing premise, that the submitted text actually contains the advertised survey, is not satisfied by the supplied material. This is a blocking concern: without a readable body, no check of coverage, citation accuracy, or taxonomy correctness can be performed. The concern is not about the survey's methodology or truth, but about whether the artifact under review is the artifact described by the abstract. If the intended paper exists and is properly retrievable, the claim may be perfectly sound; but on the supplied evidence, it is unverifiable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission presents only an abstract of a survey on safety in embodied navigation, claiming to review attack strategies, defense mechanisms, evaluation methodologies, datasets, and metrics, plus a research agenda. The full text supplied to the reviewer is unreadable mojibake, and the visible header reads 'arXiv:2508.05854v3 [quant-ph] 11 Jun 2026', which does not match the claimed paper identifier or category. Consequently, no substantive content could be audited.","tokens_in":115,"tokens_out":1583,"duration_ms":45137,"significance":"If the intended paper were properly available, a systematic survey of embodied navigation safety could be a useful contribution, particularly the proposed synthesis of attacks, defenses, and evaluation frameworks. However, because the supplied artifact's body is inaccessible and appears to belong to a different paper, the significance of the actual submission cannot be assessed. No strengths in terms of machine-checked proofs, reproducible code, or verifiable taxonomies can be identified from the available material.","major_comments":[{"comment":"The complete body of the submission is corrupted, unreadable text, and the visible header 'arXiv:2508.05854v3 [quant-ph] 11 Jun 2026' does not match the claimed identifier arXiv:2508.05855 (cs.AI). Under the reviewing rule that all manuscript passages are in-scope evidence, this is not a pipeline artifact: it means the document under review is not the claimed survey. The paper's central claim—that it provides a comprehensive analysis of embodied navigation safety—cannot be checked for coverage, citation accuracy, or correctness of the taxonomy. This is a load-bearing defect that no revision can fix short of replacing the entire artifact with the intended paper.","section":"Full text (visible header)"},{"comment":"The abstract asserts a 'comprehensive analysis' of attack strategies, defense mechanisms, evaluation methodologies, datasets, and metrics, but none of these elements appear in the readable material. Given the full-text corruption, the assertion is unverifiable and currently unsupported. If the correct manuscript is resubmitted, the authors should ensure that the abstract's claims of comprehensiveness are backed by explicit sections, tables, and citations.","section":"Abstract"}],"minor_comments":[],"recommendation":"reject","confidential_remarks":"To the editor: the supplied file is unreadable and carries a header from an unrelated arXiv submission. This appears to be a wrong or corrupted file rather than a scholarly manuscript. If the authors can provide the correct, readable paper, a fresh review may be warranted, but the current submission cannot be evaluated and should be rejected as submitted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: I can't judge this paper on the supplied artifact. The abstract is coherent and the topic is timely, but the full text is indecipherable mojibake and at the bottom the header reads 'arXiv:2508.05854v3 [quant-ph] 11 Jun 2026' — not arXiv:2508.05855 cs.AI. Reviewing rule says I shouldn't dismiss that as a pipeline artifact, so I'm treating it as evidence: the document I received is not the paper the abstract describes, and its substance cannot be audited.\n\nWhat's actually good: the abstract frames the survey around attacks, defenses, and evaluation methods for embodied navigation safety, with datasets, metrics, and a call for verification frameworks. That's the right organizational schema, and it matches a genuine need in the subfield. If the body does what the abstract promises, it would be a useful reference for people working on adversarial robustness and safe RL for navigation agents. The prose is clear and genre-appropriate.\n\nWhere it falls apart: a survey's value is coverage and accuracy, and I literally cannot check either. There are no citations, tables, or taxonomy visible in the supplied text. The mismatched header is a blocking problem. If the arXiv listing itself contains this corrupted version, the submission fails basic readability and should be rejected on that ground. If it's a transfer error on our side, then this review is moot, but I can only work with what's in front of me. The reader's low-confidence UNVERDICTED verdict is right; I'd go further and say the artifact as provided is not reviewable.\n\nBottom line: the underlying survey might be fine, and if a clean PDF of the actual 2508.05855 turns up, it deserves a serious referee. On this evidence, no. I wouldn't cite it, couldn't bring it to reading group, and wouldn't send this version out for review.","headline":"The abstract describes a useful survey, but the supplied manuscript's full text is unreadable and headed as a different paper, so nothing beyond the abstract can be checked.","tokens_in":21481,"tokens_out":1999,"would_cite":false,"duration_ms":21077,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that safety in embodied navigation—agents that perceive and move through real, unfamiliar environments—should be organized as attacks, defenses, and evaluation methods, and that verification frameworks are the field's mai","keywords":["embodied navigation","safety","adversarial attacks","defense mechanisms","evaluation methodologies","datasets and metrics","verification frameworks","LLM-driven agents"],"falsifier":"Cross-check every entry in the survey's tables of attacks, defenses, datasets, and metrics against the original cited papers: any materially misclassified or missing entry undercuts the claim to a comprehensive map. Separately, resolving the arXiv identifier mismatch—the body's header reads 2508.05854 [quant-ph] rather than the submission's 2508.05855—would determine whether the body text belongs to this survey at all.","tokens_in":20513,"feed_emoji":"🤖","tokens_out":8782,"duration_ms":87535,"temperature":0.7,"pith_summary":"The paper takes the rapid advance of LLM-driven embodied AI as motivation and focuses on navigation, where agents must perceive, interact with, and adapt to unfamiliar environments while moving toward a target. It argues that because these systems will operate in dynamic real-world settings, safety must be studied as a field of its own rather than as an afterthought. To that end, it develops a three-part map of the literature: attack strategies, defense mechanisms, and evaluation methodologies. The intended contribution is an accurate synthesis of existing challenges, mitigation technologies, datasets, and metrics, together with a research agenda pointing to better evaluation and formal verification.","feed_headline":"Find the verification gap in embodied-navigation safety","feed_subtitle":"A survey maps attacks, defenses, datasets, and metrics, and argues verification is the missing piece.","key_machinery":"The central object is the taxonomy of safety in embodied navigation, a tripartite classification into attack strategies, defense mechanisms, and evaluation methodologies. The taxonomy carries the argument by converting a scattered literature into a field-level map; the gaps it exposes—especially the absence of standardized evaluation and verification frameworks—are the paper's forward-looking findings.","core_discovery":"Building on the observation that embodied navigation systems are being deployed in dynamic, unfamiliar, real-world environments, the survey's central claim is that their safety can and should be analyzed through three lenses: attack strategies (how an adversary can compromise an agent), defense mechanisms (how to resist or detect those compromises), and evaluation methodologies (datasets and metrics that measure effectiveness and robustness). It presents this tripartite analysis as a map of the current literature and argues that the map reveals unresolved problems—new attack classes, better mitigation strategies, more reliable evaluation, and formal verification frameworks. On the paper's ow","pith_inferences":["The same attack-defense-evaluation structure could be lifted from navigation to related embodied tasks—manipulation, inspection, search and rescue—giving the survey a wider reach than its title claims.","A testable extension of the survey's gap analysis is that robustness measured in simulation will not predict robustness under physically plausible disturbances; benchmarks should include real-world or physics-grounded attacks, not only sensor-space perturbations.","Because the abstract motivates safety through LLMs, an implicit open question is whether attacks on the LLM planner are categorically different from attacks on perception; the taxonomy would be stronger if it separated them."],"forward_implications":["A common taxonomy makes it possible to compare attacks and defenses that currently live in separate papers, because each attack can be matched with the defense designed to stop it and the dataset or metric used to test it.","The survey's gap analysis implies that success rate alone is insufficient; evaluation must also count safety violations in unfamiliar environments.","The open problems named in the survey—new attack methods, better mitigations, reliable evaluation, verification frameworks—map to concrete research tasks rather than vague calls for safe AI.","If the field follows this agenda, safer embodied navigation becomes a prerequisite for critical applications, with the stated payoff of societal safety and industrial efficiency."],"supporting_citations":[],"fun_headline_variants":["Embodied navigation safety: verification is the missing piece","Survey maps attacks, defenses, and the verification void","Why embodied navigation needs formal verification","The verification gap in embodied navigation safety","Embodied navigation safety needs a verification roadmap"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The survey's central claim depends on its summaries of the cited literature being accurate and complete, and in the provided text that premise is unverifiable because the body is garbled and carries a header identifying a different paper.","fun_headline_variants_meta":{"raw":{"variants":["Embodied navigation safety: verification is the missing piece","Survey maps attacks, defenses, and the verification void","Why embodied navigation needs formal verification","The verification gap in embodied navigation safety","Embodied navigation safety needs a verification roadmap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000553,"raw_usage":{"total_tokens":2448,"prompt_tokens":693,"completion_tokens":1755,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":1688}},"tokens_in":437,"tokens_out":1755,"duration_ms":13762,"temperature":1.0,"reasoning_tokens":1688,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:06:11.429760+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Cross-check every entry in the survey's tables of attacks, defenses, datasets, and metrics against the original cited papers: any materially misclassified or missing entry undercuts the claim to a comprehensive map. Separately, resolving the arXiv identifier mismatch—the body's header reads 2508.05854 [quant-ph] rather than the submission's 2508.05855—would determine whether the body text belongs to this survey at all.","supporting_citations":[],"review_version":1}