{"id":"72a94db7-c609-4f08-81ca-634223261122","arxiv_id":"2608.03099","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A five-role taxonomy plus an evidence audit shows that many embodied-agent papers describe language as grounded while reporting evidence that only supports system-level success, not language's specific contribution.","lead":"This survey sorts the roles that language plays inside embodied robots into five categories and checks, claim by claim, whether published experiments actually support that language is the cause of the reported behavior. It finds a recurring overclaim: task success, readable language outputs, or internal replanning are often treated as proof of language grounding without the targeted evidence.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Audit frequencies rest on single-coder judgments with no reliability check; headline percentages could shift materially under independent recoding.","rationale":"The reader identifies the assumption that published descriptions are complete and accurate as the weakest premise. That is a related but distinct concern: it addresses source-document completeness, whereas the load-bearing issue here is coder reliability in applying a judgment-heavy codebook to those documents. Both concern measurement validity, so agreement is partial. The paper's own Limitations and Appendix A.3 explicitly disclaim inter-rater reliability, which is honest but leaves the magnitude of coding uncertainty unquantified. The four evidential-substitution forms are illustrated with concrete examples and do not depend on the frequencies; the taxonomy and the normative distinction between functional role and evidential warrant survive even if the percentages shift. However, the headline abstract claims—81.0% for isolation, 92.4% for targeted tests, 28.6% for closed-loop feedback—are presented as corpus statistics and support the \"recurring gap\" framing. If independent recoding produces substantially different rates, the quantitative contribution is weakened even though the conceptual contribution stands. A conditional acceptance requiring a reliability check is therefore appropriate: the paper should either report inter-rater agreement or reframe the frequencies as an author-led qualitative assessment rather than stable audit statistics. This does not impugn the authors' integrity; it reflects the standard scientific requirement that measurement uncertainty be quantified when headline numbers are reported.","tokens_in":22799,"tokens_out":3601,"duration_ms":36794,"concrete_test":"Have an independent coder, blind to the original codes, re-code a random stratified sample of 30 papers from the 105 (10 from each of the three largest joint profiles) using the Table 5 codebook. Compute per-operation Cohen's kappa and compare the sample marginal frequencies for F and I to the reported 28.6% and 81.0%. If kappa is below 0.6 or the F/I estimates differ by more than 10 percentage points, the corpus frequencies should be reported as illustrative coding outcomes with uncertainty bounds rather than as stable audit statistics.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central empirical claim—that F appears in only 28.6% of papers while I appears in 81.0%—depends on a coding protocol that is judgment-heavy. Appendix A.3 states: \"One author performed the initial coding and led the evidence recoding... this review resolved coding decisions but was not an independent second annotation.\" The Limitations paragraph repeats: \"we therefore do not claim rater-independent ground truth or inter-rater reliability.\" The codebook (Table 5) requires decisions such as whether the \"principal alternative explanation\" was addressed (I), whether an observed outcome \"changes a subsequent executable decision or physical attempt\" (F), and whether an instance-matched embodied referent \"evaluates, gates, rejects, revises, or selects\" role content (C). These are not mechanical determinations. A second coder applying the same rules could classify borderline papers differently, moving the marginal frequencies by several percentage points and either enlarging or shrinking the reported evidence gap. Because the quantitative audit is a stated contribution and the \"recurring gap\" framing leans on it, the absence of any inter-rater reliability estimate or sensitivity analysis is the most load-bearing soft spot. The authors' disclosure is honest, but disclosure does not quantify how much the headline rates could vary.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper separates two questions in the study of language grounding in embodied agents: what functional responsibility linguistic content carries, and what evidence supports that responsibility. The authors define five non-exclusive functional roles (Specification, Embodied Representation, Action Orchestration, Grounding Regulation, Execution Coupling) and five non-exclusive evidence operations (route traceability R, targeted behavioral test T, embodied-constraint check C, closed-loop feedback F, claim-relative isolation I). They apply the framework to a coded corpus of 105 papers, reporting that R appears in 100% (by inclusion), T in 92.4%, C in 72.4%, I in 81.0%, and F in only 28.6% of papers. The central claim is that the literature shows a recurring gap between functional use and evidential support, instantiated by four recurrent forms of 'evidential substitution' (inspectable language as evidence of grounding, internal revision as embodied correction, task success as language attribution, and action proximity as grounding strength). Section 6 translates the findings into role-specific evaluation and reporting recommendations. The appendices document corpus construction, the codebook, the full per-paper inventory, and joint evidence profiles; the authors explicitly limit their conclusions to what the published experiments support.","tokens_in":23124,"tokens_out":15090,"duration_ms":129266,"significance":"If the reliability caveat below is resolved, this is a genuinely useful contribution. The taxonomy is clean and the paper carefully avoids treating the evidence operations as a maturity scale; the codebook in Table 5 is specific enough to be applied by other annotators, and Appendix C provides a complete per-paper inventory with exact R/T/C/F/I profiles, making the corpus auditable. I verified that the reported marginals are internally consistent with the joint profiles in Figure 3 (T: 97, C: 76, F: 30, I: 85), a point in the paper's favor. The paper is also unusually transparent: it states in the Limitations that coding involved a single author with no inter-rater reliability, that corpus construction favored traceable mechanism coverage over exhaustive recall, that illustrative examples did not drive the coding decisions or counts, and that the findings concern reported experiments rather than unreported mechanisms. These disclosures, plus the publication-status annotation, strengthen rather than weaken the manuscript's credibility, and the paper's explicit scoping to reported experiments contains the related worry that unreported grounding mechanisms might exist.","major_comments":[{"comment":"The headline quantitative findings—F in 28.6% of papers and I in 81.0%—rest on judgment calls made by a single coder. Appendix A.3 states that 'one author performed the initial coding and led the evidence recoding' and that the subsequent team review 'was not an independent second annotation.' Table 5's rules require non-mechanical decisions for every paper–role claim: whether the principal alternative explanation was addressed (I), whether an observed outcome 'changes a subsequent executable decision or physical attempt' (F), and whether the role-bearing quantity was 'removed, replaced, perturbed, corrupted, or controlled' (T). Because the evidence audit is a stated contribution and the recurring-gap narrative is anchored in these marginals, the absence of an inter-rater reliability estimate or sensitivity analysis is load-bearing. The Limitations paragraph honestly discloses the issue, but disclosure does not quantify how much the headline rates could shift under independent recoding. I request either (a) a reliability check on a random subset of the corpus with agreement statistics, (b) a sensitivity analysis that reports how the marginals move under plausible alternative readings of borderline cases, or (c) an explicit reframing of the frequencies as indicative, with precision reduced accordingly.","section":"Section 5.1, Table 2, Appendix A.3"},{"comment":"The four forms of evidential substitution introduced in Section 1 are presented as 'recurring' in the reviewed literature, but they are never coded or counted. Section 5.1 quantifies only the five evidence operations, and Sections 5.2–5.5 support each substitution form with illustrative, paper-level examples. Consequently, the central claim of a 'recurring gap between functional use and evidential support' is not measured by the audit's own instrument, and a reader cannot determine from the data how many papers exhibit each substitution or whether the high I rate (81.0%) is compatible with the substitution narrative. The authors should either code the four forms at a coarse per-paper level or explicitly limit the 'recurring' claim to the documented examples while recasting the quantitative contribution of the audit accordingly.","section":"Section 1, Sections 5.2–5.5"}],"minor_comments":[{"comment":"The term 'primary assignment' is used for the counts in Figure 2 and organizes the inventory in Appendix C, but the main text never states how a single primary role is selected when roles are non-exclusive; the selection rule should be specified so that the primary-role counts are auditable.","section":"Figure 2, Appendix C"},{"comment":"A stray markdown fragment, '/check-circle', appears immediately before the heading 'Claim–Evidence Contract' in Section 6.4; this rendering artifact should be removed.","section":"Section 6.4"},{"comment":"The two-decimal percentages (92.4%, 81.0%, 28.6%) imply a precision that the judgment-based coding procedure cannot support; give rounded values or state a plausible uncertainty range alongside the counts.","section":"Table 2"},{"comment":"All 30 F-positive papers are also C-positive, and the paper notes that this co-occurrence is descriptive and does not imply a logical prerequisite; however, under the given definitions an observed outcome that 'revises' role content appears to satisfy the C rule, so the authors should state explicitly whether F-positive implies C-positive by construction, to avoid a misleading impression of independence between the two operations.","section":"Appendix D, Figure 3, Table 5"},{"comment":"The many-author entries in the reference list use inconsistent truncation styles ('and 1 others', 'and 26 others', 'and 5 others'); the formatting should be standardized (for example, first author plus et al.).","section":"References"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the single-coder question is the main obstacle to acceptance; I would condition acceptance on a sensitivity or reliability analysis of the coding or on a softened presentation of the quantitative claims. The paper is honest about its limitations and the internal consistency of the reported counts is good. Note that a large share of the cited systems are 2025–2026 preprints; this is appropriate for the field's pace, but it means the 'reviewed literature' coverage is provisional on those versions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's real contribution is the five-role taxonomy of what language does in embodied agents and the R/T/C/F/I evidence-operations codebook for checking what evidence supports each role. That organizing unit is new relative to the architecture- and capability-oriented surveys it cites, and it is genuinely useful: it lets you compare modular and end-to-end systems on the same footing, and the four 'evidential substitution' patterns sharpen a familiar worry into concrete, testable categories. The audit of 105 papers is a good synthetic result. The headline percentages—closed-loop feedback in only 28.6%, claim-relative isolation in 81.0%—are consistent with the qualitative examples and with the codebook definitions, and the paper is admirably transparent about its own limits: one-coder coding, no inter-rater reliability, a highest-cited 10% screening stage, and conclusions that concern only what the published experiments report. That honesty is not a substitute for rigor, but it is a real reason to trust the authors' judgment. The soft spot is exactly what the stress-test note says: the marginal frequencies rest on one person's application of a judgment-heavy codebook. Whether a borderline paper gets an F or an I depends on readings of 'changes a subsequent executable decision' or 'principal alternative explanation' that a second coder could reasonably resolve differently. The point estimates could shift by several points under independent recoding, and no sensitivity analysis or sub-sample reliability check is offered. That does not break the central claim—the gap between language attribution and evidence is visible in the concrete paper examples and in the fact that only a third of the corpus even attempts full feedback closure—but it does mean the audit should be treated as a framework plus a preliminary quantitative snapshot, not as a field-wide prevalence measurement. The companion issue, that method sections may omit interventions the coders would have credited, is real but bounded by the paper's own scope statement: it audits reported evidence, not hidden mechanisms. Who is this for? Anyone designing or evaluating embodied agents with a language component, and anyone writing the next survey on VLA models who wants a more disciplined vocabulary than 'grounding' as a black-box label. It is a review contribution, not a new result about a system, but it is a serious one. I would bring it to a reading group and would cite it in evaluation-oriented work. A serious editor should send it to peer review rather than desk reject. The referee should push for a reliability sub-sample or at least a quantified sensitivity analysis, but that is revision material, not a reason to reject.","headline":"The role–evidence taxonomy is a genuinely useful organizing device for embodied-AI evaluation, and the audit's headline finding (closed-loop feedback is rare) is credible despite the single-coder coding.","tokens_in":802,"tokens_out":809,"would_cite":true,"duration_ms":24691,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A review of 105 embodied-agent papers finds that language's contribution is frequently claimed beyond what reported experiments can support, and offers a role-by-role audit to close the gap.","keywords":["language grounding","embodied agents","functional roles","evidence audit","evidential substitution","closed-loop feedback","robotic manipulation","vision-language-action models"],"falsifier":"Take the papers coded as lacking closed-loop feedback or claim-relative isolation, inspect their released code, logs, or unreported experiment notes, or re-run their pipelines while instrumenting for whether an observed outcome actually changes a later executable attempt; if a substantial share turn out to contain unreported feedback or matched comparisons, the reported gap would shrink accordingly.","tokens_in":22588,"feed_emoji":"🤖","tokens_out":4946,"duration_ms":40026,"temperature":0.7,"pith_summary":"The paper separates two questions that are usually merged: what functional role language plays inside an embodied agent, and what evidence actually supports that role. It defines five non-exclusive roles — Specification, Embodied Representation, Action Orchestration, Grounding Regulation, and Execution Coupling — and asks, for each paper–role claim, which reported observations or interventions test it. Auditing 105 papers with five evidence operations, it finds that route traceability and targeted tests are common, but closed-loop feedback appears in only 28.6% of papers and claim-relative isolation in 81.0%. The recurring problem it identifies is evidential substitution: evidence for one property of a system is used to support a stronger claim. The paper's contribution is a transferable audit method plus concrete reporting questions, not a ranking of systems.","feed_headline":"Only 28.6% of robot-language papers trace feedback into a new attempt","feed_subtitle":"A 105-paper audit shows language is often credited beyond what experiments support.","key_machinery":"The machinery is a two-part coding scheme. The first part assigns each system a Functional Role Profile made of five non-exclusive roles defined by operational responsibility: Specification commits the agent to a task; Embodied Representation makes state available to later decisions; Action Orchestration selects and orders available capabilities; Grounding Regulation lets new evidence change downstream computation through a reported route to behavior; Execution Coupling places a language-derived quantity on the action-production path. The second part codes five non-exclusive evidence operations per paper–role claim: R route traceability, T targeted behavioral test, C embodied-constraint check, F closed-loop feedback, and I claim-relative isolation. The audit works by comparing each role claim to the operations that bear on it, so that modular planners and end-to-end policies can be compared without treating architectural proximity to action as evidence of stronger grounding.","core_discovery":"The central claim is that grounding is not a property of architecture but a property of evidence. In the reviewed literature, four forms of evidential substitution recur: explicit language does not prove grounded content, internal revision does not prove embodied correction, task success does not isolate language's contribution, and closeness to action does not strengthen grounding. Across 105 papers, targeted behavioral tests appear in 92.4% and claim-relative isolation in 81.0%, but closed-loop feedback — an observed outcome triggering a revision that changes a later executable attempt — appears in only 28.6%. The paper argues that evaluation should begin with the responsibility assigned to language and then check whether the reported evidence supports that responsibility claim by claim, rather than treating end-to-end success as evidence of language grounding.","pith_inferences":["If the audit were repeated on a random sample drawn outside the highest-cited 10% used as the screening pool, the size of the evidence gap could shrink or grow; the paper's own protocol makes this a direct check on corpus bias.","A testable extension is to use the framework prospectively: authors could preprint an R/T/C/F/I profile with each submission, turning the audit from retrospective coding into a reporting standard that reviewers can check against the methods section.","The evidential-substitution logic generalizes beyond language: the same rule that evidence strength must match claim scope applies to perception modules, world models, or any component credited with a system-level gain.","An implicit consequence is that 'grounded' should be treated as a claim about a specific responsibility in a specific loop rather than a binary property of a model, which would sharpen debates about whether language-model-based agents are genuinely grounded."],"forward_implications":["Evaluation reports should state a responsibility chain for each role, naming the language-derived quantity, its downstream consumer, and the expected behavioral consequence of altering it.","Role-specific tests should target the claimed responsibility: vary the task formulation while preserving scene and capabilities, check representations against instance-matched observations, perturb orchestration while holding capabilities fixed, trace feedback to a changed executable attempt, or perturb the action-bearing language quantity while preserving the rest of the action stack.","Reports should distinguish revision (internal change), closure (a changed executable attempt), and recovery (an improved outcome), and limit attribution to the last point directly observed.","Evidence provenance should be named, since oracle or simulator checks support consistency with modeled constraints while direct sensing and physical outcomes support different and generally stronger claims.","The R/T/C/F/I profile should be reported per role rather than collapsed into a single grounding score, and a missing operation should limit the warranted claim rather than label the system ungrounded."],"supporting_citations":[{"why":"Supplies the symbol grounding problem that motivates the requirement that linguistic content remain answerable to perception and action.","marker":"Harnad, 1990"},{"why":"Provides the semiotic-schema view of grounding language in action and perception that the review positions as a background constraint.","marker":"Roy, 2005"},{"why":"Establishes the 'experience grounds language' principle used as the inclusion boundary for the reviewed corpus.","marker":"Bisk et al., 2020"},{"why":"SayCan serves as the Action Orchestration example separating linguistic task relevance from state-conditioned feasibility.","marker":"Ahn et al., 2022"},{"why":"Inner Monologue is the canonical feedback-route example used to define Grounding Regulation and embodied closure.","marker":"Huang et al., 2023e"},{"why":"DECKARD illustrates the contrast between an inspectable world model and one corrected by environment experience.","marker":"Nottingham et al., 2023"},{"why":"RT-2 exemplifies the case where task success does not isolate the linguistic contribution from other trained components.","marker":"Zitkovich et al., 2023"},{"why":"CogACT supplies the direct example of an alternative explanation: changing the action module changes success while the language foundation stays fixed.","marker":"Li et al., 2024a"},{"why":"Statler shows an explicit state that may be updated without independent observation, illustrating why readability is not correctness.","marker":"Yoneda et al., 2024"},{"why":"ECoT provides an example where intervention on an embodied trace supports claim-relative isolation.","marker":"Zawalski et al., 2025"}],"fun_headline_variants":["Only 28.6% of robot-language papers close the feedback loop","Robot language claims often outrun the evidence in 105-paper audit","Language grounding: 92% have tests, only 28.6% have feedback loop","105-paper audit: language grounding claims exceed evidence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The audit assumes that the published method and experiment sections of the 105 papers describe the experiments fully enough that the absence of a reported operation means the operation was not performed.","fun_headline_variants_meta":{"raw":{"variants":["Only 28.6% of robot-language papers close the feedback loop","Robot language claims often outrun the evidence in 105-paper audit","Language grounding: 92% have tests, only 28.6% have feedback loop","105-paper audit: language grounding claims exceed evidence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001467,"raw_usage":{"total_tokens":5865,"prompt_tokens":878,"completion_tokens":4987,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":4908}},"tokens_in":494,"tokens_out":4987,"duration_ms":33381,"temperature":1.0,"reasoning_tokens":4908,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:51:43.270217+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the papers coded as lacking closed-loop feedback or claim-relative isolation, inspect their released code, logs, or unreported experiment notes, or re-run their pipelines while instrumenting for whether an observed outcome actually changes a later executable attempt; if a substantial share turn out to contain unreported feedback or matched comparisons, the reported gap would shrink accordingly.","supporting_citations":[],"review_version":2}