{"id":"7e5925b1-a5e1-4bc7-8c5c-008512bd7b97","arxiv_id":"2507.16398","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In an embodied Turing Test with a physical robot, 34 participants could not identify AI versus human control above chance, and human operators were misidentified as AI far more often than the reverse.","lead":"A study with 34 people tested whether they could tell if a robot was controlled by an AI or by a human operator. The participants guessed correctly only about half the time, and they more often wrongly thought a human operator was the AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Chance-level accuracy is inferred from a binomial test on non-independent trials; the load-bearing error asymmetry also lacks per-participant variance.","rationale":"The reader correctly flags the human-baseline confound (Section III.C.1 and V.B), and I agree that the two-person, guideline-driven baseline limits the scope of the conclusions. The paper's own V.B paragraph even admits that human operators 'exhibited even stronger formal tendencies than the AI' and 'strictly follow the guidelines, their behavior at times resembles that of machines', so a narrower claim about constrained operators is more defensible than the stated claim about human versus AI behavior. However, I regard the non-independence of the 136 trials as the single most load-bearing concern because it directly undermines the statistical basis of the central claim (chance-level accuracy). The binomial test and the chi-square test both assume independent observations; here the unit of randomization is the participant, not the trial. This is a standard concern in HRI repeated-measures data and in much of the Turing Test literature, so it is routeable through correctness risk. The fix is straightforward, but absent raw data or per-participant analyses it cannot be checked from the paper. The recommendation remains CONDITIONAL, because the reader's overall verdict already reflects the baseline limitations; the clustering concern strengthens the conditionality rather than overturning the verdict.","tokens_in":9647,"tokens_out":1386,"duration_ms":13927,"concrete_test":"Recompute the identification results with participant as a random effect, e.g., a mixed-effects logistic regression (glmer) predicting correct identification from operator type, task, and modality with random intercepts for participants, using the four repeated trials per participant, or, at minimum, a cluster-robust Wald test on the aggregate proportions. If the operator-type effect and the chance-level null survive the clustered analysis, the headline claims are robust; if significance disappears or the confidence intervals widen substantially, the current p-values overstate the findings.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central null result (Section IV.A) is computed as a binomial test on 136 trials (71/136 correct, p=0.334). This treats all trials as independent, but each of the 34 participants contributed four trials under the same operator (17 per operator type), with tasks and modalities as within-subject factors. The paper reports no per-participant variance, no mixed-effects model, and no clustering adjustment. Moreover, the more striking asymmetry result (50/68 human-as-AI misidentifications versus 15/68 AI-as-human; chi2=34.07, p=5.33e-9) is likewise computed on the aggregate table; a small number of participants could dominate the error counts because participants are the between-subjects unit for operator type. The reader's weakest assumption about the two-person human baseline is also real and is acknowledged by the authors in Section V.B, but the independence issue is more load-bearing because it affects the headline null result itself: if accuracy varies systematically by participant, the reported p-value is invalid, and the chance-level claim is not established. The paper drops raw data, so the magnitude of the clustering effect cannot currently be assessed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an embodied Turing Test with the IVO robot. Thirty-four participants each completed four interactions (two tasks crossed with static and dynamic modes), with operator type as a between-subjects factor: 17 interacted with an LLM (GPT-4o-mini with RAG) and 17 with one of two human teleoperators. After each interaction, participants guessed whether the operator was AI or human. The headline result is that participants identified the operator correctly in 71/136 trials (52.21%), a non-significant binomial test (p=0.334); the paper also reports a strong asymmetry in errors (human operators misclassified as AI 50 times versus AI operators misclassified as human 15 times; chi-square=34.07, p=5.33e-9), alongside analyses of confidence, reasons, demographics, and conversational data. The authors conclude that humans cannot reliably distinguish AI- from human-operated robots beyond chance and tend to over-attribute AI.","tokens_in":9970,"tokens_out":6669,"duration_ms":76233,"significance":"If the central claims are robust, the study is a meaningful extension of LLM Turing Tests from text-only settings to embodied interaction involving navigation and manipulation, and it provides a clear, falsifiable result plus an asymmetry that could inform HRI design. The study has several strengths: it uses a physical robot, two functionally different tasks, three interaction languages, a transparent prompt-design process informed by a pilot study, an explicit response-delay model (Eq. 1), and evaluation by independent participants who did not know the operator assignment. I found no circular-reasoning issue: the pilot-based prompt and delay parameter are imported design choices rather than outcome variables, and the chance-level result is not encoded in the prompt. The significance is currently conditional, however, because the principal statistical claims rest on trial-level tests that ignore the nested data structure.","major_comments":[{"comment":"The headline null result (71/136 correct, binomial p=0.334) treats the 136 judgments as independent trials. The design described in Section III.D has each of 34 participants contributing four judgments, and operator type is constant within a participant; trials are therefore clustered within participants and the operator factor is between-subjects. If accuracy varies across participants, the effective number of independent observations is closer to 34 than 136, and the reported p-value is not valid as a test of the chance-level claim. Please report per-participant accuracy, an intraclass correlation, a mixed-effects logistic regression with participant as a random effect, and/or a participant-level binomial test, or provide raw data so the clustering can be independently assessed. This is load-bearing because the abstract and title state the chance-level result as the central finding.","section":"Section IV.A"},{"comment":"The asymmetry result (50 human-as-AI misidentifications versus 15 AI-as-human, chi-square=34.07, p=5.33e-9) is computed on the aggregate 2x2 table without accounting for repeated measures. Since operator type is between-subjects, the 50 and 15 counts could be driven by a small number of participants who consistently misjudged one operator type. Please report the per-participant distribution of misidentifications (for example, counts of participants making 0-4 errors for each operator type) and provide a clustered test, such as a participant-level Mann-Whitney or Wilcoxon test on the number of misidentifications, or a mixed-effects model, before treating the asymmetry as robust.","section":"Section IV.A"},{"comment":"The human baseline is limited to two operators who were instructed to answer only from a provided document in the information task and to 'establish trust' during the handover task. The paper itself acknowledges in Section V.B that these human operators 'exhibited even stronger formal tendencies than the AI' and that when they 'strictly follow the guidelines, their behavior at times resembles that of machines.' As written, the conclusion that people cannot distinguish AI- from human-controlled robots generalizes beyond what the data support; the comparison is between an LLM and a highly constrained teleoperation protocol. Please reframe the central claim to state this boundary condition explicitly, or supplement the baseline with additional, less-constrained operators, and discuss how the null result might differ with a more natural human baseline.","section":"Section III.C.1 and Section V.B"}],"minor_comments":[{"comment":"The abstract and several other places contain typographical and grammatical errors, including 'associated to the the challenge,' 'system intelligence,' and 'participants responses'; a careful proofread is needed.","section":"Abstract and general text"},{"comment":"The notation N(0.3, 0.03) should be explicitly defined as a Gaussian random variable, and the text should state whether the delay is drawn once per response or per character and how negative or near-zero draws are handled.","section":"Section III.C.2, Eq. (1)"},{"comment":"The reported p=0.334 appears to be a one-tailed binomial probability for the direction 'better than chance'; since the text says 'no significant deviation from 50%,' the two-tailed p-value should also be reported for consistency with the wording.","section":"Section IV.A"},{"comment":"The correlation claims involving age and chatbot interaction frequency are based on small subgroups (n_ind=6, 4, 4, 20) and are presented without a correlation coefficient or confidence interval; please report the relevant statistic or use a model that accounts for participant clustering.","section":"Section IV.C"},{"comment":"The figure would be easier to interpret if the text specified whether the shaded 95% confidence interval is computed per participant or per trial and how the confidence-level groupings were formed.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The central claims depend on a reanalysis that accounts for the nested data structure; I would therefore request raw data or participant-level summary tables as part of the revision. The topic is well within the journal's scope, and I see no integrity or circularity concern, but the statistical claims as currently presented are not yet established. The manuscript also lacks a data availability statement, which is important given the reanalysis requested."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know: this is the first study I've seen that runs an embodied verbal Turing Test with a physical robot, comparing an LLM operator against a teleoperated human in real time. The central result—participants at 52% accuracy, binomial p=0.334—extends the text-only GPT-4 finding from Jones and Bergen to robots with navigation and grasping, and the asymmetry (human operators misidentified as AI in 50 of 68 trials vs 15 of 68 for AI) is a genuinely interesting pattern.\n\nWhat it does well: the design is more careful than most HRI LLM studies. Two tasks (information retrieval, package handover), two movement modes, three languages, RAG grounding, a pilot study to shape the AI prompt, and an honest discussion that acknowledges the human operators sometimes sounded machine-like. The hallucination analysis is a nice addition. The paper is transparent about its own limitations.\n\nWhere it's soft: the statistics are all aggregate. 136 trials come from 34 participants, with operator type constant per person. Treating them as independent inflates the effective N. For the null result this doesn't matter much—a clustered test would be even less significant—but the chi-square on the asymmetry needs per-participant variance before I'd trust the p=5e-9. The bigger conceptual issue is the human baseline: two operators, instructed to stick to a script and 'establish trust,' and the authors admit their behavior was at times more formal than the AI. That makes the comparison more 'LLM vs. constrained teleoperator' than 'AI vs. typical human.' The paper flags this; it doesn't solve it.\n\nAlso no raw data or code, so the clustering magnitude can't be checked. Minor: the response-delay model is imported without sensitivity analysis, and the temperature is fixed at 1. These are fixable.\n\nWho it's for: HRI and LLM-perception researchers, and anyone working on Turing Tests in embodied settings. It deserves a serious referee. My recommendation: send it out, but ask for per-participant analysis (or at least mixed-effects) and a stronger justification of the human baseline before publication.","headline":"First embodied verbal Turing Test with an LLM-driven robot against a teleoperated human; the null result and misidentification asymmetry are worth attention, but the two-person baseline and clustered statistics keep it from being definitive.","tokens_in":10377,"tokens_out":2733,"would_cite":true,"duration_ms":30650,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that people cannot reliably tell whether an embodied robot is steered by a large language model or by a human teleoperator, and that they systematically over-attribute human behavior to AI.","keywords":["embodied Turing test","human-robot interaction","large language models","teleoperation","AI perception","retrieval-augmented generation","social robots","misidentification bias"],"falsifier":"Replicate the Turing Test with unscripted human operators who answer from their own knowledge and personality. If participants then identify the operator correctly significantly more often than chance (for example, above 60%), the claim that people cannot tell AI from human operators would be falsified; if the human-misclassified-as-AI asymmetry disappears under unscripted operators, the bias interpretation would also be called into question.","tokens_in":1707,"feed_emoji":"🤖","tokens_out":2491,"duration_ms":100354,"temperature":0.7,"pith_summary":"This paper asks whether people can tell who is actually driving an embodied robot: a large language model or a human teleoperator. Across two tasks (information assistance and package handover) and two movement modes (static and dynamic), 34 participants completed 136 interactions and then guessed the operator type. The authors claim that accuracy was no better than chance, with 71 of 136 correct (52.21%, binomial p=0.334), and that errors were strongly asymmetric: human operators were mistaken for AI in 50 trials (36.76%) while AI operators were mistaken for humans in only 15 (11.03%). The paper also reports that confidence, age, and chatbot familiarity shaped judgments without rescuing overall accuracy. This matters for the design of social robots and for the wider question of whether the Turing Test transfers to embodied, task-oriented settings.","feed_headline":"People can't tell if a robot is AI- or human-driven","feed_subtitle":"Only 52% of guesses were right, and humans were more often mistaken for AI than the reverse.","key_machinery":"The load-bearing object is an embodied Turing Test built on the IVO robot with two matched operator channels. The AI channel uses GPT-4o-mini for dialogue generation plus Retrieval-Augmented Generation (RAG) with LaBSE embeddings grounded in a document about a university building; the human channel uses two remote teleoperators who receive the same document and are instructed to follow it. Both channels feed the same text-to-speech system, and the AI's response timing is artificially delayed to mimic human typing speed. Movement, person tracking, and arm control are shared and triggered by either operator, so the conversational output is the variable under test. The questionnaire, adapted from a prior LLM Turing Test study, asks participants to choose AI or human, rate confidence on a 1-10 scale, and select reasons for their judgment.","core_discovery":"The paper's central claim is that in an embodied Turing Test with an LLM-controlled robot, people cannot reliably distinguish AI from human operators. The evidence is a binomial test on 136 judgments: 52.21% correct is not significantly different from 50% ($p=0.334$). The secondary claim is that the two error types are not symmetric: participants misclassified human operators as AI far more often than the reverse, with a chi-square of $\\chi^2=34.07$ ($p=5.33\\times10^{-9}$). The authors also find that task type and robot movement mode did not affect accuracy, that RAG kept the LLM's factual errors low (8 hallucinations in 64 AI interactions), and that when hallucinations occurred, participants correctly identified the AI in 6 of 8 cases. They interpret the asymmetry as evidence that people associate AI with formality, politeness, and scripted behavior, so human operators who follow guidelines and sound formal are judged to be machines.","pith_inferences":["If the human baseline is unrepresentative—only two operators, both instructed to stick to a document and to 'establish trust'—then the result may describe how people classify machine-like behavior rather than a genuine AI–human equivalence; a naturalistic human baseline is the next test.","The confidence pattern (accuracy rising with confidence for the AI but falling for the human operator) suggests people carry overconfident stereotypes about what AI sounds like; collecting a continuous 'how AI-like is this response' rating before the binary choice would test this directly.","Because robot movement mode made no difference, the results predict that conversational content dominates in embodied Turing Tests; an extension that varies physical behavior while holding the dialogue fixed would test that prediction."],"forward_implications":["If people cannot tell the operator apart, an embodied LLM robot is already indistinguishable from a human teleoperator in short, task-oriented conversations, including when the robot is moving.","The misclassification asymmetry implies the practical risk is over-attribution of AI: human operators acting under guidelines are judged to be machines, which could affect trust in human-operated remote services.","Response timing and linguistic formality are the cues that drive judgments, so robot designers who want to be perceived as human should soften formality and vary response delays.","RAG appears to keep the LLM factually reliable, while the rare hallucinations expose the AI; further reducing hallucinations would make the AI channel harder to detect.","Accuracy did not improve across repeated interactions in the study, suggesting familiarity with the robot by itself does not teach people to spot the AI; training or performance feedback may be needed."],"supporting_citations":[{"why":"Defines the Turing Test that the study operationalizes in an embodied robotic form.","marker":"[10]"},{"why":"Supplies the premise that physical embodiment enhances natural communication, motivating the embodied setup.","marker":"[9]"},{"why":"Justifies testing intelligence jointly through language, perception, reasoning, and action in an embodied Turing Test.","marker":"[17]"},{"why":"Prior embodied nonverbal Turing Test result that this study extends to verbal, LLM-driven interaction.","marker":"[22]"},{"why":"Source of the questionnaire, the response-delay model, and the text-based baseline for LLM indistinguishability.","marker":"[26]"},{"why":"Comparison point for how LLMs and humans are perceived in constrained and displaced Turing Test settings.","marker":"[30]"},{"why":"The IVO robot platform that carries the embodied interaction in the experiment.","marker":"[37]"},{"why":"Provides the RAG method that grounds the LLM's answers in the information task document.","marker":"[43]"}],"fun_headline_variants":["AI or human? People can't tell in robot Turing test","Embodied Turing test: Humans fail to spot AI robots","Robot Turing test: Human-driven robots mistaken for AI","People can't reliably tell AI-driven robots from humans"],"cache_read_input_tokens":12544,"weakest_assumption_plain":"The central claim depends on the human operators being representative of natural human interaction, but they were only two people instructed to answer strictly from a document and to 'establish trust,' and the paper itself notes their behavior at times resembled machines; if that baseline is artificial, the result compares two scripted systems rather than AI versus human.","fun_headline_variants_meta":{"raw":{"variants":["AI or human? People can't tell in robot Turing test","Embodied Turing test: Humans fail to spot AI robots","Robot Turing test: Human-driven robots mistaken for AI","People can't reliably tell AI-driven robots from humans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000595,"raw_usage":{"total_tokens":2768,"prompt_tokens":908,"completion_tokens":1860,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":1794}},"tokens_in":524,"tokens_out":1860,"duration_ms":15042,"temperature":1.0,"reasoning_tokens":1794,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:10:08.932115+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replicate the Turing Test with unscripted human operators who answer from their own knowledge and personality. If participants then identify the operator correctly significantly more often than chance (for example, above 60%), the claim that people cannot tell AI from human operators would be falsified; if the human-misclassified-as-AI asymmetry disappears under unscripted operators, the bias interpretation would also be called into question.","supporting_citations":[{"cited_title":"Computing machinery and intelligence","cited_arxiv_id":null,"evidence_quote":"Defines the Turing Test that the study operationalizes in an embodied robotic form."},{"cited_title":"Embodiment in socially interactive robots","cited_arxiv_id":null,"evidence_quote":"Supplies the premise that physical embodiment enhances natural communication, motivating the embodied setup."},{"cited_title":"Why we need a physically embodied Turing Test and what it might look like","cited_arxiv_id":null,"evidence_quote":"Justifies testing intelligence jointly through language, perception, reasoning, and action in an embodied Turing Test."},{"cited_title":"Human-like behavioral variability blurs the distinction between a human and a machine in a nonverbal Turing Test","cited_arxiv_id":null,"evidence_quote":"Prior embodied nonverbal Turing Test result that this study extends to verbal, LLM-driven interaction."},{"cited_title":"GPT-4 is judged more human than humans in displaced and inverted Turing tests","cited_arxiv_id":"2407.08853","evidence_quote":"Comparison point for how LLMs and humans are perceived in constrained and displaced Turing Test settings."},{"cited_title":"Ivo robot: A new social robot for human- robot collaboration","cited_arxiv_id":null,"evidence_quote":"The IVO robot platform that carries the embodied interaction in the experiment."},{"cited_title":"Retrieval-augmented generation for knowledge- intensive nlp tasks","cited_arxiv_id":null,"evidence_quote":"Provides the RAG method that grounds the LLM's answers in the information task document."}],"review_version":1}