{"id":"9601da73-600a-4f54-8074-d845036134f4","arxiv_id":"2501.12128","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"For a simple industrial pick-and-place task, a scripted robot interaction matches LLM-enhanced interaction on objective efficiency and focus, while subjective ratings only marginally favor the LLM.","lead":"This paper compared a scripted robot interaction with one that used a large language model to generate responses in real time. The scripted version was about as efficient and focused for a simple industrial task, while participants rated the LLM version slightly higher but not significantly so.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'scripted is comparable' claim is undercut by Table I's significant gaze differences; the focus conclusion rests on unquantified heatmap inspection and non-significance treated as equivalence.","rationale":"The reader correctly identified the central weakness: non-significant differences are treated as evidence of equivalence without an equivalence test or power analysis. My stress-test sharpens this into a more specific, load-bearing problem: the paper's own Table I reports statistically significant gaze differences in task execution that point in the direction of more sustained LLM attention, while the abstract and discussion claim PPS is comparable or better on focus. This is not merely absence of evidence; it is an apparent conflict between the reported significant results and the qualitative conclusion drawn from heatmaps. If the significant markers are valid, the 'particularly in focus' part of the strongest claim is unsupported. If they are not valid after multiple-comparison correction, the paper needs to say so and adjust its reporting. Either way, the current manuscript requires revision, not outright rejection, because the efficiency advantage for PPS and the non-significant subjective ratings can still support a conditional version of the practical takeaway. I therefore keep the reader's CONDITIONAL verdict unchanged, while noting that the specific 'focus' wording needs correction.","tokens_in":7060,"tokens_out":6308,"duration_ms":65202,"concrete_test":"Re-run all Table I gaze comparisons as paired within-subjects tests with subject-level clustering and Holm-Bonferroni correction across the gaze metrics, and quantify spatial fixation dispersion in the task-execution phase (e.g., convex hull area or fixation-map entropy) for both conditions. If the task-execution fixation/saccade differences survive correction, the abstract's claim that PPS is comparable 'particularly in focus' is contradicted; if they do not survive, the paper must report corrected p-values and cannot cite those rows as supporting comparability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, as stated in the abstract and the reader's strongest_claim, is that objective metrics show the PPS (scripted) condition performs comparably to the LLM condition, 'particularly in efficiency and focus during simple tasks.' The efficiency part is plausible: Interval Duration (Task Execution) is significantly shorter for PPS (10.1 vs. 11.42 s, marked *). The focus part is not supported by the paper's own data. Table I marks Fixation Duration (Task Execution), Saccade Velocity (Task Execution), and Saccade Amplitude (Task Execution) with * (p < 0.05). In those rows, the LLM condition shows longer fixation duration (79.08 s vs. 70.08 s) and lower saccade velocity (420.03 vs. 434.05 °/s) and amplitude (126.47 vs. 130.64 °), which are typically interpreted as more sustained and less exploratory gaze, i.e., more focus, not less. The paper nevertheless concludes from heatmap inspection that PPS is 'more focused during task execution' (Section IV). No quantitative spatial dispersion metric is reported, so this key conclusion depends on an unquantified visual reading that appears to conflict with the significant gaze metrics. Meanwhile, the non-significant trust and Godspeed comparisons (N=15) are used to claim the LLM condition is only 'marginally higher' and PPS 'remains viable'; without an equivalence margin or power analysis, non-significant p-values cannot establish comparability. Thus the strongest claim is load-bearing on an interpretation of non-significance as equivalence and on a selective reading of heatmaps, while the paper's own significant objective results are not reconciled with the 'focus' wording.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a within-subjects user study (N=15) comparing a fully scripted ('pre-programmed schedule', PPS) interaction with an LLM-enhanced interaction for a simulated industrial pick-and-place task involving a NAO robot mounted on a forklift. The authors collect gaze-tracking metrics, Trust and Godspeed questionnaires, task-time ratios, and energy-consumption estimates. The central claim is that while subjective ratings trend toward the LLM condition, objective metrics show the scripted condition performs comparably, 'particularly in efficiency and focus during simple tasks,' and that the scripted condition may be preferable on latency and energy grounds for simple repetitive interactions.","tokens_in":7302,"tokens_out":2755,"duration_ms":29221,"significance":"If the central claim is properly supported, the paper would provide useful practical guidance for HRI designers choosing between rule-based and LLM-driven control: use the cheaper, more predictable scripted approach for low-complexity tasks and reserve LLM enhancement for tasks that genuinely require adaptation. The study has strengths worth recognizing: a realistic within-subjects setup with counterbalancing, objective eye-tracking measures rather than only self-reports, and an explicit attempt to account for energy costs including training and inference. The manuscript also openly discloses several limitations, such as the GPT-3 proxy for GPT-4o-mini energy data. However, as detailed in my major comments, the focus conclusion is not currently supported by the quantitative gaze metrics, and the equivalence interpretation of non-significant subjective results needs additional statistical backing.","major_comments":[{"comment":"The conclusion that the PPS condition is 'more focused during task execution' is not supported by the paper's own quantitative gaze data. In Table I, Fixation Duration (Task Execution) is significantly longer in the LLM condition (79.08 s vs. 70.08 s, p<0.05), and Saccade Velocity and Saccade Amplitude during task execution are significantly lower in the LLM condition (p<0.05). Longer fixations and lower saccade velocity/amplitude are conventional indicators of sustained, non-exploratory attention, i.e., the opposite of the stated conclusion. The heatmap inspection described in Section III is qualitative and appears to conflict with these metrics. Please either provide a quantitative spatial-dispersion measure (e.g., fixation-map entropy, convex hull area, or k-density spread) that supports the PPS focus claim, or revise the abstract and discussion to align with the gaze metrics actually reported.","section":"Section IV and Table I"},{"comment":"The manuscript treats non-significant differences in the Trust scale (F=1.05, p=0.32) and Godspeed subscales (p>0.05) as evidence that the PPS condition is 'a viable, simpler option' and that subjective ratings are 'marginally higher' for LLM. With N=15, a non-significant p-value does not establish equivalence or comparability; it may simply reflect low power. To support the comparability claim, the authors should either (a) report an equivalence test such as TOST with a pre-specified margin, (b) provide a post-hoc power analysis showing that the study could detect a meaningful difference in these scales, or (c) soften the language to 'no statistically significant differences were found' and avoid drawing a positive comparability conclusion from absence of significance. As written, the central 'performs comparably' conclusion partly rests on this absence-of-evidence interpretation.","section":"Section III, Trust and Godspeed results"},{"comment":"Table I reports p-values for ten pairwise gaze comparisons without any correction for multiple testing. With eight metrics and two phases, the probability of at least one false positive is substantial. The authors should apply a multiple-comparison correction (e.g., Benjamini-Hochberg within each family of gaze metrics) and re-verify which asterisked entries remain significant. This matters because the significant saccade and fixation differences are currently used to infer engagement differences, and their status could change under correction.","section":"Table I and Section III, multiple gaze comparisons"},{"comment":"The energy comparison relies on GPT-3 training and inference figures from reference [11] as a proxy for GPT-4o-mini, which the authors acknowledge. However, the per-query calculation is fragile: it divides the total training energy (1,287,000 kWh) by an estimated annual query count (71.2B) and adds a per-query inference cost derived from a daily energy figure. This treats training cost as uniformly amortized across all queries and ignores differences in model size, inference hardware, and query complexity between GPT-3 and GPT-4o-mini. Given that the title and abstract mention energy as a reason to prefer the scripted condition, the estimate should at least be presented with a sensitivity analysis (e.g., varying query counts by an order of magnitude) or explicitly labeled as an illustrative back-of-the-envelope calculation rather than a measured comparison. As written, the claim that PPS 'may have an edge' in energy is appropriately hedged, but the numerical contrast (506 Wh vs. 538–580 Wh) is not robust enough to support the emphasis it receives.","section":"Section III, Energy consumption"}],"minor_comments":[{"comment":"The notation 'GPT 4o-mini' appears without a consistent hyphen or version formatting; please unify to 'GPT-4o-mini' throughout.","section":"Section II.A"},{"comment":"The Trust scale is analyzed with a one-way ANOVA while the Godspeed subscales are analyzed with Mann-Whitney U tests because of non-normality. Please justify the parametric test for the Trust data, or report a normality check for that variable as well.","section":"Section II.C"},{"comment":"The sentence 'The lower task time ratio in the PPS condition indicates less engagement with task elements' appears to contradict the later claim that the PPS condition is 'more focused during task execution.' Please clarify what is meant by 'engagement' versus 'focus' and ensure the two statements are logically consistent.","section":"Section III"},{"comment":"The phrase 'Llama 40B' is not a standard model name; the authors likely mean 'Llama 3 70B' or 'Llama 2 70B'. Please correct the model reference.","section":"Section IV"},{"comment":"The discussion mentions 'output variability of the LLM requires prompt engineering' but does not report how many prompt iterations or what failure cases occurred. Adding a brief description of the prompt engineering process and any rejected LLM responses would strengthen the reproducibility of the LLM condition.","section":"Section IV"}],"recommendation":"major_revision","confidential_remarks":"The paper's scope fits well with a robotics/HRI venue, and the empirical comparison is useful. The main concern is that the abstract's 'focus' claim is contradicted by the quantitative gaze metrics, and the comparability claim rests on non-significance without equivalence testing. These are fixable with additional analysis and cautious rewording, so I do not recommend rejection. I would also flag that the energy estimate, while clearly labeled as a proxy, should be de-emphasized in the abstract until it is backed by a sensitivity analysis. The self-citations to the authors' prior framework [13] and PPS design [5] are legitimate building blocks and do not constitute circular evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a small within-subjects study comparing a scripted forklift interaction with a GPT-4o-mini-enhanced version. The genuinely useful finding is the efficiency and energy gap: scripted interaction is faster on task execution and uses less power, with a plausible cost estimate. That part holds up.\n\nWhat is new here is a direct empirical comparison of a scripted baseline against an LLM-controlled interaction, using gaze tracking, latency, and energy accounting on the authors' own prior framework. The data collection with mobile eye-tracking and motion capture is careful, and the energy estimate, while rough, is a sensible attempt to include training and inference costs. The paper is also honest that subjective ratings were not significantly higher for the LLM condition.\n\nThe softest spot is the 'focus' conclusion. Table I shows significantly longer total fixation duration and lower saccade velocity/amplitude for the LLM during task execution — metrics usually read as more sustained, less exploratory gaze. The paper instead claims PPS was 'more focused during task execution' based on heatmap inspection, without a quantitative dispersion metric. Maybe the heatmaps do show what the authors say, but as reported, the objective gaze metrics point the other way, and the text never reconciles this. That needs either a proper spatial dispersion measure or a softer interpretation.\n\nSecond, the 'comparable' subjective ratings rest on non-significant p-values from N=15. Without an equivalence margin or power analysis, that is absence of evidence, not evidence of equivalence. The energy comparison using GPT-3 as a proxy for GPT-4o-mini is a real but minor caveat; the direction of the result is not in doubt.\n\nOverall, the core design rule — do not add an LLM to simple repetitive tasks if you care about efficiency — is plausible and mostly supported by the data. The paper is a decent empirical contribution to HRI, not a breakthrough. It deserves a serious referee, but the authors should either quantify the focus claim properly or soften it, and add equivalence testing if they want to claim comparability in subjective measures.\n\nI would send this to peer review with a request for revision, and I would bring it to a reading group if anyone works on LLM-based HRI or energy in interactive robots. I would cite it as a reference point for the cost-benefit tradeoff of LLM control in simple industrial interactions.","headline":"A modest, useful HRI comparison whose efficiency and energy conclusions survive contact with the data, but the 'focus' claim is weaker than the abstract suggests and needs either better evidence or softer wording.","tokens_in":7933,"tokens_out":2145,"would_cite":true,"duration_ms":21984,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"For a simple pick-and-place interaction, a fully scripted robot performs comparably to an LLM-enhanced version on objective efficiency and focus, with lower latency and energy.","keywords":["human-robot interaction","large language models","scripted interaction","gaze tracking","task efficiency","engagement","energy consumption","industrial robot"],"falsifier":"A pre-registered replication with a larger sample (for example, N at least 40) and a pre-defined equivalence margin on task completion time and error rate would settle the comparability claim: if the scripted condition turns out to be significantly slower or more error-prone than the LLM condition beyond a small margin, the central conclusion fails. Separately, direct per-query power metering of the actual deployed model, rather than using GPT-3 as a proxy, could invalidate the energy edge if it shows the proxy overstates the LLM's cost.","tokens_in":6805,"feed_emoji":"🤖","tokens_out":5663,"duration_ms":56624,"temperature":0.7,"pith_summary":"This paper asks whether adding a large language model to a human-robot interaction improves measured outcomes over a plain scripted schedule. In a simple industrial pick-and-place task with a humanoid robot mounted on a forklift, the LLM condition earned slightly higher subjective trust and anthropomorphism ratings, but those differences were not statistically significant. Objective gaze-tracking metrics showed the scripted condition performing comparably on efficiency and focus, and the scripted condition had a clear edge in response latency and energy consumption. The authors conclude that interaction control should be chosen based on task complexity, with scripted control remaining a viable and cheaper default for simple, repetitive tasks.","feed_headline":"LLM-enhanced robots score higher on ratings, not on efficiency","feed_subtitle":"Gaze tracking shows scripted control is just as efficient for simple pick-and-place tasks, with lower latency and energy.","key_machinery":"The central comparison is between two interaction controllers: a pre-programmed schedule (PPS), which uses fixed timing and thresholds (for example, starting the next instruction when the participant's walking speed drops below 0.3 m/s), and an LLM-enhanced controller that feeds gaze fixations and detected objects into a large language model via an API with chain-of-thought prompting to generate robot speech, pointing, and gaze commands in real time. The metrics that carry the argument are mobile eye-tracking (fixation durations, saccade velocity and amplitude, pupil diameter), two standard questionnaires for trust and robot perception, and power measurements across local computation, network nodes, and API calls, with training and inference energy estimated from published data.","core_discovery":"The paper's central claim is that augmenting a scripted human-robot interaction with an LLM backbone does not automatically improve interaction metrics. Comparing a pre-programmed schedule (PPS) with an LLM-enhanced controller in the same task, the authors found that while the LLM condition produced slightly higher subjective ratings on trust and perceived intelligence, objective measures from gaze tracking and time allocation showed comparable or better efficiency in the scripted condition during simple task execution. The LLM condition also carried higher latency (about 2.5 s per API response) and higher estimated energy use per interaction (538–580 Wh versus 506 Wh). The authors therefore argue for aligning the interaction modality with task demands: scripted control for predictable, efficiency-driven tasks and LLM adaptation for complex, dynamic scenarios.","pith_inferences":["If the comparability claim holds, a hybrid controller that runs scripted by default and escalates to LLM reasoning only when an anomaly or unexpected state is detected could capture most of the efficiency of scripts with adaptability where it is actually needed.","The gaze differences could reflect not only engagement but also comprehension cost: longer fixations in the LLM condition might indicate that users worked harder to parse variable, generated instructions; a follow-up measuring task errors would help separate these interpretations.","Because the energy comparison substitutes GPT-3 training and inference figures for the newer model actually used, direct per-query power metering of GPT-4o-mini could shift the LLM condition's estimated energy cost in either direction.","The study's task was deliberately simple; extending the same measurement protocol to a more complex, multi-step task with unpredictable states could reveal objective benefits of LLM adaptation that the current design cannot detect."],"forward_implications":["For simple, repetitive industrial interactions, scripted control offers comparable objective efficiency with lower latency and energy consumption, making it the cheaper default.","LLM-enhanced interaction can raise subjective engagement and perceived human-likeness without producing measurable objective gains on simple tasks.","The roughly 2.5 s API response latency of the LLM condition places a practical limit on real-time interaction, which matters for time-critical robot operations.","Energy accounting that includes model training and inference shifts the cost-benefit balance against LLM deployment in resource-constrained settings.","Interaction design should be task-dependent: dynamic adaptation is valuable for complex scenarios, while predictability and focus are better for simple ones."],"supporting_citations":[{"why":"Supplies the multimodal communication design (speech, pointing, gaze) and the schedule that both conditions build on.","marker":"[5]"},{"why":"Provides the user-centric framework that combines GPT-4o-mini, eye-tracking, and object detection for the LLM condition.","marker":"[13]"},{"why":"Supplies the GPT-3 training and inference energy figures used to estimate per-query energy cost in the LLM condition.","marker":"[11]"},{"why":"Chain-of-thought prompting is the reasoning method used in the LLM condition to generate the robot's code commands.","marker":"[14]"},{"why":"The Trust in Industrial Human-Robot Collaboration scale is the trust questionnaire used to compare subjective ratings.","marker":"[15]"},{"why":"The Godspeed questionnaire measures anthropomorphism, animacy, likeability, perceived intelligence, and safety for both conditions.","marker":"[17]"},{"why":"Establishes pupil diameter as a measure of cognitive load, grounding the interpretation of that gaze metric.","marker":"[18]"}],"fun_headline_variants":["LLM robots impress users but scripted wins on efficiency","Higher ratings, not better efficiency: LLM-enhanced HRI","Scripted beats LLM on speed and energy in simple tasks","LLM-enhanced HRI: subjective win, objective tie","Scripted control matches LLM for simple pick-and-place"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that non-significant differences between conditions at N=15 can be read as evidence that the scripted condition performs comparably to the LLM condition; the study does not run an equivalence test or a power analysis to support that reading.","fun_headline_variants_meta":{"raw":{"variants":["LLM robots impress users but scripted wins on efficiency","Higher ratings, not better efficiency: LLM-enhanced HRI","Scripted beats LLM on speed and energy in simple tasks","LLM-enhanced HRI: subjective win, objective tie","Scripted control matches LLM for simple pick-and-place"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000364,"raw_usage":{"total_tokens":1943,"prompt_tokens":912,"completion_tokens":1031,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":948}},"tokens_in":528,"tokens_out":1031,"duration_ms":9151,"temperature":1.0,"reasoning_tokens":948,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:28:53.692653+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A pre-registered replication with a larger sample (for example, N at least 40) and a pre-defined equivalence margin on task completion time and error rate would settle the comparability claim: if the scripted condition turns out to be significantly slower or more error-prone than the LLM condition beyond a small margin, the central conclusion fails. Separately, direct per-query power metering of the actual deployed model, rather than using GPT-3 as a proxy, could invalidate the energy edge if it shows the proxy overstates the LLM's cost.","supporting_citations":[{"cited_title":"Advantages of Multimodal versus Verbal-Only Robot-to-Human Communication with an Anthropomorphic Robotic Mock Driver,","cited_arxiv_id":null,"evidence_quote":"Supplies the multimodal communication design (speech, pointing, gaze) and the schedule that both conditions build on."},{"cited_title":"Bidirectional Intent Communication: A Role for Large Foundation Models","cited_arxiv_id":"2408.10589","evidence_quote":"Provides the user-centric framework that combines GPT-4o-mini, eye-tracking, and object detection for the LLM condition."},{"cited_title":"The development of a scale to evaluate trust in industrial human-robot collaboration,","cited_arxiv_id":null,"evidence_quote":"The Trust in Industrial Human-Robot Collaboration scale is the trust questionnaire used to compare subjective ratings."},{"cited_title":"Measurement in- struments for the anthropomorphism, animacy, likeability, perceived intelligence, and perceived safety of robots,","cited_arxiv_id":null,"evidence_quote":"The Godspeed questionnaire measures anthropomorphism, animacy, likeability, perceived intelligence, and safety for both conditions."},{"cited_title":"Pupil diameter and load on memory,","cited_arxiv_id":null,"evidence_quote":"Establishes pupil diameter as a measure of cognitive load, grounding the interpretation of that gaze metric."}],"review_version":1}