{"id":"68f10aee-6612-47a7-81cd-a691a6afe39e","arxiv_id":"2508.05104","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In a touchscreen pointing task with the NICO humanoid, truncated arm movements plus congruent gaze let human observers predict the intended target better than pointing alone.","lead":"This paper tests whether human observers can predict a humanoid robot's intended pointing target from partial arm movements and gaze cues. It reports that both multimodal cues (pointing plus congruent gaze) and gaze alone improve legibility, supporting two long-standing hypotheses about human-robot communication.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ocular primacy results may be confounded by gaze-induced changes in arm kinematics; no kinematic control is reported in the abstract.","rationale":"The reader correctly identified that the abstract provides insufficient information to verify the claims, and I agree that the verdict should be UNVERDICTED (hence UNCHANGED). However, the reader's weakest assumption focused on ecological validity of the forced-choice task. I believe the more immediate and load-bearing concern is internal validity: the abstract's description of the conditions allows a plausible confound between gaze condition and arm kinematics. This is a sharper, more technical threat to the central claim because it could invalidate the experiment's construct even within its own controlled setting. The concrete test would settle this: if arm kinematics are unaffected, the central claim remains plausible (though still unverified due to missing statistical details). If kinematics are affected, the ocular primacy result may be an artifact. This is a good-faith concern grounded in the known mechanics of humanoid robots and the need for stimulus control in psychophysical experiments.","tokens_in":587,"tokens_out":2304,"duration_ms":31418,"concrete_test":"Obtain the recorded joint/end-effector trajectories for the three conditions that involve pointing (pointing alone, congruent, incongruent). Compare the arm trajectory features up to the truncation point (e.g., direction of the end-effector velocity vector, curvature, velocity profile) across conditions using repeated-measures ANOVA or mixed-effects models. If a significant difference is found between congruent and incongruent gaze conditions in any kinematic feature that could inform target prediction, the accuracy difference attributed to gaze is confounded; if no difference is found, the concern is dismissed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that gaze is a dominant cue for intent prediction rests on comparing prediction accuracy across conditions (gaze, pointing, pointing with congruent/incongruent gaze) using truncated arm trajectories. A load-bearing threat is that the robot's gaze direction (head orientation) might alter the arm kinematics themselves. In humanoid platforms like NICO, head movement can shift posture, starting pose, or trajectory shape due to mechanical coupling or controller state. If the arm trajectory at 60% or 80% truncation differs systematically between congruent and incongruent gaze conditions (e.g., initial direction bias or velocity changes), then participants may be responding to these kinematic differences rather than to gaze as an informational cue. The abstract does not state whether arm kinematics were held constant across gaze conditions or whether they were analyzed as a control variable. Without such a control, the multimodal superiority and ocular primacy hypotheses are not cleanly tested: the independent variable (gaze) is confounded with a potential change in the pointing signal itself. This is a testable internal-validity concern, not merely an ecological-validity caveat.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports an experiment on the legibility of humanoid robot arm movements in a pointing task. Using the NICO humanoid robot, participants observed arm movements directed at a touchscreen target under conditions that varied gaze, pointing, and combined pointing with congruent or incongruent gaze. Arm trajectories were truncated at 60% or 80% of their full length, and participants performed a forced-choice target prediction. The authors state that they tested the multimodal superiority hypothesis and the ocular primacy hypothesis, and that both were supported. This review is based on the abstract only; the full text is not available.","tokens_in":857,"tokens_out":2767,"duration_ms":35020,"significance":"If the reported findings hold, the experiment provides a useful empirical contribution to human-robot interaction, suggesting that pointing combined with congruent gaze improves human prediction of a robot's intended target and that gaze is a dominant cue. The study is clearly framed around two testable hypotheses, and the forced-choice prediction task is a straightforward operationalization of legibility. No circularity or parameter-fitting concerns arise from the abstract. However, because the abstract omits all quantitative evidence and does not address a plausible kinematic confound, the significance is conditional on the full manuscript resolving these issues.","major_comments":[{"comment":"The central claim that “both hypotheses were supported” is given with no effect sizes, confidence intervals, p-values, condition means, or participant numbers. This is not merely a stylistic omission: the claim is the entire contribution of the abstract, and without quantitative support it cannot be evaluated. If the full manuscript reports these statistics, the abstract should state at least the key effect sizes and inferential results. If the full manuscript does not, the claim is unsupported.","section":"Abstract"},{"comment":"A load-bearing internal-validity threat is not addressed. The ocular primacy hypothesis is tested by comparing pointing with congruent versus incongruent gaze. On a humanoid platform, head or gaze movement can mechanically influence arm posture, starting pose, or trajectory shape. If the arm kinematics at the 60% or 80% truncation points differ systematically between congruent and incongruent gaze conditions, participants could be responding to kinematic artifacts rather than to gaze as an informational cue. The abstract does not state whether arm trajectories were held constant across gaze conditions, whether gaze was realized independently of arm movement, or whether kinematic differences were analyzed as a control. This must be reported and, if necessary, controlled in the analysis.","section":"Abstract (experimental conditions)"},{"comment":"Legibility is operationalized as forced-choice prediction accuracy on truncated arm trajectories in a controlled touchscreen task. The manuscript should justify why this task reflects the legibility construct relevant to real human-robot coordination, including trust and perceived safety. The abstract does not provide such justification. This is not a fatal flaw, but it is a significant limitation that should be explicitly discussed if the authors intend the conclusions to generalize beyond the experimental setting.","section":"Abstract (operationalization)"}],"minor_comments":[{"comment":"The term “incongruent gaze” is used but not defined. Does it mean gaze directed at a different target than the pointing arm, or gaze toward the wrong target while pointing toward the correct one? Clarify.","section":"Abstract"},{"comment":"The truncation levels (60% and 80%) are stated, but it is not clear whether these were manipulated within participants or between participants. This information belongs in the methods section and should be included in the full text.","section":"Abstract"},{"comment":"The phrase “gaze” is ambiguous: it could refer to the robot's head orientation, eye direction, or both. Since ocular primacy is a central claim, the physical realization of gaze should be specified.","section":"Abstract"},{"comment":"Participant details (e.g., sample size, recruitment, prior experience with robots) are absent from the abstract. This is acceptable for a journal abstract, but the full manuscript must report these details, and the abstract should at least include the sample size.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The abstract is too thin to support a definitive assessment. The decisive issue for the full manuscript will be whether the authors can rule out the kinematic-confounding explanation for ocular primacy. I recommend that the editor obtain the full text and specifically check whether arm kinematics were held constant across gaze conditions or statistically controlled. The novelty relative to existing legibility literature in HRI should also be evaluated from the full text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a reasonable, incremental study—testing two established hypotheses (multimodal superiority, ocular primacy) on a new platform (NICO) with a truncated-pointing task. If the results hold, it gives robot designers a simple, actionable cue policy: point and gaze congruently when movements are only partially observable. That is practical and worth knowing.\n\nThe abstract itself is clear and honest about what was done: participants saw arm trajectories stopped at 60% or 80%, cues varied across gaze/pointing/congruent/incongruent, and both hypotheses were supported. I believe that claim could be true. But the abstract gives no effect sizes, no participant counts, no baseline condition, and no indication of variability. That makes the central claim impossible to assess from this material alone. The paper needs a full statistical report before I would trust the conclusion.\n\nThe more specific concern, which I think is the right one to raise, is the gaze-arm coupling issue. On a humanoid like NICO, head orientation or gaze direction may shift posture and therefore the arm trajectory itself. If the arm trajectories in the congruent and incongruent gaze conditions are systematically different at the truncation points, then participants could be responding to kinematic differences rather than to gaze as an informational cue. The abstract does not state whether arm trajectories were matched across conditions or analyzed as a control. This is not an ecological-validity quibble; it is a direct threat to the internal validity of the ocular primacy result. The full paper must either show the trajectories are statistically indistinguishable conditioned on the target or include them as a covariate.\n\nOne minor note: the forced-choice touchscreen task is a fine simplification, but its connection to real-world legibility (trust, safety, coordination) is asserted rather than demonstrated. That is a weaker concern and not something I would hold against the paper if the kinematic issue is handled.\n\nBottom line: this is a paper for the HRI/legibility community. It deserves peer review, not desk rejection, because the design is concrete, cheap to replicate, and addresses a real design question. But my own assessment stays open until I see the kinematic checks and the full statistics. Send it to review, and make sure the reviewers ask exactly those two questions.","headline":"A modest, plausibly useful HRI extension that is unverifiable from the abstract alone and needs a kinematic-confound check before the headline conclusions can be trusted.","tokens_in":1306,"tokens_out":1494,"would_cite":false,"duration_ms":22679,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Human observers infer a humanoid robot's target most reliably when gaze and pointing are aligned, and gaze dominates when the two cues conflict.","keywords":["human-robot interaction","legibility","gaze cue","pointing gesture","intention prediction","humanoid robot","truncated trajectories","multimodal cues"],"falsifier":"Run the same truncated-trajectory task with the robot's gaze hidden (for example, eye lights off or a head covering). If prediction accuracy with gaze-plus-pointing is no better than with pointing alone, the multimodal advantage rests entirely on visible gaze; if the advantage persists, some other cue, such as torso orientation, is doing the work.","tokens_in":575,"feed_emoji":"👀","tokens_out":4821,"duration_ms":57267,"temperature":0.7,"pith_summary":"The paper tests a practical question: when a humanoid robot reaches toward a screen, what makes its intention legible to a person watching? In the experiment, the NICO robot's arm movements toward touchscreen targets were shown as videos cut off at 60% or 80% of the full trajectory, and human participants had to predict which target the robot meant. The study found that pairing gaze with pointing improved prediction accuracy over pointing alone, and that when gaze and pointing disagreed, people predicted the target the robot was looking at. The authors read this as evidence for two principles: adding a congruent cue helps multimodal legibility, and gaze has priority over the arm.","feed_headline":"Trust a robot's gaze over its pointing arm","feed_subtitle":"Congruent gaze plus pointing beats pointing alone, and gaze wins when they conflict.","key_machinery":"The key machinery is the controlled contrast among cue conditions—pointing only, pointing plus congruent gaze, pointing plus incongruent gaze—presented as truncated arm trajectories to human raters. The truncation at 60% and 80% forces the rater to predict before seeing the full movement, and the congruent-versus-incongruent gaze contrast is what separates the added-value of gaze from its dominance.","core_discovery":"The central claim is that human onlookers can read a humanoid robot's intended target from a partial arm movement, and they do so most reliably when the robot combines its pointing arm with gaze that is aligned to the same target. In the reported forced-choice task, trajectories truncated at 60% or 80% were still legible, and the congruent gaze-plus-pointing condition produced higher prediction accuracy than pointing alone. When gaze and pointing were put in conflict, predictions followed gaze rather than pointing, supporting an ocular primacy account. The authors state that both hypotheses—multimodal superiority and ocular primacy—were supported by the experiment.","pith_inferences":["A gaze-only condition (no visible pointing arm) would let researchers size the two channels separately; the present experiment only gives the combined and conflicting cases, so the paper's ocular primacy result is comparative, not a standalone gaze-effect estimate.","Prediction accuracy probably understates real-world legibility gains, since live interaction also rewards earlier prediction and lower attentional effort; a reaction-time or gaze-following measure would likely show larger effects than the accuracy margin.","The cue-conflict design resembles standard multisensory conflict paradigms, suggesting the same logic could test whether ocular primacy persists under time pressure or with a more humanlike robot face, where social expectations might alter cue weighting."],"forward_implications":["Designers of humanoid robots can make intent clear by coordinating gaze with the pointing arm, rather than by exaggerating arm kinematics alone.","A gaze that points at a different target than the arm is an active source of misreading, not a neutral omission, since observers weight it above the arm.","Legibility can be measured as a quantitative behavioral score—prediction accuracy on truncated trajectories—which gives a comparable metric for different motion generation algorithms.","Trajectory truncation acts as a difficulty dial: the difference between 60% and 80% reveals how much of a motion must be visible before the human can commit to a prediction."],"supporting_citations":[],"fun_headline_variants":["Gaze beats pointing when robot cues conflict","Pointing plus gaze boosts robot legibility","Robot gaze outranks pointing for intent","Partial arm motion still reveals robot target"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The forced-choice accuracy score, taken from brief videos of a stationary robot with movements cut at fixed percentages, is treated as a valid measure of the legibility that matters in real human-robot collaboration, including trust and perceived safety.","fun_headline_variants_meta":{"raw":{"variants":["Gaze beats pointing when robot cues conflict","Pointing plus gaze boosts robot legibility","Robot gaze outranks pointing for intent","Partial arm motion still reveals robot target"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1225,"prompt_tokens":626,"completion_tokens":599,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":370,"completion_tokens_details":{"reasoning_tokens":545}},"tokens_in":370,"tokens_out":599,"duration_ms":6946,"temperature":1.0,"reasoning_tokens":545,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:30:57.840443+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same truncated-trajectory task with the robot's gaze hidden (for example, eye lights off or a head covering). If prediction accuracy with gaze-plus-pointing is no better than with pointing alone, the multimodal advantage rests entirely on visible gaze; if the advantage persists, some other cue, such as torso orientation, is doing the work.","supporting_citations":[],"review_version":1}