{"id":"1afb7216-6014-4699-a60f-417547d661a9","arxiv_id":"2412.20654","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"In a hybrid human-robot pyramid-building task, higher cognitive load, induced by harder visual conditions, was associated with increased self-reported and behavioral trust in the robot.","lead":"This study runs an experiment in which 54 people build block pyramids together with a robot under three difficulty levels designed to change mental workload. It reports that people trust the robot more and earn higher rewards under the hardest, inverted-camera condition, and that trust correlates with a computed failure-risk score in easier conditions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"High-load condition confounds cognitive load with inverted-camera spatial distortion; the trust increase may reflect disorientation rather than workload, and Table 1's negative within-high-load correlations contradict the proposed causal mechanism.","rationale":"The reader identified the confound between cognitive load and visual/spatial manipulation as the weakest assumption; I agree and find it the most load-bearing concern because it undermines the causal interpretation of the central claim. The paper has real strengths: a within-subject design with randomized session order, a sample of 54 participants, a post-hoc power analysis above 0.8, and a manipulation check confirming that the high-load condition elevates reported mental demand. However, the design cannot separate mental workload from the visuomotor distortion introduced by inverted camera frames, and the paper's own Table 1 provides direct evidence that the proposed mechanism does not hold within the high-load condition. The negative within-high-load regressions are not merely a failure to find a positive correlation; they are statistically significant in the opposite direction, which is internally inconsistent with the claim that higher cognitive load increases trust. This reinforces the reader's concern without requiring new external assumptions. The concrete 2x2 experiment is the cleanest way to settle whether spatial inversion alone produces the trust increase; a mediation analysis on the existing data could provide a faster, though less decisive, check if raw data are available. Given the otherwise transparent reporting and reasonable sample, the appropriate verdict remains CONDITIONAL: the authors should add a control condition or alternative load manipulation that does not alter visuomotor mapping, provide and validate the dictator-game protocol, and share data and code. I therefore see no reason to change the reader's verdict, but I want the conditions to explicitly include the disentangling of inversion from cognitive load.","tokens_in":1197,"tokens_out":865,"duration_ms":88103,"concrete_test":"Run a 2x2 within-subject experiment crossing visual modality (direct observation vs camera screen) with frame orientation (normal vs inverted), measuring NASA-TLX as the manipulation check and the same Muir/dictator-game trust measures. If the orientation main effect on trust is significant, especially when camera modality is held constant, the reported high-load effect is attributable to spatial inversion rather than cognitive load. Alternatively, re-analyze the existing raw data with a mediation analysis treating NASA-TLX change as the mediator between condition (low vs high) and trust change; the significantly negative within-high-load regressions in Table 1 predict a non-significant or negative indirect effect, which would refute the abstract's causal claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.4 manipulates cognitive load by changing task complexity in three ways: low-load uses direct observation, middle-load restricts perception to two camera feeds, and high-load inverts both camera frames. The high-load condition therefore differs from middle-load by orientation and from low-load by both modality and orientation. The independent variable is not just cognitive load; it also includes spatial inversion, which can independently reduce the participant's sense of control, alter perceived robot competence, and induce disorientation. The NASA-TLX check in Section 4.4 confirms that participants report higher mental demand in the high-load condition, but that only validates the intended manipulation; it does not rule out these alternative channels. A more direct problem appears in the paper's own Table 1: within high-load tasks, the regression of trust change on cognitive load change is significantly negative for both the Muir questionnaire (−0.7114, p<0.05) and the dictator game (−1.3687, p<0.05). This is the opposite of the between-condition effect claimed in the abstract; participants who experienced larger increases in cognitive load showed smaller increases in trust. The paper acknowledges this in Section 4.1 ('the highest levels of trust ... are associated with relatively low cognitive load states in high-load tasks') but does not resolve the contradiction. If cognitive load were the active ingredient, one would expect a positive within-condition association. The negative sign, combined with the confounded manipulation, makes the central causal claim that high cognitive load increases trust unsupported by the current design.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a 54-participant experiment on how cognitive load affects human trust in a hybrid human-robot collaboration task, in which a human and a robot jointly stack blocks into a pyramid with interdependent steps. Cognitive load is manipulated through three task-complexity conditions: direct observation (low), camera-only feedback (middle), and inverted camera feedback (high). Trust is measured with a self-report Muir questionnaire and a dictator-game allocation, and performance is measured by accumulated rewards and a newly proposed Failure Risk (FR) metric. The main claims are that high cognitive load increases trust, that rewards are higher under high load, and that trust correlates with the failure-risk metric in low- and middle-load successful tasks. The authors also report a post-hoc power analysis and discuss implications for interface design and collaboration-target selection.","tokens_in":16225,"tokens_out":3468,"duration_ms":38749,"significance":"If the central claims hold, the paper would make a useful contribution to human-robot trust research by shifting attention to hybrid, interdependent-task scenarios and by combining subjective and behavioral trust measures. The task design, with interdependent steps and a clear manipulation of perceptual-motor difficulty, is a genuine strength, and the paper offers a falsifiable prediction about delegation under high workload. However, the reported statistical evidence is fragile: the headline contrasts are supported by borderline p-values without multiple-comparison correction, the manipulation is confounded with visual modality and spatial inversion, and a within-condition analysis in the paper's own Table 1 points in the opposite direction to the between-condition claim. The failure-risk metric is also author-defined with an arbitrarily chosen discount factor. These issues undermine the causal interpretation until addressed.","major_comments":[{"comment":"The cognitive-load manipulation is confounded with visual modality and spatial orientation. The low-load condition uses direct observation, the middle-load condition uses two camera feeds, and the high-load condition inverts both camera frames. Thus high load differs from middle load by orientation, and from low load by both modality and orientation. The NASA-TLX check in §4.4 confirms that reported load differs, but it does not rule out alternative channels: inverted vision can induce disorientation, reduce perceived control, or alter expectations about the robot's competence independently of workload. The abstract's causal claim ('cognitive load exerts diverse impacts on human trust') therefore is not identified by this design. The authors should either add a control condition that varies load without inverting the visual frame, or reframe the conclusion in terms of task complexity and perceptual-motor distortion rather than cognitive load per se.","section":"§3.4 and §4.4"},{"comment":"The within-condition regressions in Table 1 contradict the proposed causal mechanism. In the high-load condition, changes in trust are significantly negatively associated with changes in cognitive load for both the Muir questionnaire (-0.7114, p<0.05) and the dictator game (-1.3687, p<0.05). The paper acknowledges this in §4.1 ('the highest levels of trust ... are associated with relatively low cognitive load states in high-load tasks') but does not reconcile it with the between-condition claim that high load increases trust. If cognitive load were the active ingredient, one would expect a positive within-condition association. At minimum, the authors need to explain this sign reversal (e.g., ceiling effects, nonlinearity, or distinct between-person and within-person processes) and show that the headline effect is not an artifact of the between-condition comparison.","section":"§4.1, Table 1"},{"comment":"The paper relies on borderline p-values without correction for multiple comparisons. The high-load vs. low-load trust differences are p=0.0494 and p=0.0262; the behavioral trust differences are p=0.0231, p=0.0209, and p=0.005; the reward difference is p=0.0389; the cognitive-load difference is p=0.041. Given three task conditions, two trust measures, multiple questionnaire subscales, and multiple performance metrics, the family-wise error rate is substantial. The post-hoc power analysis in §4.5 assumes a medium effect size (f=0.25) that does not correspond to any reported effect size and does not address the pairwise tests that support the headline claims. The authors should report adjusted p-values or a pre-specified analysis plan, and they should justify the power analysis as a sensitivity analysis rather than evidence that the pairwise contrasts were adequately powered.","section":"§4.1, §4.2, §4.4, §4.5"},{"comment":"There is a serious inconsistency in the reported sample size. The text in §3.1 and Table 1 state 54 participants, while the means and standard deviations in Appendix B and Appendix D are exactly those of an 11-person sample (e.g., 5.0909 = 56/11, 3.2727 = 36/11). If the appendices are based on a different subset, this must be stated; if the full sample is 54, the appendix means are incorrect. This discrepancy affects the credibility of all descriptive statistics and the power analysis. The authors must clarify and correct the reported N.","section":"§3.1, Table 1, Appendix B, Appendix D"},{"comment":"The failure-risk metric is an author-defined quantity whose validity is not established. The FRV in Eq. (3-2) depends on the horizontal offset between block centers and support centers, and the pyramid-level FR in Eq. (3-3) uses a discount factor gamma that is simply set to 0.8 with no empirical or theoretical justification. The claim in Table 3 that trust correlates with failure risk in low- and middle-load tasks is therefore a correlation with an unvalidated, parameter-dependent outcome. The authors should provide evidence that FRV/FR measures actual physical failure risk (e.g., convergence with observed collapses or near-falls, robustness to gamma), or at least present results across a range of gamma values to show that the conclusion is not an artifact of this particular choice.","section":"§3.6, Eq. (3-2), Eq. (3-3), Table 3"},{"comment":"The dictator game is not a standard measure of trust. As the authors themselves note in §2.2, most prior work uses questionnaires; they describe the dictator game as an 'objective measure of participants’ trust levels,' but allocations in a dictator game are typically interpreted as measures of altruism or fairness rather than trust in another agent. The paper does not validate that allocations to the robot track trust in the robot as distinct from generosity or experimenter-demand effects. The authors should either justify this measure with a pilot validation, cite evidence that dictator-game transfers are trust-sensitive in human-robot contexts, or rename the construct as behavioral delegation and temper the conclusions drawn from it.","section":"§2.2 and §3.5"}],"minor_comments":[{"comment":"The text in §4.3 says detailed regression analyses are provided in 'Appendix D,' but Appendix D contains cognitive-load descriptive statistics; the regression details are actually in Appendix C. Please correct the cross-reference.","section":"§4.3 and appendices"},{"comment":"Please report effect sizes and confidence intervals for the pairwise ANOVAs, not only p-values. This is especially important because several p-values are just below 0.05 and would not survive correction for multiple comparisons.","section":"§4.1"},{"comment":"The paper states that the order of the three sessions was randomized, but it does not report any check for order effects or learning effects. Given the repeated-measures design, a brief analysis of session order as a covariate would strengthen the results.","section":"§3.4"},{"comment":"Please clarify whether the NASA-TLX was scored with the standard weighted procedure or as an unweighted average. The appendix reports subscale means plus a 'Weighted' row, but the method does not explain how weights were obtained.","section":"§3.5"},{"comment":"The notation in Eq. (3-2) and Eq. (3-3) is difficult to parse because of subscript rendering; please use clear subscripts such as x_e and x_s, and define all variables in the text immediately after the equations.","section":"§3.6"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an interesting question and the hybrid-task setup is a strength, but the statistical and construct-validity issues are substantial. The confound between cognitive load and visual inversion is a design-level problem that cannot be fully corrected post hoc; at minimum the authors would need to reframe their claims and add sensitivity analyses. The within-condition sign reversal in Table 1 is especially concerning and should be a central focus of any revision. If the authors can address the confound with additional data or analyses, and fix the sample-size discrepancy, the paper could become publishable; otherwise the central causal claim is not supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is a genuine attempt to study trust in a hybrid HRC setting where human and robot are co-equal and steps are interdependent. The pyramid-stacking task is a real departure from the usual single-operator paradigms, and the authors deserve credit for measuring trust two ways—Muir questionnaire and a dictator game—and for checking the manipulation with NASA-TLX. The procedure is transparent enough to reproduce, and the sample of 54 is reasonable for a between-condition within-subject design. There is a real idea here: workload might push people to delegate more to a robot they otherwise don't know well. But the central causal claim is not supported as cleanly as the abstract and conclusions suggest.\n\nThe biggest problem is the manipulation. The high-load condition inverts both camera frames, so it differs from the middle-load condition by spatial orientation and from the low-load condition by both modality and orientation. Cognitive load is not isolated. The NASA-TLX check shows people feel more mentally demand in the high-load condition, but that does not rule out disorientation or perceived robot competence as alternative channels. More seriously, the paper's own Table 1 shows that within the high-load condition, changes in trust are significantly negatively correlated with changes in cognitive load for both the Muir questionnaire (-0.7114) and the dictator game (-1.3687). That is the opposite of the between-condition effect. The text acknowledges this in Section 4.1 but never resolves it. If cognitive load were the active ingredient, you would expect a positive within-condition slope. The reader's stress-test note lands: the design confounds load with visual inversion, and the within-condition data point the other way.\n\nThe statistics are also thinner than the prose. The key pairwise ANOVAs are borderline (p = 0.0494, 0.0262, 0.041, 0.0209, 0.005) with no multiple-comparison correction. The dictator game protocol is underspecified—what exactly is sent to the robot, how many tokens, and what the allocation represents in terms of trust? Failure risk (FRV with gamma = 0.8) is an author-defined, unvalidated metric; the correlation between trust change and failure risk in low/mid-load tasks is interesting but rests on that arbitrary construction. No data or code are provided.\n\nProportionately: this is not a bad paper, and it is not a malicious one. It is a decent empirical contribution with a confounded manipulation and an overstated summary. A careful revision—ideally a control condition that varies workload without spinning the camera, or at least a covariate analysis separating load from disorientation—could make the finding credible. The dictator game needs a clear protocol and validation. The abstract's claim that high cognitive load increases trust should be softened to: in this particular inverted-vision condition, participants reported more trust and delegated more, but the design does not isolate the cause.\n\nWho should read this? Researchers working on trust, workload, and autonomy allocation in HRI will want to know about it, mainly because it is a rare hybrid-teammate setup with a behavioral measure. For peer review: yes, a serious editor should send this out. It deserves referee time, but it needs major revision before acceptance. The result, if replicated with a cleaner manipulation, would be a useful data point rather than a field-changer.","headline":"Plausible but confounded: the high-load condition changes vision and orientation along with workload, and the paper's own within-condition regressions contradict its headline claim that cognitive load increases trust.","tokens_in":16774,"tokens_out":1688,"would_cite":false,"duration_ms":20315,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In hybrid human-robot collaboration with interdependent steps, high cognitive load increases trust in the robot, raises rewards, and links trust to lower failure risk in easier tasks.","keywords":["cognitive load","human trust","human-robot collaboration","hybrid human-robot collaboration","behavioral trust","dictator game","failure risk","task complexity"],"falsifier":"Run the same pyramid-stacking task with a high-load condition that raises mental workload without changing the visual feedback or the robot's apparent competence—for instance, completing a concurrent n-back task while watching normal camera views—and compare trust against the low-load condition; if the trust increase does not appear, the paper's causal interpretation is refuted. A second check is to measure perceived robot competence separately in the inverted-camera condition and test whether that perception, rather than the NASA-TLX workload score, mediates the trust effect.","tokens_in":15706,"feed_emoji":"🤖","tokens_out":7563,"duration_ms":68735,"temperature":0.7,"pith_summary":"This paper sets out to show that, in hybrid human-robot collaboration—where a human and a robot are equal teammates on a task whose steps build on one another—cognitive load does not simply damage trust the way earlier single-operator studies suggested. In a pyramid-stacking experiment with 54 participants, the authors found that the highest mental workload (watching inverted camera views) made people trust the robot more, both on a questionnaire and in a behavioral economic game where they handed money to the robot. They also report higher performance rewards in high-load tasks and, in successful low- and medium-load tasks, a significant correlation between change in trust and a computed failure-risk score. If taken at face value, the result means operators lean on their robot partners exactly when their own cognitive resources are stretched, and it gives interface designers a concrete reason to calibrate autonomy and target selection to workload rather than assuming trust falls under pressure.","feed_headline":"High cognitive load boosts trust in a robot teammate","feed_subtitle":"In a joint pyramid-building task, people delegated more and earned more when mental workload was highest.","key_machinery":"The mechanism that carries the argument is a joint pyramid-stacking task whose ten steps are interdependent: the human and robot alternately place blocks, and the placement of each block depends on what came before. Cognitive load is varied across three sessions by changing the visual channel: direct observation (low), camera feedback only (middle), and inverted camera frames (high), with each step time-limited and rewards tied to the height of the layer on which a block lands. Trust is captured twice—subjectively with the Muir questionnaire and behaviorally with a dictator game—while joint performance is scored by accumulated rewards and by a Failure Risk Value (FRV) that assigns each block a stability penalty from its horizontal offset relative to its support and discounts earlier placements geometrically (factor 0.8), so the final failure risk reflects the whole building history.","core_discovery":"On the paper's own terms, the discovery is that task-induced cognitive load has heterogeneous effects on trust in a robot teammate within interdependent joint tasks: high load increases both self-reported trust (Muir questionnaire) and behavioral trust (dictator game allocation) relative to low load, with significant differences in the changes in trust between high- and low-load sessions. The paper further reports that rewards are substantially higher in high-load tasks than in low-load tasks, and that in successful low- and middle-load tasks, increases in human trust are significantly correlated with lower failure risk, a metric that discounts each block's deviation from its support by a time factor. The conclusion is that the hybrid, interdependent scenario produces trust dynamics opposite to those often reported for independent, single-operator tasks.","pith_inferences":["The paper's high-load condition inverts the camera image, changing both workload and the perceived spatial frame; a control that raises cognitive load without distorting the robot's apparent competence (for example, a concurrent memory load) would confirm that the trust increase is driven by workload rather than perceptual confusion.","The dictator game measures willingness to give money to the robot after the task, which may reflect generalized trust or reciprocity rather than trust in this specific teammate; observing override frequencies inside the task would test whether the load effect is specific to the partnering robot.","If the failure-risk formula were calibrated against actual collapse rates, the trust-risk correlation found here could become a predictive design metric for choosing when to hand control to the robot; the paper does not test that calibration.","The results hint that adaptive autonomy—letting the robot take over when workload is high—could improve performance, but also raises the risk of overtrust; this adaptive policy is a natural next step the paper does not implement."],"forward_implications":["If high cognitive load raises trust in a robot teammate, interface designers should expect operators to delegate more and intervene less precisely when mental workload peaks.","In interdependent tasks, placing more trust under load may be adaptive: the paper's data tie high-load sessions to higher rewards and, in successful low/medium-load sessions, to lower computed failure risk.","The failure-risk metric offers a continuous, history-sensitive performance signal that can be used alongside or instead of reward counts to evaluate joint human-robot performance and to tune trust-calibrating feedback.","Because the significant trust increases appear between high and low load rather than as a smooth gradient across all three levels, workload-sensitive trust models should treat the relationship as nonlinear rather than assuming a uniform slope.","The reward advantage under high load suggests that autonomy-leaning strategies in hard conditions may improve joint outcomes, at least for manipulation tasks with interdependent steps."],"supporting_citations":[{"why":"Shows that error frequency erodes trust, the robot-factor baseline the paper's load-driven finding is contrasted with.","marker":"Desai et al., 2012"},{"why":"Supplies the trust-game paradigm and evidence that cognitive load alters trusting behavior.","marker":"Samson & Kostyszyn, 2015"},{"why":"Documents the standard finding that high load reduces trust in robot-assisted tasks, which the hybrid task set out to qualify.","marker":"Daronnat et al., 2021"},{"why":"Establishes cognitive load as a trust-influencing factor in shared-space human-robot collaboration.","marker":"Hopko et al., 2022"},{"why":"Provides evidence that people rely more on a virtual assistant under high cognitive load, used to explain the increased delegation.","marker":"Gupta et al., 2020"},{"why":"Documents automation overtrust under task load, supporting the paper's interpretation of higher trust as reliance.","marker":"Biros et al., 2004"},{"why":"Provides the validated trust-in-automation questionnaire used for subjective trust.","marker":"Muir, 1994; Muir & Moray, 1996"},{"why":"Provides the NASA-TLX workload questionnaire used to confirm the load manipulation.","marker":"Hart, 2006"},{"why":"Supplies the economic 'certain money' preference that survives cognitive load, invoked to explain higher rewards in high-load sessions.","marker":"Dominiak & Duersch, 2024"}],"fun_headline_variants":["Mental overload makes people trust robots more","Heavy thinking increases trust in robot partners","High cognitive load raises trust in robot teammates","Cognitive strain triggers higher trust in robots"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the three task-complexity conditions differ only in cognitive load and not in anything else that affects trust—in particular, the high-load condition flips the camera image, which can disorient users and change how capable the robot seems, so if the trust increase comes from that perceptual change rather than from mental workload, the causal conclusion would collapse.","fun_headline_variants_meta":{"raw":{"variants":["Mental overload makes people trust robots more","Heavy thinking increases trust in robot partners","High cognitive load raises trust in robot teammates","Cognitive strain triggers higher trust in robots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000549,"raw_usage":{"total_tokens":2591,"prompt_tokens":882,"completion_tokens":1709,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":1656}},"tokens_in":498,"tokens_out":1709,"duration_ms":12366,"temperature":1.0,"reasoning_tokens":1656,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:14:44.960325+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same pyramid-stacking task with a high-load condition that raises mental workload without changing the visual feedback or the robot's apparent competence—for instance, completing a concurrent n-back task while watching normal camera views—and compare trust against the low-load condition; if the trust increase does not appear, the paper's causal interpretation is refuted. A second check is to measure perceived robot competence separately in the inverted-camera condition and test whether that perception, rather than the NASA-TLX workload score, mediates the trust effect.","supporting_citations":[],"review_version":1}