{"id":"d946e217-15d4-40b0-abde-0b94539779c6","arxiv_id":"2506.19179","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a 32-person experiment, situational video context dominated the perceived valence of a robot's haptic feedback, while the haptic signal dominated perceived arousal.","lead":"In a 32-person study, people judged a robot's vibration feedback differently depending on what they had just seen happen to the robot, but the vibration itself set how intense the emotion felt. The result is a simple design rule for robots and wearable devices that use touch to express emotion.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The arousal-override claim is unsupported because non-significance is treated as dominance; the paper needs an equivalence or model-comparison test before this headline finding can stand.","rationale":"The reader's weakest-assumption analysis correctly identifies the same load-bearing concern: the arousal-override conclusion relies on null differences rather than a formal dominance or equivalence test. My stress-test confirms this and sharpens it by noting that the relevant comparison, video-only versus video-plus-haptic, is partially reported in Section III.D.1 and does not consistently support haptic dominance, especially for the comfort stimulus. The valence asymmetry finding is better supported by the linear mixed-effects model on absolute deviation from neutral valence, and the tactile-only validation is a useful positive control. The statistical reporting issues (Kruskal-Wallis on repeated measures, rank-sum instead of signed-rank) add uncertainty but are secondary to the main inferential gap. Because the central claim about arousal override is not established, the manuscript should remain conditional pending a reanalysis that explicitly tests modality dominance. This does not change the reader's verdict, so I recommend UNCHANGED.","tokens_in":9072,"tokens_out":4201,"duration_ms":53485,"concrete_test":"Reanalyze the arousal ratings with a linear mixed model that includes condition (video-only, haptic-only, video-plus-haptic) and situational context as fixed effects and participant as a random intercept. Then run a two one-sided tests (TOST) equivalence procedure with a pre-specified bound (e.g., ±1 SAM unit) comparing video-plus-haptic against haptic-only, and a standard significance test comparing video-plus-haptic against video-only. If the combined condition is equivalent to haptic-only and significantly different from video-only in most contexts, the override claim survives; if not, the conclusion should be weakened to 'haptic feedback shifts arousal ratings' and the word 'override' should be removed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'haptics override the participants' perception of arousal of the video' is built on a logical gap in Section III.C.1: the authors observe no significant differences between haptic-only and video-plus-haptic arousal ratings and then conclude that haptics override context. With N=32 and within-subject variability, a null p-value does not establish equivalence or dominance. The relevant test for override is whether the combined rating is equivalent to the haptic-only rating and significantly different from the video-only rating. The paper's own Section III.D.1 reports that for comfort haptics only Slap and Stroke differed significantly from video-only conditions, meaning that in most contexts the combined comfort condition was statistically indistinguishable from both video-only and haptic-only. That pattern is compatible with weak or mixed integration, not with haptic dominance. The abstract and conclusion nevertheless present arousal override as a headline finding, so this inferential gap directly threatens the paper's main contribution. A secondary but related issue is that Kruskal-Wallis tests are applied to repeated-measures data and a Wilcoxon rank-sum test is used for paired tactile-only ratings, so the supporting p-values are not on a sound inferential footing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a within-subject experiment (N=32) in which participants rated the valence and arousal of a Pepper robot's haptic feedback delivered via a wearable vibration sleeve, both alone and in combination with six video scenarios depicting the robot experiencing actions (Kiss, Slap, Eye Contact, Stroke, Flick, Cover Eyes). The authors report three main findings: (1) participants can decode valence and arousal from haptic-only stimuli, (2) situational video context dominates the perceived valence of haptic feedback, with negative haptics amplifying and positive haptics moderating valence, and (3) haptic feedback overrides the arousal perceived from the video context. The manuscript uses pairwise non-parametric tests and one linear mixed-effects model on absolute valence deviation to support these claims.","tokens_in":9269,"tokens_out":4102,"duration_ms":46841,"significance":"If the findings hold, the paper contributes to affective HRI by showing that the interpretation of robotic haptic signals is strongly context-dependent, with a dissociation between valence (context-dominant) and arousal (haptics-dominant). The use of a real robot and a vibration sleeve derived from prior human touch data grounds the stimuli in empirical behavior, and the within-subject design is appropriate for this research question. The valence amplification/moderation result is a candidate for publication after the statistical issues are addressed. However, the arousal-override claim, which is a headline finding in the abstract, rests on null pairwise differences and is not yet supported by the analysis; this issue is load-bearing and requires additional statistical work.","major_comments":[{"comment":"The central claim that haptics override the arousal of situational context is based on the absence of significant pairwise differences between haptic-only and video-plus-haptic conditions. With N=32 and within-subject variability, a non-significant p-value does not establish equivalence or dominance. The paper needs an equivalence test (e.g., TOST) with a pre-specified bound, or a Bayesian model comparison, or a direct test of whether the combined rating is closer to the haptic-only than to the video-only rating. Notably, Section III-D.1 reports that for comfort haptics, only Stroke and Slap conditions differed significantly from their video-only counterparts, meaning that in most contexts the combined condition is statistically indistinguishable from both unimodal conditions; this pattern is compatible with weak or mixed integration rather than haptic dominance. The abstract and conclusion state the override result as a firm finding, so this inferential gap directly threatens the paper's main contribution.","section":"Section III-C.1"},{"comment":"The statistical tests are not appropriate for the design. Section III-A uses a Wilcoxon rank-sum test to compare CTactileAnger and CTactileComfort, but these are paired within-subject ratings; a Wilcoxon signed-rank test (or paired t-test) is required. The same repeated-measures structure is present in the Kruskal-Wallis tests in Sections III-B, III-C, and III-D, which treat the ratings across conditions as independent. This violates the independence assumption of Kruskal-Wallis and makes the reported H values and associated p-values unreliable. The authors should use a Friedman test or linear mixed-effects models with participant random intercepts for the omnibus tests, and ensure that post-hoc tests are paired and corrected for multiple comparisons accordingly.","section":"Section III-A and throughout Section III"},{"comment":"The 'override' conclusions are not supported by a direct statistical comparison of the three condition types involved. The paper separately compares combined conditions to haptic-only (Section III-C) and to video-only (Section III-D), but no test directly quantifies whether the combined rating is significantly closer to one unisensory condition than the other. A formal approach would be to compare the absolute deviation of the combined rating from the haptic-only prediction versus from the video-only prediction, or to fit an LMM with an interaction term capturing the relative influence of haptic and visual cues. Without such a test, the language of 'override' and 'dominates' in the abstract and Section IV is stronger than the evidence justifies.","section":"Sections III-C and III-D"}],"minor_comments":[{"comment":"The linear mixed-effects model on absolute valence deviation from neutral conflates shifts toward neutral with shifts across the neutral point; a positive touch that changes valence from 6 to 3 produces the same absolute deviation as one that changes it from 6 to 5. The authors should report separate analyses for positive and negative valence directions or use a model that preserves valence sign, to confirm that comfort haptics truly moderate rather than invert valence.","section":"Section III-D.2 / Eq. (1)"},{"comment":"The description of the stimulus selection says features were clustered 'based on their acoustic features' when the context concerns tactile data; this appears to be a typo and should read 'tactile features' or 'perceptual features'.","section":"Section II-C"},{"comment":"There are formatting errors: Table II lists 'Comfort (4.16 ± 1.71' in the no-haptic column, duplicating the haptic-only value; Fig. 4 caption has 'Subfigue'; and condition labels such as 'C ContextComfort' are inconsistently spaced. Inconsistent notation (e.g., 'CContextComfort' vs 'C ContextComfort') should be unified.","section":"Tables I and II / Fig. 4"},{"comment":"The SAM response scale is not defined (e.g., whether valence and arousal each range from 1 to 9 and what anchors are used). The paper mentions neutral valence of 5 in Section III-D.2, so the scale should be stated explicitly in the methods.","section":"Section II-E / Procedure"},{"comment":"The LMM results report only beta, SE, t, and p. Please include the random intercept variance and residual variance, and consider reporting effect sizes such as Cohen's d for the fixed effects, to allow assessment of the magnitude of the reported effects.","section":"Section III-D.2 / Eq. (1)"}],"recommendation":"major_revision","confidential_remarks":"The paper's core empirical finding on valence (context dominance, amplification by negative touch, moderation by positive touch) appears plausible and is a useful contribution to affective HRI. The arousal-override claim, however, is not supported by the current analysis and is presented as a headline result. The statistical issues in the choice of tests are correctable, and the missing equivalence test for the arousal claim is a tractable but non-trivial addition. Given the scope of the required re-analysis, I recommend major revision rather than rejection, provided the authors can supply equivalence tests or an appropriate model comparison for the arousal dimension."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The valence finding is the real contribution here, and it is reasonably supported. The claim that situational context dominates the valence of a haptic signal, with negative haptics amplifying and positive haptics moderating toward neutral, is backed by the linear mixed-effects model on absolute deviation from neutral valence. That LMM is a clean, appropriate test, and the pattern across the six video contexts is consistent. This asymmetry between valence and arousal is a genuinely new empirical observation in affective HRI, even if the components are familiar. The authors also do the right thing by grounding their haptic stimuli in a prior touch data collection rather than inventing arbitrary vibration patterns.\n\nThe soft spot is the headline claim that haptics override the arousal of the situational context. That conclusion rests almost entirely on null pairwise differences between haptic-only and video-plus-haptic conditions. With N=32, a null p-value is not evidence of dominance. The paper's own Table I shows the problem: for the comfort stimulus, the combined arousal ratings hover around 4-5 across all contexts, but the video-only ratings span a similar range, and most combined conditions are statistically indistinguishable from both haptic-only and video-only. That pattern is compatible with weak or mixed integration, not with haptic override. The arousal claim needs an equivalence test or a model comparison that formally pits haptic-only equivalence against video-only equivalence. Without that, the abstract overstates what the data show.\n\nThere are also some statistical issues worth flagging. The Wilcoxon rank-sum test in Section III-A is applied to paired data; it should be a signed-rank test. The Kruskal-Wallis tests on repeated measures violate the independence assumption, so those p-values are not on solid footing. No counterbalancing of video or haptic condition order is reported, which matters for a within-subject design with only 32 participants. These are fixable, but they should be addressed before the conclusions are taken as established.\n\nWho is this for? Researchers working on affective haptic design and multimodal emotion display in HRI. The valence-context interaction is a concrete, reusable rule for designers, so the paper deserves a serious referee. I would send it to peer review, but with the expectation of substantial revision focused on the arousal claim and the stats. The valence result can carry the paper; the arousal claim should be reframed as a hypothesis needing stronger evidence.","headline":"Solid valence result, but the arousal-override claim is overreach built on null differences and needs a formal equivalence test before it can stand.","tokens_in":9783,"tokens_out":1834,"would_cite":false,"duration_ms":19925,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that the emotional valence of a robot's haptic feedback is largely determined by the situational context in which it occurs, while the haptic signal itself dominates perceived arousal.","keywords":["affective haptics","human-robot interaction","situated context","valence and arousal","vibrotactile feedback","multimodal emotion perception","robot touch"],"falsifier":"A decisive observation: if a high-arousal slap video paired with the low-arousal comfort vibration yields arousal ratings reliably above the comfort-vibration-only baseline in a larger sample, the claimed haptic override of arousal would be falsified.","tokens_in":8868,"feed_emoji":"🤖","tokens_out":8552,"duration_ms":82801,"temperature":0.7,"pith_summary":"The authors set out to establish whether the emotional meaning of a robot's touch is fixed by the vibration itself or shaped by what is happening to the robot. In a 32-person experiment, participants rated a robot's valence and arousal from vibration alone, from videos of the robot being kissed, slapped, or otherwise acted upon, and from video plus vibration. The paper reports that people can decode valence and arousal from vibration alone, that the video context usually overrides the valence of the haptic signal, and that haptic feedback overrides the video's arousal. Negative haptic cues amplified the perceived valence of the interaction, while positive haptic cues shifted it toward neutral. The work has direct consequences for designing robots that communicate emotion through touch, because it treats emotional touch as situated and co-created rather than as a fixed physical signal.","feed_headline":"Context, not vibration, sets the emotional tone of robotic touch","feed_subtitle":"In tests, video context dominated how touch valence was read, while haptics drove arousal ratings.","key_machinery":"The central machinery is a three-condition within-subject paradigm: rating haptic-only, video-only, and video-plus-haptic presentations of the same robot interaction on Self-Assessment Manikin scales of valence and arousal. The load-bearing objects are two vibrotactile stimuli selected from clustered human touch data, one labelled “anger” (negative valence, high arousal) and one labelled “comfort” (positive valence, low arousal), delivered through a wearable vibration sleeve, together with six short videos of a robot being kissed, slapped, making eye contact, stroked, flicked, and having its eyes covered. Comparing ratings across these conditions is what lets the paper attribute an effect to context or to touch.","core_discovery":"In the study's own terms, the discovery is that emotional communication through robot touch is asymmetrically dominated by modality: situational context governs how positive or negative the interaction feels (valence), while the haptic signal governs how intense it feels (arousal). The paper grounds this in three sets of comparisons: participants distinguished the anger and comfort vibrations when presented alone; adding video contexts shifted valence ratings toward the video's meaning, with the anger vibration amplifying and the comfort vibration dampening the context's valence; and arousal ratings with video plus haptics did not differ from haptic-only ratings, with haptic feedback also compressing arousal differences across contexts.","pith_inferences":["We infer that the valence/arousal asymmetry points to different underlying mechanisms: arousal judgments may track physical vibration intensity, while valence judgments draw on meaning and appraisal; a testable extension is to vary vibration amplitude versus pattern separately and see which dimension drives the arousal override.","We infer that the comfort signal's moderating effect may partly reflect its low arousal and positive valence pulling ratings toward the scale's midpoint, so a follow-up pairing comfort haptics with strongly positive contexts is needed to separate regulation from averaging.","We infer that the results would predict a practical design rule: in ambiguous or positive situations haptic feedback can steer both valence and arousal, but in strongly negative situations only arousal is steerable; this could be tested in a real-time interaction where the robot's haptic response is deliberately chosen to de-escalate or intensify.","The authors themselves note the controlled laboratory setting with video stimuli and the culturally homogeneous participant group as limitations; this means the asymmetric dominance pattern should be checked in real-world, cross-cultural interactions before it is treated as a general rule."],"forward_implications":["A vibration pattern does not carry fixed emotional meaning; designers must treat the same haptic signal as context-dependent, since a robot being slapped versus kissed changes how the touch is read.","Arousal can be communicated reliably through haptics even when the visual scene is emotionally loaded, so vibration is a usable channel for conveying urgency or intensity.","A negative haptic signal can amplify the emotional intensity of an interaction, while a positive, low-arousal signal can pull perceived valence toward neutral; this gives designers a lever for affect regulation.","In strongly negative contexts such as a slap, haptic feedback has little room to change valence, so multimodal congruency matters more in ambiguous or positive scenarios."],"supporting_citations":[{"why":"Defines the circumplex valence-arousal space used to label the comfort and anger stimuli.","marker":"[4]"},{"why":"Establishes that social and contextual factors shape the emotional interpretation of touch, motivating the research question.","marker":"[7]"},{"why":"Prior multimodal study of visual and haptic emotion perception that this work extends by adding situational context.","marker":"[16]"},{"why":"Provides the wearable vibration sleeve method used to deliver the haptic stimuli.","marker":"[20]"},{"why":"Describes the tactile sensor data collection from which the anger and comfort stimuli were selected by clustering.","marker":"[27]"},{"why":"Provided the model for designing the six situational context actions and shows context directs affective interpretation of robot expressions.","marker":"[28]"},{"why":"Supplies the emotion-regulation framework used to interpret comfort as moderating and anger as amplifying valence.","marker":"[29]"}],"fun_headline_variants":["Robot touch: context steers emotions, vibration steers intensity","In robot touch, video context wins for feel, haptics for intensity","How we read robot touch: situation sets valence, haptics set arousal","Context shapes emotional feel of robot touch; haptics shape intensity","Robot touch: video context dominates valence, vibration dominates arousal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The arousal-override result depends on treating the lack of a statistically significant difference between haptic-only and video-plus-haptic arousal ratings, in a 32-person within-subject design, as evidence that the haptic signal dominates, rather than as an absence of detectable evidence.","fun_headline_variants_meta":{"raw":{"variants":["Robot touch: context steers emotions, vibration steers intensity","In robot touch, video context wins for feel, haptics for intensity","How we read robot touch: situation sets valence, haptics set arousal","Context shapes emotional feel of robot touch; haptics shape intensity","Robot touch: video context dominates valence, vibration dominates arousal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000377,"raw_usage":{"total_tokens":1998,"prompt_tokens":925,"completion_tokens":1073,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":995}},"tokens_in":541,"tokens_out":1073,"duration_ms":8378,"temperature":1.0,"reasoning_tokens":995,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:07:15.615658+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive observation: if a high-arousal slap video paired with the low-arousal comfort vibration yields arousal ratings reliably above the comfort-vibration-only baseline in a larger sample, the claimed haptic override of arousal would be falsified.","supporting_citations":[{"cited_title":"A circumplex model of affect.,","cited_arxiv_id":null,"evidence_quote":"Defines the circumplex valence-arousal space used to label the comfort and anger stimuli."},{"cited_title":"Social touch in human–computer interac- tion,","cited_arxiv_id":null,"evidence_quote":"Establishes that social and contextual factors shape the emotional interpretation of touch, motivating the research question."},{"cited_title":"Towards multimodal affective feedback: Interaction between visual and haptic modalities,","cited_arxiv_id":null,"evidence_quote":"Prior multimodal study of visual and haptic emotion perception that this work extends by adding situational context."},{"cited_title":"Situational context directs how people affectively interpret robotic non-linguistic utterances,","cited_arxiv_id":null,"evidence_quote":"Provided the model for designing the six situational context actions and shows context directs affective interpretation of robot expressions."},{"cited_title":"The emerging field of emotion regulation: An integrative review,","cited_arxiv_id":null,"evidence_quote":"Supplies the emotion-regulation framework used to interpret comfort as moderating and anger as amplifying valence."}],"review_version":1}