REVIEW 3 major objections 5 minor 31 references
How Problematic are Suspenseful Interactions?
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A controlled replication finds the suspensefulness effect is real but small, and recommends dropping the blanket ban on suspenseful interactions.
desk verdict The replication is solid, but the practical conclusion is undermined by ceiling-saturated scales; still deserves peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a controlled four-cell replication of the visibility taxonomy: the same activity (unlocking a car) is shown with visible or hidden manipulation crossed with visible or hidden effect, producing expressive, magical, secretive, and suspenseful conditions. The suspenseful cell—visible manipulation, invisible effect—is the focus. Three single-item social acceptability scales (the original measure, plus two established alternatives) provide the dependent variables, and the decisive move is to compare each condition's absolute score against the scale midpoint as well as against the other conditions; this comparison converts a 'statistically significant' effect into an assessment of practical relevance.
What would settle it
A replication that uses a finer-grained or multi-item acceptability measure and finds the suspenseful condition's score falling well below the scale midpoint, with effect sizes above $r = .2$, would falsify the paper's claim that the effect is practically negligible; the data and preregistered procedure are available so this is directly checkable.
Extended reading notes
Core claim
The central claim is that the suspensefulness effect exists but is not practically important. Using a controlled design that holds the activity constant and varies only the visibility of manipulation and effect, the author found the suspenseful condition received significantly lower ratings than the expressive, magical, and secretive conditions on two of three social acceptability scales, with all pairwise effect sizes at or below $r = .2$; on the third scale the difference from the secretive condition was not significant. Because the suspenseful gesture's median score reached the maximum of every scale, the paper concludes that users find suspenseful interactions slightly less comfortable yet still acceptable, and that the current guideline to avoid this form of interaction is no longer justified.
Load-bearing premise
The conclusion that the suspensefulness effect is practically negligible assumes that the three single-item scales have enough headroom to reveal meaningful differences; because the suspenseful gesture's median sits at the maximum of every scale, the small effect sizes and high absolute acceptability could instead be an artifact of ceiling effects.
Editorial extensions
If this is right
- Current guidelines that tell designers to avoid suspenseful interactions should be revised, because the empirical basis for a blanket ban is no longer supported.
- Designers should weigh the social situation, audience, and location more heavily than the visibility pattern of the interaction itself.
- Suspenseful forms can stay in the design space; everyday interactions such as using a smartphone already have this visibility pattern and are widely accepted.
- Because the study used one activity and online video stimuli, further replications with other gestures and physical settings are needed before generalizing the 'still acceptable' claim.
- The non-significant difference between suspenseful and secretive on one scale suggests that form alone does not determine acceptability.
Reading between the lines
- Extending beyond the paper: if the ceiling-effect concern is real, the small effect sizes may understate the true difference; a scale with more response headroom could reveal whether the 'high acceptability' conclusion is robust.
- A practical testable extension would be a field study where bystanders rate a real suspenseful interaction (e.g., phone use in a quiet library) to see whether situation moderates the effect as the paper's situated-design argument implies.
- The paper's logic also implies that other form-based guidelines derived from the same weak evidence base deserve similar direct replication before being codified.
- One could formalize the recommendation as a two-factor model—form and situation—and predict acceptability from their interaction; that model would be falsifiable in a follow-up experiment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a preregistered online replication (n = 281) of Montero et al.'s 'suspensefulness effect', using a single car-unlocking scenario with four visibility variants (expressive, magical, secretive, suspenseful) while holding the activity constant. The study finds statistically significant differences on the original Montero scale and the Koelle scale, partial confirmation on the Pearson scale, and small rank-based effect sizes (r <= .20). The authors then argue that absolute acceptability scores are high for all conditions, including the suspenseful one, based on one-sample median tests and ALA data, and conclude that the blanket guideline against suspenseful interactions should be dropped in favor of situation-focused design.
Significance. If the practical conclusion holds, the paper is valuable: it offers a controlled, preregistered, independently powered replication of an influential but weakly grounded effect, and it challenges a widely cited design guideline. The methodological improvements over the original study—constant activity, larger sample, multiple acceptability measures, and robustness checks for comprehension-check failures—are genuine strengths, and the open data/scripts are a plus. However, the central interpretive claim that the suspensefulness effect is 'practically negligible' rests on absolute acceptability scores that may be distorted by ceiling effects in the single-item scales. The statistical replication itself is credible, but the practical conclusion is not yet established.
major comments (3)
- [Section 3.2.5 and Figure 2] The conclusion that suspenseful interactions are highly acceptable is threatened by a ceiling effect. The median of the suspenseful condition is at the scale maximum on all three measures (6 on a 6-point scale, 5 on a 5-point scale, 7 on a 7-point scale), meaning at least half of the participants chose the top category. The one-sample median tests showing scores 'significantly above scale center' are therefore nearly tautological, and the small rank-based effect sizes and 'high absolute acceptability' reading could be attenuated or even produced by the scales' inability to register differences at the top. This directly undermines the Abstract and Section 4 claim that the effect is practically negligible. Please report the full response distributions and the proportion of ceiling responses, and either use instruments with sufficient headroom or temper the practical conclusion until the ceiling artifact is ruled out.
- [Sections 3.2.1-3.2.3] The pairwise Mann-Whitney comparisons for the three main hypotheses are reported without any correction for multiple testing, whereas the exploratory confidence analysis uses the Hochberg correction. Since the paper's 'eight out of nine significant pairwise differences' count is used to argue that the effect 'mostly' replicates, this is not merely a presentational detail. After a Holm or Hochberg correction, the Pearson-scale comparison between suspenseful and magical (p < .05) may no longer be significant, reducing the count to seven of nine. Please apply a family-wise correction to the confirmatory pairwise tests or explicitly justify the unadjusted procedure.
- [Section 3.2.5, ALA analysis] The ALA analysis is also affected by ceiling saturation: median acceptability is at the scale maximum (5 on a 5-point scale) for all suitable locations and audiences, and the paper consequently restricts itself to visual inspection. This means the ALA results cannot provide independent evidence for the transferability claim that 'suspenseful gesture socially acceptable in typical locations.' In addition, the suitability criterion 'critical n = 34' is unclear without knowing the per-condition sample size, and the choice of 'at least half' is arbitrary. Please clarify the criterion and either provide distribution-level evidence or soften the transferability claim.
minor comments (5)
- [Section 3.1.2] Typo: 'Fourty-seven' should be 'Forty-seven'.
- [Section 4.1] Typo: 'spatious hand gesture' should be 'spacious hand gesture'.
- [Figure 2] The boxplots hide the mass of responses at the scale maximum; consider adding jittered raw data or the proportion of top-category responses so readers can assess the ceiling issue directly.
- [Section 2.3] The definition of 'secretive' as having an invisible manipulation and an invisible effect may still involve a visible reaching-into-purse movement; please clarify which part counts as the manipulation, since this bears on the conceptual critique the paper itself raises.
- [Section 4.3] The smartphone example is anecdotal; consider labeling it as an illustrative observation rather than as empirical support.
Circularity Check
No significant circularity: the replication result is an independent empirical test; self-citations are contextual, not load-bearing.
full rationale
The paper's central claim is an empirical replication of the suspensefulness effect (Montero et al., 2010) using new preregistered data (n=281). Hypotheses H1 and H2 are tested with Kruskal-Wallis and Mann-Whitney comparisons of the suspenseful condition against the other three conditions, using Montero et al.'s original scale and two independently published scales (Pearson et al., 2015; Koelle et al., 2018). The conclusion that the effect is small (r <= .2) and that the suspenseful interaction still has high absolute acceptability is read directly from observed distributions in Section 3.2.5; no parameter is fitted from the data and then renamed a prediction, and no equation defines the conclusion into its inputs. The paper's self-citations [26, 28, 29] appear in the background, the conceptual critique of Reeves et al.'s visibility definition, and the interpretation of the results, but they are not used to generate the replication statistics or the absolute acceptability values. The central derivation is therefore self-contained against the external benchmark of Montero et al. The manuscript itself notes a limitation relevant to the skeptical reading: all three acceptability measures are single items, and the reported medians for the suspenseful condition equal the scale maxima (6 on the 6-point Montero scale, 5 on the 5-point Pearson scale, 7 on the 7-point Koelle scale). This raises a legitimate ceiling-saturation risk for the 'small effect / high acceptability' interpretation, but that is a measurement-validity concern rather than a circularity: the paper does not argue from its own prior work or from a fitted parameter, and the effect's existence is tested against an external result. No circular steps are identified; the minor self-citations present are not load-bearing for the main empirical claim.
Assumptions & free parameters
free parameters (1)
- ALA suitability threshold (critical n = 34) =
34 participants (about 12% of the sample)
assumptions (4)
- domain assumption Perception-based definition of visibility in Reeves et al.'s taxonomy
- domain assumption Self-reported anticipated acceptability from watching videos is a proxy for real situated acceptability
- domain assumption One video per condition represents each interaction category
- domain assumption Single-item scales yield meaningful interval-level comparisons
Cite this review
Pith. "Pith review of How Problematic are Suspenseful Interactions?." pith.science (2026). https://pith.science/paper/RYZVU2CM
@misc{pith2026250601287,
author = {Pith},
title = {Pith review of: How Problematic are Suspenseful Interactions?},
year = {2026},
howpublished = {\url{https://pith.science/paper/RYZVU2CM}},
note = {Machine review of arXiv:2506.01287}
}
read the original abstract
Current "social acceptability" guidelines for interactive technologies advise against certain, seemingly problematic forms of interaction. Specifically, "suspenseful" interactions, characterized by visible manipulations and invisible effects, are generally considered be problematic. However, the empirical grounding for this claim is surprisingly weak. To test its validity, this paper presents a controlled replication study (n = 281) of the "suspensefulness effect". Although it could be statistically replicated with two out of three social acceptability measures, effect sizes were small (r =< .2), and all compared forms of interaction, including the suspenseful one, had high absolute social acceptability scores. Thus, despite the slight negative effect, suspenseful interactions seem less problematic in the overall scheme of things. We discuss alternative approaches to improve the social acceptability of interactive technology, and recommend to more closely engage with their specific social situatedness.
Figures
Reference graph
Works this paper leans on
-
[1]
David Ahlström, Khalad Hasan, and Pourang Irani. 2014. Are You Comfortable Doing That?: Acceptance Studies of Around-device Gestures In And For Public Settings. In Proceedings of the 16th International Conference on Human- Computer Interaction with Mobile Devices & Services - MobileHCI’14 . ACM, New York, NY, USA, 193–202. https: //doi.org/10.1145/2628363.2628381
arXiv 2014
-
[2]
Fouad Alallah, Ali Neshati, Yumiko Sakamoto, Khalad Hasan, Edward Lank, Andrea Bunt, and Pourang Irani. 2018. Performer vs. Observer: Whose Comfort Level Should We Consider When Examining the Social Acceptability of Input Modalities for Head-worn Display?. In Proceedings of the 24th ACM Symposium on Virtual Reality Software and Technology. ACM, New York, ...
arXiv 2018
-
[3]
Emberson, Gary Lupyan, Michael H
Lauren L. Emberson, Gary Lupyan, Michael H. Goldstein, and Michael J. Spivey. 2010. Overheard Cell-Phone Conversations: When Less Speech Is More Distracting. Psychological Science 21, 10 (2010), 1383–1388. https: //doi.org/10.1177/0956797610382126
-
[4]
Barrett Ens, Tovi Grossman, Fraser Anderson, Justin Matejka, and George Fitzmaurice. 2015. Candid Interaction: Revealing Hidden Mobile and Wearable Computing Activities. InProceedings of the 28th Annual ACM Symposium on User Interface Software & Technology - UIST’15. ACM, New York, NY, USA, 467–476. https://doi.org/10.1145/2807442.2807449
arXiv 2015
-
[5]
Franz Faul, Edgar Erdfelder, Albert-Georg Lang, and Axel Buchner. 2007. G*Power 3: A Flexible Statistical Power Analysis Program for the Social, Behavioral, and Biomedical Sciences. Behavior Research Methods 39, 2 (2007), 175–191. https://doi.org/10.3758/BF03193146
- [6]
-
[7]
Yosef Hochberg. 1988. A sharper Bonferroni procedure for multiple tests of significance. Biometrika 75, 4 (1988), 800–802. https://doi.org/10.1093/biomet/75.4.800
-
[8]
Norene Kelly and Stephen Gilbert. 2016. The WEAR Scale: Developing a Measure of the Social Acceptability of a Wearable Device. In Proceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems - CHI EA’16. ACM, New York, NY, USA, 2864–2871. https://doi.org/10.1145/2851581.2892331
arXiv 2016
Show all 31 references
-
[9]
Seoktae Kim, Minjung Sohn, Jinhee Pak, and Woohun Lee. 2006. One-key Keyboard: A Very Small QWERTY Keyboard Supporting Text Entry for Wearable Computing. InProceedings of the 2022 Australian Computer-Human Interaction Conference – OzCHI’06. ACM, New York, NY, USA, 305–308. htt...
2006
-
[10]
Marion Koelle, Swamy Ananthanarayan, and Susanne Boll. 2020. Social Acceptability in HCI: A Survey of Methods, Measures, and Design Strategies. In Proceedings of the 2020 ACM Conference on Human Factors in Computing Systems . ACM, New York, NY, USA, 1–19. https://doi.org/10.11...
2020
-
[11]
Marion Koelle, Swamy Ananthanarayan, Simon Czupalla, Wilko Heuten, and Susanne Boll. 2018. Your Smart Glasses’ Camera Bothers Me!: Exploring Opt-in and Opt-out Gestures for Privacy Mediation. In Proceedings of the 10th Nordic Conference on Human-Computer Interaction - NordiCHI...
2018
-
[12]
Theodore Kunin. 1955. The Construction of a New Type of Attitude Measure. Personnel Psychology 8, 1 (1955), 65–77. https://doi.org/10.1111/j.1744-6570.1955.tb01189.x
1955
-
[13]
Tiffany C. K. Kwok, Peter Kiefer, and Martin Raubal. 2023. Unobtrusive Interaction: a Systematic Literature Review and Expert Survey. Human-Computer Interaction (2023), 37 pages. https://doi.org/10.1080/07370024.2022.2162404
2023
-
[14]
Richard Li, Jason Wu, and Thad Starner. 2019. TongueBoard: An Oral Interface for Subtle Input. InProceedings of the 10th Augmented Human International Conference. ACM, New York, NY, USA, 1–9. https://doi.org/10.1145/3311823.3311831
2019
-
[15]
Andrew Monk, Jenni Carroll, Sarah Parker, and Mark Blythe. 2004. Why are Mobile Phones Annoying? Behaviour & Information Technology 23, 1 (2004), 33–41. https://doi.org/10.1080/01449290310001638496
2004 doi
-
[16]
Andrew Monk, Evi Fellas, and Eleanor Ley. 2004. Hearing Only One Side of Normal and Mobile Phone Conversations. Behaviour & Information Technology 23, 5 (2004), 301–305. https://doi.org/10.1080/01449290410001712744
2004 doi
-
[17]
Montero, Jason Alexander, Mark T
Calkin S. Montero, Jason Alexander, Mark T. Marshall, and Sriram Subramanian. 2010. Would You Do That?: Un- derstanding Social Acceptance of Gestural Interfaces. In Proceedings of the 12th International Conference on Hu- man Computer Interaction with Mobile Devices and Service...
2010
-
[18]
Brendan Norman and Daniel Bennett. 2014. Are Mobile Phone Conversations Always so Annoying? The ‘need-to-listen’ Effect Re-visited. Behaviour & Information Technology 33, 12 (2014), 1294–1305. https://doi.org/10.1080/0144929X.2013. 876098
2014 doi
-
[19]
Thomas Olsson, Pradthana Jarusriboonchai, Paweł Woźniak, Susanna Paasovaara, Kaisa Väänänen, and Andrés Lucero
-
[20]
Jennifer Pearson, Simon Robinson, and Matt Jones. 2015. It’s About Time: Smartwatches as Public Displays. In Proceedings of the ACM Conference on Human Factors in Computing Systems - CHI’15 . ACM, New York, NY, USA, Proc. ACM Hum.-Comput. Interact., Vol. 1, No. 1, Article . Pu...
2015
-
[21]
Halley Profita, Reem Albaghli, Leah Findlater, Paul Jaeger, and Shaun K. Kane. 2016. The AT Effect: How Disability Affects the Perceived Social Acceptability of Head-Mounted Display Use. InProceedings of the ACM Conference on Human Factors in Computing Systems - CHI’16 . ACM, ...
2016
-
[22]
Stuart Reeves, Steve Benford, Claire O’Malley, and Mike Fraser. 2005. Designing the Spectator Experience. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems - CHI’05 . ACM, New York, NY, USA, 741–750. https://doi.org/10.1145/1054972.1055074
2005
-
[23]
Jun Rekimoto. 2001. GestureWrist and GesturePad: Unobtrusive Wearable Interaction Devices. In Proceedings of the Fifth International Symposium on Wearable Computers . IEEE, Piscataway, NJ, USA, 21–27. https://doi.org/10.1109/ ISWC.2001.962092
2001
-
[24]
Julie Rico and Stephen Brewster. 2010. Usable Gestures for Mobile Interfaces: Evaluating Social Acceptability. In Proceedings of the 2010 ACM Conference on Human Factors in Computing Systems . ACM, New York, NY, USA, 887–896. https://doi.org/10.1145/1753326.1753458
2010
-
[25]
Harvey Sacks. 1992. Lectures on Conversation: Volumes I and II . Blackwell, Oxford, UK
1992
-
[26]
Alarith Uhde, Lianara Dreyer, and Marc Hassenzahl. 2025. The Witness Experience Inventory. Interacting With Computers iwaf010 (2025), 1–16. https://doi.org/10.1093/iwc/iwaf010
2025 doi
-
[27]
Alarith Uhde and Marc Hassenzahl. 2021. Towards a Better Understanding of Social Acceptability. InProceedings of the ACM CHI Conference on Human Factors in Computing Systems Extended Abstracts . ACM, New York, NY, USA, 6 pages. https://doi.org/10.1145/3411763.3451649
2021
-
[28]
Alarith Uhde, Tim zum Hoff, and Marc Hassenzahl. 2022. Obtrusive Subtleness and Why We Should Focus on Meaning, not Form, in Social Acceptability Studies. In Proceedings of the 21st International Conference on Mobile and Ubiquitous Multimedia (MUM’22). ACM, New York, NY, USA, ...
2022
-
[29]
Alarith Uhde, Tim zum Hoff, and Marc Hassenzahl. 2023. Beyond Hiding and Revealing: Exploring Effects of Visibility and Form of Interaction on the Witness Experience.Proceedings of the ACM on Human-Computer Interaction (MobileHCI) 7, 200 (2023), 23. https://doi.org/10.1145/3604247
2023 doi
-
[30]
Rand R. Wilcox. 2012.Introduction to Robust Estimation and Hypothesis Testing. Academic Press, Amsterdam, Netherlands and Boston, MA, USA. https://doi.org/10.1016/c2010-0-67044-1 Received 6 February 2025; revised 8 May 2025; accepted 29 May 2025 Proc. ACM Hum.-Comput. Interact...
2012 doi
-
[2020]
https://doi.org/10.1007/s10606-019-09345-0
Technologies for Enhancing Collocated Social Interaction: Review of Design Solutions and Approaches.Computer Supported Cooperative Work (CSCW) 29, 1-2 (2020), 29–83. https://doi.org/10.1007/s10606-019-09345-0
2020 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.