Pith. sign in

REVIEW 3 major objections 5 minor 31 references

How Problematic are Suspenseful Interactions?

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A controlled replication finds the suspensefulness effect is real but small, and recommends dropping the blanket ban on suspenseful interactions.

desk verdict The replication is solid, but the practical conclusion is undermined by ceiling-saturated scales; still deserves peer review. read the letter →

arxiv 2506.01287 v1 pith:RYZVU2CM submitted 2025-06-02 cs.HC

classification cs.HC
keywords socialacceptabilitysuspensefulnesseffectreplicationstudygestureinteractionvisibilitytaxonomydesignsituatednessHCI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tests a long-standing rule in human-computer interaction: that 'suspenseful' interactions—where the user's manipulation is visible but its effect is not, such as pointing a key fob at a car without seeing lights flash—are socially unacceptable. In a preregistered replication with 281 participants watching one of four versions of the same car-unlocking scene, the paper reproduces the effect statistically on two of three acceptability measures, but with small effect sizes ($r \leq .2$). Crucially, even the suspenseful gesture scored at or near the top of all three scales, meaning it was rated socially acceptable in absolute terms. The paper argues that the blanket recommendation to avoid suspenseful interactions should therefore be dropped, and that designers should instead consider the specific social situation.

What carries the argument

The argument is carried by a controlled four-cell replication of the visibility taxonomy: the same activity (unlocking a car) is shown with visible or hidden manipulation crossed with visible or hidden effect, producing expressive, magical, secretive, and suspenseful conditions. The suspenseful cell—visible manipulation, invisible effect—is the focus. Three single-item social acceptability scales (the original measure, plus two established alternatives) provide the dependent variables, and the decisive move is to compare each condition's absolute score against the scale midpoint as well as against the other conditions; this comparison converts a 'statistically significant' effect into an assessment of practical relevance.

What would settle it

A replication that uses a finer-grained or multi-item acceptability measure and finds the suspenseful condition's score falling well below the scale midpoint, with effect sizes above $r = .2$, would falsify the paper's claim that the effect is practically negligible; the data and preregistered procedure are available so this is directly checkable.

Watch

Extended reading notes

Core claim

The central claim is that the suspensefulness effect exists but is not practically important. Using a controlled design that holds the activity constant and varies only the visibility of manipulation and effect, the author found the suspenseful condition received significantly lower ratings than the expressive, magical, and secretive conditions on two of three social acceptability scales, with all pairwise effect sizes at or below $r = .2$; on the third scale the difference from the secretive condition was not significant. Because the suspenseful gesture's median score reached the maximum of every scale, the paper concludes that users find suspenseful interactions slightly less comfortable yet still acceptable, and that the current guideline to avoid this form of interaction is no longer justified.

Load-bearing premise

The conclusion that the suspensefulness effect is practically negligible assumes that the three single-item scales have enough headroom to reveal meaningful differences; because the suspenseful gesture's median sits at the maximum of every scale, the small effect sizes and high absolute acceptability could instead be an artifact of ceiling effects.

Editorial extensions

If this is right

  • Current guidelines that tell designers to avoid suspenseful interactions should be revised, because the empirical basis for a blanket ban is no longer supported.
  • Designers should weigh the social situation, audience, and location more heavily than the visibility pattern of the interaction itself.
  • Suspenseful forms can stay in the design space; everyday interactions such as using a smartphone already have this visibility pattern and are widely accepted.
  • Because the study used one activity and online video stimuli, further replications with other gestures and physical settings are needed before generalizing the 'still acceptable' claim.
  • The non-significant difference between suspenseful and secretive on one scale suggests that form alone does not determine acceptability.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Extending beyond the paper: if the ceiling-effect concern is real, the small effect sizes may understate the true difference; a scale with more response headroom could reveal whether the 'high acceptability' conclusion is robust.
  • A practical testable extension would be a field study where bystanders rate a real suspenseful interaction (e.g., phone use in a quiet library) to see whether situation moderates the effect as the paper's situated-design argument implies.
  • The paper's logic also implies that other form-based guidelines derived from the same weak evidence base deserve similar direct replication before being codified.
  • One could formalize the recommendation as a two-factor model—form and situation—and predict acceptability from their interaction; that model would be falsifiable in a follow-up experiment.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a preregistered online replication (n = 281) of Montero et al.'s 'suspensefulness effect', using a single car-unlocking scenario with four visibility variants (expressive, magical, secretive, suspenseful) while holding the activity constant. The study finds statistically significant differences on the original Montero scale and the Koelle scale, partial confirmation on the Pearson scale, and small rank-based effect sizes (r <= .20). The authors then argue that absolute acceptability scores are high for all conditions, including the suspenseful one, based on one-sample median tests and ALA data, and conclude that the blanket guideline against suspenseful interactions should be dropped in favor of situation-focused design.

Significance. If the practical conclusion holds, the paper is valuable: it offers a controlled, preregistered, independently powered replication of an influential but weakly grounded effect, and it challenges a widely cited design guideline. The methodological improvements over the original study—constant activity, larger sample, multiple acceptability measures, and robustness checks for comprehension-check failures—are genuine strengths, and the open data/scripts are a plus. However, the central interpretive claim that the suspensefulness effect is 'practically negligible' rests on absolute acceptability scores that may be distorted by ceiling effects in the single-item scales. The statistical replication itself is credible, but the practical conclusion is not yet established.

major comments (3)
  1. [Section 3.2.5 and Figure 2] The conclusion that suspenseful interactions are highly acceptable is threatened by a ceiling effect. The median of the suspenseful condition is at the scale maximum on all three measures (6 on a 6-point scale, 5 on a 5-point scale, 7 on a 7-point scale), meaning at least half of the participants chose the top category. The one-sample median tests showing scores 'significantly above scale center' are therefore nearly tautological, and the small rank-based effect sizes and 'high absolute acceptability' reading could be attenuated or even produced by the scales' inability to register differences at the top. This directly undermines the Abstract and Section 4 claim that the effect is practically negligible. Please report the full response distributions and the proportion of ceiling responses, and either use instruments with sufficient headroom or temper the practical conclusion until the ceiling artifact is ruled out.
  2. [Sections 3.2.1-3.2.3] The pairwise Mann-Whitney comparisons for the three main hypotheses are reported without any correction for multiple testing, whereas the exploratory confidence analysis uses the Hochberg correction. Since the paper's 'eight out of nine significant pairwise differences' count is used to argue that the effect 'mostly' replicates, this is not merely a presentational detail. After a Holm or Hochberg correction, the Pearson-scale comparison between suspenseful and magical (p < .05) may no longer be significant, reducing the count to seven of nine. Please apply a family-wise correction to the confirmatory pairwise tests or explicitly justify the unadjusted procedure.
  3. [Section 3.2.5, ALA analysis] The ALA analysis is also affected by ceiling saturation: median acceptability is at the scale maximum (5 on a 5-point scale) for all suitable locations and audiences, and the paper consequently restricts itself to visual inspection. This means the ALA results cannot provide independent evidence for the transferability claim that 'suspenseful gesture socially acceptable in typical locations.' In addition, the suitability criterion 'critical n = 34' is unclear without knowing the per-condition sample size, and the choice of 'at least half' is arbitrary. Please clarify the criterion and either provide distribution-level evidence or soften the transferability claim.
minor comments (5)
  1. [Section 3.1.2] Typo: 'Fourty-seven' should be 'Forty-seven'.
  2. [Section 4.1] Typo: 'spatious hand gesture' should be 'spacious hand gesture'.
  3. [Figure 2] The boxplots hide the mass of responses at the scale maximum; consider adding jittered raw data or the proportion of top-category responses so readers can assess the ceiling issue directly.
  4. [Section 2.3] The definition of 'secretive' as having an invisible manipulation and an invisible effect may still involve a visible reaching-into-purse movement; please clarify which part counts as the manipulation, since this bears on the conceptual critique the paper itself raises.
  5. [Section 4.3] The smartphone example is anecdotal; consider labeling it as an illustrative observation rather than as empirical support.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the replication result is an independent empirical test; self-citations are contextual, not load-bearing.

full rationale

The paper's central claim is an empirical replication of the suspensefulness effect (Montero et al., 2010) using new preregistered data (n=281). Hypotheses H1 and H2 are tested with Kruskal-Wallis and Mann-Whitney comparisons of the suspenseful condition against the other three conditions, using Montero et al.'s original scale and two independently published scales (Pearson et al., 2015; Koelle et al., 2018). The conclusion that the effect is small (r <= .2) and that the suspenseful interaction still has high absolute acceptability is read directly from observed distributions in Section 3.2.5; no parameter is fitted from the data and then renamed a prediction, and no equation defines the conclusion into its inputs. The paper's self-citations [26, 28, 29] appear in the background, the conceptual critique of Reeves et al.'s visibility definition, and the interpretation of the results, but they are not used to generate the replication statistics or the absolute acceptability values. The central derivation is therefore self-contained against the external benchmark of Montero et al. The manuscript itself notes a limitation relevant to the skeptical reading: all three acceptability measures are single items, and the reported medians for the suspenseful condition equal the scale maxima (6 on the 6-point Montero scale, 5 on the 5-point Pearson scale, 7 on the 7-point Koelle scale). This raises a legitimate ceiling-saturation risk for the 'small effect / high acceptability' interpretation, but that is a measurement-validity concern rather than a circularity: the paper does not argue from its own prior work or from a fitted parameter, and the effect's existence is tested against an external result. No circular steps are identified; the minor self-citations present are not load-bearing for the main empirical claim.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central result rests on domain assumptions about the validity of the taxonomy, the video-based measurement, and the representativeness of one gesture per category. No new entities are posited. One hand-chosen analytic threshold appears in the exploratory ALA analysis.

free parameters (1)
  • ALA suitability threshold (critical n = 34) = 34 participants (about 12% of the sample)
    Locations and audiences were excluded from the ALA analysis if at least half of the participants considered them unsuitable. The critical n was derived from this post hoc rule and was not part of the preregistered plan (Section 3.2.5).
assumptions (4)
  • domain assumption Perception-based definition of visibility in Reeves et al.'s taxonomy
    The paper defines visibility as visual perceptibility, which affects how the four interaction categories are designed and labeled (Section 2.3).
  • domain assumption Self-reported anticipated acceptability from watching videos is a proxy for real situated acceptability
    Participants watched videos and imagined performing the gesture; the paper acknowledges the ecological validity limit (Section 4.2).
  • domain assumption One video per condition represents each interaction category
    Each of the four conditions used a single video, so category differences are confounded with idiosyncratic features of that video (Section 3.1.2).
  • domain assumption Single-item scales yield meaningful interval-level comparisons
    The three acceptability scales each have a single item and show ceiling effects; the analysis treats them as suitable for hypothesis testing (Sections 3.1.3, 3.2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of How Problematic are Suspenseful Interactions?." pith.science (2026). https://pith.science/paper/RYZVU2CM

@misc{pith2026250601287,
  author       = {Pith},
  title        = {Pith review of: How Problematic are Suspenseful Interactions?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RYZVU2CM}},
  note         = {Machine review of arXiv:2506.01287}
}
read the original abstract

Current "social acceptability" guidelines for interactive technologies advise against certain, seemingly problematic forms of interaction. Specifically, "suspenseful" interactions, characterized by visible manipulations and invisible effects, are generally considered be problematic. However, the empirical grounding for this claim is surprisingly weak. To test its validity, this paper presents a controlled replication study (n = 281) of the "suspensefulness effect". Although it could be statistically replicated with two out of three social acceptability measures, effect sizes were small (r =< .2), and all compared forms of interaction, including the suspenseful one, had high absolute social acceptability scores. Thus, despite the slight negative effect, suspenseful interactions seem less problematic in the overall scheme of things. We discuss alternative approaches to improve the social acceptability of interactive technology, and recommend to more closely engage with their specific social situatedness.

Figures

Figures reproduced from arXiv: 2506.01287 by the authors.

Figure 1
Figure 1. Key elements of the interaction used in the study. The four larger pictures illustrate the interaction. [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Boxplots of the four interactions’ scores on the three social acceptability measures. The dashed line [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Scatter plot of the suspenseful gesture’s social acceptability score for the six suitable locations and all [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

31 extracted references · 13 canonical work pages

  1. [1]

    David Ahlström, Khalad Hasan, and Pourang Irani. 2014. Are You Comfortable Doing That?: Acceptance Studies of Around-device Gestures In And For Public Settings. In Proceedings of the 16th International Conference on Human- Computer Interaction with Mobile Devices & Services - MobileHCI’14 . ACM, New York, NY, USA, 193–202. https: //doi.org/10.1145/2628363.2628381

  2. [2]

    Fouad Alallah, Ali Neshati, Yumiko Sakamoto, Khalad Hasan, Edward Lank, Andrea Bunt, and Pourang Irani. 2018. Performer vs. Observer: Whose Comfort Level Should We Consider When Examining the Social Acceptability of Input Modalities for Head-worn Display?. In Proceedings of the 24th ACM Symposium on Virtual Reality Software and Technology. ACM, New York, ...

  3. [3]

    Emberson, Gary Lupyan, Michael H

    Lauren L. Emberson, Gary Lupyan, Michael H. Goldstein, and Michael J. Spivey. 2010. Overheard Cell-Phone Conversations: When Less Speech Is More Distracting. Psychological Science 21, 10 (2010), 1383–1388. https: //doi.org/10.1177/0956797610382126

  4. [4]

    Barrett Ens, Tovi Grossman, Fraser Anderson, Justin Matejka, and George Fitzmaurice. 2015. Candid Interaction: Revealing Hidden Mobile and Wearable Computing Activities. InProceedings of the 28th Annual ACM Symposium on User Interface Software & Technology - UIST’15. ACM, New York, NY, USA, 467–476. https://doi.org/10.1145/2807442.2807449

  5. [5]

    Franz Faul, Edgar Erdfelder, Albert-Georg Lang, and Axel Buchner. 2007. G*Power 3: A Flexible Statistical Power Analysis Program for the Social, Behavioral, and Biomedical Sciences. Behavior Research Methods 39, 2 (2007), 175–191. https://doi.org/10.3758/BF03193146

  6. [6]

    Kaplowitz

    Jonathan Forma and Stan A. Kaplowitz. 2012. The Perceived Rudeness of Public Cell Phone Behaviour. Behaviour & Information Technology 31, 10 (2012), 947–952. https://doi.org/10.1080/0144929X.2010.520335

  7. [7]

    Yosef Hochberg. 1988. A sharper Bonferroni procedure for multiple tests of significance. Biometrika 75, 4 (1988), 800–802. https://doi.org/10.1093/biomet/75.4.800

  8. [8]

    Norene Kelly and Stephen Gilbert. 2016. The WEAR Scale: Developing a Measure of the Social Acceptability of a Wearable Device. In Proceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems - CHI EA’16. ACM, New York, NY, USA, 2864–2871. https://doi.org/10.1145/2851581.2892331

Show all 31 references
  1. [9]

    Seoktae Kim, Minjung Sohn, Jinhee Pak, and Woohun Lee. 2006. One-key Keyboard: A Very Small QWERTY Keyboard Supporting Text Entry for Wearable Computing. InProceedings of the 2022 Australian Computer-Human Interaction Conference – OzCHI’06. ACM, New York, NY, USA, 305–308. htt...

  2. [10]

    Marion Koelle, Swamy Ananthanarayan, and Susanne Boll. 2020. Social Acceptability in HCI: A Survey of Methods, Measures, and Design Strategies. In Proceedings of the 2020 ACM Conference on Human Factors in Computing Systems . ACM, New York, NY, USA, 1–19. https://doi.org/10.11...

  3. [11]

    Marion Koelle, Swamy Ananthanarayan, Simon Czupalla, Wilko Heuten, and Susanne Boll. 2018. Your Smart Glasses’ Camera Bothers Me!: Exploring Opt-in and Opt-out Gestures for Privacy Mediation. In Proceedings of the 10th Nordic Conference on Human-Computer Interaction - NordiCHI...

  4. [12]

    Theodore Kunin. 1955. The Construction of a New Type of Attitude Measure. Personnel Psychology 8, 1 (1955), 65–77. https://doi.org/10.1111/j.1744-6570.1955.tb01189.x

  5. [13]

    Tiffany C. K. Kwok, Peter Kiefer, and Martin Raubal. 2023. Unobtrusive Interaction: a Systematic Literature Review and Expert Survey. Human-Computer Interaction (2023), 37 pages. https://doi.org/10.1080/07370024.2022.2162404

  6. [14]

    Richard Li, Jason Wu, and Thad Starner. 2019. TongueBoard: An Oral Interface for Subtle Input. InProceedings of the 10th Augmented Human International Conference. ACM, New York, NY, USA, 1–9. https://doi.org/10.1145/3311823.3311831

  7. [15]

    Andrew Monk, Jenni Carroll, Sarah Parker, and Mark Blythe. 2004. Why are Mobile Phones Annoying? Behaviour & Information Technology 23, 1 (2004), 33–41. https://doi.org/10.1080/01449290310001638496

  8. [16]

    Andrew Monk, Evi Fellas, and Eleanor Ley. 2004. Hearing Only One Side of Normal and Mobile Phone Conversations. Behaviour & Information Technology 23, 5 (2004), 301–305. https://doi.org/10.1080/01449290410001712744

  9. [17]

    Montero, Jason Alexander, Mark T

    Calkin S. Montero, Jason Alexander, Mark T. Marshall, and Sriram Subramanian. 2010. Would You Do That?: Un- derstanding Social Acceptance of Gestural Interfaces. In Proceedings of the 12th International Conference on Hu- man Computer Interaction with Mobile Devices and Service...

  10. [18]

    Brendan Norman and Daniel Bennett. 2014. Are Mobile Phone Conversations Always so Annoying? The ‘need-to-listen’ Effect Re-visited. Behaviour & Information Technology 33, 12 (2014), 1294–1305. https://doi.org/10.1080/0144929X.2013. 876098

  11. [19]

    Thomas Olsson, Pradthana Jarusriboonchai, Paweł Woźniak, Susanna Paasovaara, Kaisa Väänänen, and Andrés Lucero

  12. [20]

    Jennifer Pearson, Simon Robinson, and Matt Jones. 2015. It’s About Time: Smartwatches as Public Displays. In Proceedings of the ACM Conference on Human Factors in Computing Systems - CHI’15 . ACM, New York, NY, USA, Proc. ACM Hum.-Comput. Interact., Vol. 1, No. 1, Article . Pu...

  13. [21]

    Halley Profita, Reem Albaghli, Leah Findlater, Paul Jaeger, and Shaun K. Kane. 2016. The AT Effect: How Disability Affects the Perceived Social Acceptability of Head-Mounted Display Use. InProceedings of the ACM Conference on Human Factors in Computing Systems - CHI’16 . ACM, ...

  14. [22]

    Stuart Reeves, Steve Benford, Claire O’Malley, and Mike Fraser. 2005. Designing the Spectator Experience. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems - CHI’05 . ACM, New York, NY, USA, 741–750. https://doi.org/10.1145/1054972.1055074

  15. [23]

    Jun Rekimoto. 2001. GestureWrist and GesturePad: Unobtrusive Wearable Interaction Devices. In Proceedings of the Fifth International Symposium on Wearable Computers . IEEE, Piscataway, NJ, USA, 21–27. https://doi.org/10.1109/ ISWC.2001.962092

  16. [24]

    Julie Rico and Stephen Brewster. 2010. Usable Gestures for Mobile Interfaces: Evaluating Social Acceptability. In Proceedings of the 2010 ACM Conference on Human Factors in Computing Systems . ACM, New York, NY, USA, 887–896. https://doi.org/10.1145/1753326.1753458

  17. [25]

    Harvey Sacks. 1992. Lectures on Conversation: Volumes I and II . Blackwell, Oxford, UK

  18. [26]

    Alarith Uhde, Lianara Dreyer, and Marc Hassenzahl. 2025. The Witness Experience Inventory. Interacting With Computers iwaf010 (2025), 1–16. https://doi.org/10.1093/iwc/iwaf010

  19. [27]

    Alarith Uhde and Marc Hassenzahl. 2021. Towards a Better Understanding of Social Acceptability. InProceedings of the ACM CHI Conference on Human Factors in Computing Systems Extended Abstracts . ACM, New York, NY, USA, 6 pages. https://doi.org/10.1145/3411763.3451649

  20. [28]

    Alarith Uhde, Tim zum Hoff, and Marc Hassenzahl. 2022. Obtrusive Subtleness and Why We Should Focus on Meaning, not Form, in Social Acceptability Studies. In Proceedings of the 21st International Conference on Mobile and Ubiquitous Multimedia (MUM’22). ACM, New York, NY, USA, ...

  21. [29]

    Alarith Uhde, Tim zum Hoff, and Marc Hassenzahl. 2023. Beyond Hiding and Revealing: Exploring Effects of Visibility and Form of Interaction on the Witness Experience.Proceedings of the ACM on Human-Computer Interaction (MobileHCI) 7, 200 (2023), 23. https://doi.org/10.1145/3604247

  22. [30]

    Rand R. Wilcox. 2012.Introduction to Robust Estimation and Hypothesis Testing. Academic Press, Amsterdam, Netherlands and Boston, MA, USA. https://doi.org/10.1016/c2010-0-67044-1 Received 6 February 2025; revised 8 May 2025; accepted 29 May 2025 Proc. ACM Hum.-Comput. Interact...

  23. [2020]

    https://doi.org/10.1007/s10606-019-09345-0

    Technologies for Enhancing Collocated Social Interaction: Review of Design Solutions and Approaches.Computer Supported Cooperative Work (CSCW) 29, 1-2 (2020), 29–83. https://doi.org/10.1007/s10606-019-09345-0

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.