{"id":"44e85ca6-2eb6-48ac-8dc9-0b1071d689cd","arxiv_id":"2506.06774","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"All four dark patterns tested in simulated MR videos significantly reduced comfort and adoption intent while increasing reactance and perceived system darkness.","lead":"A video-based study with 74 participants tested four manipulative design tricks (dark patterns) in simulated mixed-reality city walks, applied to a restaurant, a pair of shoes, and a person. All four patterns lowered comfort and intention to use MR glasses and increased user resistance, with effects varying by target.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The one-scenario-per-pattern design makes the dark-pattern category inseparable from its concrete visual design, so the central category-level claim is not identifiable from the reported data.","rationale":"The reader's weakest assumption names the single-scenario design; I agree and treat it as the most load-bearing threat. The paper's headline empirical contribution is a category-level generalization: dark patterns in MR are harmful. That generalization is only as strong as the mapping from one concrete visualization to the abstract pattern. Because every cell has exactly one stimulus, the design is saturated at the stimulus level; pattern and exemplar are the same variable. This is not fixable by additional analyses of the current dataset (e.g., no variance component can be estimated). The limitation section is transparent about this, and that transparency is a strength. Nevertheless, the causal language in Section 4.1 ('We found that all dark patterns significantly impact...') goes beyond what the design can identify. The same issue also threatens the secondary 'similar impact' claim, which additionally lacks an equivalence test; the 'similar' pattern in Figures 3a and 4b is an informal visual assessment. However, since the reader's verdict was already CONDITIONAL and explicitly cited the single-scenario limitation, my read does not move the verdict. The empirical directional effects are plausible, the statistics appear appropriate, and the authors' own caveats are honest. No additional concern—such as video-based presentation or unvalidated SDS use in MR—changes the assessment, as these are acknowledged and affect confidence rather than validity. Thus verdict_should_be: UNCHANGED.","tokens_in":24495,"tokens_out":5557,"duration_ms":64304,"concrete_test":"Run a stimulus-sampling replication: have 3–4 independent designers create alternate implementations of each of the four dark patterns for one target (e.g., product), producing 12–16 videos matched in length, camera path, and baseline augmentations. Have a fresh sample (N≈100) rate them with the same reactance, SDS, comfort, and intention-to-use items. Fit a mixed-effects model with dark pattern as a fixed effect and exemplar as a random effect. If the exemplar variance component is large relative to the pattern effect—or if pattern-level estimates differ in sign or magnitude across designers—the category-level attribution fails. As a minimal variant, compare two visually distinct 'Urgency' countdown designs: if their SDS/comfort scores differ substantially, scenario-specific effects are present.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—'all dark patterns significantly impact our quantitative measures negatively' (Section 4.1)—requires the dark-pattern category, not the specific video design, to be the causal driver. But each Dark Pattern × Target cell contains exactly one scenario (Table 1, Section 3.1.1), so the pattern factor is perfectly confounded with the concrete stimulus: e.g., Forced Registration on Person is a heart-and-lock overlay obscuring a face, Hiding Information on Product is an 'added to your cart' banner, and Urgency is a countdown timer. Ratings could therefore be driven by surface features (text content, occlusion, animation, layout) rather than by the defining deceptive mechanism. The authors acknowledge this in Section 4.4: 'we explored only one scenario per dark pattern,' but the consequence is stronger than a generalizability caveat: with one exemplar per cell, scenario-level variance is unidentifiable, so no statistical analysis of the present data can separate pattern effects from exemplar effects. The same confound undermines the similarity claim for Emotional or Sensory Manipulation and Hiding Information (Section 4.2), since those two exemplar sets share non-occluding text-based overlays that the other two patterns' exemplars do not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an online within-subject video study (N=74) in which four dark patterns taken from Gray et al.'s meso-level taxonomy were adapted to Mixed Reality and applied to three augmentation targets (place, product, person), plus a baseline condition. The authors measured reactance, System Darkness Scale, comfort, and intention to use, and report that all dark patterns significantly worsen these measures relative to baseline, with Forced Registration on a person and Hiding Information on a product being the most disruptive. They further argue that Emotional or Sensory Manipulation and Hiding Information produce similar user effects and that dark-pattern classifications should be reconsidered on the basis of user impact rather than design technique alone.","tokens_in":24645,"tokens_out":5215,"duration_ms":51311,"significance":"If the pattern-level claims are valid, this is a useful quantitative contribution to a literature that has so far largely relied on qualitative or speculative methods for dark patterns in Mixed Reality. The study uses a clear within-subject design, an a-priori power analysis, appropriate non-parametric repeated-measures analyses (Friedman test with Durbin-Conover post-hoc and Holm correction, ART ANOVA), and reports effect sizes. The inclusion of a baseline condition and attention checks is methodologically sound. The main weakness is that each Dark Pattern × Target cell contains exactly one video, which makes the dark-pattern category inseparable from the concrete visual design; this limits the central generalization to specific stimuli rather than dark-pattern types.","major_comments":[{"comment":"The design assigns exactly one video to each Dark Pattern × Augmentation Target cell, so the dark-pattern factor is perfectly confounded with the concrete stimulus design (e.g., Hiding Information on Product is an \"added to cart\" banner, Urgency is a countdown timer, and Forced Registration on Person is a heart-and-lock occlusion). Consequently, the statement in Section 4.1 that \"all dark patterns significantly impact our quantitative measures negatively, regardless of whether they target a place, product, or person\" cannot be attributed to the dark-pattern categories. The data support only that these particular 12 video stimuli differ from the baseline. The Section 4.4 caveat that \"we explored only one scenario per dark pattern\" understates the issue: exemplar-level variance is unidentifiable, so no statistical analysis of the present data can separate pattern effects from exemplar effects. This is load-bearing for the central claim.","section":"Section 4.1 and Table 1 / Section 3.1.1"},{"comment":"The claim that Emotional or Sensory Manipulation and Hiding Information have similar impacts and that current classifications should be re-evaluated is based on visual inspection of the interaction plots rather than a statistical equivalence test or a targeted interaction contrast. Furthermore, the exemplars of these two patterns share surface features that the other two patterns' exemplars do not (non-occluding text-based overlays), so the observed similarity is also confounded with the concrete designs. The conclusion that these two categories should be reconsidered is therefore not yet established by the reported analyses.","section":"Section 4.2 and Figures 3a/4b"},{"comment":"For intention to use, only a significant main effect of Dark Pattern is reported; no interaction effect or target-specific post-hoc contrasts are given. The Section 4.1 claim that effects hold \"regardless of whether they target a place, product, or person\" is thus stronger than the reported evidence for this dependent variable. The authors should either report the missing factorial results or soften the generalization to the measures for which target-level effects were actually examined.","section":"Section 3.4.5 / Section 4.1"}],"minor_comments":[{"comment":"The text says \"we recruited 79 participants but excluded five due to incorrect attention checks\" and then refers to \"These seven nonsensical attention checks\"; the relationship between the seven checks and five excluded participants should be clarified.","section":"Section 3.4.1"},{"comment":"The Matsuda Hyper-Reality reference is dated \"2026\", which appears to be a typo; the correct access/publication date should be verified.","section":"Reference [49]"},{"comment":"The pattern name is inconsistently written as \"Emotional or Sensory Manipulation\", \"Emotional Manipulation\", and \"Emotional Manipulation or Sensory\" across the text and tables; the terminology should be standardized.","section":"Section 3.1.1 / Table 1 / Table 3"},{"comment":"The phrase \"clearly recognized by users\" overstates what the System Darkness Scale measures: the SDS captures perceived system darkness, not recognition or correct identification of dark patterns as in Mildner et al.'s recognition task.","section":"Section 4.1"},{"comment":"The design is described as a \"4 × 3 + 1 baseline factorial design\" and later as a \"two-factorial\" analysis; it would be clearer to state explicitly that the baseline is not a factor level in the ART analysis but was used only for the separate baseline comparisons.","section":"Section 3 / Abstract"}],"recommendation":"major_revision","confidential_remarks":"The reviewer concern about the single-exemplar confound genuinely lands: the study's central category-level claims are not identifiable from the current design. I recommend major revision rather than rejection because the empirical materials, procedure, and measurements are valuable and the claims can be reframed to the level of the concrete stimuli or complemented with additional exemplars. The paper's publication venue appears to be a peer-reviewed venue, so the bar for the claims in the final version should match the evidence actually provided."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my read. The headline: this is a genuinely useful first cut at measuring dark patterns in mixed reality, with clean directional results, but the specific stimulus designs and the pattern-level conclusions are more tangled than the abstract admits.\n\nWhat is actually new: most prior MR dark-pattern work is speculative or qualitative (co-design workshops, app audits, scenario construction). This paper runs a 74-person within-subject video study, crossing four patterns from Gray et al.'s ontology with three targets, and measures reactance, perceived system darkness, comfort, and intention to use, each compared against a dark-pattern-free baseline. The baseline comparison is clean: all 12 dark-pattern videos produce statistically significant negative shifts on all four measures. The statistical toolkit (Friedman, ART, Durbin-Conover with Holm, a priori power analysis) fits the non-normal repeated-measures data, and effect sizes are reported. The limitations section is unusually honest, which earns real credit.\n\nSoft spots, in order of severity. First, the pattern factor is perfectly confounded with the single concrete visual per cell. Forced Registration on a person is a heart-and-lock overlay obscuring a face; Hiding Information on a product is an “added to your cart” banner; Urgency is a countdown timer. With one exemplar per Dark Pattern × Target cell, you can attribute the reactions to those exact stimuli, but you cannot statistically separate “Forced Registration as a category” from “that particular overlay.” The authors flag this in Section 4.4, but the stress-test note is right that it goes beyond a generalizability caveat: the category-level claim is not identifiable from these data. Second, the “similar impact” claim for Emotional and Sensory Manipulation and Hiding Information rests on visual inspection of interaction plots, not an equivalence test, and the two patterns’ exemplars share surface features (text overlays, no occlusion) that the other patterns’ exemplars lack. That part of the Discussion is speculative. Third, the System Darkness Scale was validated for web shops; the authors acknowledge this but still build recommendations on it. Minor point: stimuli are only available on request, which makes independent replication harder than necessary.\n\nWho is this for? Researchers working on deceptive design, MR ethics, and regulation will get value from having this quantitative baseline, with the caveat that the specific visualizations did the work. It deserves a serious referee: the design is coherent, the analysis is sound, and the limitations are stated. My recommendation: engage with it, but treat the category-level conclusions as hypotheses for future multi-exemplar studies, not as established effects.","headline":"Solid first quantitative study of dark patterns in MR, but the one-stimulus-per-cell design means the category-level claims are weaker than the abstract suggests.","tokens_in":25251,"tokens_out":3503,"would_cite":true,"duration_ms":38829,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that dark patterns—deceptive overlays in mixed reality—significantly reduce user comfort and intention to use MR glasses while increasing reactance and perceived system darkness, across all four patterns and all three…","keywords":["dark patterns","deceptive design","mixed reality","user study","reactance","system darkness","intention to use","ethical design"],"falsifier":"Re-run the study with several independently produced designs per dark-pattern-target pair; if the within-pattern variance in comfort and reactance is as large as the between-pattern differences, the claim that dark patterns themselves—rather than their specific renderings—drive the effects is falsified.","tokens_in":24249,"feed_emoji":"🕶️","tokens_out":8186,"duration_ms":86575,"temperature":0.7,"pith_summary":"Dark patterns—deceptive interface techniques that push people into actions they did not intend—are usually studied in websites and apps; this paper asks whether they do measurable harm when transplanted into mixed reality. The authors try to establish that they do: in a two-factorial video study with 74 participants, four dark patterns applied to three real-world targets (a place, a product, a person) all significantly lowered comfort and intention to use MR glasses while raising reactance and perceived system darkness relative to a dark-pattern-free baseline. The most disruptive combinations involved personal interference (obscuring a person's face behind a paywall) and monetary manipulation (sneaking a product into the shopping cart). If the claim holds, deceptive design is not just a web problem but an emerging MR problem that designers and regulators should address before the hardware becomes widespread.","feed_headline":"Dark patterns in mixed reality push users away","feed_subtitle":"All four manipulative tactics cut comfort and intent to use and boosted reactance, whatever the target","key_machinery":"The study's engine is a two-factorial within-subject video study: 13 two-minute videos shot as a first-person city walk, structured as 4 dark patterns (Emotional or Sensory Manipulation, Forced Registration, Hiding Information, Urgency) times 3 augmentation targets (place, product, person), plus one dark-pattern-free baseline. Each pattern was adapted from the meso-level of a published dark-pattern ontology, chosen because that level is context-agnostic and can be interpreted for a specific application type. Participants rated every video on four instruments: a reactance questionnaire, the System Darkness Scale (a validated measure of how manipulative a system feels), a single comfort item, and a two-item intention-to-use scale, allowing the authors to separate the effect of the pattern, the effect of the target, and their interaction against the same baseline of ordinary MR aids such as navigation and weather overlays.","core_discovery":"The paper's central claim is that all four dark patterns it tested—Emotional or Sensory Manipulation, Forced Registration, Hiding Information, and Urgency—significantly harm the mixed-reality user experience no matter what real-world object they are attached to. Across 74 participants who each watched 13 videos of a simulated city walk through MR glasses, every pattern scored worse than the baseline on all four measures: higher reactance, higher system darkness, lower comfort, and lower intention to use the glasses. The authors further claim that the impact is most severe when the pattern touches personal identity or money, and that two patterns from different high-level families, Emotional or Sensory Manipulation and Hiding Information, produced similar effects, suggesting that dark-pattern taxonomies should incorporate measured user impact rather than only design technique.","pith_inferences":["My inference: the 'She likes you' overlay produced relatively low reactance, so a dark pattern that flatters may slip past users' defenses more easily than an openly obstructive one; an impact-based measure built only on felt resistance could therefore underrate the most effective manipulations.","My inference: because the effect clusters track the consequence the user experiences (losing money, losing control of personal appearance) rather than the UI tactic, a consequence-based taxonomy would likely generalize to augmented-reality advertising and diminished-reality interfaces, not just the four scenarios tested.","My inference: the video method probably understates or shifts attention effects compared with a worn headset, so a headset-based replication using the same stimulus designs would be a direct test of whether the reported effect sizes transfer to real MR use."],"forward_implications":["Designers of MR applications now have quantitative evidence that these four manipulative tactics reduce users' comfort and willingness to adopt the product, so deploying them is likely to backfire commercially.","The severe reactions to face-obscuring registration and cart-sneaking suggest that applications targeting personal identity or direct monetary loss will encounter the strongest user resistance.","The similar user impact of Emotional or Sensory Manipulation and Hiding Information supports moving toward classifications that group dark patterns by how they affect people, not only by how they are built.","Regulation and automated detection tools for dark patterns, already debated for websites, can now be extended to MR on the basis of measurable user responses rather than speculation."],"supporting_citations":[{"why":"Supplies the working definition of dark patterns as tricks that make users do things they did not mean to do.","marker":"[5]"},{"why":"Provides the reactance questionnaire used to measure perceived threats to the user's freedom of choice.","marker":"[19]"},{"why":"Supplies the ontology and the four meso-level dark patterns that the study adapts to mixed reality.","marker":"[29]"},{"why":"Defines the appropriate baseline for measuring dark-pattern effects, which the study follows for its control condition.","marker":"[48]"},{"why":"Provides the neutral baseline augmentations (navigation, weather, contextual information) used in every video.","marker":"[70]"},{"why":"Supplies the single-item comfort measure and prior evidence that augmentation targets affect comfort.","marker":"[73]"},{"why":"Provides the System Darkness Scale used to measure how manipulative or dark participants perceived each MR scenario.","marker":"[83]"},{"why":"Supplies the two-item intention-to-use scale from the technology acceptance literature.","marker":"[84]"}],"fun_headline_variants":["Dark patterns in MR backfire, cutting comfort and usage intent","Manipulative MR designs push users away, study finds","Mixed reality dark patterns harm user experience across the board","Forced registration and hidden info hurt MR users equally","Dark patterns in mixed reality lower trust and comfort"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results are interpreted as effects of the dark pattern category, but each category was shown through one concrete visual design, so a different design might have produced different ratings.","fun_headline_variants_meta":{"raw":{"variants":["Dark patterns in MR backfire, cutting comfort and usage intent","Manipulative MR designs push users away, study finds","Mixed reality dark patterns harm user experience across the board","Forced registration and hidden info hurt MR users equally","Dark patterns in mixed reality lower trust and comfort"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1257,"prompt_tokens":852,"completion_tokens":405,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":468,"completion_tokens_details":{"reasoning_tokens":327}},"tokens_in":468,"tokens_out":405,"duration_ms":4270,"temperature":1.0,"reasoning_tokens":327,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:48:47.170658+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the study with several independently produced designs per dark-pattern-target pair; if the within-pattern variance in comfort and reactance is as large as the between-pattern differences, the claim that dark patterns themselves—rather than their specific renderings—drive the effects is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the reactance questionnaire used to measure perceived threats to the user's freedom of choice."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the neutral baseline augmentations (navigation, weather, contextual information) used in every video."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the System Darkness Scale used to measure how manipulative or dark participants perceived each MR scenario."}],"review_version":1}