{"id":"8d387f16-adc7-4fe5-9856-96dc2bdb0ba8","arxiv_id":"1908.10048","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A highlighted 'select all' button in consent dialogs increases accepted purposes, reduces accurate recall, and increases regret and perceived deception after users learn what they consented to, though the deception effect is not robust across sites.","lead":"A controlled experiment with 150 university students finds that a highlighted 'select all and confirm' button in a cookie consent dialog makes people accept more data collection purposes and remember their choice less accurately. The findings matter for regulators and designers because they provide behavioral evidence on consent-dialog features that may fall afoul of the GDPR's informed and freely given consent requirements.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Robustness table contradicts the perceived-deception claim: H2b is absent in Germany (p=0.564), so the abstract's deception result is not supported in both sites.","rationale":"The paper is a carefully conducted classroom experiment with random assignment, pre-tested instruments, honest reporting of the incomplete 2x2 design, and a behavioral effect (H1) that is large and statistically credible. The recall and regret findings are also plausible and consistent with the mechanism. My concern is narrower: the perceived-deception construct is one of three headline outcomes in the abstract, yet the site breakdown shows the effect does not replicate in Germany. The paper's claim that Table 6 confirms H1-H3 in both populations is demonstrably inaccurate for H2b, since Germany's two-sided p=0.564 is not rescued by the one-sided convention. Because the pooled test is only just below 0.05, the deception result is fragile. This is a correctness risk for the generalization claim, not a challenge to the central H1 finding. The reader's conditional verdict is appropriate; I would not change it, but the authors should either restrict the deception claim to the Austrian subsample or add a replication before publication.","tokens_in":120,"tokens_out":5743,"duration_ms":124552,"concrete_test":"Re-analyze PDE with a treatment-by-site interaction, e.g., a two-way ANOVA or robust regression with PDE ~ treatment + site + treatment:site on the T1 and control participants. If the interaction is significant, report per-site Cohen's d and 95% CIs; if the Germany-only CI for the T1-control difference includes zero while Austria's excludes it, the pooled H2b result is site-specific. As a simpler check, recompute the one-sided H2b test on the German subsample from the reported data; p=0.282, which directly falsifies the 'supported in both populations' claim and requires the abstract to be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing weakness is the generalized perceived-deception claim (H2b). The pooled t-test is only marginally significant (t(96.83)=2.24, p<0.05, d=0.44), and the paper's own Table 6 shows the effect is present in Austria (p=0.019) but absent in Germany (p=0.564). Section 6.2 states 'Table 6 confirms that H1-H3 are supported in both populations,' but this is false for H2b: even applying the authors' own remedy of halving the two-sided p-value for the smaller German subsample gives p=0.282. The abstract presents perceived deception as one of the headline outcomes, so this is an internal inconsistency in the robustness argument rather than a minor caveat. The effect could be driven by one site or by chance among the many tests reported. H1 and the recall/regret findings are more secure, so the paper's main behavioral result survives; but the deception-perception result needs qualification or replication before it can be stated as a general finding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a controlled classroom experiment (N=150 German-speaking university students in Austria and Germany) testing how two design features of GDPR consent dialogs affect users' consent decisions and perceptions. The three conditions are: a \"deceptive\" dialog with a highlighted \"Select all and confirm\" default button (T1), a reduced-choice dialog with only one purpose (T2), and a control dialog with three purposes and no highlighted default button. The authors hypothesize that the default button increases effective consent (H1), increases regret and perceived deception after users are informed of their choice (H2a, H2b), that multiple purposes increase response time (H3), and that multiple purposes increase perceived difficulty (H4). Results support H1 (T1 participants consented to significantly more purposes), H2a (regret increased after being informed in T1), and H3 (response time was longer for three purposes than one), while H4 is rejected. The paper interprets these findings as evidence that a highlighted default button can mislead users into accepting more data processing than intended, with policy implications for GDPR consent design.","tokens_in":21982,"tokens_out":4053,"duration_ms":39277,"significance":"The study addresses a timely and practically important question: whether specific design elements in GDPR consent dialogs distort users' consent decisions. The main behavioral result (H1) is important and appears well supported: participants who saw the highlighted default button accepted significantly more purposes, and the effect size is large relative to prior default-effect results. The paper also contributes a detailed, transparently described instrument with an attempt at cross-site replication, and it appropriately discusses ethical considerations and limitations. However, the headline claim about perceived deception (H2b) is fragile: the pooled effect is only marginally significant and the site-specific robustness analysis contradicts the paper's own claim of replication. Because the abstract and Section 6.1 present perceived deception as a central outcome, the generalization of this particular result needs to be qualified or substantiated with additional analysis before the paper can be considered robust.","major_comments":[{"comment":"The sentence \"Table 6 confirms that H1–H3 are supported in both populations\" is incorrect for H2b. Table 6 shows the perceived-deception effect is significant in Austria (p=0.019) but clearly absent in Germany (p=0.564). Even applying the authors' stated remedy of halving the two-sided p-value for the smaller German subsample yields p=0.282, not a significant effect. The abstract and Section 6.1 list perceived deception as a headline finding, so this is an internal inconsistency in the robustness argument rather than a minor caveat. The paper should either remove the generalizable deception claim, report the site-specific nature of the effect, or provide a formal location-by-treatment interaction test to demonstrate that the site difference is not meaningful. At minimum, the claim that Table 6 confirms support for H1–H3 in both populations must be corrected.","section":"Section 6.2, Table 6"},{"comment":"The pooled support for H2b rests on a marginally significant t-test (t(96.8279)=2.24, p<0.05, d=0.44) conducted after several other hypothesis tests on the same data. The paper does not apply any multiple-comparison correction, and the normality justification for PDE is the weakest among the constructs (Kolmogorov–Smirnov p=0.10 in Table 4). Given the site-specific null result in Germany, the perceived-deception finding is fragile. I recommend reporting a non-parametric Mann-Whitney U test as a robustness check and explicitly discussing the false-positive risk, or softening the conclusion to indicate the effect is suggestive and requires replication.","section":"Section 5.1, H2b"},{"comment":"The H3 comparison between T1 and T2 confounds the number of purposes with the specific purpose content: T2 presents only \"personalization\", which the authors themselves note is the most sensitive purpose, whereas T1 presents three purposes including personalization. The observed response-time difference (median 5.36s vs 3.16s) may therefore be due to the particular purpose shown rather than to choice proliferation per se. This is acknowledged only indirectly by calling T2 \"somewhat artificial\". Since H4 is rejected and the choice-proliferation contribution is secondary to the main H1 result, the paper should explicitly acknowledge this confound as a limitation of the H3 evidence.","section":"Section 4.1, H3"}],"minor_comments":[{"comment":"The t-test for H2b uses unpooled (Welch) degrees of freedom (t(96.8279)); please state explicitly that Welch's correction is used, as it affects the interpretation of the test.","section":"Section 5.1, H2b"},{"comment":"The Q-Q plots do not render as figures in the text; consider replacing them with a single combined figure or a verbal summary, since the current table is not informative to readers.","section":"Table 4"},{"comment":"The footnote stating that some p-values for Germany are above 5% only because two-sided tests are used is itself a form of one-sided reasoning applied post hoc. The one-sided correction is not applied consistently (it was not pre-registered), and for H2b it does not rescue the null result. Please revise the footnote for accuracy.","section":"Section 6.2, footnote 6"},{"comment":"In the description of pretests, the sentence \"During the test, we observed that several test subjects glanced at their neighbors' screens\" is informal; \"several\" would be better expressed as a count or proportion for precision.","section":"Section 4.2"},{"comment":"The phrase \"our experimental results confirm the common conjecture\" is a bit strong given that one of the four hypotheses was rejected and another is site-dependent; consider \"provide evidence consistent with\" instead.","section":"Section 6.1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: the main behavioral result is real and worth citing; the perceived-deception claim is softer than the abstract implies. The paper does something genuinely new: a controlled classroom experiment on a real-world GDPR consent-dialog pattern—a highlighted 'select all and confirm' button that overrides manual checkbox choices—with behavioral, recall, regret, and perception measures. The H1 effect is large and clean: 54% of T1 accepted all three purposes versus 23% of controls, and 34.5% of default-button clickers had also manually selected at least one purpose, suggesting they did not understand the button's semantics. The instrument is well documented and the two-country replication is a nice touch.\n\nThe main soft spot is H2b (perceived deception). The pooled t-test is only marginally significant (p<0.05), and the paper's own Table 6 shows the effect holds in Austria (p=0.019) but not in Germany (p=0.564). Yet Section 6.2 states that Table 6 'confirms that H1–H3 are supported in both populations.' That is false for H2b, and the footnote about one-sided tests does not fix it: halving the German p-value still gives 0.282. So the abstract's claim that participants 'perceive the consent dialog as more deceptive' is not robust across sites and needs qualification or replication.\n\nThe choice-proliferation tests (H3/H4) are weaker for a separate reason: T2 reduces three purposes to one, but the one kept purpose (personalization) is also the most sensitive, so count is confounded with content. The response-time result and the null for perceived difficulty cannot cleanly separate number-of-options from which-option. This does not threaten H1, but it limits what H3/H4 can say.\n\nThe paper is for usable-privacy and policy researchers. It deserves a serious referee: it is a real experiment with real stakes and mostly honest reporting. I would send it to peer review but require the authors to correct the H2b generalization and explicitly acknowledge the count/content confound. The central finding—deceptive default buttons increase effective consent and reduce accurate recall—stands.","headline":"Solid behavioral evidence on deceptive 'select all' buttons; the perceived-deception claim is not supported across both sites.","tokens_in":22575,"tokens_out":3405,"would_cite":true,"duration_ms":31786,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A highlighted 'select all and confirm' default button in a GDPR consent dialog makes users accept significantly more purposes, recall fewer of their choices, and later regret and distrust the dialog.","keywords":["web privacy","consent dialogs","GDPR","cookies","user study","choice proliferation","perceived deception","dark patterns"],"falsifier":"Run the same three-dialog experiment with balanced group sizes at several independent sites and pre-registered analyses that report the treatment-versus-control comparison separately for each site. The central consent effect would fail if the treatment group's count of consented purposes is not significantly higher than control ($\\chi^2(1)$, $p<0.05$); the deception claim would fail specifically if the perceived-deception difference is significant only in a minority of sites.","tokens_in":21612,"feed_emoji":"🍪","tokens_out":10324,"duration_ms":88403,"temperature":0.7,"pith_summary":"The paper asks whether the design of a GDPR cookie consent dialog can push users into agreeing to more data collection than they intend, and tests this with a controlled classroom experiment in which 150 university students searched for flights after encountering one of three dialogs. The central claim is that a highlighted default button labeled 'Select all and confirm' — which grants all three listed purposes no matter which boxes are checked — makes users effectively consent to significantly more purposes than a control dialog without such a button (54% accepted all three, versus 23.1%; non-parametric rank test $\\chi^2(1)=7.2$, $p<0.01$). Users who saw the button were also less able to correctly recall what they had consented to, and, once informed, reported more regret and perceived the website as more deceptive. Presenting one purpose instead of three had no significant effect on decisions or perceived difficulty, although it shortened response time. If the result is right, a common design element can undermine the 'freely given and informed' consent the GDPR requires, and regulators have concrete evidence to treat such dialogs as suspect.","feed_headline":"Highlighted cookie button lifts consent to 54 percent","feed_subtitle":"Users who click it recall less, regret more, and may not give the informed consent the GDPR requires.","key_machinery":"The load-bearing object is the experimental consent dialog itself: a blocking pop-up, copied from a real airline website, that lists three purposes (statistics, comfort, personalization) with initially unchecked checkboxes, plus a highlighted yellow button 'Select all and confirm' that grants all three purposes regardless of checkbox state and a colorless 'Confirm selection' button that grants only what was actively selected. The paper's central comparison is between this T1 dialog, a T2 dialog with only the personalization purpose, and a control dialog with the same three purposes but no highlighted button. The outcome is 'effective consent,' a score from 0 to 3 counting the purposes the site records as agreed to, independent of the user's intention; this is measured alongside free recall of the choice and multi-item scales for perceived deception, perceived difficulty, regret (measured before and after the user is informed of the effective choice), and privacy attitudes.","core_discovery":"The paper's central discovery is that a highlighted 'select all and confirm' button in a blocking multi-purpose consent dialog is not a neutral shortcut: it changes what users effectively consent to and how they feel about it afterward. In the treatment group, 54% ended up consenting to all three purposes versus 23.1% in the control group, a significant difference in the count of consented purposes ($\\chi^2(1)=7.2$, $p<0.01$). The accuracy of recall also dropped: 73.5% of the treatment group could correctly report their choice, versus 90.0% of the control group; among those who actually clicked the default button, correct recall fell to 55.6%. After being shown what they had effectively agreed to, treatment-group participants reported a significant increase in regret (paired $t(49)=2.81$, $p<0.01$, effect size $d=0.40$), and their perceived-deception score was significantly higher than control ($t(96.83)=2.24$, $p<0.05$, effect size $d=0.44$). In contrast, reducing the dialog from three purposes to one produced no significant difference in consented purposes or perceived difficulty, only a shorter response time.","pith_inferences":["The paper's own per-site breakdown leaves the perceived-deception effect open: because it reached significance in only one of the two locations, a pre-registered multi-site replication with per-site power is needed before treating 'users feel deceived' as a settled consequence of this dialog design.","The recall measure could be developed into an audit test for consent dialogs: if users who click the main button cannot correctly state what they agreed to, the dialog likely fails the GDPR's informed-consent requirement; validating that test against actual browsing behavior is the natural next step.","The study's convenience sample of computer-literate students makes the consent effect a lower-bound estimate for the general public; a field experiment on a live website with a non-student population could test whether susceptibility is even larger.","A design change already observed in practice — relabeling the button once a checkbox is selected — can be tested as a remedy: a follow-up experiment could compare uninformed consent and post-hoc regret between the static label and the dynamic label."],"forward_implications":["Designs that pair a highlighted select-all button with initially unchecked checkboxes can record consent that users themselves do not accurately remember, so such recorded consent should not be treated as evidence of informed, freely given agreement.","If regulators adopt recall as a proxy for consent quality, consent dialogs with accurate recall rates around 90% (control) versus 55–74% (treatment) would be distinguishable in audits.","Presenting up to three purposes in one dialog does not produce the choice-overload effects predicted by choice proliferation: perceived difficulty was not significantly higher, so multi-purpose dialogs can be designed without unavoidable overload.","Response time does increase with the number of purposes, so extending the result to dialogs with many more than three purposes is not supported by this study.","The default-button effect in a multi-purpose dialog is about four times larger than the previously reported binary default effect, so bundling purposes amplifies the nudge."],"supporting_citations":[{"why":"This concurrent field experiment on cookie banners provides the real-world baseline and the field-versus-lab comparison against which the paper positions its results.","marker":"[9]"},{"why":"This experimental study of choice proliferation in privacy settings supplies the paradigm and the satisfaction/regret constructs the paper adapts.","marker":"[13]"},{"why":"This field experiment on binary consent dialogs gives the prior default-effect estimate that the paper compares with its four-times-larger effect.","marker":"[31]"},{"why":"This scale-development work is the source of the perceived-deception items used to test the deception hypothesis.","marker":"[44]"},{"why":"This measurement study documents the prevalence and granularity of post-GDPR consent notices and motivates the design space explored here.","marker":"[8]"},{"why":"This ruling that pre-selected checkboxes do not constitute valid consent sets the legal baseline for the opt-in dialogs tested.","marker":"[15]"},{"why":"This systematization of dark patterns classifies designs such as the highlighted accept button as privacy-invasive and frames the policy interpretation.","marker":"[56]"},{"why":"This study of purpose-specific willingness to share data supports the choice of the three distinct purposes and the 0–3 scoring of effective consent.","marker":"[17]"}],"fun_headline_variants":["Highlighted consent button lifts clicks, sinks recall","Default cookie button: more consent, less memory","One button boosts consent to 54%, regret to match","Consent dialog: highlighted 'select all' cuts informed choice","GDPR consent: default button changes minds, not just clicks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result that users see the dialog as more deceptive depends on pooling the two study sites: in the site-by-site breakdown the effect is significant in Austria ($p=0.019$) but not in Germany ($p=0.564$), so if the samples are not combined, the deception claim does not robustly generalize.","fun_headline_variants_meta":{"raw":{"variants":["Highlighted consent button lifts clicks, sinks recall","Default cookie button: more consent, less memory","One button boosts consent to 54%, regret to match","Consent dialog: highlighted 'select all' cuts informed choice","GDPR consent: default button changes minds, not just clicks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000701,"raw_usage":{"total_tokens":3197,"prompt_tokens":1013,"completion_tokens":2184,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":2105}},"tokens_in":629,"tokens_out":2184,"duration_ms":19553,"temperature":1.0,"reasoning_tokens":2105,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:54:13.552884+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same three-dialog experiment with balanced group sizes at several independent sites and pre-registered analyses that report the treatment-versus-control comparison separately for each site. The central consent effect would fail if the treatment group's count of consented purposes is not significantly higher than control ($\\chi^2(1)$, $p<0.05$); the deception claim would fail specifically if the perceived-deception difference is significant only in a minority of sites.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This concurrent field experiment on cookie banners provides the real-world baseline and the field-versus-lab comparison against which the paper positions its results."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This experimental study of choice proliferation in privacy settings supplies the paradigm and the satisfaction/regret constructs the paper adapts."},{"cited_title":"Böhme, S","cited_arxiv_id":null,"evidence_quote":"This field experiment on binary consent dialogs gives the prior default-effect estimate that the paper compares with its four-times-larger effect."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This scale-development work is the source of the perceived-deception items used to test the deception hypothesis."},{"cited_title":"Degeling, C","cited_arxiv_id":null,"evidence_quote":"This measurement study documents the prevalence and granularity of post-GDPR consent notices and motivates the design space explored here."},{"cited_title":"Judgement in Case C-673/17 in the proceedings Bundesverband der Verbraucherzentralen und Verbraucherverbände vs Planet49 GmbH (2019)","cited_arxiv_id":null,"evidence_quote":"This ruling that pre-selected checkboxes do not constitute valid consent sets the legal baseline for the opt-in dialogs tested."},{"cited_title":"Bösch, B","cited_arxiv_id":null,"evidence_quote":"This systematization of dark patterns classifies designs such as the highlighted accept button as privacy-invasive and frames the policy interpretation."},{"cited_title":"Ackerman, L","cited_arxiv_id":null,"evidence_quote":"This study of purpose-specific willingness to share data supports the choice of the three distinct purposes and the 0–3 scoring of effective consent."}],"review_version":1}