REVIEW 4 major objections 4 minor 43 references
The paper claims that, in a high-trust civic AI setting, neither transparency nor control mechanisms shift data donation behavior, with donation rates uniformly high at 91.7% and Bayesian checks supporting genuine null effects.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 08:37 UTC pith:GG7DRMIW
load-bearing objection A clean, honestly reported null result that is more interesting for what the control condition reveals than for what the headline claims; the Bayesian 'evidence of absence' is prior-sensitive and overstates the case. the 4 major comments →
Trust by Context, Not by Design? A Quantitative Study of Data Donation Willingness for Open-Source Civic AI in Switzerland
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central discovery is a null result with positive content: in a high-trust civic setting, data donation decisions are driven by context rather than by interface design. In a 2x2 experiment with 205 Swiss residents, a Data Nutrition Label (transparency) and a granular consent dashboard (control) produced no significant differences in donation rates, which hovered between 90.0% and 93.3% across conditions (overall 91.7%). Logistic regressions gave odds ratios near 1 (transparency OR=1.02; control OR=1.46; interaction OR=0.89), and beta-binomial Bayes factors provided strong to extreme evidence for the null. The control manipulation did increase perceived control, a
What carries the argument
The experimental apparatus consists of two stimuli plus a behavioral outcome. The Data Nutrition Label is a visual, icon-based summary of model provenance, data sources, privacy safeguards, and output generation—designed to trigger heuristic transparency. The granular consent dashboard is a sequential configuration interface (data scope, purpose, storage location, retention period) operationalizing dynamic consent. The decision context is a functional civic AI chatbot, grounded in official Swiss ballot data, which gives the donation decision ecological validity. Analysis uses chi-square tests, nested logistic regressions, and beta-binomial Bayes factors, with manipulation checks (perceived t
Load-bearing premise
The load-bearing premise is that the transparency manipulation actually worked—that the Data Nutrition Label increased perceived transparency—but the manipulation check found no such increase (p=.480), so the paper's conclusion that transparency does not affect donation rests on an untested manipulation.
What would settle it
Run the same 2x2 experiment in a low-trust or commercial frame—say, a private company or a foreign-hosted service—with real (non-simulated) data transfer and a graded willingness scale; if donation drops well below 91.7% or the label/dashboard show measurable effects, the 'context, not design' conclusion would be falsified. A simpler check: measure donation rates when the identical interface is paired with commercial rather than civic framing.
If this is right
- If the paper is right, adding richer transparency and control widgets to civic AI consent screens will not increase donation rates in settings with high institutional trust and salient public benefit.
- The value of granular consent shifts from persuading users to donate to letting them negotiate scope, purpose, storage, and retention; designers should measure configuration choices, not just yes/no donation.
- Because the control manipulation worked while donation did not change, high donation and boundary-setting can coexist; privacy-preserving defaults matter even at near-universal donation rates.
- The Bayesian evidence (BF01 of 7.99–428.76) indicates the null effects are not merely underpowered; they are genuine absences of treatment effects in this context.
- The findings imply that institutional trust and framing do more work than interface design, so consent-design experiments in low-trust or commercial contexts may yield different results.
Where Pith is reading between the lines
- Editorial inference: Because the transparency manipulation check failed (MC-T p=.480), the paper does not actually test whether transparency can move donation; it tests a label that participants did not perceive as more transparent. A stronger or more novel transparency stimulus could still show an effect.
- Editorial inference: The simulated donation and civic academic framing may create demand characteristics; a replication with real data transfer or a neutral, non-civic frame would reveal how much of the 91.7% ceiling is an artifact of the setting.
- Editorial inference: The dashboard data suggest sovereignty is the dimension that matters: 68.6% chose Swiss or EU storage and 67.6% academic-only purpose. A direct test would vary server location or purpose while holding the interface identical and measure donation willingness.
- Editorial inference: The privacy-calculus explanation could be probed by varying query sensitivity (e.g., innocuous vs. politically charged topics) to see whether the ceiling breaks when perceived risk rises; that would convert the post-hoc interpretation into a testable prediction.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a 2x2 between-subjects experiment (N=205) in which Swiss residents interacted with a chatbot based on the Apertus-70B model and then chose whether to donate their anonymized conversation for training an open-source civic AI. The two factors were presence of a Data Nutrition Label and a granular consent dashboard. Donation rates were uniformly high (91.7% overall) with no significant main or interaction effects; the dashboard increased perceived control but not donation. Bayesian robustness checks (Beta(1,1) priors) were interpreted as strong evidence of absence. Open-ended responses emphasized public benefit and low perceived sensitivity. The authors conclude that in high-trust civic contexts, context matters more than interface design, and control functions mainly to let donors define terms of use.
Significance. If the null effects were robust, the paper would be a useful contribution to consent-mechanism research in civic AI, showing ceiling effects in high-trust settings and documenting how users exercise granular restrictions. The strengths include the embedded functional chatbot with sovereign infrastructure, the structured analysis plan, the reproducible codebook for qualitative coding, and the use of a real ballot-information task. However, the primary interpretive claims go beyond what the design can support: the transparency manipulation failed its manipulation check, the Bayesian absence evidence is prior-sensitive under ceiling conditions, and the study is underpowered for plausible small effects. The descriptive results on dashboard configuration choices and qualitative themes remain informative despite these limitations.
major comments (4)
- [Section 4.2, Table 5; Section 4.3; Abstract] The transparency manipulation failed its manipulation check (U=5518.0, p=.480, rank-biserial r=0.06). Consequently, H1 is not merely unsupported; it is untested. The null behavioral effect for the Data Nutrition Label cannot be interpreted as evidence that transparency has no effect because participants in the label conditions did not perceive greater transparency. The abstract and conclusion nevertheless state that 'neither transparency nor control significantly affected donation' and that 'Bayesian checks confirmed the absence of treatment effects.' The limitation in Section 5.2 partially acknowledges the failed manipulation, but the central claim is not qualified accordingly. This is load-bearing for the paper's conclusion that interface design has little leverage.
- [Section 4.3, 'Bayesian robustness check'] The Bayes factors are computed with Beta(1,1) priors on cell probabilities. With observed rates near 0.92 and cell sizes of 45–57, the independent-rates alternative under Beta(1,1) places most prior mass on very large probability differences (rates near 0 or 1), making the common-rate model look overwhelmingly better. BF01 values of 10.30, 7.99, and 428.76 are therefore prior-sensitive artifacts, not calibrated evidence about effect sizes of interest. A sensitivity analysis with informative priors centered on small effects, or a prior directly on the odds ratio, is required before claiming 'genuine absence.' This is load-bearing because the 'no leverage' conclusion depends on converting non-significance into evidence of absence.
- [Section 3.5 and Table 2] The power analysis assumed OR>=2.0, but with a baseline donation rate of 90%, OR=2.0 corresponds to roughly a 4.7 percentage-point increase. With n≈50 per arm, the design has low power to detect such effects, and the authors acknowledge power is only about 0.75 for this assumed size. Thus the experiment cannot rule out practically meaningful smaller effects. The Bayesian claim does not repair this problem because it uses the same low-information data. The conclusion that 'interface design has little leverage' exceeds what this design can establish.
- [Title, Section 5.1, Section 6] The study uses a single high-trust civic context (Swiss academic, non-commercial, Apertus infrastructure) and does not manipulate context. The claim encapsulated in the title, 'Trust by Context, Not by Design?', is therefore not directly tested. The high baseline rate could be due to social desirability or the simulated donation (both acknowledged in Section 5.2), and the single-cell design cannot compare context against design. The conclusion should be reframed as a descriptive finding within this particular high-trust context, not as a comparative claim that context matters more than design.
minor comments (4)
- [Section 3.4, Figure 4] The stimulus is described as an independently designed artifact following the Data Nutrition Project framework, but the text sometimes refers to it simply as 'the Data Nutrition Label.' Clarify early that this is an adapted, context-specific version, not the original MIT label, to avoid confusion for readers familiar with the original.
- [Table 12] The two storage rows ('Swiss servers only' and 'Swiss or EU servers') are not mutually exclusive, as the note states. Please clarify the response-option structure: is 'Swiss or EU' a separate option, or a collapsed category? The current presentation makes the 68.6% figure ambiguous.
- [Section 4.4, qualitative coding] The deterministic keyword-matching approach is reproducible, but it may miss paraphrases and context-dependent meanings. The paper states this only implicitly; a sentence acknowledging this limitation and its effect on theme frequencies would be helpful.
- [Abstract and Introduction] The phrase 'Bayesian checks confirmed the absence of treatment effects' is too strong given the prior sensitivity described above. Even after revision, the abstract should be worded to reflect that these checks are consistent with null effects under diffuse priors, not that absence is established.
Circularity Check
No significant circularity: the analysis is an empirical null-result study with no fitted predictions, no self-citation chain, and no result that reduces to its inputs by construction.
full rationale
The paper's claims are empirical and self-contained: donation rates are directly measured, hypotheses are tested with chi-square tests and logistic regression, and no parameter is fitted to one part of the data and then 'predicted' for a closely related quantity. The Bayesian robustness check uses a stated Beta(1,1) prior and is therefore a transparent statistical assumption, not a hidden fit; whether that prior is well chosen is a statistical validity question, not circularity. The manipulation-check failure for transparency (MC-T p=.480) is explicitly acknowledged as limiting interpretation of H1, which weakens the paper's claim but does not constitute circular reasoning. The 'context, not design' interpretation is an explanatory narrative built around the null result, not a derivation from the hypotheses. There are no self-citations among the references and no imported uniqueness theorems. The study's ceiling effect is a real limitation, but it does not make the analysis circular: the outcome was not defined in terms of the explanation, nor was the conclusion forced by a fitted input.
Axiom & Free-Parameter Ledger
free parameters (2)
- Beta(1,1) prior in Bayesian robustness check
- Assumed effect size OR >= 2.0 in power analysis
axioms (4)
- domain assumption Simulated donation decisions measure the same latent construct as real data donation
- domain assumption Privacy calculus (risk/benefit tradeoff) is the operative decision model in this civic context
- domain assumption The chatbot interaction was sufficiently realistic to evoke genuine donation intentions
- standard math Beta-binomial conjugate model is an appropriate statistical model for computing the reported Bayes factors
read the original abstract
Civic AI systems increasingly support democratic participation, yet interactions with them may reveal sensitive political views, creating tension between improving AI models and residents' expectations of privacy and consent. This study examines the conditions of transparency and user control under which Swiss residents are willing to donate their anonymized chatbot conversations to train an open-source AI model. A 2x2 between-subjects factorial design evaluated how a Data Nutrition Label and a granular consent dashboard influence donation decisions. The experiment was delivered via a multilingual online survey featuring a custom chatbot powered by the Apertus-70B model. Analysis of the 205 participants revealed that neither transparency nor control significantly affected donation behavior. Rates were uniformly high (91.7% overall), producing a ceiling effect, and Bayesian checks confirmed the absence of treatment effects. The dashboard raised perceived control but not donation, and high-control participants actively restricted their data-use settings. A qualitative analysis of 120 open-ended responses indicates that residents framed donation as a contribution to the public good, motivated by democratic participation, an open-source model, and research, while many regarded their anonymized queries as non-personal and therefore low in risk. Interpreted through the privacy calculus, a high perceived benefit coincided with a low perceived risk under high institutional trust, so both sides of the trade-off aligned and interface design had little leverage. Offering control served less to raise donation than to let residents define the terms of their contribution.
Figures
Reference graph
Works this paper leans on
-
[1]
Phi-4 technical report , year =
Abdin, Marah and Aneja, Jyoti and Behl, Harkirat and Bubeck, S. Phi-4 technical report , year =
-
[2]
Science , year =
Acquisti, Alessandro and Brandimarte, Laura and Loewenstein, George , title =. Science , year =
-
[3]
IEEE Security & Privacy , year =
Acquisti, Alessandro and Grossklags, Jens , title =. IEEE Security & Privacy , year =
-
[4]
and Adaji, Ifeoma and Ituma, Chidinma and Ricciardelli, Rosemary , title =
Andreotta, Adam J. and Adaji, Ifeoma and Ituma, Chidinma and Ricciardelli, Rosemary , title =. Ethics and Information Technology , year =
-
[5]
2025 , publisher =
Codebuch f. 2025 , publisher =
2025
-
[6]
and Busby, Ethan C
Argyle, Lisa P. and Busby, Ethan C. and Fulda, Nancy and Gubler, Joshua R. and Rytting, Christopher and Wingate, David , title =. Political Analysis , year =
-
[7]
, title =
Baum, Howell S. , title =. International Encyclopedia of the Social & Behavioral Sciences , edition =. 2001 , doi =
2001
-
[8]
Social Psychological and Personality Science , year =
Brandimarte, Laura and Acquisti, Alessandro and Loewenstein, George , title =. Social Psychological and Personality Science , year =
-
[9]
Journal of Personality and Social Psychology , year =
Chaiken, Shelly , title =. Journal of Personality and Social Psychology , year =
-
[10]
2022 , howpublished =
Chmielinski, Kasia and Newman, Sarah and Taylor, Matt and Joseph, Josh and Thomas, Kemi and Yurkofsky, Jessica and Holland, Sarah , title =. 2022 , howpublished =
2022
-
[11]
and Armstrong, Pamela K
Culnan, Mary J. and Armstrong, Pamela K. , title =. Organization Science , year =
-
[12]
2021 , howpublished =
Delgado, Fernando and Yang, Stephen and Madaio, Michael and Yang, Qian , title =. 2021 , howpublished =
2021
-
[13]
Information Systems Research , year =
Dinev, Tamara and Hart, Paul , title =. Information Systems Research , year =
-
[14]
Regulation (EU) 2016/679 (General Data Protection Regulation) , year =
2016
-
[15]
2023 , howpublished =
Gao, Yunfan and Xiong, Yun and Gao, Xinyu and Jia, Kangxiang and Pan, Jinliu and Bi, Yuxi and Dai, Yi and Sun, Jiawei and Wang, Meng and Wang, Haofen , title =. 2023 , howpublished =
2023
-
[16]
Apertus: Democratizing open and compliant
Hern. Apertus: Democratizing open and compliant. 2025 , howpublished =
2025
-
[17]
and Lutz, Christoph , title =
Hoffmann, Christian P. and Lutz, Christoph , title =. European Journal of Communication , year =
-
[18]
2018 , howpublished =
Holland, Sarah and Hosny, Ahmed and Newman, Sarah and Joseph, Joshua and Chmielinski, Kasia , title =. 2018 , howpublished =
2018
-
[19]
Jeffreys, Harold , title =
-
[20]
Nature Machine Intelligence , year =
Jobin, Anna and Ienca, Marcello and Vayena, Effy , title =. Nature Machine Intelligence , year =
-
[21]
and Lund, David and Morrison, Michael and Teare, Harriet and Melham, Karen , title =
Kaye, Jane and Whitley, Edgar A. and Lund, David and Morrison, Michael and Teare, Harriet and Melham, Karen , title =. European Journal of Human Genetics , year =
-
[22]
Information Systems Journal , year =
Kehr, Flavius and Kowatsch, Tobias and Wentzel, Daniel and Fleisch, Elgar , title =. Information Systems Journal , year =
-
[23]
, title =
Lee, Min Kyung and Kusbit, Daniel and Kahng, Anson and Kim, Ji Tae and Yuan, Xinran and Chan, Allissa and See, Daniel and Noothigattu, Ritesh and Lee, Siheon and Psomas, Alexandros and Procaccia, Ariel D. , title =. Proceedings of the ACM on Human-Computer Interaction , year =
-
[24]
, title =
Leonard, Peter G. , title =. 2018 , howpublished =
2018
-
[25]
Ethics in Design and Communication: Critical Perspectives , edition =
Madaio, Michael and Martin, Shari , title =. Ethics in Design and Communication: Critical Perspectives , edition =
-
[26]
and Cranor, Lorrie F
McDonald, Aleecia M. and Cranor, Lorrie F. , title =. I/S: A Journal of Law and Policy for the Information Society , year =
-
[27]
2025 , howpublished =
Murray-Rust, Dave and Alfrink, Kars and Zaga, Cristina , title =. 2025 , howpublished =
2025
-
[28]
MixtureVitae: Open web-scale pre-training dataset with high quality instruction and reasoning data built from permissive-first text sources , year =
Nguyen, Huu and May, Victor and Raj, Harsh and Nezhurina, Marianna and Wang, Yishan and Luo, Yanqi and Vu, Minh Chien and Nakamura, Taishi and Tsui, Ken and Nguyen, Van Khue and Salinas, David and Krasnod. MixtureVitae: Open web-scale pre-training dataset with high quality instruction and reasoning data built from permissive-first text sources , year =
-
[29]
and Horne, Daniel R
Norberg, Patricia A. and Horne, Daniel R. and Horne, David A. , title =. Journal of Consumer Affairs , year =
-
[30]
Governing with artificial intelligence: The state of play and way forward in core government functions , year =
-
[31]
Ouyang, Long and Wu, Jeffrey and Jiang, Xu and Almeida, Diogo and Wainwright, Carroll L. and Mishkin, Pamela and Zhang, Chong and Agarwal, Sandhini and Slama, Katarina and Ray, Alex and Schulman, John and Hilton, Jacob and Kelton, Fraser and Miller, Luke and Simens, Maddie and Askell, Amanda and Welinder, Peter and Christiano, Paul and Leike, Jan and Lowe...
-
[32]
2025 , howpublished =
Overney, Cassandra , title =. 2025 , howpublished =
2025
-
[33]
, title =
Peck, Joann and Shu, Suzanne B. , title =. Journal of Consumer Research , year =
-
[34]
and Kostova, Tatiana and Dirks, Kurt T
Pierce, Jon L. and Kostova, Tatiana and Dirks, Kurt T. , title =. Review of General Psychology , year =
-
[35]
Shklovski, Irina and Mainwaring, Scott D. and Sk. Leakiness and creepiness in app economies: Investigating experiences of surveillance and control in mobile checking-in , journal =. 2014 , pages =
2014
-
[36]
International Journal of Human-Computer Interaction , year =
Shneiderman, Ben , title =. International Journal of Human-Computer Interaction , year =
-
[37]
Environment and Planning B: Urban Analytics and City Science , year =
Sieber, Renee and Brandusescu, Ana and Sangiambut, Suthee and Adu-Daako, Anthony , title =. Environment and Planning B: Urban Analytics and City Science , year =
-
[38]
PLOS ONE , year =
Skatova, Anya and Goulding, James , title =. PLOS ONE , year =
-
[39]
IEEE Data Engineering Bulletin , year =
Stoyanovich, Julia and Howe, Bill , title =. IEEE Data Engineering Bulletin , year =
-
[40]
and Kim, Jang Hyun , title =
Sun, Seungjong and Lee, Eungu and Nan, Dongyan and Zhao, Xiangying and Lee, Wonbyung and Jansen, Bernard J. and Kim, Jang Hyun , title =. 2024 , howpublished =
2024
-
[41]
Federal Act on Data Protection (FADP) of 25 September 2020, as revised 1 September 2023 (SR 235.1) , year =
2020
-
[42]
, title =
Thomson, Ian and Boutilier, Robert G. , title =. SME Mining Engineering Handbook , edition =
-
[43]
Proceedings of the 11th International Conference on Communities and Technologies (C&T '23) , year =
Drobotowicz, Karolina and Truong, Nghiep Lucy and Ylipulli, Johanna and Gonzalez Torres, Ana Paula and Sawhney, Nitin , title =. Proceedings of the 11th International Conference on Communities and Technologies (C&T '23) , year =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.