Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Are Data Experts Buying into Differentially Private Synthetic Data? Gathering Community Perspectives

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Data experts see private synthetic data as a last resort until real-data validation is the norm.

desk verdict A credible, well-conducted qualitative study with one internal inconsistency ("all participants" vs. 14/17) and a sampling frame that is narrower than the framing implies; the recommendations are useful and the paper deserves review. read the letter →

arxiv 2412.13030 v1 pith:YCVFJXGH submitted 2024-12-17 cs.HC cs.CRcs.DB

classification cs.HCcs.CRcs.DB
keywords differentialprivacysyntheticdataqualitativeinterviewstudyexpertstrustandvalidationutilitytrade-offtieredaccess
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the people who actually work with sensitive data—researchers, data scientists, policy analysts, and medical professionals—are willing to use differentially private synthetic data as a stand-in for the original. Through 17 semi-structured interviews, it finds that most data experts do not currently use such data and would only reach for it as a last resort, even though they clearly see benefits in broader research access. The central barrier is trust: every participant said synthetic data must be validated against real data, but there is little consensus on what that validation should look like. From these findings the authors derive three recommendations: partner-vetted evidence of validation, published discipline-specific standards of evidence, and a tiered 'driver's license' model for access to sensitive data. The study matters because DP synthetic data is being promoted as a privacy-preserving way to share data, yet adoption depends on the evidentiary standards of the experts who would use it.

What carries the argument

The central mechanism is a semi-structured interview protocol built around a hypothetical medical tabular dataset and a differentially private synthetic version of it, analyzed with two-tier thematic coding. Participants were shown histograms and correlations from both the real and fake data and asked whether the comparisons were convincing, which elicited their evidentiary standards in concrete terms. The authors also construct a Privacy Prior score, a numeric index of each participant's familiarity with privacy concepts derived from their definitions of de-identified data, k-anonymity, and differential privacy, to position each response on a familiarity spectrum. The protocol carries the argument because it converts an abstract technology-acceptance question into specific judgments about evidence.

What would settle it

A representative survey of data experts drawn from outside privacy-focused channels that found most already use differentially private synthetic data in routine work and accept benchmark-based evidence without real-data validation would contradict the paper's central claim.

Watch

Extended reading notes

Core claim

The authors claim that data experts are broadly skeptical of adopting differentially private synthetic data and treat it as a last resort, and that this skepticism is rooted in epistemological concerns about generalizability and the risk of drawing wrong conclusions about individuals or underrepresented groups. They claim that validation against real data is a universal requirement among data experts, but the form of that validation remains contested. They further claim that current quantitative DP benchmarks, built on sanitized public datasets and proxy tasks, are insufficient to ground trust, and that the path to adoption runs through concrete, context-aware evidence and governance: partner-vetted use cases, published standards of evidence, and tiered access to sensitive data.

Load-bearing premise

The 17 interviewees, recruited through privacy-focused mailing lists and the authors' professional networks, stand in for the entire population of data experts the conclusions address.

Editorial extensions

If this is right

  • If the central claim is right, benchmark evaluations of DP synthetic data that report only proxy-task performance on sanitized datasets will not persuade actual users; evidence must come from a partner-vetted real use case.
  • Organizations that release DP synthetic data should publish their standards of evidence, tailored to each discipline's shared training, such as statisticians demanding precise error characterization and empiricists building application-specific benchmarks.
  • A tiered 'driver's license' access model, where researchers start with high-privacy, low-fidelity synthetic data and earn access to richer data, would align with the iterative and exploratory reality of research.
  • Trust, not technical privacy guarantees, is the deciding factor for uptake, so communication and community standards deserve as much attention as mechanism quality.
  • DP synthetic data is currently positioned for lower-stakes uses like testing and tinkering, while mission-critical applications will wait until validation norms mature.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's 17-person, US-centric, privacy-leaning sample likely overstates skepticism for the broader population of data experts; a representative sample could find more routinized adoption among some practitioner groups.
  • The 'chicken and the egg' problem named by one participant suggests a collective-action dynamic: a single visible partner-vetted success could shift the equilibrium faster than many more benchmark papers.
  • The same trust logic likely applies to LLM-era synthetic text data, where validation against real data is even harder, so a tiered access or sandbox model may become more necessary there.
  • A testable extension would be to ask participants which specific validation artifacts, such as confidence intervals on query answers, joint-distribution diagnostic plots, or full replication studies, would change their willingness to publish on synthetic data.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper reports a qualitative interview study of 17 U.S.-based data experts on their perspectives toward differentially private (DP) synthetic data. Through semi-structured interviews, the authors elicit participants' general privacy views, their familiarity with privacy concepts (via a 'Privacy Prior' score), and their reactions to a hypothetical DP synthetic medical dataset. Thematic analysis yields findings that most participants do not currently use DP synthetic data, view it as a last resort, and demand validation against real data before trusting it. These findings motivate three recommendations: partner-vetted evidence of validation (Recommendation 1), published discipline-specific standards of evidence (Recommendation 2), and a tiered 'driver's license' data access model (Recommendation 3). The paper includes the interview protocol, a hierarchical codebook with example quotations, and a participant table.

Significance. If the central claims hold, the paper makes a useful contribution to the growing literature on the practical adoption of differentially private synthetic data. Its strengths include a transparent qualitative methodology: the full interview guide is in Appendix B, the two-tier codebook with representative quotes is in Appendix C, and participant characteristics are tabulated in Table 2. The paper also engages seriously with prior interview and survey work (Table 1) and situates its recommendations in participants' own words. The main value is in identifying a specific evidentiary barrier to adoption—participants' desire for real-data validation—and translating it into concrete, actionable recommendations for DP researchers and deploying organizations. However, the paper's prevalence-style claims ('most participants,' 'all participants') go beyond what a convenience sample of 17 can support, and one such claim is internally contradicted by the paper's own Figure 2. These issues are fixable within the manuscript's scope, but they are load-bearing for the paper's central generalization.

major comments (3)
  1. [1.2 and Section 6, versus Figure 2] The text claims in Section 1.2 and again in Section 6 that 'all participants expressed that validation against real data is required' and that 'all participants agreed on the necessity of validating synthetic data against real data.' Figure 2, however, reports that only 14 of 17 participants (82%) expressed a need or desire for validating DP outputs against real data. This is a direct internal contradiction, not a matter of interpretation: the paper's own quantitative summary of the coded data contradicts the universal claim. Because Recommendation 1 and Recommendation 2 are presented as grounded in unanimity ('all participants agreed'), the recommendations currently rest on a stronger evidentiary foundation than the data provide. Please correct the claim to reflect the actual count, and adjust the corresponding framing in the abstract, Section 1.2, and Section 6.
  2. [3, Recruitment; Table 2; Section 5.5] The recruitment strategy in Section 3 relies on public postings to data-privacy and synthetic-data listserves and Slack channels, plus snowball sampling through the authors' professional network. Table 2 shows that all 17 participants were U.S.-based, 11 of 17 were from academic or academic/government roles, the median Privacy Prior was 6/10, and 11 of 17 scored above the median. This is a convenience sample selected for interest in and connection to privacy topics, which is problematic for the paper's prevalence-like claims ('most participants reported that they do not currently use...', 'respondents considered it more as a last resort') in Section 1.2 and Section 6. The paper's own Section 5.5 acknowledges limited generalizability, but that limitation is not carried through to the headline findings and recommendations. Please temper the prevalence claims to the interviewed sample, or provide additional justification for why this sample is representative of the broader population of 'data experts' as defined in Section 1.
  3. [Section 1.2 and Section 4.2, 'last resort' claim] The paper's central assertion that participants view DP synthetic data as 'a last resort' is supported by illustrative quotes but is not quantified in Figure 2 or in any summary table. Figure 2 reports counts for related constructs (e.g., skepticism of existing DP methods, desire for real-data validation) but not for the 'last resort' theme. The manuscript would be stronger if the authors reported how many of the 17 participants actually expressed this view, and how that count is distributed across the high-PP and low-PP groups. As written, the 'last resort' finding is presented as a general result but lacks the same transparent accounting that the authors provide for other, less central claims.
minor comments (4)
  1. [Section 3, 'Table ??' references] In the Participants and Recruitment subsection, the paper refers to 'Table ?? in Section ?? of the appendix' for aggregate race and demographic information. This is a broken cross-reference; no such table appears in the appendix included with the manuscript.
  2. [Table 2 caption and Section 4.1] The Privacy Prior scoring (2/4/2/4 points for de-identified data, k-anonymity, and differential privacy) is presented as a 'general indicator of familiarity,' but no validation of this scoring is provided. Please add a brief note that the score is an unvalidated ad-hoc index, or acknowledge this in the limitations.
  3. [Throughout] There are several typographical errors, including 'heirarchical' in Section 3 (Coding and Analysis), 'Decenial' in Section 2, and inconsistent use of 'data' vs. 'dataset.' These do not affect the technical content but should be corrected in a revision.
  4. [Figure 2 caption] The caption describes participants as sorted by 'above average privacy priors,' but the cutoff is described elsewhere as 'above median' (Section 4.1). Please clarify whether the split is at the mean (5.8) or median (6), since Table 2 shows a median of 6 and the split is 11/6, which is consistent with a median split but not with an above/below-average split.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: interview findings are grounded in participant responses rather than constructed from fitted inputs or self-referential definitions.

full rationale

No circularity identified. This is a qualitative interview study, not a formal derivation with fitted parameters, equations, or definitional equivalences. The load-bearing claims — that most participants do not currently use differentially private synthetic data, view it as a last resort, and demand validation against real data — are anchored in direct interview responses and coded transcripts (e.g., P17: “We [need to] see that the outcomes of our analyses would be the same across real and fake data sets,” and Figure 2). The real-data-validation theme is not definitionally forced by the interview protocol: participants were asked whether the shown real-versus-synthetic comparisons were convincing and could have accepted them, but instead asked for additional validation, which is an empirical response rather than a tautology. The paper’s self-citations (e.g., [63] and [64]) appear in literature reviews and lists of existing tools and evaluation approaches, but they are not load-bearing: the recommendations are supported by participant quotations and by comparison with external prior work such as Garrido et al., Dwork et al., and boyd and Sarathy. The discrepancy between the text’s “all participants” claim and Figure 2’s 14/17 count for validation demand is an internal consistency and quantification concern, not a circularity. No equation or fitted parameter is renamed as a prediction, and no uniqueness theorem or imported ansatz is used to force a conclusion. The study is self-contained as a qualitative investigation, so the appropriate circularity verdict is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no mathematical free parameters and no newly postulated entities. Its load-bearing assumptions are qualitative: saturation, the privacy-prior proxy, the single-analyst coding process, and the transferability of a medical-data scenario. These are domain assumptions a reader must accept for the findings to generalize.

assumptions (4)
  • domain assumption Thematic saturation at 17 interviews is sufficient to draw stable conclusions about RQ1 and RQ2.
    Section 3, 'Saturation and Sample Size', stops data collection at 17 based on recurring themes; saturation is a judgment call with no formal bound, and the authors acknowledge variability across participant backgrounds.
  • ad hoc to paper The Privacy Prior score computed from self-provided definitions of three terms is a valid proxy for data privacy expertise.
    The numeric PP index (2/4/4 points for complete definitions) is constructed for this paper and used to split participants into High/Low PP groups in Figure 2; its reliability and validity are not independently established.
  • domain assumption A single researcher who conducted the interviews and initial coding can produce themes representative of the whole group.
    Section 3, 'Coding and Analysis', relies on one researcher as ethnographer and initial coder, with multiple researchers reviewing codes but no inter-rater reliability reported; this creates potential for analyst bias.
  • domain assumption The hypothetical medical tabular data scenario is an appropriate probe for how data experts would evaluate DP synthetic data in their own domains.
    Section 3, Step 5, presents all participants with the same synthetic medical table and plots, but their actual work domains differ; the conclusions assume their responses transfer to economics, education, policy, and other fields.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Are Data Experts Buying into Differentially Private Synthetic Data? Gathering Community Perspectives." pith.science (2026). https://pith.science/paper/YCVFJXGH

@misc{pith2026241213030,
  author       = {Pith},
  title        = {Pith review of: Are Data Experts Buying into Differentially Private Synthetic Data? Gathering Community Perspectives},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YCVFJXGH}},
  note         = {Machine review of arXiv:2412.13030}
}
read the original abstract

Data privacy is a core tenet of responsible computing, and in the United States, differential privacy (DP) is the dominant technical operationalization of privacy-preserving data analysis. With this study, we qualitatively examine one class of DP mechanisms: private data synthesizers. To that end, we conducted semi-structured interviews with data experts: academics and practitioners who regularly work with data. Broadly, our findings suggest that quantitative DP benchmarks must be grounded in practitioner needs, while communication challenges persist. Participants expressed a need for context-aware DP solutions, focusing on parity between research outcomes on real and synthetic data. Our analysis led to three recommendations: (1) improve existing insufficient sanitized benchmarks; successful DP implementations require well-documented, partner-vetted use cases, (2) organizations using DP synthetic data should publish discipline-specific standards of evidence, and (3) tiered data access models could allow researchers to gradually access sensitive data based on demonstrated competence with high-privacy, low-fidelity synthetic data.

Figures

Figures reproduced from arXiv: 2412.13030 by the authors.

Figure 1
Figure 1. Materials used in Steps 3–5 of the interview. In the first part of the slide prompts, participants were presented with sample [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Absolute count of participants discussing general topics, out of a total of 17 participants. Presented overall, as well as sorted [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Full slides used for participant prompting during interviews, abbreviated in Figure 1 [PITH_FULL_IMAGE:figures/full_fig_p024_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy

    cs.CR 2025-12 conditional novelty 2.0 of 10

    A practical, extremely thorough survey of differentially private synthetic data generation: methods, privacy units, evaluation metrics, and end-to-end system components across four data modalities.

Reference graph

Works this paper leans on

84 extracted references · 55 canonical work pages · cited by 1 Pith paper

  1. [1]

    Christian Arnold and Marcel Neunhoeffer. 2020. Really Useful Synthetic Data–A Framework to Evaluate the Quality of Differentially Private Synthetic Data. arXiv preprint arXiv:2004.07740 (2020)

  2. [2]

    ATLAS.ti Scientific Software Development GmbH. 2023. ATLAS.ti Mac (version 23.2.1). Qualitative data analysis software. https://atlasti.com [Online; accessed 20-Jan-2024]

  3. [3]

    Sergul Aydore, William Brown, Michael Kearns, Krishnaram Kenthapadi, Luca Melis, Aaron Roth, and Ankit A Siva. 2021. Differentially private query release through adaptive projection. In International Conference on Machine Learning . PMLR, 457–467

  4. [4]

    Jane Bailey and Ian Kerr. 2007. Seizing control?: The experience capture experiments of Ringley & Mann. Ethics and information technology 9 (2007), 129–139

  5. [5]

    Rebecca Balebako, Abigail Marsh, Jialiu Lin, Jason Hong, and Lorrie Faith Cranor. 2014. The privacy and security behaviors of smartphone app developers. In Workshop on Usable Security. Citeseer, 1–10

  6. [6]

    Andrés F Barrientos, Aaron R Williams, Joshua Snoke, and CM Bowen. 2021. Differentially Private Methods for Validation Servers . Technical Report. Urban Institute research report

  7. [7]

    Subhajit Basu et al. 2012. Privacy protection: a tale of two cultures. Masaryk University Journal of Law and Technology 6, 1 (2012), 1–34

  8. [8]

    Barry Becker and Ronny Kohavi. 1996. Adult. UCI Machine Learning Repository. DOI: https://doi.org/10.24432/C5XW20

Show all 84 references
  1. [9]

    Kathrin Bednar, Sarah Spiekermann, and Marc Langheinrich. 2019. Engineering Privacy by Design: Are engineers ready to live up to the challenge? The Information Society 35, 3 (2019), 122–142

  2. [10]

    Joseph R Biden. 2023. Executive order on the safe, secure, and trustworthy development and use of artificial intelligence. (2023)

  3. [11]

    Alberto Blanco-Justicia, David Sánchez, Josep Domingo-Ferrer, and Krishnamurty Muralidhar. 2022. A critical review on the use (and misuse) of differential privacy in machine learning. Comput. Surveys 55, 8 (2022), 1–16

  4. [12]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101

  5. [13]

    Mir, and Evan M

    Brooke Bullek, Stephanie Garboski, Darakhshan J. Mir, and Evan M. Peck. 2017. Towards Understanding Differential Privacy: When Do People Trust Randomized Response Technique?. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems, Denver, CO, USA, May ...

  6. [14]

    Leonard E Burman, Alex Engler, Surachai Khitatrakun, James R Nunns, Sarah Armstrong, John Iselin, Graham MacDonald, and Philip Stallworth

  7. [15]

    Miranda Christ, Sarah Radway, and Steven M Bellovin. 2022. Differential Privacy and Swapping: Examining De-Identification’s Impact on Minority Representation and Privacy Preservation in the US Census. In2022 IEEE Symposium on Security and Privacy (SP) . IEEE Computer Society, ...

  8. [16]

    I need a better description

    Rachel Cummings, Gabriel Kaptchuk, and Elissa M. Redmiles. 2021. "I need a better description": An Investigation Into User Expectations For Differential Privacy. In CCS ’21: 2021 ACM SIGSAC Conference on Computer and Communications Security, Virtual Event, Republic of Korea, N...

  9. [17]

    Asmita Dalela, Saverio Giallorenzo, Oksana Kulyk, Jacopo Mauro, and Elda Paja. 2022. A study on security and privacy practices in Danish companies. In Usable Security and Privacy (USEC) Symposium 2022 . Internet society

  10. [18]

    danah boyd and Jayshree Sarathy. 2022. Differential perspectives: Epistemic disconnects surrounding the US Census Bureau’s use of differential privacy. Harvard Data Science Review (2022)

  11. [19]

    Fida Kamal Dankar and Khaled El Emam. 2013. Practicing differential privacy in health care: A review. Trans. Data Priv. 6, 1 (2013), 35–67

  12. [20]

    Jiahao Ding, Xinyue Zhang, Xiaohuan Li, Junyi Wang, Rong Yu, and Miao Pan. 2020. Differentially private and fair classification via calibrated functional mechanism. In AAAI, Vol. 34. 622–629

  13. [21]

    Cynthia Dwork. 2008. Differential privacy: A survey of results. In International conference on theory and applications of models of computation . Springer, 1–19

  14. [22]

    Cynthia Dwork, Nitin Kohli, and Deirdre Mulligan. 2019. Differential privacy in practice: Expose your epsilons!Journal of Privacy and Confidentiality 9, 2 (2019)

  15. [23]

    Cynthia Dwork, Aaron Roth, et al. 2014. The algorithmic foundations of differential privacy. Foundations and Trends® in Theoretical Computer Science 9, 3–4 (2014), 211–407

  16. [24]

    Pardis Emami-Naeini, Yuvraj Agarwal, Lorrie Faith Cranor, and Hanan Hibshi. 2020. Ask the experts: What should be on an IoT privacy and security label?. In 2020 IEEE Symposium on Security and Privacy (SP) . IEEE, 447–464

  17. [25]

    Daniele Fanelli. 2018. Is science really facing a reproducibility crisis, and do we need it to? Proceedings of the National Academy of Sciences 115, 11 (2018), 2628–2631

  18. [26]

    Joseph Ficek, Wei Wang, Henian Chen, Getachew Dagne, and Ellen Daley. 2021. Differential privacy in health research: A scoping review. Journal of the American Medical Informatics Association 28, 10 (2021), 2269–2276

  19. [27]

    AJ Flanagan, J King, and S Warren. [n. d.]. Redesigning data privacy: reimagining notice & consent for human-technology interaction (2020)

  20. [28]

    Lorenzo Frigerio, Anderson Santana de Oliveira, Laurent Gomez, and Patrick Duverger. 2019. Differentially private generative adversarial networks for time series, continuous, and discrete open data. In ICT Systems Security and Privacy Protection: 34th IFIP TC 11 International ...

  21. [29]

    Patricia I Fusch Ph D and Lawrence R Ness. 2015. Are we there yet? Data saturation in qualitative research. (2015)

  22. [30]

    Georgi Ganev, Bristena Oprisanu, and Emiliano De Cristofaro. 2021. Robin Hood and Matthew Effects–Differential Privacy Has Disparate Impact on Synthetic Data. arXiv preprint arXiv:2109.11429 (2021)

  23. [31]

    Gonzalo Munilla Garrido, Xiaoyuan Liu, Florian Matthes, and Dawn Song. 2023. Lessons Learned: Surveying the Practicality of Differential Privacy in the Industry. Proc. Priv. Enhancing Technol. 2023, 2 (2023), 151–170. https://doi.org/10.56553/POPETS-2023-0045

  24. [32]

    Greg Guest, Arwen Bunce, and Laura Johnson. 2006. How many interviews are enough? An experiment with data saturation and variability. Field methods 18, 1 (2006), 59–82

  25. [33]

    Greg Guest, Emily Namey, and Mario Chen. 2020. A simple method to assess and report thematic saturation in qualitative research. PloS one 15, 5 (2020), e0232076

  26. [34]

    Marco Gutfleisch, Jan H Klemmer, Niklas Busch, Yasemin Acar, M Angela Sasse, and Sascha Fahl. 2022. How does usable security (not) end up in software products? results from a qualitative interview study. In 2022 IEEE Symposium on Security and Privacy (SP) . IEEE, 893–910

  27. [35]

    Irit Hadar, Tomer Hasson, Oshrat Ayalon, Eran Toch, Michael Birnhack, Sofia Sherman, and Arod Balissa. 2018. Privacy by designers: software developers’ privacy mindset. Empirical Software Engineering 23 (2018), 259–289

  28. [36]

    Moritz Hardt, Katrina Ligett, and Frank McSherry. 2010. A simple and practical algorithm for differentially private data release. arXiv preprint arXiv:1012.4763 (2010)

  29. [37]

    Michael Hay, Ashwin Machanavajjhala, Gerome Miklau, Yan Chen, and Dan Zhang. 2016. Principled evaluation of differentially private algorithms using dpbench. In Proceedings of the 2016 International Conference on Management of Data . 139–154

  30. [38]

    Raquel Hill. 2015. Evaluating the Utility of Differential Privacy: A Use Case Study of a Behavioral Science Dataset. InMedical Data Privacy Handbook. Springer, 59–82

  31. [39]

    Shlomi Hod and Ran Canetti. 2024. Differentially Private Release of Israel’s National Registry of Live Births. arXiv preprint arXiv:2405.00267 (2024)

  32. [40]

    Naoise Holohan, Stefano Braghin, Pól Mac Aonghusa, and Killian Levacher. 2019. Diffprivlib: the IBM differential privacy library. ArXiv e-prints 1907.02444 [cs.CR] (July 2019)

  33. [41]

    Eric Horvitz and Deirdre Mulligan. 2015. Data, privacy, and the greater good. Science 349, 6245 (2015), 253–255

  34. [42]

    Leonardo Horn Iwaya, Muhammad Ali Babar, and Awais Rashid. 2023. Privacy Engineering in the Wild: Understanding the Practitioners’ Mindset, Organisational Aspects, and Current Practices. IEEE Transactions on Software Engineering (2023)

  35. [43]

    Priyank Jain, Manasi Gyanchandani, and Nilay Khare. 2016. Big data privacy: a technological perspective and review. Journal of Big Data 3 (2016), 1–25. Are Data Experts Buying into Differentially Private Synthetic Data? 21

  36. [44]

    Zhanglong Ji, Zachary C Lipton, and Charles Elkan. 2014. Differential privacy and machine learning: a survey and review. arXiv preprint arXiv:1412.7584 (2014)

  37. [45]

    Bailey Kacsmar, Vasisht Duddu, Kyle Tilbury, Blase Ur, and Florian Kerschbaum. 2023. Comprehension from Chaos: Towards Informed Consent for Private Computation. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security . 210–224

  38. [46]

    Farzaneh Karegar, Ala Sarah Alaqra, and Simone Fischer-Hübner. 2022. Exploring User-Suitable Metaphors for Differentially Private Data Analyses. In Eighteenth Symposium on Usable Privacy and Security (SOUPS 2022) . USENIX Association, Boston, MA, 175–193

  39. [47]

    We are a startup to the core

    Dilara Keküllüoğlu and Yasemin Acar. 2023. " We are a startup to the core": A qualitative interview study on the security and privacy development practices in Turkish software startups. In 2023 IEEE Symposium on Security and Privacy (SP) . IEEE, 2015–2031

  40. [48]

    Christopher T Kenny, Shiro Kuriwaki, Cory McCartan, Evan TR Rosenman, Tyler Simko, and Kosuke Imai. 2021. The use of differential privacy for census data and its impact on redistricting: The case of the 2020 US Census. Science advances 7, 41 (2021), eabk3283

  41. [49]

    Patrick Kühtreiber, Viktoriya Pak, and Delphine Reinhardt. 2022. Replication: The Effect of Differential Privacy Communication on German Users’ Comprehension and Data Sharing Attitudes. In Eighteenth Symposium on Usable Privacy and Security, SOUPS 2022, Boston, MA, USA, August...

  42. [50]

    It doesn’t tell me anything about how my data is used

    Lin Kyi, Abraham Mhaidli, Cristiana Teixeira Santos, Franziska Roesner, and Asia J Biega. 2024. “It doesn’t tell me anything about how my data is used”: User Perceptions of Data Collection Purposes. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–12

  43. [51]

    Tianshi Li, Yuvraj Agarwal, and Jason I Hong. 2018. Coconut: An IDE plugin for developing privacy-friendly apps. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 2, 4 (2018), 1–35

  44. [52]

    Bo Liu, Ming Ding, Sina Shaham, Wenny Rahayu, Farhad Farokhi, and Zihuai Lin. 2021. When machine learning meets privacy: A survey and outlook. ACM Computing Surveys (CSUR) 54, 2 (2021), 1–36

  45. [53]

    Terrance Liu, Giuseppe Vietri, and Steven Z Wu. 2021. Iterative methods for private synthetic data: Unifying framework and new methods.Advances in Neural Information Processing Systems 34 (2021), 690–702

  46. [54]

    Ashwin Machanavajjhala, Xi He, and Michael Hay. 2017. Differential privacy in the wild: A tutorial on current practices & open challenges. In Proceedings of the 2017 ACM International Conference on Management of Data . 1727–1730

  47. [55]

    Moira Maguire and Brid Delahunt. 2017. Doing a thematic analysis: A practical, step-by-step guide for learning and teaching scholars. All Ireland Journal of Higher Education 9, 3 (2017)

  48. [56]

    Nora McDonald, Sarita Schoenebeck, and Andrea Forte. 2019. Reliability and inter-rater reliability in qualitative research: Norms and guidelines for CSCW and HCI practice. Proceedings of the ACM on human-computer interaction 3, CSCW (2019), 1–23

  49. [57]

    Ryan McKenna, Gerome Miklau, and Daniel Sheldon. 2021. Winning the NIST Contest: A scalable and general approach to differentially private synthetic data. arXiv preprint arXiv:2108.04978 (2021)

  50. [58]

    Ryan McKenna, Brett Mullins, Daniel Sheldon, and Gerome Miklau. 2022. Aim: An adaptive and iterative mechanism for differentially private synthetic data. arXiv preprint arXiv:2201.12677 (2022)

  51. [59]

    Jack Murtagh, Kathryn Taylor, George Kellaris, and Salil Vadhan. 2018. Usable differential privacy: A case study with psi. arXiv preprint arXiv:1809.04103 (2018)

  52. [60]

    Priyanka Nanayakkara, Mary Anne Smart, Rachel Cummings, Gabriel Kaptchuk, and Elissa M Redmiles. 2023. What are the chances? explaining the epsilon parameter in differential privacy. In 32nd USENIX Security Symposium (USENIX Security 23) . 1613–1630

  53. [61]

    Ngong, Brad Stenger, Joseph P

    Ivoline C. Ngong, Brad Stenger, Joseph P. Near, and Yuanyuan Feng. 2023. Evaluating the Usability of Differential Privacy Tools with Data Practitioners. CoRR abs/2309.13506 (2023). https://doi.org/10.48550/ARXIV.2309.13506 arXiv:2309.13506

  54. [62]

    Daniel L Oberski and Frauke Kreuter. 2020. Differential privacy and social science: An urgent puzzle. Harvard Data Science Review 2, 1 (2020), 1–21

  55. [63]

    Lucas Rosenblatt, Anastasia Holovenko, Taras Rumezhak, Andrii Stadnik, Bernease Herman, Julia Stoyanovich, and Bill Howe. 2023. Epistemic Parity: Reproducibility as an Evaluation Metric for Differential Privacy. Proceedings of the VLDB Endowment (2023)

  56. [64]

    Lucas Rosenblatt, Xiaoyan Liu, Samira Pouyanfar, Eduardo de Leon, Anuj Desai, and Joshua Allen. 2020. Differentially private synthetic data: Applied evaluations and enhancements. arXiv preprint arXiv:2011.05537 (2020)

  57. [65]

    Steven Ruggles, Catherine Fitch, Diana Magnuson, and Jonathan Schroeder. 2019. Differential privacy and census data: Implications for social and economic research. In AEA papers and proceedings , Vol. 109. American Economic Association 2014 Broadway, Suite 305, Nashville, TN 3...

  58. [66]

    Shruti Sannon, Natalya N Bazarova, and Dan Cosley. 2018. Privacy lies: Understanding how, when, and why people lie to protect their privacy in multiple online contexts. In Proceedings of the 2018 CHI conference on human factors in computing systems . 1–13

  59. [67]

    Jayshree Sarathy, Sophia Song, Audrey Haque, Tania Schlatter, and Salil Vadhan. 2023. Don’t look at the data! how differential privacy reconfigures the practices of data science. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–19

  60. [68]

    Jeremy Seeman and Daniel Susser. 2024. Between privacy and utility: On differential privacy in theory and practice. ACM Journal on Responsible Computing 1, 1 (2024), 1–18

  61. [69]

    Michael Shoemate, Andrew Vyrros, Chuck McCallum, Raman Prasad, Philip Durbin, Sílvia Casacuberta Puig, Ethan Cowan, Vicki Xu, Zachary Ratliff, Nicolás Berrios, Alex Whitworth, Michael Eliot, Christian Lebeda, Oren Renard, and Claire McKay Bowen. [n. d.]. OpenDP Library . https...

  62. [70]

    Sean Sirur, Jason RC Nurse, and Helena Webb. 2018. Are we there yet? Understanding the challenges faced in complying with the General Data Protection Regulation (GDPR). In Proceedings of the 2nd International Workshop on Multimedia Privacy and Security . 88–95. 22 Rosenblatt et al

  63. [71]

    Mary Anne Smart, Dhruv Sood, and Kristen Vaccaro. 2022. Understanding risks of privacy theater with differential privacy. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2 (2022), 1–24

  64. [72]

    Ryan Steed and Alessandro Acquisti. 2024. Adoption of’Privacy-Preserving’Analytics: Drivers, Designs, & Decoupling. SSRN (2024)

  65. [73]

    Mohammad Tahaei, Alisa Frik, and Kami Vaniea. 2021. Privacy champions in software teams: Understanding their motivations, strategies, and challenges. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–15

  66. [74]

    Shun Takagi, Tsubasa Takahashi, Yang Cao, and Masatoshi Yoshikawa. 2021. P3gm: Private high-dimensional data release via privacy preserving phased generative model. In 2021 IEEE 37th International Conference on Data Engineering (ICDE) . IEEE, 169–180

  67. [75]

    Yuchao Tao, Ryan McKenna, Michael Hay, Ashwin Machanavajjhala, and Gerome Miklau. 2021. Benchmarking Differentially Private Synthetic Data Generation Algorithms. arXiv preprint arXiv:2112.09238 (2021)

  68. [76]

    Robin Daniël van Hoorn. 2024. On the Acceptance, Adoption, and Utility of Synthetic Data for Healthcare Innovation. (2024)

  69. [77]

    Williams, Joshua Snoke, Claire McKay Bowen, and Andrés F

    Aaron R. Williams, Joshua Snoke, Claire McKay Bowen, and Andrés F. Barrientos. 2023. Disclosing Economists’ Privacy Perspectives: A Survey of American Economic Association Members’ Views on Differential Privacy and the Usability of Noise-Infused Data. (2023)

  70. [78]

    Aiping Xiong, Tianhao Wang, Ninghui Li, and Somesh Jha. 2020. Towards effective differential privacy communication for users’ data sharing decision and comprehension. In 2020 IEEE Symposium on Security and Privacy (SP) . IEEE, 392–410

  71. [79]

    Aiping Xiong, Chuhao Wu, Tianhao Wang, Robert W Proctor, Jeremiah Blocki, Ninghui Li, and Somesh Jha. 2022. Using Illustrations to Communicate Differential Privacy Trust Models: An Investigation of Users’ Comprehension, Perception, and Data Sharing Decision.arXiv preprint arXi...

  72. [80]

    Aiping Xiong, Chuhao Wu, Tianhao Wang, Robert W Proctor, Jeremiah Blocki, Ninghui Li, and Somesh Jha. 2023. Exploring Use of Explanative Illustrations to Communicate Differential Privacy Models. In Proceedings of the Human Factors and Ergonomics Society Annual Meeting . SAGE P...

  73. [81]

    SIGMOD 2014

    Jun Zhang, Graham Cormode, Cecilia M Procopiuc, Divesh Srivastava, and Xiaokui Xiao. SIGMOD 2014. PrivBayes: Private Data Release via Bayesian Networks. (SIGMOD 2014)

  74. [82]

    Ying Zhao and Jinjun Chen. 2022. A survey on differential privacy for unstructured data content. ACM Computing Surveys (CSUR) 54, 10s (2022), 1–28

  75. [83]

    Public (Government, etc.)

    Zhiru Zhu and Raul Castro Fernandez. 2023. Making Differential Privacy Easier to Use for Data Controllers and Data Analysts using a Privacy Risk Indicator and an Escrow-Based Platform. arXiv preprint arXiv:2310.13104 (2023). Are Data Experts Buying into Differentially Private ...

  76. [2019]

    Technical report US, Internal Revenue Service (2019)

    Safely expanding research access to administrative tax data: creating a synthetic public use file and a validation server. Technical report US, Internal Revenue Service (2019). 20 Rosenblatt et al

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.