{"id":"f29d1854-11b1-48cc-baf0-38d048fe64d3","arxiv_id":"2412.13030","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Interviews with 17 data experts show skepticism toward differentially private synthetic data, a last-resort stance, and a demand for validation against real data.","lead":"The paper presents 17 interviews with data experts about differentially private synthetic data, finding that most do not use it, view it as a last resort, and want proof that it gives the same results as real data. The findings matter because they challenge the utility of current sanitized benchmarks and propose concrete changes: partner-vetted use cases, public standards of evidence, and tiered data access.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sampling via privacy-specific listserves and the authors' network selects on the outcome; with 17 US-based experts and an 'all participants' claim that contradicts the paper's own 14/17 figure, the central generalization to data experts is not yet established.","rationale":"Good-faith reading: this is a carefully reported qualitative study with an IRB-approved protocol, a published codebook, transparent recruitment details, saturation rationale, and a limitations section. It does not pretend to be a statistical survey; the issue is not the absence of random sampling per se, but the alignment between the sampled population and the population named in the central claim. The title and abstract address 'data experts' broadly, and the recommendations are written as universal normative statements ('organizations should publish standards'), so the sample must carry more weight than a purely internal qualitative description. The strongest reason to worry is selection: recruiting from privacy-focused listserves and the authors' network plausibly over-represents people who are already thinking about DP, and the high median Privacy Prior (6/10) and academic weighting (11/17) confirm the sample skews toward privacy-engaged academics. Because the outcome of interest - willingness to use DP synthetic data and demand for validation - is exactly what privacy-engaged people are likely to have opinions about, the observed 'last resort' stance may be an artifact of who volunteered. The paper modestly acknowledges generalizability limits in Section 5.5, but the abstract and conclusion do not carry that caveat with the same strength. Additionally, the 'all participants' phrasing is literally contradicted by the paper's own Figure 2 (14/17), which indicates the narrative has drifted from the evidence. That drift matters because Recommendations 1 and 2 are explicitly grounded in the universality of the validation demand. None of this impugns the authors' honesty; it is a correctable evidentiary gap. The reader's conditional verdict already captures most of this, so I do not recommend moving the verdict. The concrete check - a stratified replication, or at minimum a within-sample stratification by recruitment source and Privacy Prior - would determine whether the concern lands.","tokens_in":1137,"tokens_out":2547,"duration_ms":74900,"concrete_test":"Conduct a pre-registered replication with a stratified sample of at least 24 data experts recruited through general data-science/ML and domain-specific professional venues (economics, medicine, policy) rather than privacy listserves, with quotas for sector (industry/academic/government) and geography; administer the same interview protocol and code transcripts with the published codebook. Report prevalence for three codes: 'currently uses DP synthetic data', 'describes it as last resort', and 'expresses need for validation against real data.' If the validation rate falls below 80% or the 'last resort' characterization is not a majority, the central claim and recommendations do not survive generalization. As a cheaper first pass, stratify the existing 17 by recruitment source and Privacy Prior and re-check Figure 2 and the last-resort theme within the low-PP and industry subgroups.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Participants were recruited through public postings on data-privacy and synthetic-data listserves/Slack channels, supplemented by snowball sampling through the authors' professional network (Section 3). All 17 were US-based at interview time, 11/17 had above-median Privacy Prior scores, and only 6/17 were from industry (Table 2, Figure 2). This convenience sample is selected on interest in the very technology the study asks about: people who respond to privacy-themed calls and are connected to the authors are more likely to have formed opinions about DP synthetic data, and those opinions may be more skeptical or more specialized than the population of 'data experts' defined in Section 1. The headline findings - most do not currently use it, would only use it as a last resort, and require real-data validation - are prevalence-like claims, yet the sampling strategy supplies no basis for estimating population prevalences. The internal contradiction sharpens the point: Section 1.2 and the Conclusion say 'all participants' expressed that validation against real data is required, but Figure 2 reports 14/17 (82%), not 17/17. Thus the strongest quantitative anchor for Recommendations 1 and 2 is weaker than the text claims, and the recommendations are presented as universal 'should' statements grounded in that universality. This is the load-bearing concern because if a broader, less privacy-engaged sample shows lower validation demand or more routine use, the paper's central 'not buying in' narrative and its recommendations would need to be substantially reframed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative interview study of 17 U.S.-based data experts on their perspectives toward differentially private (DP) synthetic data. Through semi-structured interviews, the authors elicit participants' general privacy views, their familiarity with privacy concepts (via a 'Privacy Prior' score), and their reactions to a hypothetical DP synthetic medical dataset. Thematic analysis yields findings that most participants do not currently use DP synthetic data, view it as a last resort, and demand validation against real data before trusting it. These findings motivate three recommendations: partner-vetted evidence of validation (Recommendation 1), published discipline-specific standards of evidence (Recommendation 2), and a tiered 'driver's license' data access model (Recommendation 3). The paper includes the interview protocol, a hierarchical codebook with example quotations, and a participant table.","tokens_in":25480,"tokens_out":2170,"duration_ms":22287,"significance":"If the central claims hold, the paper makes a useful contribution to the growing literature on the practical adoption of differentially private synthetic data. Its strengths include a transparent qualitative methodology: the full interview guide is in Appendix B, the two-tier codebook with representative quotes is in Appendix C, and participant characteristics are tabulated in Table 2. The paper also engages seriously with prior interview and survey work (Table 1) and situates its recommendations in participants' own words. The main value is in identifying a specific evidentiary barrier to adoption—participants' desire for real-data validation—and translating it into concrete, actionable recommendations for DP researchers and deploying organizations. However, the paper's prevalence-style claims ('most participants,' 'all participants') go beyond what a convenience sample of 17 can support, and one such claim is internally contradicted by the paper's own Figure 2. These issues are fixable within the manuscript's scope, but they are load-bearing for the paper's central generalization.","major_comments":[{"comment":"The text claims in Section 1.2 and again in Section 6 that 'all participants expressed that validation against real data is required' and that 'all participants agreed on the necessity of validating synthetic data against real data.' Figure 2, however, reports that only 14 of 17 participants (82%) expressed a need or desire for validating DP outputs against real data. This is a direct internal contradiction, not a matter of interpretation: the paper's own quantitative summary of the coded data contradicts the universal claim. Because Recommendation 1 and Recommendation 2 are presented as grounded in unanimity ('all participants agreed'), the recommendations currently rest on a stronger evidentiary foundation than the data provide. Please correct the claim to reflect the actual count, and adjust the corresponding framing in the abstract, Section 1.2, and Section 6.","section":"1.2 and Section 6, versus Figure 2"},{"comment":"The recruitment strategy in Section 3 relies on public postings to data-privacy and synthetic-data listserves and Slack channels, plus snowball sampling through the authors' professional network. Table 2 shows that all 17 participants were U.S.-based, 11 of 17 were from academic or academic/government roles, the median Privacy Prior was 6/10, and 11 of 17 scored above the median. This is a convenience sample selected for interest in and connection to privacy topics, which is problematic for the paper's prevalence-like claims ('most participants reported that they do not currently use...', 'respondents considered it more as a last resort') in Section 1.2 and Section 6. The paper's own Section 5.5 acknowledges limited generalizability, but that limitation is not carried through to the headline findings and recommendations. Please temper the prevalence claims to the interviewed sample, or provide additional justification for why this sample is representative of the broader population of 'data experts' as defined in Section 1.","section":"3, Recruitment; Table 2; Section 5.5"},{"comment":"The paper's central assertion that participants view DP synthetic data as 'a last resort' is supported by illustrative quotes but is not quantified in Figure 2 or in any summary table. Figure 2 reports counts for related constructs (e.g., skepticism of existing DP methods, desire for real-data validation) but not for the 'last resort' theme. The manuscript would be stronger if the authors reported how many of the 17 participants actually expressed this view, and how that count is distributed across the high-PP and low-PP groups. As written, the 'last resort' finding is presented as a general result but lacks the same transparent accounting that the authors provide for other, less central claims.","section":"Section 1.2 and Section 4.2, 'last resort' claim"}],"minor_comments":[{"comment":"In the Participants and Recruitment subsection, the paper refers to 'Table ?? in Section ?? of the appendix' for aggregate race and demographic information. This is a broken cross-reference; no such table appears in the appendix included with the manuscript.","section":"Section 3, 'Table ??' references"},{"comment":"The Privacy Prior scoring (2/4/2/4 points for de-identified data, k-anonymity, and differential privacy) is presented as a 'general indicator of familiarity,' but no validation of this scoring is provided. Please add a brief note that the score is an unvalidated ad-hoc index, or acknowledge this in the limitations.","section":"Table 2 caption and Section 4.1"},{"comment":"There are several typographical errors, including 'heirarchical' in Section 3 (Coding and Analysis), 'Decenial' in Section 2, and inconsistent use of 'data' vs. 'dataset.' These do not affect the technical content but should be corrected in a revision.","section":"Throughout"},{"comment":"The caption describes participants as sorted by 'above average privacy priors,' but the cutoff is described elsewhere as 'above median' (Section 4.1). Please clarify whether the split is at the mean (5.8) or median (6), since Table 2 shows a median of 6 and the split is 11/6, which is consistent with a median split but not with an above/below-average split.","section":"Figure 2 caption"}],"recommendation":"major_revision","confidential_remarks":"This is a well-conducted qualitative study with transparent appendices, and the topic is timely for CHI/CSCW and privacy-oriented venues. The main fixable issue is the gap between the sample evidence and the universal or prevalence-style claims in the abstract, introduction, and conclusion. If the authors recalibrate the claims to the sample and correct the 'all participants' contradiction, the paper would be a solid contribution. I do not see grounds for rejection; the methodological limitations are normal for an interview study and are partially acknowledged in Section 5.5."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a genuine contribution—the first interview study I know of that centers differentially private synthetic data as the object of inquiry—but the text stretches the evidence in one place and the recruitment strategy limits how far the conclusions can generalize. The core direction is still credible.\n\nWhat's new and good: the paper interviews 17 data experts and gives them a concrete medical data scenario, then derives three concrete recommendations: partner-vetted use cases, published standards of evidence, and a tiered \"driver's license\" access model. The method is transparent: full interview protocol in the appendix, a codebook with example quotes, a participant table with privacy-prior scores, a positionality statement, and a saturation rationale. The related-work table is genuinely useful. The authors engage prior qualitative studies in this space rather than pretending to be the first to talk to practitioners.\n\nSoft spots: the \"all participants\" claim appears twice—Section 1.2 and the Conclusion—but Figure 2 reports 14/17 for real-data validation. That is a small but real internal inconsistency, and it matters because the recommendations are pitched as universal. The sampling is also a legitimate concern: recruitment started on privacy/SD listserves and snowballed through the authors' network; all participants were US-based; only 6 of 17 come from industry. The paper's own limitations section admits generalizability is limited, but the abstract and intro frame the findings as something like \"data experts think X.\" That framing overreaches. A careful revision should replace \"all participants\" with \"most participants\" and add a sentence or two making clear this is a convenience sample, not a representative survey.\n\nThere's one production issue: Section 3 references \"Table ?? in Section ??\" for demographic data. That's an obvious placeholder, but easy to fix.\n\nMy overall read: the direction of the findings is believable—practitioners demand validation against real data and are skeptical of drop-in replacement—and the paper provides a helpful vocabulary for thinking about adoption. It deserves a serious referee. I'd send it to review, with a request to fix the overclaim and soften the generalization.","headline":"A credible, well-conducted qualitative study with one internal inconsistency (\"all participants\" vs. 14/17) and a sampling frame that is narrower than the framing implies; the recommendations are useful and the paper deserves review.","tokens_in":25961,"tokens_out":2843,"would_cite":true,"duration_ms":28772,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Data experts see private synthetic data as a last resort until real-data validation is the norm.","keywords":["differential privacy","synthetic data","qualitative interview study","data experts","trust and validation","privacy utility trade-off","tiered data access"],"falsifier":"A representative survey of data experts drawn from outside privacy-focused channels that found most already use differentially private synthetic data in routine work and accept benchmark-based evidence without real-data validation would contradict the paper's central claim.","tokens_in":24987,"feed_emoji":"💬","tokens_out":4399,"duration_ms":41607,"temperature":0.7,"pith_summary":"This paper asks whether the people who actually work with sensitive data—researchers, data scientists, policy analysts, and medical professionals—are willing to use differentially private synthetic data as a stand-in for the original. Through 17 semi-structured interviews, it finds that most data experts do not currently use such data and would only reach for it as a last resort, even though they clearly see benefits in broader research access. The central barrier is trust: every participant said synthetic data must be validated against real data, but there is little consensus on what that validation should look like. From these findings the authors derive three recommendations: partner-vetted evidence of validation, published discipline-specific standards of evidence, and a tiered 'driver's license' model for access to sensitive data. The study matters because DP synthetic data is being promoted as a privacy-preserving way to share data, yet adoption depends on the evidentiary standards of the experts who would use it.","feed_headline":"Data experts see private synthetic data as a last resort","feed_subtitle":"Interviews with 17 data experts find trust depends on real-data validation; authors urge partner-vetted evidence and tiered access.","key_machinery":"The central mechanism is a semi-structured interview protocol built around a hypothetical medical tabular dataset and a differentially private synthetic version of it, analyzed with two-tier thematic coding. Participants were shown histograms and correlations from both the real and fake data and asked whether the comparisons were convincing, which elicited their evidentiary standards in concrete terms. The authors also construct a Privacy Prior score, a numeric index of each participant's familiarity with privacy concepts derived from their definitions of de-identified data, k-anonymity, and differential privacy, to position each response on a familiarity spectrum. The protocol carries the argument because it converts an abstract technology-acceptance question into specific judgments about evidence.","core_discovery":"The authors claim that data experts are broadly skeptical of adopting differentially private synthetic data and treat it as a last resort, and that this skepticism is rooted in epistemological concerns about generalizability and the risk of drawing wrong conclusions about individuals or underrepresented groups. They claim that validation against real data is a universal requirement among data experts, but the form of that validation remains contested. They further claim that current quantitative DP benchmarks, built on sanitized public datasets and proxy tasks, are insufficient to ground trust, and that the path to adoption runs through concrete, context-aware evidence and governance: partner-vetted use cases, published standards of evidence, and tiered access to sensitive data.","pith_inferences":["The paper's 17-person, US-centric, privacy-leaning sample likely overstates skepticism for the broader population of data experts; a representative sample could find more routinized adoption among some practitioner groups.","The 'chicken and the egg' problem named by one participant suggests a collective-action dynamic: a single visible partner-vetted success could shift the equilibrium faster than many more benchmark papers.","The same trust logic likely applies to LLM-era synthetic text data, where validation against real data is even harder, so a tiered access or sandbox model may become more necessary there.","A testable extension would be to ask participants which specific validation artifacts, such as confidence intervals on query answers, joint-distribution diagnostic plots, or full replication studies, would change their willingness to publish on synthetic data."],"forward_implications":["If the central claim is right, benchmark evaluations of DP synthetic data that report only proxy-task performance on sanitized datasets will not persuade actual users; evidence must come from a partner-vetted real use case.","Organizations that release DP synthetic data should publish their standards of evidence, tailored to each discipline's shared training, such as statisticians demanding precise error characterization and empiricists building application-specific benchmarks.","A tiered 'driver's license' access model, where researchers start with high-privacy, low-fidelity synthetic data and earn access to richer data, would align with the iterative and exploratory reality of research.","Trust, not technical privacy guarantees, is the deciding factor for uptake, so communication and community standards deserve as much attention as mechanism quality.","DP synthetic data is currently positioned for lower-stakes uses like testing and tinkering, while mission-critical applications will wait until validation norms mature."],"supporting_citations":[{"why":"Provides the formal definition of differential privacy and its guarantees, the technical baseline the study assumes throughout.","marker":"[23]"},{"why":"Documents the epistemic disagreements around the US Census Bureau's DP deployment, the motivating case for studying expert perspectives.","marker":"[18]"},{"why":"Industry practitioner interviews showing that most organizations do not deploy DP, giving the comparison point for practical gaps.","marker":"[31]"},{"why":"Interview study of DP tool users showing workflow disruption and evidentiary risk, which this study builds on for the synthetic-data case.","marker":"[67]"},{"why":"Survey of economists' familiarity with and attitudes toward DP, establishing the expert-side evidence base this study extends.","marker":"[77]"},{"why":"Interviews with DP practitioners showing lack of consensus on epsilon, which the authors contrast with their standards-of-evidence recommendation.","marker":"[22]"},{"why":"Study of how epsilon is explained to people, cited to show that communication gaps persist and to motivate the call for publishing standards of evidence.","marker":"[60]"},{"why":"Interview study of medical professionals on synthetic data uptake, the closest prior work on synthetic-data adoption this study distinguishes itself from.","marker":"[76]"}],"fun_headline_variants":["For data experts, synthetic data is plan B","Trust in synthetic data hinges on real-data checks","Experts urge tiered access to private synthetic data","DP synthetic data seen as last resort by experts","Data pros want context-aware privacy solutions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 17 interviewees, recruited through privacy-focused mailing lists and the authors' professional networks, stand in for the entire population of data experts the conclusions address.","fun_headline_variants_meta":{"raw":{"variants":["For data experts, synthetic data is plan B","Trust in synthetic data hinges on real-data checks","Experts urge tiered access to private synthetic data","DP synthetic data seen as last resort by experts","Data pros want context-aware privacy solutions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000527,"raw_usage":{"total_tokens":2493,"prompt_tokens":844,"completion_tokens":1649,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":460,"completion_tokens_details":{"reasoning_tokens":1580}},"tokens_in":460,"tokens_out":1649,"duration_ms":13888,"temperature":1.0,"reasoning_tokens":1580,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:28:44.400923+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A representative survey of data experts drawn from outside privacy-focused channels that found most already use differentially private synthetic data in routine work and accept benchmark-based evidence without real-data validation would contradict the paper's central claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the formal definition of differential privacy and its guarantees, the technical baseline the study assumes throughout."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the epistemic disagreements around the US Census Bureau's DP deployment, the motivating case for studying expert perspectives."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Industry practitioner interviews showing that most organizations do not deploy DP, giving the comparison point for practical gaps."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Interview study of DP tool users showing workflow disruption and evidentiary risk, which this study builds on for the synthetic-data case."},{"cited_title":"Williams, Joshua Snoke, Claire McKay Bowen, and Andrés F","cited_arxiv_id":null,"evidence_quote":"Survey of economists' familiarity with and attitudes toward DP, establishing the expert-side evidence base this study extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Interviews with DP practitioners showing lack of consensus on epsilon, which the authors contrast with their standards-of-evidence recommendation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Study of how epsilon is explained to people, cited to show that communication gaps persist and to motivate the call for publishing standards of evidence."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Interview study of medical professionals on synthetic data uptake, the closest prior work on synthetic-data adoption this study distinguishes itself from."}],"review_version":1}