REVIEW 4 major objections 6 minor 1 cited by
Differentially private text generators lose most utility on specialized data
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Applying a new benchmark to five specialized domains, the paper shows current privacy-preserving text generators lose much of their utility and fidelity, especially at strict privacy levels and on gated datasets.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection Useful benchmark, credible qualitative degradation, but the 'open-domain overestimation' claim goes beyond the evidence and Section 4's numbers need a rewrite. the 4 major comments →
Evaluating Differentially Private Generation of Domain-Specific Text
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that current differentially private text generators are far less capable on specialist domains than their reported open-domain results suggest. Concretely, the authors benchmark DP-Gen (DP-SGD fine-tuning) and AUG-PE (private distribution-aligned evolution) on five datasets spanning clinical notes (HoC, N2C2'08), patient-reported drug effects (PsyTAR), financial news (DMSAFN), and legal refugee-case analysis (AsyLax), at privacy budgets epsilon in {0.5, 1, 2, 4} and infinity. Averaged over datasets, downstream classification utility under strict privacy (epsilon <= 4) sits at about 50 percent of real-data performance, nearly independent of epsilon; the best model
What carries the argument
The benchmark itself is the central mechanism: five domain-specific datasets, two DP generation paradigms—DP-SGD fine-tuning (DP-Gen) and private distribution alignment via a privatized evolution scheme (AUG-PE)—and a two-axis evaluation. Utility is measured by training downstream classifiers on synthetic data and testing them on a held-out real test set; fidelity is measured by surface overlap (BLEU, METEOR), semantic similarity (BERTScore, USE, MAUVE), and distributions of recognized entities and text lengths. The load-bearing design decisions are the use of gated-access datasets to reduce prior exposure in the model's pre-training data, and scoring improvements over random/majority baseli
Load-bearing premise
The results assume the gated-access datasets were not already memorized by the Llama-3 backbone during pre-training or closed post-training; if they were, the measured drop conflates memory leakage with domain difficulty.
What would settle it
Search Llama-3 outputs (or run a membership/canary audit) for verbatim or near-verbatim strings from N2C2'08 and PsyTAR; if such strings appear even without DP training, the prior-exposure control fails and the comparison against open-domain results is confounded. Alternatively, run the identical protocol on an open-domain dataset and show high utility, confirming the domain-specificity interpretation.
If this is right
- Utility of DP synthetic text in realistic medical, financial, or legal settings is far lower than previous open-domain numbers imply; deployments should expect near-baseline performance at strict privacy budgets.
- Gated-access or otherwise non-public data is essential for realistic evaluation, because public datasets inflate results through prior exposure and memorization.
- Generated text can look superficially plausible while losing domain-specific entities and structure, so surface-level quality metrics are insufficient for high-stakes domains.
- Averaging across classifier baselines can hide the privacy-utility trade-off; evaluation protocols should report best- and per-model performance separately.
- Current state-of-the-art methods need domain-specific adaptation—such as domain-aware embeddings in AUG-PE or handling the noise-to-signal ratio of long, jargon-heavy texts in DP-SGD—before they can support real data sharing.
- If these results are right, existing positive results on open-domain datasets should be re-derived on gated domains before being used as evidence of deployability.
Where Pith is reading between the lines
- A natural extension is to run the same benchmark with a backbone pre-trained or fine-tuned on public in-domain text before DP training; if utility recovers, domain mismatch rather than DP noise is the main bottleneck.
- The near-zero MAUVE values across all methods suggest generated text is detectably different from real domain text in embedding space; a human-preference study of which failures matter—entities, style, or factuality—would sharpen the metric set.
- The prior-exposure caveat implies that the effective privacy guarantee may be weaker than epsilon suggests when the underlying model has already seen similar data during closed post-training; this should be part of any deployment decision.
- The paper's protocol could serve as a template for adversarial auditing: the same gated datasets and relative-gain scoring could be used to test whether a new DP generator improves over the baselines reported here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a benchmark for evaluating differentially private (DP) synthetic text generation on domain-specific datasets. It evaluates two DP generators, DP-Gen (DP-SGD fine-tuning) and AUG-PE (private evolution), across five datasets (HoC, N2C2'08, PsyTAR, DMSAFN, AsyLax) with privacy budgets ε ∈ {0.5, 1, 2, 4} and ε = ∞. Utility is measured by downstream classification F1 of models trained on synthetic data; fidelity via MAUVE, BLEU/METEOR/BERTScore/USE, entity-overlap, and text-length divergences. The paper reports substantial degradation under DP, especially on gated-access, long-document datasets, and claims that open-domain evaluations overestimate real-world performance.
Significance. The benchmark addresses an important gap: prior evaluations used open-domain datasets and simple metrics. Strengths are the use of gated-access datasets to mitigate prior exposure, realistic ε choices, multiple utility and fidelity metrics, and public release of code and extended results. If supported, the central finding that current DP text generators degrade sharply on domain-specific data is practically important. The main limitations—only two methods, no variance estimates, and the unverified prior-exposure assumption—mean the quantitative claims should be read as preliminary rather than definitive. The paper is a useful first step but does not yet establish the general overestimation claim.
major comments (4)
- [Section 4 / Table 2] The baseline-adjusted summary is internally inconsistent and not reproducible from Table 2. The text says 'lowers this score to 55%/21% without privacy guarantees for DP-Gen/Aug-PE, and to an improvement of at most 28%/52% with ε≤4', then later reports '28%, 26%, 23%, 21%, 15% retention for AUG-PE at ε∈{inf,4,2,1,0.5}'. These numbers cannot both be 'improvement over baselines' and 'retention' without a clear statement of the formula. Table 2 values (e.g., AUG-PE on DMSAFN: avg F1 51.0 at ε=4 vs random 30.5, majority 41.1, original 76.8) do not obviously yield the cited figures. Moreover, no error bars, confidence intervals, or number of seeds are reported for any table or figure, so the claim that average performance is 'strikingly independently of ε' is not substantiated.
- [Abstract; Section 4, last paragraph] The central generalization—'open-domain, simple datasets overestimate their performance for real use-cases'—rests entirely on comparing the present numbers to prior work [29,45,46] under different evaluation protocols (different datasets, downstream classifiers, synthetic corpus sizes, and metrics). No open-domain dataset is evaluated under the same protocol, so any observed difference could be due to protocol variation rather than domain specificity. A controlled within-protocol comparison (adding a non-sensitive, open-domain corpus) or a substantial weakening of the claim is needed.
- [Section 3, 'Addressing the challenges' and footnote 2] The prior-exposure control is central to the domain-specificity argument, but the paper concedes it cannot verify that the gated-access datasets were absent from Llama-3's pre-training or closed-source post-training. This admitted limitation means the interpretation of N2C2'08/PsyTAR as 'more realistic' is confounded: degradation could reflect memorization or distributional mismatch rather than domain-specific difficulty. Please make this limitation more prominent, and ideally include a diagnostic (e.g., perplexity/canary checks or a synthetic control) or present the claim as conditional.
- [Section 3, 'Methods & Datasets'] The evaluation covers only two generators, one per paradigm. While the paper states this explicitly, the abstract and conclusion refer to 'state-of-the-art privacy-preserving generation methods' and 'current approaches' in general; the evidence base is too narrow for those general claims. Adding at least one more method per paradigm or explicitly restricting the conclusions would strengthen the paper.
minor comments (6)
- [Section 4, first paragraph] Typo: 'theeval-uated' should be 'the evaluated'.
- [Figure 2] The caption says 'many reference metrics' but does not specify which metrics are shown or how they are aggregated. Please clarify the legend and the exact computation of 'relative decrease'.
- [Section 4] The dataset name is inconsistent: 'AsyLax' in Table 1 and 'AsyLex' in the text (e.g., 'AsyLex again mirrors this pattern'). Use one spelling throughout.
- [References] References [31] and [32] are the same work (The Canary's Echo); please merge them or distinguish them clearly.
- [Section 3, 'Evaluation Protocol'] The protocol is described as rigorous and reproducible, but no details are given for the downstream classifiers (architecture, hyperparameters, training procedure). Please specify these, even if briefly, or cite the repository location where they are defined.
- [Section 4, footnote 4] The footnote 'which demands larger parameter updates that are clipped, noised; long contexts force smaller batch size' is grammatically awkward and would benefit from rewriting.
Circularity Check
No significant circularity: the paper is an empirical benchmark with held-out evaluations; the open-domain overestimation claim is confounded but not circular.
full rationale
This paper is an empirical evaluation study, not a derivation. It applies two externally proposed DP text generators (DP-Gen from Yue et al. [46] and AUG-PE from Xie et al. [45]) to five domain-specific datasets and measures utility via downstream classifiers trained on synthetic data and evaluated on held-out real test sets, plus fidelity via MAUVE, entity/length divergences, and reference-based metrics. No parameter is fitted to the evaluation data and then renamed as a prediction; the reported numbers are direct measurements. The methods under test are external, and the benchmark does not introduce a theoretical claim whose conclusion is built into its definitions. The central interpretive claim—that open-domain evaluations overestimate real-world performance—is based on comparing this paper's results to previously reported numbers [29,45,46]. That comparison is indeed confounded by different datasets, classifiers, and protocols, but confounding is a validity threat, not circularity: the claim does not assume its own conclusion, and the prior results are external evidence rather than a self-citation chain. The paper's self-citations (e.g., [28,33,34,35,41]) are background and survey references and are not load-bearing for the empirical measurements. The footnote admitting that gated-access data 'might have still encountered the data during closed-source post-training' weakens the prior-exposure control, but again this is an acknowledged empirical limitation, not a definitional circularity. Consequently, no circular step meeting the required evidentiary standard can be identified; the appropriate finding is no significant circularity.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption The two evaluated methods, DP-Gen and AUG-PE, correctly implement their claimed differential privacy guarantees.
- domain assumption Downstream classification F1 on held-out real test data is a valid proxy for synthetic data utility.
- domain assumption The gated datasets are sufficiently absent from Llama-3 pre-training.
- domain assumption Standard reference-based and distributional metrics (MAUVE, entity divergence, BLEU, etc.) capture text fidelity.
Cite this review
Pith. "Pith review of Evaluating Differentially Private Generation of Domain-Specific Text." pith.science (2026). https://pith.science/paper/QGKYJY5D
@misc{pith2026250820452,
author = {Pith},
title = {Pith review of: Evaluating Differentially Private Generation of Domain-Specific Text},
year = {2026},
howpublished = {\url{https://pith.science/paper/QGKYJY5D}},
note = {Machine review of arXiv:2508.20452}
}
read the original abstract
Generative AI offers transformative potential for high-stakes domains such as healthcare and finance, yet privacy and regulatory barriers hinder the use of real-world data. To address this, differentially private synthetic data generation has emerged as a promising alternative. In this work, we introduce a unified benchmark to systematically evaluate the utility and fidelity of text datasets generated under formal Differential Privacy (DP) guarantees. Our benchmark addresses key challenges in domain-specific benchmarking, including choice of representative data and realistic privacy budgets, accounting for pre-training and a variety of evaluation metrics. We assess state-of-the-art privacy-preserving generation methods across five domain-specific datasets, revealing significant utility and fidelity degradation compared to real data, especially under strict privacy constraints. These findings underscore the limitations of current approaches, outline the need for advanced privacy-preserving data sharing methods and set a precedent regarding their evaluation in realistic scenarios.
Figures
Forward citations
Cited by 1 Pith paper
-
SynBench: A Benchmark for Differentially Private Text Generation
SynBench benchmarks DP text generators across nine datasets and uses a new MIA to show that public pre-training on portions of private data overestimates synthetic text quality and breaks DP privacy bounds.
Reference graph
Works this paper leans on
-
[1]
Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang
Martin Abadi, Andy Chu, Ian Goodfellow, H. Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016. Deep Learning with Differential Privacy. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, Vol. 24-28-October-2016. ACM, New York, NY, USA, 308–318. doi:10. 1145/2976749.2978318
arXiv 2016
-
[2]
Kareem Amin, Alex Bie, Weiwei Kong, Alexey Kurakin, Natalia Ponomareva, Umar Syed, Andreas Terzis, and Sergei Vassilvitskii. 2024. Private prediction for large-scale synthetic text generation. arXiv preprint arXiv:2407.12108 (2024)
Pith/arXiv arXiv 2024
-
[3]
Simon Baker, Ilona Silins, Yufan Guo, Imran Ali, Johan Högberg, Ulla Stenius, and Anna Korhonen. 2016. Automatic semantic classification of scientific literature according to the hallmarks of cancer. Bioinformatics 32, 3 (2016), 432–440. doi:10. 1093/bioinformatics/btv585
work page 2016
-
[4]
Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. 65–72 pages. https://aclanthology.org/W05-0909/
work page 2005
-
[5]
Claire Barale, Michael Rovatsos, and Nehal Bhuta. 2023. Automated Refugee Case Analysis: A NLP Pipeline for Supporting Legal Practitioners. InFindings of the As- sociation for Computational Linguistics: ACL 2023 . Association for Computational Linguistics, Toronto, Canada, 2992–3005. doi:10.18653/v1/2023.findings-acl.187
-
[6]
Colin Bellinger, Christopher Drummond, and Nathalie Japkowicz. 2016. Beyond the Boundaries of SMOTE: A Framework for Manifold-Based Synthetically Over- sampling. In Machine Learning and Knowledge Discovery in Databases (ECML PKDD 2016) (Lecture Notes in Computer Science, vol. 9851) . Springer, Cham, 248–
work page 2016
-
[7]
Liu, Vi- jay Prakash Dwivedi, Thanh-Tung Nguyen, Xiaoxue Gao, Nancy F
Kuluhan Binici, Abhinav Ramesh Kashyap, Viktor Schlegel, Andy T. Liu, Vi- jay Prakash Dwivedi, Thanh-Tung Nguyen, Xiaoxue Gao, Nancy F. Chen, and Stefan Winkler. 2025. MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic Dialogues. Proceed- ings of the AAAI Conference on Artificial Intelligence 39, 22 ...
-
[8]
John Blitzer, Mark Dredze, and Fernando Pereira. 2007. Biographies, Bollywood, Boom-boxes and Blenders: Domain Adaptation for Sentiment Classification. 440– 447 pages. https://aclanthology.org/P07-1056/
work page 2007
-
[9]
Kathi Canese and Sarah Weis. 2013. PubMed: the bibliographic database. The NCBI handbook 2, 1 (2013)
work page 2013
-
[10]
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert- Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, Alina Oprea, and Colin Raffel. 2021. Extracting Training Data from Large Lan- guage Models. In 30th USENIX Security Symposium (USENIX Security 21) . 2633–
work page 2021
-
[11]
John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil. 2018. Universal Sentence Encoder for English. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for Com...
-
[12]
Vikram S Chundawat, Ayush K Tarun, Murari Mandal, Mukund Lahoti, and Pratik Narang. 2022. TabSynDex: A Universal Metric for Robust Evaluation of Synthetic Tabular Data. In Proceedings of the IEEE Conference on Artificial Intelligence. arXiv:2207.05295v2
Pith/arXiv arXiv 2022
-
[13]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training Verifiers to Solve Math Word Problems. arXiv:2110.14168 (10 2021). https://arxiv.org/abs/2110.14168v2
Pith/arXiv arXiv 2021
-
[14]
Personal Data Protection Commission. 2023. Proposed Guide on Synthetic Data Generation. Personal Data Protection Commission, Singapore (2023)
work page 2023
-
[15]
Daniel-ML. 2023. Sentiment Analysis for Financial News v2. https://huggingface. co/datasets/Daniel-ML/sentiment-analysis-for-financial-news-v2. Accessed: 2025-05-22
work page 2023
-
[16]
Fabrizio Dell’Acqua, Edward McFowland, Ethan R. Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2023. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. SSRN Electronic Journal (2023). doi:10.213...
-
[17]
Gelei Deng, Yi Liu, Víctor Mayoral-Vilches, Peng Liu, Yuekang Li, Yuan Xu, Tianwei Zhang, Yang Liu, Martin Pinzger, and Stefan Rass. 2024. PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetra- tion Testing. In 33rd USENIX Security Symposium (USENIX Security 24) . USENIX Association, Philadelphia, PA, 847–864. https://www.useni...
work page 2024
-
[18]
Cynthia Dwork. 2006. Differential Privacy. In Automata, Languages and Pro- gramming, Michele Bugliesi, Bart Preneel, Vladimiro Sassone, and Ingo Wegener (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 1–12
work page 2006
-
[19]
Cynthia Dwork, Nitin Kohli, and Deirdre Mulligan. 2019. Differential Privacy in Practice: Expose your Epsilons! Journal of Privacy and Confidentiality 9, 2 (10 2019), 2019. doi:10.29012/jpc.689
-
[20]
James Flemings, Meisam Razaviyayn, and Murali Annavaram. 2024. Differentially Private Next-Token Prediction of Large Language Models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) . Association for Computational Linguistics, Stroudsb...
work page 2024
-
[21]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)
Pith/arXiv arXiv 2024
-
[22]
Florent Guépin, Matthieu Meeus, Ana Maria Creţu, and Yves Alexandre de Montjoye. 2023. Synthetic is all you need: removing the auxiliary data as- sumption for membership inference attacks against synthetic data. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial In- telligence and Lecture Notes in Bioinformatics) 14398 LNCS...
-
[23]
Kenneth Hudson. 1978. The jargon of the professions . Springer
work page 1978
-
[24]
Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J
Alistair E.W. Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J. Pollard, Benjamin Moody, Brian Gow, Li wei H. Lehman, Leo A. Celi, and Roger G. Mark. 2023. MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data 10, 1 (12 2023). doi:10.1038/S41597-022- 01899-X
-
[25]
Alistair E.W. Johnson, Tom J. Pollard, Seth J. Berkowitz, Nathaniel R. Greenbaum, Matthew P. Lungren, Chih ying Deng, Roger G. Mark, and Steven Horng. 2019. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data 2019 6:1 6, 1 (12 2019), 1–8. doi:10.1038/s41597- 019-0322-0
doi:10.1038/s41597- 2019
-
[26]
Tatsuki Koga, Ruihan Wu, and Kamalika Chaudhuri. 2024. Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy. arXiv preprint arXiv:2412.04697 (2024)
arXiv 2024
-
[27]
Jaewoo Lee and Chris Clifton. 2011. How Much Is Enough? Choosing 𝜖 for Differential Privacy. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 7001 LNCS (2011), 325–340. doi:10.1007/978-3-642-24861-0-22
-
[28]
Hao Li, Yuping Wu, Viktor Schlegel, Riza Batista-Navarro, Thanh-Tung Nguyen, Abhinav Ramesh Kashyap, Xiao-Jun Zeng, Daniel Beck, Stefan Winkler, and Goran Nenadic. 2023. Team:PULSAR at ProbSum 2023:PULSAR: Pre-training with Extracted Healthcare Terms for Summarising Patients’ Problems and Data Augmentation with Black-box Large Language Models. In The 22nd...
2023
-
[29]
Justus Mattern, Zhijing Jin, Benjamin Weggenmann, Bernhard Schoelkopf, and Mrinmaya Sachan. 2022. Differentially Private Language Models for Secure Data Sharing. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Stroudsburg, PA, USA, 4860–4873. doi:10.18653/v1/2022.emnlp-m...
-
[30]
Justus Mattern, Benjamin Weggenmann, and Florian Kerschbaum. 2022. The limits of word level differential privacy. arXiv preprint arXiv:2205.02130 (2022)
Pith/arXiv arXiv 2022
-
[32]
Matthieu Meeus, Lukas Wutschitz, Santiago Zanella-Béguelin, Shruti Tople, and Reza Shokri. 2025. The Canary’s Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text. (2 2025). https://arxiv.org/abs/2502.14921v1
arXiv 2025
-
[33]
uMedSum: A Unified Framework for Advancing Medical Abstractive Summarization
Aishik Nagar, Yutong Liu, Andy T. Liu, Viktor Schlegel, Vijay Prakash Dwivedi, Arun-Kumar Kaliya-Perumal, Guna Pratheep Kalanchiam, Yili Tang, and Robby T. Tan. 2024. uMedSum: A Unified Framework for Advancing Medical Abstractive Summarization. (8 2024). https://arxiv.org/abs/2408.12095v2
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[34]
Aishik Nagar, Viktor Schlegel, Thanh-Tung Nguyen, Hao Li, Yuping Wu, Ku- luhan Binici, and Stefan Winkler. 2025. LLMs are not Zero-Shot Reasoners for Biomedical Information Extraction. 106–120 pages. https://aclanthology.org/ 2025.insights-1.11/
work page 2025
-
[35]
Thanh Tung Nguyen, Viktor Schlegel, Abhinav Kashyap, and Stefan Winkler
-
[36]
Shiwen Ni, Xiangtao Kong, Chengming Li, Xiping Hu, Ruifeng Xu, Jia Zhu, and Min Yang. 2025. Training on the Benchmark Is Not All You Need. Proceedings of the AAAI Conference on Artificial Intelligence 39, 23 (4 2025), 24948–24956. doi:10.1609/aaai.v39i23.34678
-
[37]
Sebastian Ochs and Ivan Habernal. 2025. Private Synthetic Text Generation with Diffusion Models. 10612–10626 pages. https://aclanthology.org/2025.naacl- long.532/
work page 2025
-
[38]
Nicolas Papernot, Shuang Song, Ilya Mironov, Ananth Raghunathan, Kunal Tal- war, and Úlfar Erlingsson. 2018. Scalable Private Learning with PATE. In Interna- tional Conference on Learning Representations . http://arxiv.org/abs/1802.08908
Pith/arXiv arXiv 2018
-
[39]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-jing Zhu. 2001. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics - ACL ’02. Association for Computational Linguistics, Morristown, NJ, USA, 311–318. doi:10.3115/1073083.1073135
arXiv 2001
-
[40]
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, Zaid Harchaoui, and Paul G Allen. 2021. MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers. In Advances in Neural Information Processing Systems , Vol. 34. 4816–4828. https: //github.com/krishnap25/mauve
work page 2021
-
[41]
Viktor Schlegel, Anil A Bharath, Zilong Zhao, and Kevin Yee. 2025. Generating Synthetic Data with Formal Privacy Guarantees: State of the Art and the Road Ahead. (3 2025). https://arxiv.org/abs/2503.20846v1
Pith/arXiv arXiv 2025
-
[42]
Latanya Sweeney. 2002. k-anonymity: A model for protecting privacy. Interna- tional journal of uncertainty, fuzziness and knowledge-based systems 10, 05 (2002), 557–570
work page 2002
-
[43]
Florian Tramèr, Gautam Kamath, and Nicholas Carlini. 2024. Position: Consider- ations for Differentially Private Learning with Large-Scale Public Pretraining. 48453–48467 pages. https://proceedings.mlr.press/v235/tramer24a.html
work page 2024
-
[44]
Ozlem Uzuner. 2009. Recognizing Obesity and Comorbidities in Sparse Data. Journal of the American Medical Informatics Association 16, 4 (07 2009), 561–
work page 2009
-
[45]
Chulin Xie, Zinan Lin, Arturs Backurs, Sivakanth Gopi, Da Yu, Huseyin A Inan, Harsha Nori, Haotian Jiang, Huishuai Zhang, Yin Tat Lee, Bo Li, and Sergey Yekhanin. 2024. Differentially Private Synthetic Data via Foundation Model APIs 2: Text. In Proceedings of the 41st International Conference on Machine Learning . PMLR, 54531–54560. http://arxiv.org/abs/2...
Pith/arXiv arXiv 2024
-
[46]
Xiang Yue, Huseyin A. Inan, Xuechen Li, Girish Kumar, Julia McAnallen, Hoda Shajari, Huan Sun, David Levitan, and Robert Sim. 2023. Synthetic Text Gen- eration with Differential Privacy: A Simple and Practical Recipe. Proceedings of the Annual Meeting of the Association for Computational Linguistics 1 (2023), 1321–1342. doi:10.18653/V1/2023.ACL-LONG.74
-
[47]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi
-
[48]
Patrick, Paul Fontelo, Hadi Kharrazi, Anthony Faiola, Yi Shuan Shirley Wu, Christina E
Maryam Zolnoori, Kin Wah Fung, Timothy B. Patrick, Paul Fontelo, Hadi Kharrazi, Anthony Faiola, Yi Shuan Shirley Wu, Christina E. Eldredge, Jake Luo, Mike Conway, Jiaxi Zhu, Soo Kyung Park, Kelly Xu, Hamideh Moayyed, and Somaieh Goudarzvand. 2019. A systematic approach for developing a corpus of patient reported adverse drug events: A case study for SSRI ...
-
[263]
doi:10.1007/978-3-319-46128-1_16
-
[570]
doi:10.1197/jamia.M3115 arXiv:https://academic.oup.com/jamia/article- pdf/16/4/561/2302602/16-4-561.pdf
-
[2019]
BERTScore: Evaluating Text Generation with BERT. (4 2019). doi:10.48550/ arxiv.1904.09675
-
[2023]
Proceedings of the Annual Meeting of the Association for Computational Linguistics (2023), 4658–4665
A Two-Stage Decoder for Efficient ICD Coding. Proceedings of the Annual Meeting of the Association for Computational Linguistics (2023), 4658–4665. doi:10. 18653/V1/2023.FINDINGS-ACL.285
work page 2023
-
[2650]
http://arxiv.org/abs/2012.07805
Pith/arXiv arXiv 2012
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.