REVIEW 2 major objections 4 minor 1 cited by
Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities
T0 review · 2 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read The paper argues that privacy-preserving techniques for mental-health AI currently fail to protect longitudinal, multimodal therapy data while preserving diagnostic utility, and proposes a pipeline to fix that.
desk verdict A solid, useful survey of privacy in mental health AI with a localized but real error in the DP-SGD epsilon definition that should be fixed before acceptance. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying framework is the proposed privacy-aware pipeline: data collection, anonymization or synthetic data generation, data-level privacy-utility evaluation, privacy-aware model training, and model-level privacy-utility evaluation. It is organized by the paper's taxonomy of threats (identification versus impersonation, data-level versus model-level leakage) and its per-modality map of solutions (text PII removal, voice anonymization, face anonymization, synthetic generation), joined with a common metric set that includes equal error rate, membership-inference accuracy, the differential-privacy guarantee epsilon, and downstream diagnostic performance. This pipeline does the argumentative work: it turns scattered technique reviews into a sequence with explicit decision points, such as choosing synthetic augmentation for small datasets and requiring cross-modal re-identification tests.
What would settle it
Run the paper's recommended pipeline on a real longitudinal therapy dataset and measure whether any anonymization or differential-privacy configuration reaches the same downstream diagnostic F1 as the non-private baseline; if one does at a defensible epsilon, the paper's claim that current methods fall short is falsified for that configuration.
Extended reading notes
Core claim
The central claim, stated on the paper's own terms, is that no existing privacy technique is ready for mental-health AI because therapy data are longitudinal, multimodal, and full of indirect identifiers. For text, NER-based PII removal misses implicit contextual disclosures and cross-session links; for audio, anonymization lacks multi-speaker or differential-privacy guarantees suited to therapy conversations; for video, face obfuscation leaks age, gender, and body-language cues; and for models, differential privacy and federated learning degrade diagnostic utility, especially on small realistic datasets. The paper's proposed answer is a workflow—consent-based recording, local transcription, anonymization or synthetic generation chosen by dataset scale, privacy-utility evaluation with re-identification and downstream-task metrics, then DP-based training and model-level attack testing—which it presents as the feasible path to clinically useful, privacy-aware systems.
Load-bearing premise
The whole pipeline assumes that privacy-preserving methods can eventually keep enough diagnostic signal in longitudinal, multimodal therapy data; the paper itself documents that current methods degrade utility, so if that trade-off cannot be overcome the pipeline does not produce clinically usable models.
Editorial extensions
If this is right
- If the paper is correct, any privacy-preserving mental-health AI system should evaluate privacy and utility twice—once on the protected data and once on the trained model—and should report downstream diagnostic performance, not just privacy metrics.
- Multi-speaker anonymization with strong threat models becomes a prerequisite for audio privacy in therapy, since real sessions are dialogues with overlapping speech and informed attackers.
- Differential privacy should be applied to fine-tuning and fusion layers, while federated learning is used only with local differential privacy, because federated learning alone leaks through gradients.
- Synthetic therapy data must become long-form, multimodal, and grounded in clinical frameworks before it can substitute for real data.
- Cross-modal leakage becomes a standard evaluation target, for example lip movements revealing names that were redacted from the transcript.
Reading between the lines
- A natural next step the paper leaves implicit is a hybrid strategy: use synthetic data for demographic and diagnostic coverage and anonymized real data for fidelity, with a measured information-retention budget.
- An adversarial cross-modal linking model—trained to match anonymized audio or video to a text transcript—would turn the paper's cross-modal leakage concern into a concrete benchmark.
- If differential privacy's documented disparate impact holds for mental-health tasks, deployment may require group-specific utility reporting before these models are used clinically.
- A direct head-to-head comparison of DP-SGD and LDP-FL on the same diagnostic benchmarks, which the paper notes is missing, would be the decisive experiment for choosing a training paradigm.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a survey of privacy challenges, solutions, and evaluation methods for AI models in mental health, with emphasis on multimodal (text, audio, video) therapy data. It reviews privacy leakage from datasets and from trained models, then surveys data anonymization, synthetic data generation, and privacy-aware training (DP, FL, TEE, autoencoders). It also catalogues privacy and utility evaluation metrics and proposes a concrete pipeline for developing privacy-aware mental health AI systems, together with a list of open research directions.
Significance. If the survey is accurate, it provides a valuable structured map of privacy risks and mitigation strategies in a high-stakes application domain, with a practical pipeline that could guide future work. Its strengths include broad modality coverage (text, audio, video), attention to longitudinal and multi-speaker settings, and explicit discussion of fairness and utility degradation (e.g., DP-SGD utility loss, FIVA demographic bias). The paper also explicitly identifies gaps such as cross-modal privacy leakage and the lack of DP guarantees for multimodal anonymization, which are timely and actionable research targets.
major comments (2)
- [Privacy-aware training, 'Privacy evaluation' paragraph] The sentence defining the DP-SGD ε-value is incorrect: 'For models trained with DP-SGD, the privacy guarantee is quantified by the ε-value38 (which determines the distance within which errors are considered to be zero in Stochastic Gradient Descent).' In differential privacy, ε bounds the log-likelihood ratio of the algorithm's output on adjacent datasets (i.e., the privacy loss); it is not a distance threshold within which errors vanish. Since DP-SGD is one of the paper's principal recommended solutions and the paper explicitly aims to guide privacy-aware training, this misdefinition is misleading for readers and must be corrected, with proper citations (e.g., Abadi et al. 2016 or Dwork & Roth).
- [Threats, PII leakage] The claim that 'age, address, and gender ... can uniquely identify most Americans' is attributed to ref 69 (Krishnamurthy & Wills 2009), but that reference is about PII leakage in online social networks, not the well-known Sweeney re-identification result. The authors should cite the correct source (e.g., Sweeney 2000, 'Simple demographics often identify people uniquely') or rephrase the claim to match what ref 69 actually supports.
minor comments (4)
- [General] The terms 'privacy-utility trade-off' and 'privacy-utility evaluation' are used frequently; a one-sentence definition at first use would improve clarity.
- [Figure 1] The text in Figure 1 is very small and dense, especially in the evaluation columns; consider enlarging or splitting the figure for readability.
- [Threats] When stating that 'LLMs trained on therapy data are prone to privacy breaches', the cited references (22–25) generally concern neural models and memorization; the attribution to LLMs specifically could be made more precise.
- [Prospects, LDP-FL] There are minor grammatical issues, e.g., 'a LDP-FL setup' should be 'an LDP-FL setup', and some sentences are long and could be split for readability.
Circularity Check
No circularity found: the paper is a literature review whose proposed pipeline is a synthesis of external results, not a derivation from its own assumptions.
full rationale
This manuscript is a survey and position paper. It does not derive new results from first principles, fit any parameter, or make quantitative predictions that could reduce to its own inputs. The proposed privacy-aware pipeline (Figure 2) is presented as a recommendation based on the surveyed literature; it is not claimed to be proven by the paper's own equations, and no load-bearing self-citation is used to justify it. The skeptical observation that DP-SGD's epsilon value is misdefined as 'the distance within which errors are considered to be zero' is a technical correctness error in the survey's exposition, but it is not circularity: the discussion of differential privacy is not used to derive a result that presupposes that definition. Similarly, the attribution of the 'age, address, and gender... uniquely identify most Americans' claim to reference 69 is a citation-accuracy issue, not a circular-dependency issue. There are no fitted inputs renamed as predictions, no uniqueness theorems imported from the authors' prior work, and no ansatz smuggled in via self-citation. The central contribution is an organized review and a set of research directions, so the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Differential privacy and other privacy-preserving training methods provide meaningful privacy guarantees when applied to mental health data.
- domain assumption Multimodal AI models can support mental health diagnosis.
- domain assumption Therapy data cannot be publicly released and patient consent is required.
Cite this review
Pith. "Pith review of Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities." pith.science (2026). https://pith.science/paper/ID5CUW27
@misc{pith2026250200451,
author = {Pith},
title = {Pith review of: Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities},
year = {2026},
howpublished = {\url{https://pith.science/paper/ID5CUW27}},
note = {Machine review of arXiv:2502.00451}
}
read the original abstract
Mental health disorders create profound personal and societal burdens, yet conventional diagnostics are resource-intensive and limit accessibility. Advances in artificial intelligence, particularly natural language processing and multimodal methods, offer promise for detecting and addressing mental disorders, but raise critical privacy risks. This paper examines these challenges and proposes solutions, including anonymization, synthetic data, and privacy-preserving training, while outlining frameworks for privacy-utility trade-offs, aiming to advance reliable, privacy-aware AI tools that support clinical decision-making and improve mental health outcomes.
Figures
Forward citations
Cited by 1 Pith paper
-
A Comprehensive Review of Datasets for Clinical Mental Health AI Systems
A systematic catalog of 89 clinical mental health datasets and 16 synthetic datasets, with a gap analysis on access, culture, and modality.
Reference graph
Works this paper leans on
-
[1]
Slonim, D. A. et al. Facing change: using automated facial expression analysis to examine emotional flexibility in the treatment of depression. Adm. Policy Mental Heal. Mental Heal. Serv. Res. 1–8 (2023)
2023
-
[2]
Cohn, J. F. et al. Detecting depression from facial actions and vocal prosody. In Affective Computing and Intelligent Interaction, Third International Conference and Workshops, ACII 2009, Amsterdam, The Netherlands, September 10-12, 2009, Proceedings, 1–7, DOI: 10.1109/ACII.2009.5349358 (IEEE Computer Society, 2009)
arXiv 2009
-
[3]
Scherer, S. et al. Automatic audiovisual behavior descriptors for psychological disorder analysis. Image Vis. Comput. 32, 648–658, DOI: https://doi.org/10.1016/j.imavis.2014.06.001 (2014). Best of Automatic Face and Gesture Recognition 2013
-
[4]
Cummins, N. et al. A review of depression and suicide risk assessment using speech analysis. Speech Commun. 71, 10–49 (2015)
2015
-
[5]
Chim, J. et al. Overview of the CLPsych 2024 shared task: Leveraging large language models to identify evidence of suicidality risk in online posts. In Yates, A. et al. (eds.) Proceedings of the 9th Workshop on Computational Linguistics and Clinical Psychology (CLPsych 2024), 177–190 (Association for Computational Linguistics, St. Julians, Malta, 2024)
2024
-
[6]
K., Lim, M
Langer, J. K., Lim, M. H., Fernandez, K. C. & Rodebaugh, T. L. Social anxiety disorder is associated with reduced eye contact during conversation primed for conflict. Cogn. therapy research 41, 220–229 (2017)
2017
-
[7]
Shafique, S. et al. Towards automatic detection of social anxiety disorder via gaze interaction. Appl. Sci. 12, DOI: 10.3390/app122312298 (2022)
-
[8]
Kathan, A. et al. The effect of clinical intervention on the speech of individuals with PTSD: features and recognition performances. In Harte, N., Carson-Berndsen, J. & Jones, G. (eds.) 24th Annual Conference of the International Speech Communication Association, Interspeech 2023, Dublin, Ireland, August 20-24, 2023, 4139–4143, DOI: 10.21437/ INTERSPEECH....
2023
Show all 131 references
-
[9]
& Ren, Z
Hu, J., Zhao, C., Shi, C., Zhao, Z. & Ren, Z. Speech-based recognition and estimating severity of ptsd using machine learning. J. Affect. Disord. 362, 859–868, DOI: https://doi.org/10.1016/j.jad.2024.07.015 (2024)
2024 doi
-
[10]
Gideon, J., Provost, E. M. & McInnis, M. G. Mood state prediction from speech of varying acoustic quality for individuals with bipolar disorder. In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2016, Shanghai, China, March 20-25, 2016, 2...
2016
-
[11]
Gilanie, G. et al. A robust method of bipolar mental illness detection from facial micro expressions using machine learning methods. Intell. Autom. & Soft Comput. 39 (2024)
2024
- [12]
-
[13]
R., Warnita, T., Uto, K
Makiuchi, M. R., Warnita, T., Uto, K. & Shinoda, K. Multimodal fusion of BERT-CNN and gated CNN representations for depression detection. In Ringeval, F. et al. (eds.) Proceedings of the 9th International on Audio/Visual Emotion Challenge and Workshop, AVEC@MM 2019, Nice, Fran...
2019
-
[14]
& Morency, L
Baltrusaitis, T., Robinson, P. & Morency, L. Openface: An open source facial behavior analysis toolkit. In 2016 IEEE Winter Conference on Applications of Computer Vision, WACV 2016, Lake Placid, NY, USA, March 7-10, 2016, 1–10, DOI: 10.1109/W ACV .2016.7477553 (IEEE Computer S...
2016
-
[15]
Sadeghi, M. et al. Harnessing multimodal approaches for depression detection using large language models and facial expressions. npj Mental Heal. Res. 3, 66 (2024)
2024
-
[16]
& Mahmoud, M
Zhang, Z., Lin, W., Liu, M. & Mahmoud, M. Multimodal deep learning framework for mental disorder recognition. In 15th IEEE International Conference on Automatic Face and Gesture Recognition, FG 2020, Buenos Aires, Argentina, November 16-20, 2020, 344–350, DOI: 10.1109/FG47880....
2020
-
[17]
Multimodal assessment of schizophrenia symptom severity from linguistic, acoustic and visual cues
Chuang, C.-Y .et al. Multimodal assessment of schizophrenia symptom severity from linguistic, acoustic and visual cues. IEEE Transactions on Neural Syst. Rehabil. Eng. 31, 3469–3479, DOI: 10.1109/TNSRE.2023.3307597 (2023)
2023
-
[18]
& Wang, K
Zhao, Z. & Wang, K. Unaligned multimodal sequences for depression assessment from speech. In 44th Annual International Conference of the IEEE Engineering in Medicine & Biology Society, EMBC 2022, Glasgow, Scotland, United Kingdom, July 11-15, 2022, 3409–3413, DOI: 10.1109/EMBC...
2022
-
[19]
Qin, J. et al. Mental-perceiver: Audio-textual multi-modal learning for estimating mental disorders. In Walsh, T., Shah, J. & Kolter, Z. (eds.) AAAI-25, Sponsored by the Association for the Advancement of Artificial Intelligence, February 25 - March 4, 2025, Philadelphia, PA, ...
2025 doi
-
[20]
Regulation (eu) 2016/679 of the european parliament and of the council of 27 april 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing directive 95/46/ec (general data protection regulati...
2016
-
[21]
Health insurance portability and accountability act of 1996
Act, A. Health insurance portability and accountability act of 1996. Public law 104, 191 (1996)
1996
-
[22]
& Shmatikov, V
Shokri, R., Stronati, M., Song, C. & Shmatikov, V . Membership inference attacks against machine learning models. In 2017 IEEE Symposium on Security and Privacy, SP 2017, San Jose, CA, USA, May 22-26, 2017 , 3–18, DOI: 10.1109/SP.2017.41 (IEEE Computer Society, 2017)
2017 doi
-
[23]
& Raghunathan, A
Song, C. & Raghunathan, A. Information leakage in embedding models. In Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security, CCS ’20, 377–390, DOI: 10.1145/3372297.3417270 (Association for Computing Machinery, New York, NY , USA, 2020)
2020
-
[24]
& Song, D
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J. & Song, D. The secret sharer: Evaluating and testing unintended memorization in neural networks. In Heninger, N. & Traynor, P. (eds.)28th USENIX Security Symposium, USENIX Security 2019, Santa Clara, CA, USA, August 14-16, 2019, 26...
2019
- [25]
-
[26]
Tang, B. et al. De-identification of clinical text via bi-lstm-crf with neural language models. In AMIA 2019, American Medical Informatics Association Annual Symposium, Washington, DC, USA, November 16-20, 2019 (AMIA, 2019)
2019
-
[27]
& Zhou, S
Yue, X. & Zhou, S. PHICON: improving generalization of clinical text de-identification models via data augmentation. In Rumshisky, A., Roberts, K., Bethard, S. & Naumann, T. (eds.) Proceedings of the 3rd Clinical Natural Language Processing Workshop, ClinicalNLP@EMNLP 2020, On...
2020 doi
-
[28]
Liu, Z. et al. Deid-gpt: Zero-shot medical text de-identification by GPT-4. CoRR abs/2303.11032, DOI: 10.48550/ ARXIV .2303.11032 (2023). 2303.11032
2023 doi
-
[29]
& Skala, P
Flechl, M., Yin, S., Park, J. & Skala, P. End-to-end speech recognition modeling from de-identified data. In Ko, H. & Hansen, J. H. L. (eds.) 23rd Annual Conference of the International Speech Communication Association, Interspeech 2022, Incheon, Korea, September 18-22, 2022, ...
2022 doi
-
[30]
Panariello, M. et al. The voiceprivacy 2022 challenge: Progress and perspectives in voice anonymisation. IEEE/ACM Transactions on Audio, Speech, Lang. Process.32, 3477–3491, DOI: 10.1109/TASLP.2024.3430530 (2024)
2024
-
[31]
Tomashenko, N. et al. The voice privacy 2024 challenge evaluation plan. arXiv preprint arXiv:2404.02677 (2024)
2024 arXiv
-
[32]
& Kankanhalli, M
Singh, A., Fan, S. & Kankanhalli, M. Human attributes prediction under privacy-preserving conditions. In Proceedings of the 29th ACM International Conference on Multimedia , MM ’21, 4698–4706, DOI: 10.1145/3474085.3475687 (Association for Computing Machinery, New York, NY , USA, 2021)
-
[33]
SoulChat: Improving LLMs’ empathy, listening, and comfort abilities through fine-tuning with multi-turn empathy conversations
Chen, Y .et al. SoulChat: Improving LLMs’ empathy, listening, and comfort abilities through fine-tuning with multi-turn empathy conversations. In Bouamor, H., Pino, J. & Bali, K. (eds.) Findings of the Association for Computational Linguistics: EMNLP 2023 , 1170–1183, DOI: 10....
2023 doi
-
[34]
& Zhang, Y
Wu, Y ., Chen, J., Mao, K. & Zhang, Y . Automatic post-traumatic stress disorder diagnosis via clinical transcripts: A novel text augmentation with large language models. In IEEE Biomedical Circuits and Systems Conference, BioCAS 2023, Toronto, ON, Canada, October 19-21, 2023,...
2023
-
[35]
Patient-ψ: Using large language models to simulate patients for training mental health professionals
Wang, R.et al. Patient-ψ: Using large language models to simulate patients for training mental health professionals. In Al- Onaizan, Y ., Bansal, M. & Chen, Y . (eds.)Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL,...
2024
-
[36]
Lee, S. et al. Cactus: Towards psychological counseling conversations using cognitive behavioral theory. In Al-Onaizan, Y ., Bansal, M. & Chen, Y .-N. (eds.)Findings of the Association for Computational Linguistics: EMNLP 2024, 14245– 14274, DOI: 10.18653/v1/2024.findings-emnl...
2024 doi
- [37]
-
[38]
Abadi, M. et al. Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security , CCS ’16, 308–318, DOI: 10.1145/2976749.2978318 (Association for Computing Machinery, New York, NY , USA, 2016)
2016
- [39]
-
[40]
& Giuffrida, V
Plant, R., Gkatzia, D. & Giuffrida, V . CAPE: Context-aware private embeddings for private language learning. In Moens, M.-F., Huang, X., Specia, L. & Yih, S. W.-t. (eds.)Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 7970–7978, DOI: 10...
2021 doi
-
[41]
Yu, D. et al. Differentially private fine-tuning of language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 (OpenReview.net, 2022)
2022
-
[42]
& Tuyls, J
Kerrigan, G., Slack, D. & Tuyls, J. Differentially private language models benefit from public pre-training. In Feyisetan, O., Ghanavati, S., Malmasi, S. & Thaine, P. (eds.)Proceedings of the Second Workshop on Privacy in NLP, 39–45, DOI: 10.18653/v1/2020.privatenlp-1.5 (Assoc...
- [43]
-
[44]
& Karypis, G
Bu, Z., Wang, Y .-X., Zha, S. & Karypis, G. Differentially private bias-term only fine-tuning of foundation models. In Workshop on Trustworthy and Socially Responsible Machine Learning, NeurIPS 2022(2022)
2022
-
[45]
X., Chiu, J
Morris, J. X., Chiu, J. T., Zabih, R. & Rush, A. M. Unsupervised text deidentification. In Goldberg, Y ., Kozareva, Z. & Zhang, Y . (eds.) Findings of the Association for Computational Linguistics:EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, 4777–4788, DOI...
2022 doi
-
[46]
Yang, X. et al. A study of deep learning methods for de-identification of clinical notes in cross-institute settings. BMC Med. Informatics Decis. Mak. 19-S, 6, DOI: 10.1186/S12911-019-0935-4 (2019)
2019 doi
-
[47]
Maouche, M. et al. A comparative study of speech anonymization metrics. In Meng, H., Xu, B. & Zheng, T. F. (eds.)21st Annual Conference of the International Speech Communication Association, Interspeech 2020, Virtual Event, Shanghai, China, October 25-29, 2020, 1708–1712, DOI:...
2020 doi
-
[48]
& Jha, S
Rosenberg, H., Tang, B., Fawaz, K. & Jha, S. Fairness properties of face recognition and obfuscation systems. In Calandrino, J. A. & Troncoso, C. (eds.) 32nd USENIX Security Symposium, USENIX Security 2023, Anaheim, CA, USA, August 9-11, 2023, 7231–7248 (USENIX Association, 2023)
2023
-
[49]
Osorio-Marulanda, P. A. et al. Privacy mechanisms and evaluation metrics for synthetic data generation: A systematic review. IEEE Access 12, 88048–88074, DOI: 10.1109/ACCESS.2024.3417608 (2024)
2024
-
[50]
Murtaza, H. et al. Synthetic data generation: State of the art in health care domain. Comput. Sci. Rev. 48, 100546, DOI: https://doi.org/10.1016/j.cosrev.2023.100546 (2023)
2023
- [51]
- [52]
- [53]
-
[55]
& Osman, H
Ghanadian, H., Nejadgholi, I. & Osman, H. A. Socially aware synthetic data generation for suicidal ideation detection using large language models. IEEE Access 12, 14350–14363, DOI: 10.1109/ACCESS.2024.3358206 (2024)
2024
-
[56]
Mehta, S. et al. Fake it to make it: Using synthetic data to remedy the data shortage in joint multimodal speech-and-gesture synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 1952–1964 (2024)
2024
-
[57]
Mughal, M. H. et al. Convofusion: Multi-modal conversational diffusion for co-speech gesture synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1388–1398 (2024)
2024
-
[58]
& Shmatikov, V
Narayanan, A. & Shmatikov, V . Robust de-anonymization of large sparse datasets. In2008 IEEE Symposium on Security and Privacy (SP 2008), 18-21 May 2008, Oakland, California, USA, 111–125, DOI: 10.1109/SP.2008.33 (IEEE Computer Society, 2008)
2008 doi
-
[59]
Nautsch, A. et al. Preserving privacy in speaker and speech characterisation. Comput. Speech & Lang. 58, 441–480, DOI: https://doi.org/10.1016/j.csl.2019.06.001 (2019)
2019 doi
-
[60]
& Verkholyak, O
Markitantov, M. & Verkholyak, O. Automatic recognition of speaker age and gender based on deep neural networks. In Salah, A. A., Karpov, A. & Potapova, R. (eds.)Speech and Computer - 21st International Conference, SPECOM 2019, Istanbul, Turkey, August 20-25, 2019, Proceedings,...
2019 doi
-
[61]
& Milner, B
Shao, X. & Milner, B. Pitch prediction from MFCC vectors for speech reconstruction. In 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2004, Montreal, Quebec, Canada, May 17-21, 2004, 97–100, DOI: 10.1109/ICASSP.2004.1325931 (IEEE, 2004)
2004 arXiv
-
[62]
& Kim, K
Lim, J. & Kim, K. Wav2vec-vc: V oice conversion via hidden representations of wav2vec 2.0. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2024, Seoul, Republic of Korea, April 14-19, 2024, 10326–10330, DOI: 10.1109/ICASSP48485.2024.10447984...
2024
-
[63]
& Sun, J
He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV , USA, June 27-30, 2016 , 770–778, DOI: 10.1109/CVPR.2016.90 (IEEE Computer Society, 2016)
2016 doi
-
[64]
& Philbin, J
Schroff, F., Kalenichenko, D. & Philbin, J. Facenet: A unified embedding for face recognition and clustering. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015, 815–823, DOI: 10.1109/CVPR.2015.7298682 (IEEE Computer Soci...
2015
-
[65]
& Hinton, G
Krizhevsky, A., Sutskever, I. & Hinton, G. E. Imagenet classification with deep convolutional neural networks. In Bartlett, P. L., Pereira, F. C. N., Burges, C. J. C., Bottou, L. & Weinberger, K. Q. (eds.)Advances in Neural Information Processing Systems 25: 26th Annual Confer...
2012
-
[66]
& Brox, T
Dosovitskiy, A. & Brox, T. Inverting visual representations with convolutional networks. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV , USA, June 27-30, 2016, 4829–4837, DOI: 10.1109/CVPR.2016.522 (IEEE Computer Society, 2016)
2016 doi
-
[67]
Mai, G., Cao, K., Yuen, P. C. & Jain, A. K. On the reconstruction of face images from deep face templates. IEEE Trans. Pattern Anal. Mach. Intell. 41, 1188–1202, DOI: 10.1109/TPAMI.2018.2827389 (2019)
2019
-
[68]
3d face reconstruction with dense landmarks
Wood, E.et al. 3d face reconstruction with dense landmarks. InComputer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XIII, 160–177, DOI: 10.1007/978-3-031-19778-9_10 (Springer- Verlag, Berlin, Heidelberg, 2022)
2022 doi
-
[69]
& Wills, C
Krishnamurthy, B. & Wills, C. E. On the leakage of personally identifiable information via online social networks. In Proceedings of the 2nd ACM Workshop on Online Social Networks, WOSN ’09, 7–12, DOI: 10.1145/1592665.1592668 (Association for Computing Machinery, New York, NY ...
-
[70]
& Dua, M
Joshi, S. & Dua, M. Noise robust automatic speaker verification systems: review and analysis. Telecommun. Syst. 87, 845–886, DOI: 10.1007/s11235-024-01212-8 (2024)
2024 doi
-
[71]
& Kasak, P
Jakubec, M., Jarina, R., Lieskovska, E. & Kasak, P. Deep speaker embeddings for speaker verification: Review and experimental comparison. Eng. Appl. Artif. Intell. 127, 107232, DOI: https://doi.org/10.1016/j.engappai.2023.107232 (2024)
2024
-
[72]
& Atri, M
Kortli, Y ., Jridi, M., Al Falou, A. & Atri, M. Face recognition systems: A survey.Sensors 20, DOI: 10.3390/s20020342 (2020)
2020 doi
-
[73]
T., Terhörst, P
Huber, M., Luu, A. T., Terhörst, P. & Damer, N. Efficient explainable face verification based on similarity score argument backpropagation. In IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2024, Waikoloa, HI, USA, January 3-8, 2024, 4724–4733, DOI: 10.110...
2024
-
[74]
& Elgibreen, H
Almutairi, Z. & Elgibreen, H. A review of modern audio deepfake detection methods: Challenges and future directions. Algorithms 15, DOI: 10.3390/a15050155 (2022)
2022 doi
-
[75]
A., Yildirim, R
Shaaban, O. A., Yildirim, R. & Alguttar, A. A. Audio deepfake approaches. IEEE Access 11, 132652–132682, DOI: 10.1109/ACCESS.2023.3333866 (2023)
2023
-
[76]
W., McCarthy, I
Kietzmann, J., Lee, L. W., McCarthy, I. P. & Kietzmann, T. C. Deepfakes: Trick or treat?Bus. Horizons 63, 135–146, DOI: https://doi.org/10.1016/j.bushor.2019.11.006 (2020). ARTIFICIAL INTELLIGENCE AND MACHINE LEARNING
2020 doi
-
[77]
Masood, M. et al. Deepfakes generation and detection: state-of-the-art, open challenges, countermeasures, and way forward. Appl. Intell. 53, 3974–4026, DOI: 10.1007/S10489-022-03766-Z (2023)
2023 doi
-
[78]
Gu, H. et al. Utilizing speaker profiles for impersonation audio detection. In Proceedings of the 32nd ACM International Conference on Multimedia, MM ’24, 1961–1970, DOI: 10.1145/3664647.3681602 (Association for Computing Machinery, New York, NY , USA, 2024)
1961
-
[79]
& Ortega-Garcia, J
Tolosana, R., Vera-Rodriguez, R., Fierrez, J., Morales, A. & Ortega-Garcia, J. Deepfakes and beyond: A survey of face manipulation and fake detection. Inf. Fusion 64, 131–148, DOI: https://doi.org/10.1016/j.inffus.2020.06.014 (2020)
2020 doi
-
[80]
& Woo, S
Khalid, H., Tariq, S., Kim, M. & Woo, S. S. Fakeavceleb: A novel audio-video multimodal deepfake dataset. In Vanschoren, J. & Yeung, S. (eds.) Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, Dec...
2021
-
[81]
& Dwivedi, Y
Mustak, M., Salminen, J., Mäntymäki, M., Rahman, A. & Dwivedi, Y . K. Deepfakes: Deceptions, mitigations, and opportunities. J. Bus. Res. 154, 113368, DOI: https://doi.org/10.1016/j.jbusres.2022.113368 (2023)
2023
- [82]
-
[83]
& Zhang, Y
Li, Q., Zhang, Y ., Ren, J., Li, Q. & Zhang, Y . You can use but cannot recognize: Preserving visual privacy in deep neural networks. In 31st Annual Network and Distributed System Security Symposium, NDSS 2024, San Diego, California, USA, February 26 - March 1, 2024(The Intern...
2024
-
[84]
Y ., Uzuner, O
Dernoncourt, F., Lee, J. Y ., Uzuner, O. & Szolovits, P. De-identification of patient notes with recurrent neural networks. J. Am. Med. Informatics Assoc. 24, 596–606, DOI: 10.1093/jamia/ocw156 (2016). https://academic.oup.com/jamia/ article-pdf/24/3/596/34945844/ocw156.pdf
2016 doi
-
[85]
& Lee, J
Kim, W., Hahm, S. & Lee, J. Generalizing clinical de-identification models by privacy-safe data augmentation using GPT-4. In Al-Onaizan, Y ., Bansal, M. & Chen, Y .-N. (eds.)Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 21204–21218, DO...
2024 doi
-
[86]
Dhingra, P. et al. Speech de-identification data augmentation leveraging large language model. In Liu, R. et al. (eds.) International Conference on Asian Language Processing, IALP 2024, Hohhot, China, August 4-6, 2024, 97–102, DOI: 10.1109/IALP63756.2024.10661176 (IEEE, 2024)
2024
-
[87]
& Roller, R
Baroud, I., Raithel, L., Möller, S. & Roller, R. Beyond de-identification: A structured approach for defining and detecting indirect identifiers in medical texts. In Habernal, I., Ghanavati, S., Jain, V ., Igamberdiev, T. & Wilson, S. (eds.)Proceedings of the Sixth Workshop on...
-
[88]
& Dusek, O
Balloccu, S., Schmidtová, P., Lango, M. & Dusek, O. Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs. In Graham, Y . & Purver, M. (eds.)Proceedings of the 18th Conference of the European Chapter of the Association for Computational Ling...
2024
-
[89]
Cohn, I. et al. Audio de-identification - a new entity recognition task. In Loukina, A., Morales, M. & Kumar, R. (eds.) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Ind...
2019 doi
-
[90]
S., Dhingra, P., Wang, Z
Veerappan, C. S., Dhingra, P., Wang, Z. & Tong, R. SpeeDF - A Speech De-identification Framework, DOI: 10.25447/sit. 27013936.v1 (2024)
2024 doi
-
[91]
& Tomashenko, N
Miao, X., Wang, X., Cooper, E., Yamagishi, J. & Tomashenko, N. Speaker anonymization using orthogonal householder neural network. IEEE/ACM Transactions on Audio, Speech, Lang. Process.31, 3681–3695, DOI: 10.1109/TASLP.2023. 3313429 (2023)
2023 doi
-
[92]
& Larcher, A
Champion, P., Jouvet, D. & Larcher, A. Are disentangled representations all you need to build speaker anonymization systems? In INTERSPEECH 2022 - Human and Humanizing Speech Technology(incheon, South Korea, 2022)
2022
- [93]
-
[94]
Face-off: Adversarial face obfuscation
Chandrasekaran, V .et al. Face-off: Adversarial face obfuscation. Proc. Priv. Enhancing Technol.2021, 369–390, DOI: 10.2478/POPETS-2021-0032 (2021)
2021 doi
-
[95]
Lowkey: Leveraging adversarial attacks to protect social media users from facial recognition
Cherepanova, V .et al. Lowkey: Leveraging adversarial attacks to protect social media users from facial recognition. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 (OpenReview.net, 2021)
2021
-
[96]
& Kohno, T
Evtimov, I., Sturmfels, P. & Kohno, T. Foggysight: A scheme for facial lookup privacy.Proc. Priv. Enhancing Technol. 2021, 204–226, DOI: 10.2478/POPETS-2021-0044 (2021)
2021 doi
-
[97]
Shan, S. et al. Fawkes: Protecting privacy against unauthorized deep learning models. In Capkun, S. & Roesner, F. (eds.) 29th USENIX Security Symposium, USENIX Security 2020, August 12-14, 2020, 1589–1604 (USENIX Association, 2020)
2020
-
[98]
& Tramèr, F
Radiya-Dixit, E., Hong, S., Carlini, N. & Tramèr, F. Data poisoning won’t save you from facial recognition. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022(OpenReview.net, 2022)
2022
-
[99]
N., Utz, V
Yalçın, O. N., Utz, V . & DiPaola, S. Empathy through aesthetics: Using ai stylization for visual anonymization of interview videos. In Proceedings of the 3rd Empathy-Centric Design Workshop: Scrutinizing Empathy Beyond the Individual, EmpathiCH ’24, 63–68, DOI: 10.1145/366179...
-
[100]
E., Englund, C
Rosberg, F., Aksoy, E. E., Englund, C. & Alonso-Fernandez, F. FIV A: facial image and video anonymization and anonymization defense. In IEEE/CVF International Conference on Computer Vision, ICCV 2023 - Workshops, Paris, France, October 2-6, 2023, 362–371, DOI: 10.1109/ICCVW607...
2023
-
[101]
& Hernandez-Matamoros, A
Kikuchi, H., Miyoshi, S., Mori, T. & Hernandez-Matamoros, A. A vulnerability in video anonymization - privacy disclosure from face-obfuscated video. In 19th Annual International Conference on Privacy, Security & Trust, PST 2022, Fredericton, NB, Canada, August 22-24, 2022, 1–1...
2022
-
[102]
& Chen, J
Zhao, Y . & Chen, J. A survey on differential privacy for unstructured data content. ACM Comput. Surv. 54, DOI: 10.1145/3490237 (2022)
2022 doi
-
[103]
& Song, L
Wen, Y ., Liu, B., Ding, M., Xie, R. & Song, L. Identitydp: Differential private identification protection for face images. Neurocomputing 501, 197–211, DOI: https://doi.org/10.1016/j.neucom.2022.06.039 (2022)
2022 doi
-
[104]
Lozoya, D. et al. Generating mental health transcripts with SAPE (Spanish adaptive prompt engineering). In Duh, K., Gomez, H. & Bethard, S. (eds.) Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language ...
2024 doi
-
[105]
& Lan, Z
Qiu, H., He, H., Zhang, S., Li, A. & Lan, Z. SMILE: single-turn to multi-turn inclusive language expansion via chatgpt for mental health support. In Al-Onaizan, Y ., Bansal, M. & Chen, Y . (eds.)Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Flor...
2024
-
[106]
Yue, X. et al. Synthetic text generation with differential privacy: A simple and practical recipe. In Rogers, A., Boyd- Graber, J. & Okazaki, N. (eds.) Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1321–1342, D...
-
[107]
Nahid, M. M. H. & Hasan, S. B. Safesynthdp: Leveraging large language models for privacy-preserving synthetic data generation using differential privacy. arXiv preprint arXiv:2412.20641 (2024)
2024 arXiv
-
[108]
Multichallenge: A realistic multi-turn conversation evaluation benchmark challenging to frontier llms
Sirdeshmukh, V .et al. Multichallenge: A realistic multi-turn conversation evaluation benchmark challenging to frontier llms. CoRR abs/2501.17399, DOI: 10.48550/ARXIV .2501.17399 (2025). 2501.17399
-
[109]
A., Jiang, X., Darefsky, J., Zhu, G
Li, Y . A., Jiang, X., Darefsky, J., Zhu, G. & Mesgarani, N. Styletalker: Finetuning audio language model and style-based text-to-speech model for fast spoken dialogue generation. In First Conference on Language Modeling (2024)
2024
-
[110]
Ng, E. et al. From audio to photoreal embodiment: Synthesizing humans in conversations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1001–1010 (2024)
2024
-
[111]
& Roth, A
Dwork, C. & Roth, A. The algorithmic foundations of differential privacy. Found. Trends Theor. Comput. Sci.9, 211–407, DOI: 10.1561/0400000042 (2014)
2014 doi
-
[112]
& y Arcas, B
McMahan, B., Moore, E., Ramage, D., Hampson, S. & y Arcas, B. A. Communication-efficient learning of deep networks from decentralized data. In Singh, A. & Zhu, X. J. (eds.) Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, AISTATS 2017...
2017
-
[113]
When the curious abandon honesty: Federated learning is not private
Boenisch, F.et al. When the curious abandon honesty: Federated learning is not private. In8th IEEE European Symposium on Security and Privacy, EuroS&P 2023, Delft, Netherlands, July 3-7, 2023, 175–199, DOI: 10.1109/EUROSP57164. 2023.00020 (IEEE, 2023)
2023
-
[114]
A., Mdhaffar, S., Tommasi, M., Estève, Y
Tomashenko, N. A., Mdhaffar, S., Tommasi, M., Estève, Y . & Bonastre, J. Privacy attacks for automatic speech recognition acoustic models in A federated learning framework. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022, Virtual and Si...
2022
-
[115]
Kariyappa, S. et al. Cocktail party attack: Breaking aggregation-based privacy in federated learning using independent component analysis. In Krause, A. et al. (eds.) International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, vol. 202 of P...
2023
-
[116]
Nagy, B. et al. Privacy-preserving federated learning and its application to natural language processing.Knowledge-Based Syst. 268, 110475, DOI: https://doi.org/10.1016/j.knosys.2023.110475 (2023)
2023
-
[117]
& Chen, X
Sun, L., Qian, J. & Chen, X. LDP-FL: practical private aggregation in federated learning with local differential privacy. In Zhou, Z. (ed.) Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI 2021, Virtual Event / Montreal, Canada, 19-...
2021 doi
-
[118]
Basu, P. et al. Benchmarking differential privacy and federated learning for BERT models. CoRR abs/2106.13973 (2021). 2106.13973. 17/18
2021 arXiv
-
[119]
& Chaspari, T
Ravuri, V ., Gutierrez-Osuna, R. & Chaspari, T. Preserving mental health information in speech anonymization. In10th International Conference on Affective Computing and Intelligent Interaction, ACII 2022 - Workshops and Demos, Nara, Japan, October 17-21, 2022, 1–8, DOI: 10.110...
2022
-
[120]
Pranjal, R. et al. Toward privacy-enhancing ambulatory-based well-being monitoring: Investigating user re-identification risk in multimodal data. In IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP 2023, Rhodes Island, Greece, June 4-10, 2023, 1–...
2023
-
[121]
& Viejo, A
Sánchez, D., Batet, M. & Viejo, A. Utility-preserving privacy protection of textual healthcare documents. J. Biomed. Informatics 52, 189–198, DOI: https://doi.org/10.1016/j.jbi.2014.06.008 (2014). Special Section: Methods in Clinical Research Informatics
2014 doi
-
[122]
Goncalves, A. et al. Generation and evaluation of synthetic patient data. BMC medical research methodology 20, 1–40 (2020)
2020
-
[123]
Kuppa, A., Aouad, L. M. & Le-Khac, N. Towards improving privacy of synthetic datasets. In Gruschka, N., Antunes, L. F. C., Rannenberg, K. & Drogkaris, P. (eds.) Privacy Technologies and Policy - 9th Annual Privacy Forum, APF 2021, Oslo, Norway, June 17-18, 2021, Proceedings, v...
2021 doi
-
[124]
& Norgeot, B
Shi, J., Wang, D., Tesei, G. & Norgeot, B. Generating high-fidelity privacy-conscious synthetic patient data for causal effect estimation with multiple treatments. Front. Artif. Intell. 5, DOI: 10.3389/frai.2022.918813 (2022)
2022
-
[125]
& Chua, T
Wu, S., Fei, H., Qu, L., Ji, W. & Chua, T. Next-gpt: Any-to-any multimodal LLM. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024(OpenReview.net, 2024)
2024
-
[126]
& Shmatikov, V
Bagdasaryan, E., Poursaeed, O. & Shmatikov, V . Differential privacy has disparate impact on model accuracy. In Wallach, H. M. et al. (eds.) Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, Dec...
2019
-
[127]
S., Richter, D
Kolekar, S. S., Richter, D. J., Bappi, M. I. & Kim, K. Advancing AI voice synthesis: Integrating emotional expression in multi-speaker voice generation. In International Conference on Artificial Intelligence in Information and Communication , ICAIIC 2024, Osaka, Japan, Februar...
2024
- [128]
-
[129]
& Dumontier, M
Sun, C., van Soest, J. & Dumontier, M. Generating synthetic personal health data using conditional generative adversarial networks combining with differential privacy. J. Biomed. Informatics 143, 104404, DOI: https://doi.org/10.1016/j.jbi. 2023.104404 (2023)
2023
-
[130]
Qian, Z. et al. Synthetic data for privacy-preserving clinical risk prediction. Sci. Reports 14, 25676 (2024)
2024
-
[131]
Ziegler, J. D. et al. Multi-modal conditional GAN: Data synthesis in the medical domain. In NeurIPS 2022 Workshop on Synthetic Data for Empowering ML Research (2022)
2022
-
[132]
Xin, B. et al. Private FL-GAN: differential privacy synthetic data generation based on federated learning. In 2020 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2020, Barcelona, Spain, May 4-8, 2020, 2927–2931, DOI: 10.1109/ICASSP40776.2020.9...
2020
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.