Pith. sign in

REVIEW 4 cited by

Real-time Detection of AI-Generated Speech for DeepFake Voice Conversion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.12734 v1 pith:67CRTHQV submitted 2023-08-24 cs.SD cs.CLcs.HCcs.LGeess.AS

classification cs.SDcs.CLcs.HCcs.LGeess.AS
keywords speechvoiceconversionreal-timeai-generateddetectionthereanother
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There are growing implications surrounding generative AI in the speech domain that enable voice cloning and real-time voice conversion from one individual to another. This technology poses a significant ethical threat and could lead to breaches of privacy and misrepresentation, thus there is an urgent need for real-time detection of AI-generated speech for DeepFake Voice Conversion. To address the above emerging issues, the DEEP-VOICE dataset is generated in this study, comprised of real human speech from eight well-known figures and their speech converted to one another using Retrieval-based Voice Conversion. Presenting as a binary classification problem of whether the speech is real or AI-generated, statistical analysis of temporal audio features through t-testing reveals that there are significantly different distributions. Hyperparameter optimisation is implemented for machine learning models to identify the source of speech. Following the training of 208 individual machine learning models over 10-fold cross validation, it is found that the Extreme Gradient Boosting model can achieve an average classification accuracy of 99.3% and can classify speech in real-time, at around 0.004 milliseconds given one second of speech. All data generated for this study is released publicly for future research on AI speech detection.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Towards Trustworthy Audio Deepfake Detection: A Systematic Framework for Diagnosing and Mitigating Gender Bias

    cs.SD 2026-05 unverdicted novelty 7.0 of 10

    A diagnosis-first framework for gender bias in audio deepfake detection identifies acoustic representation differences and feature leakage as sources, with per-gender threshold adjustment reducing unfairness by 54-75%...

  2. What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Underrepresented gender in training suffers higher deepfake-detection error; WavLM gaps stay large under balance, and all post-hoc calibrations leave the EER gap fixed at 1.317 pp.

  3. Spectro-Temporal Modulation Representation Framework for Human-Imitated Speech Detection

    cs.SD 2026-04 unverdicted novelty 5.0 of 10

    STM representations from auditory filterbanks detect human-imitated speech at or above human listener accuracy levels.

  4. Gender Fairness in Audio Deepfake Detection: Performance and Disparity Analysis

    cs.SD 2026-03 unverdicted novelty 5.0 of 10

    Fairness metrics uncover gender disparities in audio deepfake detection error distributions that standard Equal Error Rate metrics obscure.

Pith tools