REVIEW 2 major objections 5 minor 7 cited by
From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview
T0 review · 2 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This review claims to be the first comprehensive overview of AI-generated music detection, arguing that detectors should be built on intrinsic musical features—melody, harmony, rhythm, and lyrics—rather than surface artifacts like…
desk verdict Genuinely the first map of AI-generated music detection, useful for newcomers, but the composition/arrangement 'essence' assumption is unproven and the reference list needs cleanup before it can serve as a trusted entry point. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the five-stage music production pipeline—composition, arrangement, sound design, mixing, and mastering—with the first two stages treated as the creative core that defines a song's essence. Around this the review builds a feature taxonomy: content-based features (melody, harmony, rhythm, lyrics) and decoration-based features (timbre, instrument, emotion tone, genre), matched against a detection taxonomy of end-to-end versus feature-based classifiers. This machinery organizes the survey: generation models condition on these musical features, detectors are judged by whether they capture them, and datasets are evaluated by whether they include them, making musicological feature understanding the proposed anchor for transferring foundation models from audio deepfake detection to AIGM detection.
What would settle it
Take songs composed and arranged by humans, pass only the synthesis, mixing, and mastering stages through AI tools, then test whether detectors built on intrinsic musical features can still separate them from fully human productions; if artifact-focused detectors succeed while intrinsic-feature detectors fail, the paper's central recommendation would be refuted.
Extended reading notes
Core claim
The paper's central claim is that AIGM detection is an emerging but neglected field, and that its first comprehensive review shows a workable direction: intrinsic musical features carry the detectable signature of AI authorship, while surface-level artifacts such as watermarks do not. It surveys the two dedicated AIGM datasets (FakeMusicCaps, an audio-only collection of human and AI-generated clips, and SONICS, a multimodal audio-and-lyrics set), along with the few existing detectors such as SpecTTTra, a transformer that tokenizes spectro-temporal features, and a convolutional music deepfake detector. The authors ground this in a five-stage music production model in which composition and arrangement establish the creative core, so detectors should classify on content-based features and decoration-based features rather than on post-production polish. They conclude that direct transfer from speech deepfake detection is insufficient and that future detectors need music-specific, multimodal, explainable designs.
Load-bearing premise
The review assumes that what makes music creative and detectable resides in the composition and arrangement stages, with sound design, mixing, and mastering acting only as aesthetic polish, so if AI artifacts enter mainly during those later production stages the proposed intrinsic-feature focus would miss the signal.
Editorial extensions
If this is right
- If intrinsic musical features are the right signal, future AIGM detectors should be benchmarked on whether they capture melody, harmony, rhythm, and lyrics, not just on in-domain accuracy against a single generator.
- Audio deepfake detection models can serve as starting points for AIGM detection, but the paper argues direct transfer or fine-tuning is insufficient because music carries musicological structure and subjective qualities that speech audio lacks.
- Because lyrics are a separate text modality that shapes a song's meaning and emotion, the paper concludes that multimodal audio-and-lyrics detectors are necessary for accurate AIGM detection.
- The scarcity of dedicated datasets—only FakeMusicCaps and SONICS exist—means that building comprehensive, accessible benchmarks is a prerequisite for progress, the paper states.
- Surface-level cues such as watermarking are unreliable because generators can avoid them, so the paper prioritizes intrinsic features as the core detection signal.
Reading between the lines
- A testable extension is the review's composition-and-arrangement focus: if AI artifacts concentrate in mixing and mastering, detectors built on intrinsic features will miss them, and artifact-focused methods would win in that scenario.
- The paper leaves open whether self-supervised speech representations or music-specific pretraining should anchor the transfer pathway; a likely next experiment is to compare both on the same AIGM benchmark.
- The adversarial relationship between generation and detection implies that once detectors learn intrinsic-feature signatures, generators will be optimized against those signatures, so benchmarks will need periodic refreshment to stay meaningful.
- A concrete editorial extension is that the survey's gap analysis suggests a shared out-of-domain evaluation protocol, where detectors are tested on unseen generators and unseen production styles, would serve the field better than any single accuracy number.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey/overview of AI-generated music (AIGM) detection. It begins with an introduction to music features and their role in generation, then describes existing AIGM detection datasets and detectors, reviews audio deepfake detection methods and multimodal fusion techniques, and proposes a pathway for adapting deepfake audio foundation models to AIGM detection. The paper argues that intrinsic musicological features (melody, harmony, rhythm, lyrics, timbre, etc.) should be the core signals for detection, and it closes with challenges and future directions, including the need for benchmarks, domain-specific models, and explainability.
Significance. If the paper's claims hold, it would provide a useful early map of a small but emerging field, and its proposed focus on musicological features is a distinctive position that could guide subsequent research. The paper also gathers the two known AIGM-specific datasets (FakeMusicCaps and SONICS) and connects them to the broader audio deepfake literature, which is valuable for researchers entering the area. However, the paper's scholarly value depends on accurate citations and on the defensibility of its scoping decision to prioritize composition and arrangement while downplaying later production stages; both of these currently need attention.
major comments (2)
- [Section I] The claim that composition and arrangement carry 'the most creativity and foundational stages' while sound design, mixing, and mastering 'primarily serve as aesthetic enhancements rather than altering the essence of the music' is load-bearing for the review's scope and for the Section V recommendation that 'intrinsic features unique to music are essential and should be prioritised as core detection features.' No evidence or citation is provided for this empirical claim, and it stands in tension with the paper's own discussion in Section III.B: the BEAR framework of Shih et al. shows that audio deepfake detectors can key on surface-level, noise-like artifacts, and the paper itself criticizes AIGM detectors for 'dependence on surface-level features.' Such artifacts can plausibly be introduced in rendering, synthesis, or mastering stages, and modern end-to-end generators such as MusicLM or Suno do not separate composition from production. The authors should either support the scoping claim with evidence or explicitly expand the review to include artifact-based and production-stage detection signals; without this, the proposed pathway rests on an unsupported assumption.
- [Section II; References] The reference list contains several defects that undermine the verifiability expected of a review: Suno AI is cited with a placeholder '[?]' in the 'Commercial tools' paragraph of Section II; reference [73] is labeled 'Image transformers' with an author list that does not match arXiv:2010.11929, which is in fact the same paper as reference [104] ('An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale'); and the SongComposer entry appears twice as references [40] and [59]. In addition, references [110] and [114] duplicate the same AlexNet paper. For a survey whose contribution is explicitly to provide a reliable map of the literature, these errors are not merely cosmetic; they need to be corrected before publication.
minor comments (5)
- [Section II] There is a typo in 'catagorised' (should be 'categorised') and the phrase 'thought it is a bit out-dated' is grammatically awkward and should be revised.
- [Section I, Figure 1] The figure caption contains scattered music symbols and a somewhat unstructured sentence; a cleaner description of the five production steps would improve readability.
- [Section III.A and Table I] The paper uses 'Afchar [60]' both for the proposed dataset and for the detector described in the same section; please disambiguate the dataset reference from the detection method reference.
- [Section III.B] The dataset name 'sound8k' appears with inconsistent capitalization (elsewhere as 'Sound8K'); please standardize the naming.
- [Table II] The column heading 'Compared Baseline' is unclear: some entries list model names (e.g., RawNet2, ResNet) rather than a baseline comparison, and the column should be more precisely titled or supplemented with a description in the table notes.
Circularity Check
Review paper with no derived predictions or fitted parameters; no circular reasoning found.
full rationale
This manuscript is a literature review and research roadmap, not a derivation or empirical study. It contains no fitted parameters, no equations that reduce to their own inputs, and no predictive claims that are forced by construction. The paper's central contribution is the claim to provide the first comprehensive AIGM detection review; while this novelty claim is not independently verified, it is an assertion about the literature, not a derived result. The proposed future pathway (focusing on intrinsic music features) is presented as a suggestion, explicitly grounded in the authors' qualitative judgment that composition and arrangement define music's essence, and it is not derived from the reviewed material in a way that would constitute circularity. The authors' self-citations are limited to routine references to their own affiliations and prior work in the context of broader literature; none of these citations is load-bearing for a specific claim that reduces to itself. The review is self-contained as a survey and does not make predictions that could be circular. Score 0 is appropriate: no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption AIGM detection can be treated as a binary classification task (human vs. AI generated music).
- domain assumption Music production can be decomposed into five stages (composition, arrangement, sound design, mixing, mastering), and the first two carry the detectable creative content.
- domain assumption The reviewed papers and datasets correctly represent the state of the art.
Cite this review
Pith. "Pith review of From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview." pith.science (2026). https://pith.science/paper/RUY7D5JJ
@misc{pith2026241200571,
author = {Pith},
title = {Pith review of: From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview},
year = {2026},
howpublished = {\url{https://pith.science/paper/RUY7D5JJ}},
note = {Machine review of arXiv:2412.00571}
}
read the original abstract
As Artificial Intelligence (AI) technologies continue to evolve, their use in generating realistic, contextually appropriate content has expanded into various domains. Music, an art form and medium for entertainment, deeply rooted into human culture, is seeing an increased involvement of AI into its production. However, despite the effective application of AI music generation (AIGM) tools, the unregulated use of them raises concerns about potential negative impacts on the music industry, copyright and artistic integrity, underscoring the importance of effective AIGM detection. This paper provides an overview of existing AIGM detection methods. To lay a foundation to the general workings and challenges of AIGM detection, we first review general principles of AIGM, including recent advancements in deepfake audios, as well as multimodal detection techniques. We further propose a potential pathway for leveraging foundation models from audio deepfake detection to AIGM detection. Additionally, we discuss implications of these tools and propose directions for future research to address ongoing challenges in the field.
Figures
Forward citations
Cited by 7 Pith papers
-
MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing
MixAssist is the first audio-grounded, multi-turn conversational dataset for co-creative music mixing instruction, and fine-tuning Qwen-Audio on it yields human-comparable mixing advice.
-
Assessing AI-generated music detection in real-world broadcast monitoring
Two CNN detectors that are near-perfect on clean AI music drop to F1 0.19 to 0.47 on real TV broadcast recordings from the new BAMM dataset.
-
Finding the noise: Zero-shot AI Music Detection
A zero-shot method based on fakeprints, NMF and a blur-based reconstruction error detects unknown AI-music generators in one-class and clustering setups, working for most services but missing Mubert and pre-v9 Mureka.
-
Echoes: A semantically-aligned music deepfake detection dataset
A semantically aligned, multi-provider music deepfake dataset is harder for detectors and trains models that transfer better than prior AI-music datasets.
-
A Fourier Explanation of AI-music Artifacts
Zero-upsampling in audio generators replicates the low-frequency spectrum at fixed intervals, creating architecture-dependent spectral peaks that enable a simple, interpretable detector for AI-generated music.
-
Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
A late-fusion model that combines ASR-transcribed lyrics and speech embeddings detects AI-written lyrics from audio alone, achieving 94.9% recall in-domain and staying robust to attacks.
-
Improved Robustness in AI-Generated Music Detection
Log-frequency remapping plus a single cross-correlation filter makes AI-music artifact detection invariant to speed change by design, matching clean-audio SOTA while recovering the speed factor.
Reference graph
Works this paper leans on
-
[104]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” 2021. [Online]. Available: https://arxiv.org/abs/2010.11929
arXiv 2021
-
[106]
Wav2vec 2.0: A framework for self-supervised learning of speech representations,
Z. Y . Baevski, A. et al., “Wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proceedings of the NeurIPS Conference. NeurIPS, 2020, pp. 12 449–12 460
2020
-
[59]
Songcomposer: A large language model for lyric and melody composition in song generation,
S. Ding, Z. Liu, X. Dong, P. Zhang, R. Qian, C. He, D. Lin, and J. Wang, “Songcomposer: A large language model for lyric and melody composition in song generation,” 2024. [Online]. Available: https://arxiv.org/abs/2402.17645
arXiv 2024
-
[114]
Imagenet classification with deep convolutional neural networks,
S. I. Krizhevsky, A. and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in Neural Information Processing Systems (NeurIPS) , vol. 25, pp. 1097–1105, 2012
2012
-
[1]
Gpt-4 technical report,
OpenAI, “Gpt-4 technical report,” 2023
2023
-
[2]
Robust speech recognition via large-scale weak supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” 2022. [Online]. Available: https://github.com/openai/ whisper
2022
-
[3]
MuseNet: Ai-generated music with deep learning models,
OpenAI, “MuseNet: Ai-generated music with deep learning models,” https://openai.com/research/musenet, 2019, capable of composing mu- sic with up to 10 different instruments in various styles and genres, MuseNet uses deep learning models to generate harmonies, melodies, and full compositions
2019
-
[4]
Jukedeck: Background music generation,
Jukedeck, “Jukedeck: Background music generation,” https://jukedeck. com, 2020, focuses on generating background music for videos, allow- ing users to specify genre, mood, and tempo, with auto-composition tools to streamline music production. Acquired by TikTok
2020
Show all 154 references
-
[5]
Aiva: Artificial intelligence virtual artist,
A. Technologies, “Aiva: Artificial intelligence virtual artist,” https:// www.aiva.ai, 2016, aIV A is a tool for creating music in various genres and can compose pieces for soundtracks or game development. It allows users to select instruments and provides sheet music
2016
-
[6]
Impacts of ai on music consumption and fairness,
A. Henry, V . Wiratama, A. Afilipoaie, H. Ranaivoson, and E. Arrivé, “Impacts of ai on music consumption and fairness,” Emerging Media, p. 27523543241269047, 2024
2024
-
[7]
An adaptive meta-heuristic for music plagiarism detection based on text similarity and clustering,
D. Malandrino, R. De Prisco, M. Ianulardo, and R. Zaccagnino, “An adaptive meta-heuristic for music plagiarism detection based on text similarity and clustering,” Data Mining and Knowledge Discovery , vol. 36, no. 4, pp. 1301–1334, 2022
2022
-
[8]
Audio deepfake detection: A survey,
J. Yi, C. Wang, J. Tao, X. Zhang, C. Y . Zhang, and Y . Zhao, “Audio deepfake detection: A survey,” arXiv preprint arXiv:2308.14970, 2023
2023 arXiv
-
[9]
Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild,
X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kinnunen, M. Todisco, J. Yamagishi, N. Evans, A. Nautsch et al. , “Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp...
2021
-
[10]
Add 2023: the second audio deepfake detection challenge,
J. Yi, J. Tao, R. Fu, X. Yan, C. Wang, T. Wang, C. Y . Zhang, X. Zhang, Y . Zhao, Y . Renet al., “Add 2023: the second audio deepfake detection challenge,” arXiv preprint arXiv:2305.13774 , 2023
2023 arXiv
-
[11]
Deepfakes generation and detection: State-of-the-art, open challenges, countermeasures, and way forward,
M. Masood, M. Nawaz, K. M. Malik, A. Javed, A. Irtaza, and H. Malik, “Deepfakes generation and detection: State-of-the-art, open challenges, countermeasures, and way forward,”Applied intelligence, vol. 53, no. 4, pp. 3974–4026, 2023
2023
-
[12]
Detecting multimedia generated by large ai models: A survey,
L. Lin, N. Gupta, Y . Zhang, H. Ren, C.-H. Liu, F. Ding, X. Wang, X. Li, L. Verdoliva, and S. Hu, “Detecting multimedia generated by large ai models: A survey,” arXiv preprint arXiv:2402.00045 , 2024
2024 arXiv
-
[13]
Review of audio deepfake detection techniques: Issues and prospects,
A. Dixit, N. Kaur, and S. Kingra, “Review of audio deepfake detection techniques: Issues and prospects,” Expert Systems , vol. 40, no. 8, p. e13322, 2023
2023
-
[14]
Deepfake generation and detection: Case study and challenges,
Y . Patel, S. Tanwar, R. Gupta, P. Bhattacharya, I. E. Davidson, R. Nyameko, S. Aluvala, and V . Vimal, “Deepfake generation and detection: Case study and challenges,” IEEE Access, 2023
2023
-
[15]
A review of modern audio deepfake de- tection methods: challenges and future directions,
Z. Almutairi and H. Elgibreen, “A review of modern audio deepfake de- tection methods: challenges and future directions,” Algorithms, vol. 15, no. 5, p. 155, 2022
2022
-
[16]
Taylor, Music, Subjectivity, and Schumann
B. Taylor, Music, Subjectivity, and Schumann . Cambridge University Press, 2022
2022
-
[17]
Music composition with deep learning: A review,
B. Sturm, E. Benetos et al. , “Music composition with deep learning: A review,” arXiv preprint arXiv:2108.12290 , 2021
2021 arXiv
-
[18]
Poetic discourse in popular song lyrics: An analysis of the linguistic features of a selection of popular songs,
D. Irvine, “Poetic discourse in popular song lyrics: An analysis of the linguistic features of a selection of popular songs,” International Journal of Language and Linguistics , vol. 2, no. 2, pp. 74–81, 2015
2015
-
[19]
Technical, musical, and legal aspects of an ai-aided algorithmic music production system,
J. Kwiecie ´n, P. Skrzy ´nski, W. Chmiel, A. D ˛ abrowski, B. Szadkowski, and M. Pluta, “Technical, musical, and legal aspects of an ai-aided algorithmic music production system,” Applied Sciences, vol. 14, no. 9, p. 3541, 2024
2024
-
[20]
Foundation models for music: A survey,
Y . Ma, A. Øland, A. Ragni, B. M. D. Sette, C. Saitis, C. Donahue, C. Lin, C. Plachouras, E. Benetos, E. Shatri, F. Morreale, G. Zhang, G. Fazekas, G. Xia, H. Zhang, I. Manco, J. Huang, J. Guinot, L. Lin, L. Marinelli, M. W. Y . Lam, M. Sharma, Q. Kong, R. B. Dannenberg, R. Yu...
2024 arXiv
-
[21]
Deep learning for music genre recognition,
A. Roy, N. Ippili, and D. K. Teja, “Deep learning for music genre recognition,” in 2023 International Conference on Communication, Security and Artificial Intelligence (ICCSAI) , 2023, pp. 744–748
2023
-
[22]
musif: a python package for symbolic music feature extraction,
A. Llorens, F. Simonetta, M. Serrano, and Á. Torrente, “musif: a python package for symbolic music feature extraction,” arXiv preprint arXiv:2307.01120, 2023
2023
-
[23]
Benward and M
B. Benward and M. Saker, Music in Theory and Practice , 8th ed. Boston: McGraw-Hill, 2009, vol. 1
2009
-
[24]
Twinkle, twinkle, little star,
W. A. Briggs, “Twinkle, twinkle, little star,” Retrieved from the Library of Congress, Boston, 1880, [Notated Music]. [Online]. Available: https://www.loc.gov/item/2023838232/
-
[25]
Melody extraction from polyphonic music by deep learning approaches: A review,
K. S. Rao, P. P. Das et al. , “Melody extraction from polyphonic music by deep learning approaches: A review,” arXiv preprint arXiv:2202.01078, 2022
2022 arXiv
-
[26]
Frequency-temporal attention network for singing melody extraction,
S. Yu, X. Sun, Y . Yu, and W. Li, “Frequency-temporal attention network for singing melody extraction,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 251–255
2021
-
[27]
Tonet: Tone-octave network for singing melody extraction from poly- phonic music,
K. Chen, S. Yu, C.-i. Wang, W. Li, T. Berg-Kirkpatrick, and S. Dubnov, “Tonet: Tone-octave network for singing melody extraction from poly- phonic music,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 621–625
2022
-
[28]
Melody is all you need for music generation,
S. Wei, M. Wei, H. Wang, Y . Zhao, and G. Kou, “Melody is all you need for music generation,” arXiv preprint arXiv:2409.20196 , 2024
2024 arXiv
-
[29]
Towards automatic extraction of harmony information from music signals,
C. Harte, “Towards automatic extraction of harmony information from music signals,” Ph.D. dissertation, 2010
2010
-
[30]
Functional harmony ontology: Musical harmony analysis with description logics,
S. Kantarelis, E. Dervakos, N. Kotsani, and G. Stamou, “Functional harmony ontology: Musical harmony analysis with description logics,” Journal of Web Semantics , vol. 75, p. 100754, 2023
2023
-
[31]
Structure-enhanced pop music generation via harmony-aware learning,
X. Zhang, J. Zhang, Y . Qiu, L. Wang, and J. Zhou, “Structure-enhanced pop music generation via harmony-aware learning,” in Proceedings of the 30th ACM International Conference on Multimedia , ser. MM ’22. ACM, Oct. 2022, p. 1204–1213. [Online]. Available: http://dx.doi.org/10...
2022
-
[32]
Emotional harmony through deep learning: A facial expression- based music therapy,
A. Gobinath, A. K. SNS, M. Athithyan, M. Anandan, P. Rajeswari et al., “Emotional harmony through deep learning: A facial expression- based music therapy,” in 2024 Third International Conference on Intelligent Techniques in Control, Optimization and Signal Processing (INCOS). ...
2024
-
[33]
Emotion-driven melody harmonization via melodic variation and functional representation,
J. Huang and Y .-H. Yang, “Emotion-driven melody harmonization via melodic variation and functional representation,” arXiv preprint arXiv:2407.20176, 2024
2024 arXiv
-
[34]
Rhythm pattern representations for tempo detection in music,
S. Gulati and P. Rao, “Rhythm pattern representations for tempo detection in music,” in Proceedings of the First International Conference on Intelligent Interactive Technologies and Multimedia , ser. IITM ’10. New York, NY , USA: Association for Computing Machinery, 2011, p. 2...
-
[35]
Cadence detection in symbolic classical music using graph neural networks,
E. Karystinaios and G. Widmer, “Cadence detection in symbolic classical music using graph neural networks,” 2022. [Online]. Available: https://arxiv.org/abs/2208.14819
2022 arXiv
-
[36]
Mrbert: Pre-training of melody and rhythm for automatic music generation,
S. Li and Y . Sung, “Mrbert: Pre-training of melody and rhythm for automatic music generation,” Mathematics, vol. 11, no. 4, p. 798, 2023
2023
-
[37]
Musicongen: Rhythm and chord control for transformer-based text-to-music gener- ation,
Y .-H. Lan, W.-Y . Hsiao, H.-C. Cheng, and Y .-H. Yang, “Musicongen: Rhythm and chord control for transformer-based text-to-music gener- ation,” arXiv preprint arXiv:2407.15060 , 2024
2024 arXiv
-
[38]
Lyemobert: Classification of lyrics’ emotion and recommendation using a pre-trained model,
V . Revathy, A. S. Pillai, and F. Daneshfar, “Lyemobert: Classification of lyrics’ emotion and recommendation using a pre-trained model,” Procedia Computer Science , vol. 218, pp. 1196–1208, 2023. JOURNAL OF LATEX CLASS FILES 11
2023
-
[39]
Songcreator: Lyrics-based universal song generation,
S. Lei, Y . Zhou, B. Tang, M. W. Lam, F. Liu, H. Liu, J. Wu, S. Kang, Z. Wu, and H. Meng, “Songcreator: Lyrics-based universal song generation,” arXiv preprint arXiv:2409.06029 , 2024
2024 arXiv
-
[41]
Controllable lyrics-to-melody genera- tion,
Z. Zhang, Y . Yu, and A. Takasu, “Controllable lyrics-to-melody genera- tion,” Neural Computing and Applications, vol. 35, no. 27, pp. 19 805– 19 819, 2023
2023
-
[42]
Innovations in cover song detection: A lyrics-based approach,
M. Balluff, P. Mandl, and C. Wolff, “Innovations in cover song detection: A lyrics-based approach,” arXiv preprint arXiv:2406.04384, 2024
2024 arXiv
-
[43]
Song authorship attribution: a lyrics and rhyme based approach,
T. Yılmaz and T. Scheffler, “Song authorship attribution: a lyrics and rhyme based approach,” International Journal of Digital Humanities , vol. 5, no. 1, pp. 29–44, 2023
2023
-
[44]
Detecting synthetic lyrics with few-shot inference,
Y . Labrak, G. Meseguer-Brocal, and E. V . Epure, “Detecting synthetic lyrics with few-shot inference,”arXiv preprint arXiv:2406.15231, 2024
2024 arXiv
-
[45]
Musical timbre style transfer with diffusion model,
H. Huang, J. Man, L. Li, and R. Zeng, “Musical timbre style transfer with diffusion model,” PeerJ Computer Science , vol. 10, p. e2194, 2024
2024
-
[46]
Gtr-ctrl: instrument and genre conditioning for guitar-focused music generation with transformers,
P. Sarmento, A. Kumar, Y .-H. Chen, C. Carr, Z. Zukowski, and M. Bar- thet, “Gtr-ctrl: instrument and genre conditioning for guitar-focused music generation with transformers,” in International Conference on Computational Intelligence in Music, Sound, Art and Design (Part of E...
2023
-
[47]
Emopia: A multi-modal pop piano dataset for emotion recognition and emotion-based music generation,
H.-T. Hung, J. Ching, S. Doh, N. Kim, J. Nam, and Y .-H. Yang, “Emopia: A multi-modal pop piano dataset for emotion recognition and emotion-based music generation,” arXiv preprint arXiv:2108.01374 , 2021
2021 arXiv
-
[48]
Generative adversarial nets,
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems (NeurIPS) , 2014, pp. 2672–2680. [Online]. Available: https://papers.nips.cc/paper/ 20...
2014
-
[49]
Musenet,
C. Payne, “Musenet,” https://openai.com/research/musenet, 2019, https: //openai.com/research/musenet
2019
-
[50]
Music transformer,
C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, N. Shazeer, I. Simon, C. Hawthorne, A. M. Dai, M. D. Hoffman, M. Dinculescu, and D. Eck, “Music transformer,” 2018. [Online]. Available: https://arxiv.org/abs/1809.04281
2018 arXiv
-
[51]
A hierarchical latent vector model for learning long-term structure in music,
A. Roberts, J. Engel, C. Raffel, C. Hawthorne, and D. Eck, “A hierarchical latent vector model for learning long-term structure in music,” in Proceedings of the 35th International Conference on Machine Learning (ICML) , 2018. [Online]. Available: https: //proceedings.mlr.press...
2018
-
[52]
Musegan: Multi-track sequential generative adversarial networks for symbolic music generation and accompaniment,
H.-W. Dong, W.-Y . Hsiao, L.-C. Yang, and Y .-H. Yang, “Musegan: Multi-track sequential generative adversarial networks for symbolic music generation and accompaniment,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018. [Online]. Available: ...
2018
-
[53]
Auto-encoding variational bayes,
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114 , 2013. [Online]. Available: https: //arxiv.org/abs/1312.6114
2013 arXiv
-
[54]
Diffwave: A versatile diffusion model for audio synthesis,
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “Diffwave: A versatile diffusion model for audio synthesis,” 2021. [Online]. Available: https://arxiv.org/abs/2009.09761
2021 arXiv
-
[55]
Melody-conditioned lyrics generation with seqgans,
Y . Chen and A. Lerch, “Melody-conditioned lyrics generation with seqgans,” in 2020 IEEE International Symposium on Multimedia (ISM), 2020, pp. 189–196
2020
-
[56]
Conditional lstm-gan for melody generation from lyrics,
Y . Yu, A. Srivastava, and S. Canales, “Conditional lstm-gan for melody generation from lyrics,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) , vol. 17, no. 1, pp. 1–20, 2021
2021
-
[57]
Melody generation sys- tem based on a theory of melody sequence,
S. Yazawa, M. Hamanaka, and T. Utsuro, “Melody generation sys- tem based on a theory of melody sequence,” in 2014 International Conference of Advanced Informatics: Concept, Theory and Application (ICAICTA). IEEE, 2014, pp. 336–341
2014
-
[58]
Continuous melody generation via disentangled short-term representations and structural conditions,
K. Chen, G. Xia, and S. Dubnov, “Continuous melody generation via disentangled short-term representations and structural conditions,” in 2020 IEEE 14th International Conference on Semantic Computing (ICSC). IEEE, 2020, pp. 128–135
2020
-
[60]
Detecting music deep- fakes is easy but actually hard,
D. Afchar, G. M. Brocal, and R. Hennequin, “Detecting music deep- fakes is easy but actually hard,”arXiv preprint arXiv:2405.04181, 2024
2024 arXiv
-
[61]
As good as a coin toss human detection of ai-generated images, videos, audio, and audiovisual stimuli,
D. Cooke, A. Edwards, S. Barkoff, and K. Kelly, “As good as a coin toss human detection of ai-generated images, videos, audio, and audiovisual stimuli,” arXiv preprint arXiv:2403.16760 , 2024
2024 arXiv
-
[62]
Fakemusiccaps: a dataset for detection and attribution of synthetic music generated via text-to- music models,
L. Comanducci, P. Bestagini, and S. Tubaro, “Fakemusiccaps: a dataset for detection and attribution of synthetic music generated via text-to- music models,” arXiv preprint arXiv:2409.10684 , 2024
2024 arXiv
-
[63]
Sonics: Synthetic or not–identifying counterfeit songs,
M. A. Rahman, Z. I. A. Hakim, N. H. Sarker, B. Paul, and S. A. Fattah, “Sonics: Synthetic or not–identifying counterfeit songs,” arXiv preprint arXiv:2408.14080, 2024
2024 arXiv
-
[64]
Musiclm: Generating music from text,
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, M. Sharifi, N. Zeghidour, and C. Frank, “Musiclm: Generating music from text,”
-
[65]
Genius song lyrics with language information,
C. G. D. C. J., “Genius song lyrics with language information,” 2023, accessed: 2024-07-05. [Online]. Available: https://www.kaggle.com/ datasets/carlosgdcj/genius-song-lyrics-with-language-information
2023
-
[66]
Dali: A large dataset of synchronized audio, lyrics and notes, automatically created using teacher-student machine learning paradigm
G. Meseguer-Brocal, A. Cohen-Hadria, and G. Peeters, “Dali: A large dataset of synchronized audio, lyrics and notes, automatically created using teacher-student machine learning paradigm.” 2018. [Online]. Available: https://zenodo.org/record/1492443
2018
-
[67]
Fma: A dataset for music analysis,
M. Defferrard, K. Benzi, P. Vandergheynst, and X. Bresson, “Fma: A dataset for music analysis,” 2017. [Online]. Available: https://arxiv.org/abs/1612.01840
2017 arXiv
-
[68]
Enabling factorized piano music modeling and generation with the MAESTRO dataset,
C. Hawthorne, A. Stasyuk, A. Roberts, I. Simon, C.-Z. A. Huang, S. Dieleman, E. Elsen, J. Engel, and D. Eck, “Enabling factorized piano music modeling and generation with the MAESTRO dataset,” in International Conference on Learning Representations, 2019. [Online]. Available: ...
2019
-
[69]
The million song dataset
T. Bertin-Mahieux, D. Ellis, B. Whitman, and P. Lamere, “The million song dataset.” 01 2011, pp. 591–596
2011
-
[70]
Learning features of music from scratch,
J. Thickstun, Z. Harchaoui, and S. M. Kakade, “Learning features of music from scratch,” in International Conference on Learning Representations (ICLR), 2017
2017
-
[71]
The mtg-jamendo dataset for automatic music tagging,
D. Bogdanov, M. Won, P. Tovstogan, A. Porter, and X. Serra, “The mtg-jamendo dataset for automatic music tagging,” in Machine Learning for Music Discovery Workshop, International Conference on Machine Learning (ICML 2019) , Long Beach, CA, United States,
2019
-
[72]
Sunocaps: A novel dataset of text-prompt based ai-generated music with emotion annotations,
M. Civit, V . Drai-Zerbib, D. Lizcano, and M. J. Escalona, “Sunocaps: A novel dataset of text-prompt based ai-generated music with emotion annotations,” Data in Brief , vol. 55, p. 110743, 2024
2024
-
[74]
Convnext: Revisiting convolutional neural networks for vision,
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, J. Cao, Q. Lu, F. Wei et al. , “Convnext: Revisiting convolutional neural networks for vision,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2022, pp. 5578– 5588, https://arxiv.or...
2022 arXiv
-
[75]
Random forests,
L. Breiman, “Random forests,” Machine learning , vol. 45, no. 1, pp. 5–32, 2001. [Online]. Available: https://doi.org/10.1023/A: 1010933404324
2001 doi
-
[76]
Melodic deception: Exploring the complexities of deepfakes of music generated by generative adversarial networks (gans)
A. Ahuja et al. , “Melodic deception: Exploring the complexities of deepfakes of music generated by generative adversarial networks (gans).” Sangeet Galaxy, vol. 13, no. 2, 2024
2024
-
[77]
Wavefake: A data set to facilitate audio deepfake detection,
J. Frank and L. Schönherr, “Wavefake: A data set to facilitate audio deepfake detection,” arXiv preprint arXiv:2111.02813 , 2021
2021 arXiv
-
[78]
For: A dataset for synthetic speech detection,
R. Reimao and V . Tzerpos, “For: A dataset for synthetic speech detection,” in 2019 International Conference on Speech Technology and Human-Computer Dialogue (SpeD) . IEEE, 2019, pp. 1–10
2019
-
[79]
Fakeavceleb: A novel audio-video multimodal deepfake dataset,
H. Khalid, S. Tariq, M. Kim, and S. S. Woo, “Fakeavceleb: A novel audio-video multimodal deepfake dataset,” arXiv preprint arXiv:2108.05080, 2021
2021 arXiv
-
[80]
Add 2022: The first audio deepfake detection challenge,
J. Yi et al., “Add 2022: The first audio deepfake detection challenge,” in ICASSP, 2022
2022
-
[81]
Add 2023: Audio deepfake detection challenge,
H. Zhang et al. , “Add 2023: Audio deepfake detection challenge,” in ICASSP, 2023
2023
-
[82]
Asvspoof 2015: The first automatic speaker verification spoofing and countermeasures challenge,
Z. Wu et al., “Asvspoof 2015: The first automatic speaker verification spoofing and countermeasures challenge,” in Interspeech, 2015
2015
-
[83]
Asvspoof 2019: Automatic speaker verification spoofing and countermeasures challenge evaluation,
M. Todisco et al. , “Asvspoof 2019: Automatic speaker verification spoofing and countermeasures challenge evaluation,” in Interspeech, 2019. JOURNAL OF LATEX CLASS FILES 12
2019
-
[84]
Asvspoof 2021: Automatic speaker verification spoofing and countermeasures challenge evaluation,
J. Yamagishi et al. , “Asvspoof 2021: Automatic speaker verification spoofing and countermeasures challenge evaluation,” in ASVspoof Workshop, 2021
2021
-
[85]
Safeear: Content privacy-preserving audio deepfake detection,
X. Li, K. Li, Y . Zheng, C. Yan, X. Ji, and W. Xu, “Safeear: Content privacy-preserving audio deepfake detection,” arXiv preprint arXiv:2409.09272, 2024
2024 arXiv
-
[86]
The ami meeting corpus: A pre-announcement,
J. Carletta et al., “The ami meeting corpus: A pre-announcement,” in International Workshop on Machine Learning for Multimodal Interac- tion, 2006
2006
-
[87]
Deep speech 2: End-to-end speech recognition in english and mandarin,
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen et al. , “Deep speech 2: End-to-end speech recognition in english and mandarin,” in International conference on machine learning . PMLR, 2016, pp. 173–182
2016
-
[88]
The voice conversion challenge 2018: Pro- moting development of parallel and non-parallel methods,
J. Lorenzo-Trueba et al., “The voice conversion challenge 2018: Pro- moting development of parallel and non-parallel methods,” in Odyssey: The Speaker and Language Recognition Workshop , 2018
2018
-
[89]
Mlaad: The multi-language audio anti-spoofing dataset,
N. M. Müller, P. Kawa, W. H. Choong, E. Casanova, E. Gölge, T. Müller, P. Syga, P. Sperl, and K. Böttinger, “Mlaad: The multi-language audio anti-spoofing dataset,” arXiv preprint arXiv:2401.09512, 2024
2024 arXiv
-
[90]
A dataset and taxonomy for urban sound research,
J. Salamon, C. Jacoby, and J. P. Bello, “A dataset and taxonomy for urban sound research,” in 22nd ACM International Conference on Multimedia (ACM-MM’14), Orlando, FL, USA, Nov. 2014, pp. 1041– 1044
2014
-
[91]
Does audio deepfake detection generalize?
N. M. Müller, P. Czempin, F. Dieckmann, A. Froghyar, and K. Böt- tinger, “Does audio deepfake detection generalize?” Interspeech, 2022
2022
-
[92]
Attention is all you need,
A. Vaswani, “Attention is all you need,” Advances in Neural Informa- tion Processing Systems , 2017
2017
-
[93]
Detection of ai-synthesized speech using cepstral & bispectral statistics,
A. K. Singh and P. Singh, “Detection of ai-synthesized speech using cepstral & bispectral statistics,” in 2021 IEEE 4th International Con- ference on Multimedia Information Processing and Retrieval (MIPR) . IEEE, 2021, pp. 412–417
2021
-
[94]
Gaussian mixture models
D. A. Reynolds et al. , “Gaussian mixture models.” Encyclopedia of biometrics, vol. 741, no. 659-663, 2009
2009
-
[95]
Multi-path gmm- mobilenet based on attack algorithms and codecs for synthetic speech and deepfake detection
Y . Wen, Z. Lei, Y . Yang, C. Liu, and M. Ma, “Multi-path gmm- mobilenet based on attack algorithms and codecs for synthetic speech and deepfake detection.” in INTERSPEECH, 2022, pp. 4795–4799
2022
-
[96]
An introduction to convolutional neural networks,
K. O’Shea and R. Nash, “An introduction to convolutional neural networks,” 2015. [Online]. Available: https://arxiv.org/abs/1511.08458
2015 arXiv
-
[97]
Lcnn: Lookup- based convolutional neural network,
H. Bagherinezhad, M. Rastegari, and A. Farhadi, “Lcnn: Lookup- based convolutional neural network,” 2017. [Online]. Available: https://arxiv.org/abs/1611.06473
2017 arXiv
-
[98]
End-to-end anti-spoofing with rawnet2,
H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, “End-to-end anti-spoofing with rawnet2,” 2021. [Online]. Available: https://arxiv.org/abs/2011.01108
2021 arXiv
-
[99]
Deep residual learning for image recognition,
Z. X. R. S. He, K. and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 2015, pp. 770–778
2015
-
[100]
Detecting synthetic audio with resnet architecture,
Z. Li et al. , “Detecting synthetic audio with resnet architecture,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020. [Online]. Available: https: //ieeexplore.ieee.org
2020
-
[101]
Squeeze-and-excitation networks,
J. Hu, L. Shen, S. Albanie, G. Sun, and E. Wu, “Squeeze-and-excitation networks,” 2019. [Online]. Available: https://arxiv.org/abs/1709.01507
2019 arXiv
-
[102]
Improving the robustness of deepfake audio detection through confidence calibra- tion
Y . Zhang, J. Lu, Z. Li, Z. Shang, W. Wang, and P. Zhang, “Improving the robustness of deepfake audio detection through confidence calibra- tion.” in DADA@ IJCAI, 2023, pp. 70–75
2023
-
[103]
Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks,
J. weon Jung, H.-S. Heo, H. Tak, H. jin Shim, J. S. Chung, B.-J. Lee, H.-J. Yu, and N. Evans, “Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks,” 2021. [Online]. Available: https://arxiv.org/abs/2110.01200
2021 arXiv
-
[105]
Deepfake video detection using convolutional vision transformer,
D. Wodajo and S. Atnafu, “Deepfake video detection using convolutional vision transformer,” 2021. [Online]. Available: https: //arxiv.org/abs/2102.11126
2021 arXiv
-
[107]
The regression analysis of binary sequences,
D. R. Cox, “The regression analysis of binary sequences,” Journal of the Royal Statistical Society: Series B (Methodological) , vol. 20, no. 2, pp. 215–232, 1958
1958
-
[108]
F. J. O. R. Breiman, L. and C. Stone, Classification and Regression Trees. Wadsworth & Brooks, 1986
1986
-
[109]
Support vector machines,
M. Hearst, S. Dumais, E. Osuna, J. Platt, and B. Scholkopf, “Support vector machines,” IEEE Intelligent Systems and their Applications , vol. 13, no. 4, pp. 18–28, 1998
1998
-
[111]
Long short-term memory,
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997
1997
-
[112]
Res2net: A new multi-scale backbone ar- chitecture,
L. W. S. X. Gao, X. et al., “Res2net: A new multi-scale backbone ar- chitecture,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, 2019, pp. 1–9
2019
-
[113]
Very deep convolutional networks for large-scale image recognition,
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Proceedings of the International Conference on Neural Information Processing Systems (NeurIPS) . MIT Press, 2014, pp. 1–9
2014
-
[115]
End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detection,
H. Tak, J. weon Jung, J. Patino, M. Kamble, M. Todisco, and N. Evans, “End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detection,” in Proc. 2021 Edition of the Automatic Speaker Verification and Spoofing Counterme...
2021
-
[116]
Deep speech: Scaling up end-to-end speech recognition,
A. Hannun et al. , “Deep speech: Scaling up end-to-end speech recognition,” Proceedings of the International Conference on Machine Learning (ICML) , pp. 1–12, 2014. [Online]. Available: http: //arxiv.org/abs/1412.5567
2014 arXiv
-
[117]
Does audio deepfake detection generalize?
N. M. Müller, P. Czempin, F. Dieckmann, A. Froghyar, and K. Böt- tinger, “Does audio deepfake detection generalize?” arXiv preprint arXiv:2203.16263, 2022
2022
-
[118]
What to remember: Self-adaptive continual learning for audio deepfake detec- tion,
X. Zhang, J. Yi, C. Wang, C. Y . Zhang, S. Zeng, and J. Tao, “What to remember: Self-adaptive continual learning for audio deepfake detec- tion,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 17, 2024, pp. 19 569–19 577
2024
-
[119]
Abc-capsnet: Attention based cascaded capsule network for audio deepfake detection,
T. M. Wani, R. Gulzar, and I. Amerini, “Abc-capsnet: Attention based cascaded capsule network for audio deepfake detection,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 2464–2472
2024
-
[120]
A hybrid cnn- lstm approach for deepfake audio detection,
M. Chitale, A. Dhawale, M. Dubey, and S. Ghane, “A hybrid cnn- lstm approach for deepfake audio detection,” in 2024 3rd International Conference on Artificial Intelligence For Internet of Things (AIIoT) . IEEE, 2024, pp. 1–6
2024
-
[121]
Deepfake audio detection by speaker verification,
A. Pianese, D. Cozzolino, G. Poggi, and L. Verdoliva, “Deepfake audio detection by speaker verification,” 12 2022, pp. 1–6
2022
-
[122]
Bts-e: Audio deepfake detection using breathing-talking-silence encoder,
T.-P. Doan, L. Nguyen-Vu, S. Jung, and K. Hong, “Bts-e: Audio deepfake detection using breathing-talking-silence encoder,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
-
[123]
Domain generalization via aggregation and separation for audio deepfake detection,
Y . Xie, H. Cheng, Y . Wang, and L. Ye, “Domain generalization via aggregation and separation for audio deepfake detection,” IEEE Transactions on Information Forensics and Security , 2023
2023
-
[124]
Slim: Style-linguistics mismatch model for generalized audio deepfake detection,
Y . Zhu, S. Koppisetti, T. Tran, and G. Bharaj, “Slim: Style-linguistics mismatch model for generalized audio deepfake detection,” arXiv preprint arXiv:2407.18517, 2024
2024 arXiv
-
[125]
Audio deepfake detection based on a combination of f0 information and real plus imaginary spectrogram features,
J. Xue, C. Fan, Z. Lv, J. Tao, J. Yi, C. Zheng, Z. Wen, M. Yuan, and S. Shao, “Audio deepfake detection based on a combination of f0 information and real plus imaginary spectrogram features,” in Proceedings of the 1st international workshop on deepfake detection for audio mult...
2022
-
[126]
Harmonet: Partial deepfake detection network based on multi-scale harmof0 feature fusion,
L. Liu, H. Wei, D. Liu, and Z. Fu, “Harmonet: Partial deepfake detection network based on multi-scale harmof0 feature fusion,” in Proc. Interspeech 2024 , 2024, pp. 2255–2259
2024
-
[127]
A comparative study on physical and perceptual features for deepfake audio detection,
M. Li, Y . Ahmadiadli, and X.-P. Zhang, “A comparative study on physical and perceptual features for deepfake audio detection,” in Proceedings of the 1st International Workshop on Deepfake Detection for Audio Multimedia , 2022, pp. 35–41
2022
-
[128]
A robust audio deepfake detection system via multi-view feature,
Y . Yang, H. Qin, H. Zhou, C. Wang, T. Guo, K. Han, and Y . Wang, “A robust audio deepfake detection system via multi-view feature,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 13 131– 13 135
2024
-
[129]
Exploring green ai for audio deepfake detection,
S. Saha, M. Sahidullah, and S. Das, “Exploring green ai for audio deepfake detection,” arXiv preprint arXiv:2403.14290 , 2024
2024 arXiv
-
[130]
Deepfake audio detection: a deep learning based JOURNAL OF LATEX CLASS FILES 13 solution for group conversations,
R. Wijethunga, D. Matheesha, A. Al Noman, K. De Silva, M. Tissera, and L. Rupasinghe, “Deepfake audio detection: a deep learning based JOURNAL OF LATEX CLASS FILES 13 solution for group conversations,” in 2020 2nd International conference on advancements in computing (ICAC) , ...
2020
-
[131]
Convolutional sequence to sequence learning,
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y . N. Dauphin, “Convolutional sequence to sequence learning,” 2017. [Online]. Available: https://arxiv.org/abs/1705.03122
2017 arXiv
-
[132]
Going deeper with convolutions,
C. Szegedy, W. Liu, Y . Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V . Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, 2015, pp. 1–9
2015
-
[133]
Conformer: Convolution-augmented transformer for speech recognition,
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y . Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y . Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” 2020. [Online]. Available: https://arxiv.org/abs/2005.08100
2020 arXiv
-
[134]
Fighting ai with ai: fake speech detection using deep learning,
H. Malik, “Fighting ai with ai: fake speech detection using deep learning,” in 2019 AES INTERNATIONAL CONFERENCE ON AUDIO FORENSICS (June 2019) , 2019
2019
-
[135]
Audio-deepfake detection: Adversarial attacks and countermeasures,
M. Rabhi, S. Bakiras, and R. Di Pietro, “Audio-deepfake detection: Adversarial attacks and countermeasures,” Expert Systems with Appli- cations, vol. 250, p. 123941, 2024
2024
-
[136]
Does audio deepfake detection rely on artifacts?
T.-H. Shih, C.-Y . Yeh, and M.-S. Chen, “Does audio deepfake detection rely on artifacts?” in ICASSP 2024-2024 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 12 446–12 450
2024
-
[137]
Parameter- efficient transfer learning of audio spectrogram transformers,
U. Cappellazzo, D. Falavigna, A. Brutti, and M. Ravanelli, “Parameter- efficient transfer learning of audio spectrogram transformers,” 2024. [Online]. Available: https://arxiv.org/abs/2312.03694
2024 arXiv
-
[138]
Gated multimodal units for information fusion,
J. Arevalo, T. Solorio, M. M. y Gómez, and F. A. González, “Gated multimodal units for information fusion,” 2017. [Online]. Available: https://arxiv.org/abs/1702.01992
2017 arXiv
-
[139]
Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing,
J. Ao, R. Wang, L. Zhou, C. Wang, S. Ren, Y . Wu, S. Liu, T. Ko, Q. Li, Y . Zhang, Z. Wei, Y . Qian, J. Li, and F. Wei, “Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing,” 2022. [Online]. Available: https://arxiv.org/abs/2110.07205
2022 arXiv
-
[140]
St-bert: Cross- modal language model pre-training for end-to-end spoken language understanding,
M. Kim, G. Kim, S.-W. Lee, and J.-W. Ha, “St-bert: Cross- modal language model pre-training for end-to-end spoken language understanding,” 2021. [Online]. Available: https://arxiv.org/abs/2010. 12283
2021
-
[141]
Multimodal transformer for unaligned multimodal language sequences,
Y .-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal transformer for unaligned multimodal language sequences,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Fl...
2019
-
[142]
Speechbert: An audio-and-text jointly learned language model for end-to- end spoken question answering,
Y .-S. Chuang, C.-L. Liu, H.-Y . Lee, and L. shan Lee, “Speechbert: An audio-and-text jointly learned language model for end-to- end spoken question answering,” 2020. [Online]. Available: https: //arxiv.org/abs/1910.11559
2020 arXiv
-
[143]
A multimodal music emotion classification method based on multifeature combined network classifier,
C. Chen and Q. Li, “A multimodal music emotion classification method based on multifeature combined network classifier,” Mathematical Problems in Engineering , vol. 2020, no. 1, p. 4606027, 2020
2020
-
[144]
Multimodal music mood classification frame- work for kokborok music,
S. Das and S. Satpathy, “Multimodal music mood classification frame- work for kokborok music,” Solid State Technology, vol. 63, no. 6, pp. 9209–9219, 2020
2020
-
[145]
MMD- MII Model: A Multilayered Analysis and Multimodal Integration Interaction Approach Revolutionizing Music Emotion Classification,
J. Wang, A. Sharifi, T. R. Gadekallu, and A. Shankar, “MMD- MII Model: A Multilayered Analysis and Multimodal Integration Interaction Approach Revolutionizing Music Emotion Classification,” International Journal of Computational Intelligence Systems , vol. 17, p. 99, 2024. [On...
2024
-
[146]
Multimodal metric learning for tag-based music retrieval,
M. Won, S. Oramas, O. Nieto, F. Gouyon, and X. Serra, “Multimodal metric learning for tag-based music retrieval,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP). IEEE, 2021, pp. 591–595
2021
-
[147]
A novel multimodal music genre classifier using hierarchical attention and convolutional neural network,
M. Agrawal and A. Nandy, “A novel multimodal music genre classifier using hierarchical attention and convolutional neural network,” 2020. [Online]. Available: https://arxiv.org/abs/2011.11970
2020 arXiv
-
[148]
Muchomusic: Evaluating music understanding in multimodal audio-language models,
B. Weck, I. Manco, E. Benetos, E. Quinton, G. Fazekas, and D. Bog- danov, “Muchomusic: Evaluating music understanding in multimodal audio-language models,” arXiv preprint arXiv:2408.01337 , 2024
2024 arXiv
-
[149]
On the development and practice of ai technology for contemporary popular music production
E. Deruty, M. Grachten, S. Lattner, J. Nistal, and C. Aouameur, “On the development and practice of ai technology for contemporary popular music production.” Transactions of the International Society for Music Information Retrieval, vol. 5, no. 1, pp. 35–50, 2022
2022
-
[150]
Largest report on ai music reveals potentially dev- astating impact on australian and new zealand music creators,
APRA AMCOS, “Largest report on ai music reveals potentially dev- astating impact on australian and new zealand music creators,” 2024
2024
-
[151]
Generative ai and the future of music: Quality and authenticity concerns,
T. Atlantic, “Generative ai and the future of music: Quality and authenticity concerns,” 2024, accessed: 2024-11-13. [Online]. Available: https://www.theatlantic.com/technology/archive/ 2024/07/generative-ai-music-suno-udio/679114/
2024
-
[152]
Detecting machine-generated text: An arms race with the advancements of large language models,
U. of Pennsylvania School of Engineering and A. Science, “Detecting machine-generated text: An arms race with the advancements of large language models,” 2023
2023
-
[153]
Deepbach: a steerable model for bach chorales generation,
G. Hadjeres, F. Pachet, and F. Nielsen, “Deepbach: a steerable model for bach chorales generation,” 2017. [Online]. Available: https://arxiv.org/abs/1612.01010
2017 arXiv
-
[154]
Music inter- ventions for improving psychological and physical outcomes in people with cancer,
J. Bradt, C. Dileo, K. Myers-Coffman, and J. Biondo, “Music inter- ventions for improving psychological and physical outcomes in people with cancer,” Cochrane Database of Systematic Reviews , no. 10, 2021
2021
-
[155]
The promise of personalisation: Exploring how music streaming platforms are shaping the performance of class identities and distinction,
J. Webster, “The promise of personalisation: Exploring how music streaming platforms are shaping the performance of class identities and distinction,” New Media & Society , vol. 25, no. 8, pp. 2140–2162, 2023
2023
-
[2019]
Available: http://hdl.handle.net/10230/42015
[Online]. Available: http://hdl.handle.net/10230/42015
-
[2023]
Available: https://arxiv.org/abs/2301.11325
[Online]. Available: https://arxiv.org/abs/2301.11325
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.