REVIEW 2 major objections 5 minor 54 references
IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups
T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read IdolSongsJp is a new, freely distributable corpus of 15 commissioned idol-style songs with exact stems, dry vocals, and expert chord annotations.
desk verdict A genuinely useful corpus that ships with a license clause contradicting its own advertised MSS training use—worth refereeing, but the license needs to be fixed before anyone can build on it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the corpus's multi-layer production design: every song exists as individual instrument stems, linear stem sums, stems processed with mastering effects, and fully mastered versions at two loudness targets, so that ground-truth audio can be compared against realistic commercial-style mixes. The utawari structure, in which singers alternate solo lines and sing together in sections, is the organizing musical feature that makes the corpus suited to multi-singer tasks such as singer diarization. Expert chord annotations in Harte's shorthand supply symbolic ground truth, and the 414 dry vocal tracks provide clean sources for singing voice synthesis and vocal processing.
What would settle it
A quantitative distributional comparison between the 15 corpus tracks and the 4,483 real idol tracks (e.g., a two-sample test on CLAP embeddings or on loudness and instrument-count statistics) that shows the corpus tracks come from a different distribution, or a listener study in which participants do not identify the commissioned tracks as stylistically similar to idol songs, would falsify the realistic-resource claim.
Extended reading notes
Core claim
The central claim is that IdolSongsJp provides a realistic, distributable benchmark for music information processing on multi-singer idol-style songs. The 15 songs were created by professional or semi-professional musicians, with each song assigned a unique combination of singers among 10 female and 8 male vocalists, designed with utawari song-division structures, varied genres and lyrical themes, and mastered to a loudness of -7 LUFS (with a -9 LUFS alternative). The corpus bundles linear stem sums, stems for four-category source separation, 414 individually recorded dry vocal tracks, unmastered master-bus signals, off-vocal and minus-one versions, 95 solo-song mixes, and chord annotations in Harte's shorthand agreed by at least two expert annotators. Application experiments show the corpus is demanding for current systems: HT Demucs separation accuracy drops on mastered tracks compared with unmastered sums, chord estimation exceeds 80 percent on roots and major/minor quality but falls below 60 percent on tetrads and below 30 percent on four-note MIREX chords, and lyrics transcription with Whisper is more robust than with a HuBERT-Conformer system. These results support the claim that the corpus can serve as a challenging resource for general and song-specific music information processing tasks.
Load-bearing premise
The load-bearing premise is that the commissioned songs genuinely resemble real Japanese idol songs in the properties that matter — loudness, arrangement density, utawari structure, and stylistic coverage — which the paper supports only with a qualitative UMAP visualization and no statistical test.
Editorial extensions
If this is right
- HT Demucs source separation performs noticeably worse on the mastered -7 LUFS tracks than on unmastered stem sums, especially for drums and vocals, showing that mastering effects remain a real obstacle for separation systems.
- Chord estimation systems exceed 80 percent accuracy for roots and major/minor chords but stay below 60 percent for tetrads and below 30 percent for MIREX4 chords, so extended chord vocabulary is still an open problem.
- Whisper-based lyrics transcription keeps similar character error rates across mastered tracks, Demucs-separated vocals, and dry lead vocals, while a HuBERT-Conformer system degrades on mastered audio and produces different error patterns.
- Because the corpus includes utawari structures and dry vocals for each singer, it is directly reusable for singer diarization, multi-pitch detection, and singing voice synthesis, beyond the three tasks the paper evaluates.
Reading between the lines
- The paired -7 and -9 LUFS masters could be used as a controlled experiment to isolate how limiter gain alone changes separation and chord estimation performance, something the paper does not investigate.
- Adding a quantitative distributional test to the UMAP visualization would strengthen the claim that the corpus matches real idol-song style; the current evidence is purely qualitative.
- The license structure, which permits research use but forbids sampling instrumental stems to train generative models, offers a model for sharing music data while protecting sample-library rights.
- The unusually poor separation result on the UK Garage track suggests genre-specific bass and sound design may need specialized models, pointing toward genre-conditioned source separation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces IdolSongsJp, a corpus of 15 commissioned songs in the style of Japanese idol groups, produced by professional creators. The corpus provides mastered audio at −7 and −9 LUFS, instrument stems, four-stem versions for music source separation, 414 dry vocal tracks, 95 solo versions, off-vocal/minus-one tracks, and expert chord annotations. The authors report a qualitative embedding-space comparison with 4,483 real idol-group tracks and present benchmark evaluations for music source separation, automatic chord estimation, and automatic lyrics transcription. The central claim is that the corpus is a realistic, distributable resource for tasks such as source separation, singer diarization, and chord estimation under commercial-like loudness and arrangement conditions.
Significance. If the licensing contradictions are resolved, IdolSongsJp could be a valuable addition to MIR resources: it is, to my knowledge, a rare corpus with exact stems, dry vocal tracks, and expert chord annotations for multi-singer songs at commercial loudness, and the construction details (loudness targets, mixing paths, consensus annotations) are concrete and checkable. The authors are also transparent about the limitations of their MSS evaluation (nonlinear mastering effects) and about the fact that their system evaluations use pretrained models rather than training on the corpus. The main weaknesses are that the advertised usability for MSS training is currently undercut by the license text, and the realism/diversity evidence is only qualitative.
major comments (2)
- [Section 2, final paragraph; Section 9] There is an internal contradiction in the license. The text states that 'Users may rearrange, parody, and apply machine learning techniques to the corpus,' but immediately adds that 'sampling the instrumental tracks to create unrelated content or to train machine learning models is prohibited.' Because Section 1 motivates the corpus by noting that supervised MSS 'requires a corpus consisting of ground-truth stems,' and Section 2 lists 'Stems for MSS' as a core data type, this prohibition removes the central advertised use of the instrumental stems. The phrase 'apply machine learning techniques' is unqualified and appears to allow what the next sentence forbids. Please clarify which data types (vocal vs instrumental; stems vs mastered tracks) may be used for training, and align the Hugging Face license with the text; the current wording makes the resource unusable for the main task it is designed to support.
- [Section 3, Figure 2] The claim that the corpus songs are 'broadly distributed across the embedding space' and include both distinctive and typical idol-style songs is based solely on visual inspection of a UMAP plot. The comparison against 4,483 real-world tracks lacks any quantitative support, such as a coverage statistic, nearest-neighbor distances between corpus and real tracks, or a statistical test of distribution overlap. Since the conclusion that the corpus is a 'realistic resource' for real-world conditions depends in part on this comparison, please add a quantitative diversity/coverage analysis or explicitly weaken the claim to avoid overstating what Figure 2 demonstrates.
minor comments (5)
- [Section 4, conditions 2 and 3] In the MSS evaluation, the reference stems for the mastered conditions are produced by applying nonlinear mastering effects to individual stems, and the paper acknowledges that simply summing the processed stems does not reproduce the final mastered tracks. This means the SDR values in Figure 3 are computed against approximate references whose relation to the true sources is not quantified; please state whether the approximation affects all stems equally or report the reconstruction error of the mastered mix from the processed stems.
- [Section 5, Figure 4] The statement that the proportion of major and minor chords is 'significantly lower' than in the McGill Billboard corpus would benefit from a statistical test or at least a confidence interval; as written, the comparison is qualitative and the denominator (after excluding non-chorded sections and tensions) is not fully specified.
- [Section 2, bullet list] The terms 'calls and mixes' and 'utawari' are used without a brief definition for non-Japanese readers; a one-sentence explanation in Section 2 would improve accessibility.
- [First-page footnote vs Section 2] The first-page footnote states that the paper is 'Licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0),' while Section 2 says the corpus is 'available free of charge for non-commercial research and entertainment purposes.' Please clarify which license applies to the paper and which to the corpus, because CC BY 4.0 permits commercial use and is at odds with the corpus's non-commercial restriction.
- [Table 1, column 'No. of stems'] The stem counts in Table 1 (ranging from 8 to 17) refer to the instrument stems, whereas the MSS data use four categories; please make this distinction explicit in the table caption or in Section 2 to avoid confusion.
Circularity Check
No circularity found: the corpus is externally constructed and evaluated with off-the-shelf models, so the reported results do not reduce to their inputs.
full rationale
The paper's central claim is the construction of a new corpus (IdolSongsJp) by commissioning professional creators, plus descriptive evaluations with existing pretrained or open-source systems. There is no fitted parameter that is later renamed as a prediction, and no equation in the paper derives a target quantity from a quantity that was defined in terms of it. The loudness targets, stem definitions, and chord annotations are corpus-construction choices, not predicted outputs. The application sections evaluate external models (HT Demucs, Whisper, HuBERT+Conformer, and two chord estimators) on the new corpus; these models were not trained on IdolSongsJp, so the evaluation scores cannot feed back into corpus construction. Self-citations to the authors' earlier singer-diarization work and FruitsMusic corpus appear only as motivation and context in Section 1, not as load-bearing evidence for the present corpus's properties. The style-diversity comparison in Section 3 is a qualitative UMAP visualization and is not a derivation; its limitation is weak statistical support, not circularity. The license clause restricting ML training on instrumental stems, while it conflicts with some advertised MSS uses, is a consistency concern about usability, not a circular-reasoning issue. Therefore, no significant circularity is present.
Assumptions & free parameters
assumptions (4)
- domain assumption Commissioned professional composers and singers produced tracks representative of Japanese idol group style.
- domain assumption Fine-tuned Hybrid Transformer Demucs removes vocals accurately enough that accompaniment-only embeddings represent musical style.
- domain assumption The CLAP/UMAP embedding space of accompaniment signals is a valid proxy for musical style diversity.
- domain assumption A target loudness near -7 LUFS matches commercial idol group songs.
Cite this review
Pith. "Pith review of IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups." pith.science (2026). https://pith.science/paper/WR2C56TL
@misc{pith2026250701349,
author = {Pith},
title = {Pith review of: IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups},
year = {2026},
howpublished = {\url{https://pith.science/paper/WR2C56TL}},
note = {Machine review of arXiv:2507.01349}
}
read the original abstract
Japanese idol groups, comprising performers known as "idols," are an indispensable part of Japanese pop culture. They frequently appear in live concerts and television programs, entertaining audiences with their singing and dancing. Similar to other J-pop songs, idol group music covers a wide range of styles, with various types of chord progressions and instrumental arrangements. These tracks often feature numerous instruments and employ complex mastering techniques, resulting in high signal loudness. Additionally, most songs include a song division (utawari) structure, in which members alternate between singing solos and performing together. Hence, these songs are well-suited for benchmarking various music information processing techniques such as singer diarization, music source separation, and automatic chord estimation under challenging conditions. Focusing on these characteristics, we constructed a song corpus titled IdolSongsJp by commissioning professional composers to create 15 tracks in the style of Japanese idol groups. This corpus includes not only mastered audio tracks but also stems for music source separation, dry vocal tracks, and chord annotations. This paper provides a detailed description of the corpus, demonstrates its diversity through comparisons with real-world idol group songs, and presents its application in evaluating several music information processing techniques.
Reference graph
Works this paper leans on
-
[1]
IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups
INTRODUCTION Processing music audio signals is one of the major direc- tions in the music information retrieval (MIR) field [1, 2]. The techniques include beat tracking [3, 4], fundamental frequency estimation [5, 6], automatic musical chord esti- mation [7–9], music source separation (MSS) [10,11], and automatic lyrics transcription (ALT) [12, 13]. Train...
work page Pith review arXiv 2025
-
[2]
IDOLSONGSJP CORPUS: A MULTI-SINGER CORPUS IN THE JAPANESE IDOL GROUP STYLE We constructed a novel corpus titled IdolSongsJp, com- prising 15 multi-singer songs in the style of Japanese idol groups. This corpus features 10 female and 8 male singers, all of whom are professionals or semi- professionals, and each of the 15 songs features a unique combination...
-
[3]
COMPARISON WITH REAL-WORLD IDOL GROUP SONGS The IdolSongsJp corpus is designed to capture the diverse styles of idol group music. This section demonstrates the style diversity of the songs in our corpus by comparing their music embeddings to those derived from real-world idol group songs. We processed 4,483 publicly available preview tracks performed by 2...
-
[4]
APPLICATION 1: MUSIC SOURCE SEPARATION As mentioned in Section 2, the IdolSongsJp corpus in- cludes stem signals, which help the evaluation of MSS techniques. In this section, we evaluate the performance of the fine-tuned model of HT Demucs [11, 17], which sep- arates music signals into four stems: bass, drums, vocals, and other. We evaluated the performa...
-
[5]
The stems are summed linearly without any additional mastering processing
Linear summation of stems without mastering ef- fects. The stems are summed linearly without any additional mastering processing
-
[6]
Mastered tracks at −9 LUFS. For these tracks, the same mastering effects as those used in produc- ing the mastered tracks were applied to the ground- truth stems
-
[7]
Mastered tracks at −7 LUFS. The only differ- ence from condition 2 is the gain parameter applied to the final limiter. Since mastering effects include not only maximizers and limiters but also equalizers, stereo imaging plug-ins, and other processing plug-ins, the final mixed signals differ considerably from the raw stems. To address this discrep- ancy, w...
-
[8]
APPLICATION 2: AUTOMATIC CHORD ESTIMATION As mentioned in Section 2, the IdolSongsJp corpus con- tains musical chord annotations provided by expert an- notators. Figure 4 shows the occurrence rates of chord qualities, excluding chord roots, inversions, tensions, and non-chorded sections. The proportion of major and minor chords is 50%, which is significan...
Show all 54 references
-
[9]
thank you for your watching
APPLICATION 3: AUTOMATIC LYRICS TRANSCRIPTION As described in Sections 2 and 3, the songs in the Idol- SongsJp corpus cover a wide range of musical styles with a variety of genres, tempos, and lyrical themes. Therefore, the corpus serves as a benchmark dataset for various nat-...
-
[10]
CONCLUSION We constructed the IdolSongsJp corpus, a novel multi- singer song corpus in the style of Japanese idol groups. The corpus includes not only mastered tracks but also stems for evaluating MSS techniques, dry vocal tracks, solo versions for all song–singer pairs, and c...
-
[11]
R&D on Generative AI Foundation Models for the Physical Domain
ACKNOWLEDGMENTS This research was partially supported by the AIST policy- based budget project “R&D on Generative AI Foundation Models for the Physical Domain.” The authors would like to acknowledge Mr. Takizawa (AIST) for his support in the speech recognition evaluation
-
[12]
Instrument stems may be utilized to create new musical content through music generation techniques or sampling, potentially infringing on the rights of these sources
ETHICS STATEMENT This corpus includes instrument stems and dry vocal tracks, which may be used for music generation and related applications. Instrument stems may be utilized to create new musical content through music generation techniques or sampling, potentially infringing ...
-
[13]
Signal processing for music analysis,
M. Müller, D. P. W. Ellis, A. Klapuri, and G. Richard, “Signal processing for music analysis,” IEEE journal of selected topics in signal processing , vol. 5, no. 6, pp. 1088–1110, 2011
2011
-
[14]
Deep learning for audio signal pro- cessing,
H. Purwins, B. Li, T. Virtanen, J. Schluter, S.-Y . Chang, and T. Sainath, “Deep learning for audio signal pro- cessing,” IEEE journal of selected topics in signal pro- cessing, vol. 13, no. 2, pp. 206–219, 2019
2019
-
[15]
An audio-based real-time beat tracking sys- tem for music with or without drum-sounds,
M. Goto, “An audio-based real-time beat tracking sys- tem for music with or without drum-sounds,” Journal of new music research , vol. 30, no. 2, pp. 159–171, 2001
2001
-
[16]
Efficient tempo and beat tracking in audio recordings,
L. Jean, “Efficient tempo and beat tracking in audio recordings,” Journal of the Audio Engineering Society. Audio Engineering Society, vol. 51, pp. 226–233, 2003
2003
-
[17]
A robust predominant-F0 estimation method for real-time detection of melody and bass lines in CD recordings,
M. Goto, “A robust predominant-F0 estimation method for real-time detection of melody and bass lines in CD recordings,” in Proc. 2000 IEEE International Con- ference on Acoustics, Speech, and Signal Processing (ICASSP), vol. 2, 2000, pp. 757–760
2000
-
[18]
Melody extraction from polyphonic music signals: Approaches, applications, and challenges,
J. Salamon, E. Gomez, D. P. W. Ellis, and G. Richard, “Melody extraction from polyphonic music signals: Approaches, applications, and challenges,” IEEE Sig- nal Processing Magazine, vol. 31, no. 2, pp. 118–134, 2014
2014
-
[19]
Automatic chord recognition for music classification and retrieval,
H.-T. Cheng, Y .-H. Yang, Y .-C. Lin, I.-B. Liao, and H. H. Chen, “Automatic chord recognition for music classification and retrieval,” in Proc. 2008 IEEE Inter- national Conference on Multimedia and Expo , 2008, pp. 1505–1508
2008
-
[20]
Audio chord recognition with recurrent neural networks,
N. Boulanger-Lewandowski, Y . Bengio, and P. Vin- cent, “Audio chord recognition with recurrent neural networks,” in Proc. 14th International Society for Mu- sic Information Retrieval Conference (ISMIR 2013) , 2013
2013
-
[21]
A fully convolu- tional deep auditory model for musical chord recogni- tion,
F. Korzeniowski and G. Widmer, “A fully convolu- tional deep auditory model for musical chord recogni- tion,” in Proc. 2016 IEEE 26th International Workshop on Machine Learning for Signal Processing (MLSP) , 2016, pp. 1–6
2016
-
[22]
Monaural sound source separation by nonnegative matrix factorization with temporal conti- nuity and sparseness criteria,
T. Virtanen, “Monaural sound source separation by nonnegative matrix factorization with temporal conti- nuity and sparseness criteria,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 3, pp. 1066–1074, 2007
2007
-
[23]
Hybrid spectrogram and waveform source separation,
A. Défossez, “Hybrid spectrogram and waveform source separation,” in Proc. ISMIR 2021 Music Demix- ing Workshop, 2021, pp. 1–11
2021
-
[24]
Automatic recognition of lyrics in singing,
A. Mesaros and T. Virtanen, “Automatic recognition of lyrics in singing,”EURASIP Journal on Audio, Speech, and Music Processing , vol. 2010, no. 1, pp. 1–11, 2010
2010
-
[25]
Automatic lyrics tran- scription of polyphonic music with lyrics-chord multi- task learning,
X. Gao, C. Gupta, and H. Li, “Automatic lyrics tran- scription of polyphonic music with lyrics-chord multi- task learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 2280– 2294, 2022
2022
-
[26]
The MUSDB18 corpus for music separation,
Z. Rafii, A. Liutkus, F.-R. Stöter, S. I. Mimilakis, and R. Bittner, “The MUSDB18 corpus for music separation,” 2017. [Online]. Available: https://doi.org/ 10.5281/zenodo.1117372
2017 doi
-
[27]
MUSDB18-HQ - an uncompressed version of MUSDB18,
——, “MUSDB18-HQ - an uncompressed version of MUSDB18,” 2019. [Online]. Available: https: //doi.org/10.5281/zenodo.3338373
2019 doi
-
[28]
MoisesDB: A dataset for source separation beyond 4- stems,
I. Pereira, F. Araújo, F. Korzeniowski, and R. V ogl, “MoisesDB: A dataset for source separation beyond 4- stems,” in Proc. 24th International Conference on Mu- sic Information Retrieval Conference (ISMIR 2023) , 2023, pp. 619–626
2023
-
[29]
Hybrid trans- formers for music source separation,
S. Rouard, F. Massa, and A. Défossez, “Hybrid trans- formers for music source separation,” in Proc. 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5
2023
-
[30]
Towards robust music source separation on loud commercial music,
C.-B. Jeon and K. Lee, “Towards robust music source separation on loud commercial music,” in Proc. 23rd International Society for Music Information Retrieval Conference (ISMIR 2022), 2022
2022
-
[31]
Deep learning based source separation ap- plied to choir ensembles,
D. Petermann, P. Chandna, H. Cuesta, J. Bonada, and E. Gomez, “Deep learning based source separation ap- plied to choir ensembles,” in Proc. 21st International Society for Music Information Retrieval Conference (ISMIR 2020), 2020
2020
-
[32]
jaCappella Corpus: A Japanese a cappella vocal ensemble corpus,
T. Nakamura, S. Takamichi, N. Tanji, S. Fukayama, and H. Saruwatari, “jaCappella Corpus: A Japanese a cappella vocal ensemble corpus,” in Proc. 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5
2023
-
[33]
MedleyV ox: An evaluation dataset for multi- ple singing voices separation,
C.-B. Jeon, H. Moon, K. Choi, B. S. Chon, and K. Lee, “MedleyV ox: An evaluation dataset for multi- ple singing voices separation,” in Proc. 2023 IEEE In- ternational Conference on Acoustics, Speech and Sig- nal Processing (ICASSP), 2023, pp. 1–5
2023
-
[34]
Okada, Ed., Living with Idols (in Japanese)
Y . Okada, Ed., Living with Idols (in Japanese) . Pot Publishing, 2013
2013
-
[35]
Singer diarization for polyphonic music with unison singing,
H. Suda, D. Saito, S. Fukayama, T. Nakano, and M. Goto, “Singer diarization for polyphonic music with unison singing,” IEEE/ACM Transactions on Au- dio, Speech, and Language Processing , vol. 30, pp. 1531–1545, 2022
2022
-
[36]
FruitsMusic: A real-world corpus of Japanese idol-group songs,
H. Suda, S. Yoshida, T. Nakamura, S. Fukayama, and J. Ogata, “FruitsMusic: A real-world corpus of Japanese idol-group songs,” inProc. 25th International Society for Music Information Retrieval Conference (ISMIR 2024), 2024
2024
-
[37]
Singer diarization: Application to ethnomusicologi- cal recordings,
M. Thlithi, C. Barras, J. Pinquier, and T. Pellegrini, “Singer diarization: Application to ethnomusicologi- cal recordings,” in Proc. 5th International Workshop on F olk Music Analysis (FMA 2015) , 2015, pp. 124– 125
2015
-
[38]
Active music listening interfaces based on signal processing,
M. Goto, “Active music listening interfaces based on signal processing,” in Proc. 2007 IEEE International Conference on Acoustics, Speech and Signal Process- ing (ICASSP), vol. 4, 2007, pp. IV–1441–IV–1444
2007
-
[39]
RWC music database: Popular, classical and jazz mu- sic databases,
M. Goto, H. Hashiguchi, T. Nishimura, and R. Oka, “RWC music database: Popular, classical and jazz mu- sic databases,” in Proc. 3rd International Conference on Music Information Retrieval Conference (ISMIR 2002), 2002
2002
-
[40]
Japanese “idols
W. Xie, “Japanese “idols” in trans-cultural reception: the case of AKB48,” inThe Art of Reception, J. Bracker and A.-K. Hubrich, Eds., 2021, pp. 371–399
2021
-
[41]
Recommendation ITU-R BS.1770-3: Algorithms to measure audio programme loudness and true-peak audio level,
Radiocommunication Sector of International Telecom- munication Union (ITU-R), “Recommendation ITU-R BS.1770-3: Algorithms to measure audio programme loudness and true-peak audio level,” 2012
2012
-
[42]
Symbolic representation of musical chords: A pro- posed syntax for text annotations,
C. Harte, M. B. Sandler, S. A. Abdallah, and E. Gómez, “Symbolic representation of musical chords: A pro- posed syntax for text annotations,” in Proc. 6th Inter- national Conference on Music Information Retrieval (ISMIR 2005), 2005, pp. 66–71
2005
-
[43]
Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Y . Wu, K. Chen, T. Zhang, Y . Hui, M. Nezhurina, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,”arXiv [cs.SD], 2022
2022
-
[44]
UMAP: Uniform manifold approximation and projec- tion,
L. McInnes, J. Healy, N. Saul, and L. Großberger, “UMAP: Uniform manifold approximation and projec- tion,” Journal of open source software , vol. 3, no. 29, p. 861, 2018
2018
-
[45]
Separate what you de- scribe: Language-queried audio source separation,
X. Liu, H. Liu, Q. Kong, X. Mei, J. Zhao, Q. Huang, M. D. Plumbley, and W. Wang, “Separate what you de- scribe: Language-queried audio source separation,” in Proc. Interspeech 2022, 2022, pp. 1801–1805
2022
-
[46]
An expert ground truth set for audio chord recognition and music analysis,
J. A. Burgoyne, J. Wild, and I. Fujinaga, “An expert ground truth set for audio chord recognition and music analysis,” in Proc. 12th International Society for Music Information Retrieval Conference (ISMIR 2011), 2011
2011
-
[47]
Large- vocabulary chord transcription via chord structure decomposition,
J. Jiang, K. Chen, W. Li, and G. Xia, “Large- vocabulary chord transcription via chord structure decomposition,” in Proc. 20th International Society for Music Information Retrieval Conference (ISMIR 2019), 2019
2019
-
[48]
A bi-directional Transformer for musical chord recogni- tion,
J. Park, K. Choi, S. Jeon, D. Kim, and J. Park, “A bi-directional Transformer for musical chord recogni- tion,” in Proc. 20th International Society for Music In- formation Retrieval Conference (ISMIR 2019) , 2019
2019
-
[49]
Robust speech recog- nition via large-scale weak supervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recog- nition via large-scale weak supervision,” in Proc. 40th International Conference on Machine Learning (ICML’23), 2022, pp. 28 492–28 518
2022
-
[50]
Conformer: Convolution-augmented transformer for speech recognition,
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y . Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y . Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. Interspeech 2020 , 2020, pp. 5036–5040
2020
-
[51]
HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
-
[52]
Automatic speech recognition of Japanese dialects using large-scale self-supervised learning models (in Japanese),
D. Takizawa, T. Nakamura, H. Suda, and S. Fukayama, “Automatic speech recognition of Japanese dialects using large-scale self-supervised learning models (in Japanese),” in Proc. 2025 Spring Meeting of the Acous- tical Society of Japan, 2025
2025
-
[53]
Construction of a large-scale Japanese ASR corpus on TV recordings,
S. Ando and H. Fujihara, “Construction of a large-scale Japanese ASR corpus on TV recordings,” inProc. 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 6948– 6952
2021
-
[54]
Lyrics transcription for humans: A readability- aware benchmark,
O. Cífka, H. Schreiber, L. Miner, and F.-R. Stöter, “Lyrics transcription for humans: A readability- aware benchmark,” in Proc. 25th International Soci- ety for Music Information Retrieval Conference (IS- MIR 2024), 2024
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.