Pith. sign in

REVIEW 4 major objections 6 minor 80 references

CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning

T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Continual pre-training on 650 hours of Greek, Turkish, and Indian music adapts a Western-trained music foundation model to non-Western traditions, improving auto-tagging by 4.9% on average with almost no forgetting on Western benchmarks.

desk verdict Useful, honest adaptation study whose 'SOTA on all tasks' claim is not supported by its own Table 2. read the letter →

arxiv 2506.17818 v1 pith:GMVBVWP7 submitted 2025-06-21 cs.SD cs.AIcs.LGeess.AS

classification cs.SDcs.AIcs.LGeess.AS
keywords musicfoundationmodelscontinualpre-trainingcross-culturalrepresentationtaskarithmeticmodelmergingMERTnon-Westernauto-tagging
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that a music foundation model trained mostly on Western music can be adapted to other musical traditions without retraining from scratch and without losing its Western abilities. The authors continually pre-train MERT-v1-95M on a 650-hour mix of Greek, Turkish, and Indian classical and traditional music using a two-stage schedule with learning-rate re-warming. The resulting CultureMERT beats the original model on every non-Western auto-tagging task they test, with an average 4.9% improvement in ranking metrics (ROC-AUC and average precision), and loses only about 0.05% on Western benchmarks. They also show that merging single-culture adapted models with task arithmetic performs about as well, offering a training-free alternative.

What carries the argument

The engine is the MERT-95M encoder (a HuBERT-style CNN plus 12-layer transformer) trained with a masked-language-modeling objective whose targets come from two frozen teachers: EnCodec's eight 1024-codeword residual vector-quantization codebooks for acoustic tokens, and constant-Q transform spectrogram reconstruction for pitch and harmonic content. The paper's contribution is the two-stage continual pre-training recipe around that engine: Stage 1 freezes the transformer and trains only the CNN feature extractor and codeword embeddings on 100 hours (20% Western replay) with a warm-up to $\eta_{\max}=5\times10^{-4}$; Stage 2 unfreezes everything and trains on the full 650 hours with a gentler schedule peaking at $\eta_{\max}=5\times10^{-5}$. This staging is what lets a small effective batch size adapt a large model without the stability gap or catastrophic forgetting. Task arithmetic supplies the alternative mechanism: each single-culture adapted model defines a task vector $\tau_i=\theta_i-\theta_{\mathrm{pre}}$, and merging sums scaled task vectors $\lambda\tau_i$ onto the base weights.

What would settle it

Take an unseen non-Western tradition absent from the 650-hour mix (for example gamelan or Arabic maqam) and compare CultureMERT against MERT-v1 on its auto-tagging benchmark: if there is no improvement, the model adapts only to the four trained cultures, not to world music generally. Alternatively, retrain or fine-tune EnCodec on the same non-Western audio and repeat the two-stage continual pre-training: if the gains disappear or shrink sharply, the frozen tokenizer was the limiting component.

Watch

Extended reading notes

Core claim

The core claim is that staged continual pre-training, not architectural change, is enough to make a Western-biased music encoder useful across cultures. Using the same masked pseudo-label objective as MERT—predicting EnCodec acoustic tokens and reconstructing constant-Q transform (CQT) spectrograms—the authors first adapt the low-level feature extractor and codeword embeddings on 100 hours of multi-cultural audio with 20% Western replay, then unfreeze the transformer and train on all 650 hours. On Turkish makam, Hindustani, Carnatic, and Greek Lyra auto-tagging, the adapted model consistently outperforms the base MERT-v1 and also exceeds the previous state of the art on the non-Western tasks, while Western tagging performance stays essentially flat. A companion result is that task arithmetic—adding task vectors from separately adapted single-culture models back to the base with scaling $\lambda=0.2$—matches the multi-culturally trained model on non-Western tasks and does even better on Western benchmarks.

Load-bearing premise

The load-bearing premise is that the frozen audio tokenizer (EnCodec), trained mostly on Western music, still turns Greek, Turkish, and Indian audio into discrete pseudo-labels useful enough for the model to learn from; if those codebook symbols miss microtonal pitches or unfamiliar rhythms, the adaptation signal is weak or biased.

Editorial extensions

If this is right

  • Non-Western auto-tagging improves on average 4.9% in ROC-AUC and average precision over MERT-v1, while Western benchmarks drop only 0.05%.
  • A 50-hour low-resource tradition (Lyra) still benefits from adaptation, and only the multi-cultural models beat the base model on Lyra.
  • Cross-cultural transfer is asymmetric: Carnatic adaptation transfers most consistently, and Turkish-makam and Carnatic models transfer strongly to each other.
  • Task arithmetic with scaling factor 0.2 gives a training-free route to a multi-cultural model that also slightly improves Western benchmarks over the base model.
  • Token distribution similarity between cultures predicts positive transfer, giving a data-selection signal for future continual pre-training mixes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the token-similarity result holds, the same EnCodec token statistics could be used prospectively to choose which cultures to add to a pre-training mix before training, not just to explain transfer afterward.
  • The method's ceiling is likely set by the frozen Western-trained EnCodec tokenizer; fine-tuning or replacing that tokenizer on non-Western audio is a natural test of whether the 4.9% gain is a floor or a starting point.
  • The two-stage recipe should transfer to other music foundation models with masked-token objectives, including larger ones, but the hyperparameters and scaling factor would need re-tuning; the paper only demonstrates one model scale.
  • Because evaluation is limited to auto-tagging with a frozen probe, the gains may not carry over to generation, retrieval, or fine-tuned downstream tasks; that is an open extension, not something the paper establishes.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces CultureMERT-95M, a continual pre-training (CPT) adaptation of the MERT-v1-95M music foundation model to Greek, Turkish, and Indian music traditions, using 650 hours of multi-cultural data. The method is a two-stage CPT recipe with learning-rate re-warming/re-decaying and a Western replay fraction in Stage 1. The authors also construct CultureMERT-TA by merging single-culture adapted models with task arithmetic. Evaluation on six auto-tagging benchmarks (four non-Western, two Western) reports consistent gains over MERT-v1 on non-Western tasks, an average relative improvement of 4.9%, minimal forgetting on Western benchmarks, and a claim of surpassing prior state-of-the-art on all non-Western tasks. The paper also analyzes cross-cultural transfer and token-level similarity. The authors release both model checkpoints.

Significance. If the empirical results survive scrutiny, this is a useful contribution to the under-studied problem of cross-cultural music representation learning. The paper provides a concrete, computationally cheap adaptation recipe, a systematic comparison against the base MERT-v1 model on five seeds, and a comparison of continual pre-training against weight-space task arithmetic. The public release of the adapted models is a practical asset for the community. However, the headline claim of surpassing prior state-of-the-art on all non-Western tasks is contradicted by the paper's own Table 2 on Carnatic average precision, and several hyperparameters are selected using the same evaluation benchmarks that later appear in the main table. These issues currently weaken the central narrative, though they are fixable with a careful rewriting of the claims and an explicit separation of model selection from evaluation.

major comments (4)
  1. [§5, Table 2] The claim in the Abstract, Contribution 4, and Section 5 that the adapted models 'surpass previous state-of-the-art results on all evaluated non-Western music tagging tasks' is contradicted by the Carnatic AP row of Table 2. The cited prior SOTA is 43.9 [46], while CultureMERT obtains 43.1±0.22 and CultureMERT-TA obtains 43.3±0.13. With standard errors of roughly 0.1–0.2 points, these are not ties within noise; on this metric both models are below the cited SOTA. The claim must be restricted to ROC-AUC, or to a defined aggregate, or the evaluation protocol against [46] must be documented and justified if a difference is claimed.
  2. [§5.3, Figure 4 and §4.3, Table 1] The final models are selected using the same benchmarks on which they are subsequently evaluated. The task-arithmetic scaling factor λ=0.2 is chosen from Figure 4, which plots ROC-AUC across all six test tasks, and the two-stage CPT recipe (stage budgets, Western replay fraction, warm-up ratios) is selected using Turkish-makam and MTAT in Table 1, both of which later appear in Table 2. For these datasets, the comparison against MERT-v1 is in-sample. Please introduce a validation split for recipe selection, or explicitly report which results are model-selection results and which are independent hold-out evaluations.
  3. [Abstract and §5] The '4.9% average improvement' is not defined precisely. It appears to be the relative change in the overall all-dataset average in Table 2 (66.1 to 69.3), not an average of per-dataset relative improvements (which is about 5.4% across all twelve cells and about 8% across the non-Western cells). Please state the exact aggregation formula, and report the non-Western and Western improvements separately so that the reader can see how the headline number relates to the 'diverse non-Western music auto-tagging tasks' mentioned in the Abstract.
  4. [§6 and §3.1, Eq. (2)] The paper's own limitation that the frozen EnCodec tokenizer 'may be suboptimal for encoding culturally diverse musical languages' directly concerns the training targets in Eq. (2), because all continual pre-training optimizes prediction of these discrete tokens. The empirical gains mitigate this concern, but the manuscript would be stronger with a diagnostic showing that the tokenizer provides meaningful targets for the four non-Western traditions (for example, codebook utilization, token entropy, or reconstruction error on held-out non-Western audio). As written, the central adaptation mechanism rests on an assumption that is acknowledged but not tested.
minor comments (6)
  1. [§3.1, Eq. (3)] The notation c∈C_k is confusing: C is defined as the number of codewords (C=1024) and K as the number of codebooks. Please write c∈{1,...,C} for a codebook index and clarify that the sum in Eq. (3) runs over codewords in the k-th codebook.
  2. [Table 1] The column header 'Turkish-makam MTA T' contains a spacing typo; it should read 'Turkish-makam MTAT'.
  3. [§3.2] The phrase 'we strike to balance plasticity' should read 'we strike a balance between plasticity'.
  4. [Figure 1] The flame and snowflake symbols (🔥/❄) used to denote trainable and frozen components are not defined in the caption; please define them or use standard textual labels.
  5. [§4.2] The evaluation is said to follow the MARBLE protocol 'under constrained settings', but the constraints are not specified. Please state which MARBLE elements are adopted and which are modified, since these details affect comparability with prior results.
  6. [§2] The paper states that 30-second segments are extracted from the training split of each non-Western dataset, but it does not explicitly confirm that the evaluation split is disjoint from the continual pre-training segments. This should be stated for reproducibility.

Circularity Check

1 steps flagged · score 2.0 of 10

Central continual pre-training result is an independent empirical comparison; only the task-arithmetic scaling factor is selected on the evaluation set, and the SOTA claim is overstated as a correctness issue, not a circularity issue.

  1. fitted input called prediction [Section 5.3 'Task Arithmetic Scaling Factor' and Section 5, Table 2]
    "We systematically evaluate different values of a shared scaling factor λ∈{0.1, 0.2, 0.25, 0.3, 0.5, 0.75, 1.0}, applied uniformly across all task vectors... Finally, CultureMERT and CultureMERT-TA surpass previous state-of-the-art (SOTA) results on all non-Western music tagging tasks, with the best task arithmetic variant obtained using λ = 0.2 (see Figure 4)."

    The reported CultureMERT-TA row in Table 2 is not an out-of-sample prediction: λ was chosen as the best-performing value on the very six benchmark tasks (Turkish-makam, Hindustani, Carnatic, Lyra, MTAT, FMA-medium) that are then used to conclude that 'task arithmetic performs on par' and 'surpasses previous SOTA.' Selecting the maximum over seven λ values on the evaluation set makes the favorable TA comparison partly an artifact of test-set selection rather than a fixed-model result. The choice is consequential: with the selected λ=0.2, CultureMERT-TA still does not beat the cited Carnatic AP SOTA (43.3 vs 43.9 in Table 2), so the 'surpasses SOTA on all non-Western tasks' conclusion is not robust to the hyperparameter-selection procedure.

full rationale

The central claim—that continual pre-training on 650 hours of Greek, Turkish, and Indian music improves non-Western tagging over MERT-v1—is an empirical training comparison against an external base checkpoint (MERT-v1-95M). The reported 4.9% average relative improvement is computed directly from Table 2 and does not reduce to any fitted parameter or self-citation. The two-stage CPT strategy is ablated against single-stage variants in Table 1, so its contribution is not definitional. The previous-SOTA baseline cites the authors' own ISMIR 2023 paper [46], but the cited numbers are external benchmark results, and using the same splits is a legitimate evaluation choice. That becomes a correctness problem, not a circularity problem: Table 2 shows CultureMERT's Carnatic AP (43.1±0.22) and CultureMERT-TA's (43.3±0.13) below the cited SOTA of 43.9, so the abstract and Contribution 4 claim of surpassing SOTA 'across all evaluated non-Western music tagging tasks' is overstated. The only mild circular element is the task-arithmetic λ: Section 5.3 selects the scaling factor by ROC-AUC on the same six benchmark tasks, and Section 5 then reports the best-λ CultureMERT-TA as 'on par' and 'SOTA-surpassing.' That is test-set hyperparameter selection, statistically forcing an optimistic comparison, rather than an independent prediction. Section 5.2's token-similarity/transfer correlation is post hoc and non-load-bearing. Overall, derivation-level circularity is minimal; score 2.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim rests mainly on domain assumptions about the transferability of the MERT objective, the EnCodec tokenizer, and the validity of probing evaluation. These are standard in the SSL-for-audio literature but are not independently verified here. The task arithmetic result additionally rests on a scaling factor selected on the benchmark itself.

free parameters (2)
  • Task arithmetic scaling factor lambda = 0.2
    Selected by sweeping lambda in {0.1, 0.2, 0.25, 0.3, 0.5, 0.75, 1.0} on the same six evaluation benchmarks (Section 5.3, Figure 4), with no held-out validation set reported. The final CultureMERT-TA numbers are the best point of this sweep.
  • Two-stage CPT recipe (stage budgets, warmup ratios, Western replay fraction) = Stage 1: 100h (20% Music4All), 2,250 steps, 10% warmup; Stage 2: 650h, 14,625 steps, 1% warmup; single-culture stages…
    Chosen by hand and supported by a limited ablation on two datasets (Table 1). These choices directly determine the reported gains, and no grid search or sensitivity analysis is given.
assumptions (4)
  • domain assumption MERT's masked acoustic-token prediction plus CQT reconstruction objective remains an effective learning signal for non-Western music.
    The CPT reuses the exact MERT objective (Eq. 1) on culturally shifted audio. If the objective does not capture modal or microtonal structures, the adaptation signal is weak.
  • domain assumption The EnCodec tokenizer provides faithful discrete tokens for non-Western audio even though it was trained on Western music.
    All CPT targets come from EnCodec (Section 3.1). The paper itself flags this as a limitation in Section 6.
  • domain assumption A probing MLP on frozen representations is an adequate measure of representation quality for cross-cultural music understanding.
    All conclusions rest on the MARBLE probing protocol (Section 4.2), which does not test generation or fine-grained sequence modeling.
  • domain assumption The Lyra and CompMusic corpora, with their tag sets, are representative of the target traditions.
    The ethics statement (Section 7.1) admits potential limitations in the representativeness and coverage of the corpora used.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning." pith.science (2026). https://pith.science/paper/GMVBVWP7

@misc{pith2026250617818,
  author       = {Pith},
  title        = {Pith review of: CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GMVBVWP7}},
  note         = {Machine review of arXiv:2506.17818}
}
read the original abstract

Recent advances in music foundation models have improved audio representation learning, yet their effectiveness across diverse musical traditions remains limited. We introduce CultureMERT-95M, a multi-culturally adapted foundation model developed to enhance cross-cultural music representation learning and understanding. To achieve this, we propose a two-stage continual pre-training strategy that integrates learning rate re-warming and re-decaying, enabling stable adaptation even with limited computational resources. Training on a 650-hour multi-cultural data mix, comprising Greek, Turkish, and Indian music traditions, results in an average improvement of 4.9% in ROC-AUC and AP across diverse non-Western music auto-tagging tasks, surpassing prior state-of-the-art, with minimal forgetting on Western-centric benchmarks. We further investigate task arithmetic, an alternative approach to multi-cultural adaptation that merges single-culture adapted models in the weight space. Task arithmetic performs on par with our multi-culturally trained model on non-Western auto-tagging tasks and shows no regression on Western datasets. Cross-cultural evaluation reveals that single-culture models transfer with varying effectiveness across musical traditions, whereas the multi-culturally adapted model achieves the best overall performance. To support research on world music representation learning, we publicly release CultureMERT-95M and CultureMERT-TA-95M, fostering the development of more culturally aware music foundation models.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

80 extracted references · 60 canonical work pages

  1. [46]

    Model merging in llms, mllms, and be- yond: Methods, theories, applications and opportuni- ties,

    E. Yang, L. Shen, G. Guo, X. Wang, X. Cao, J. Zhang, and D. Tao, “Model merging in llms, mllms, and be- yond: Methods, theories, applications and opportuni- ties,” CoRR, vol. abs/2408.07666, 2024

  2. [1]

    CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning

    INTRODUCTION Foundation models have recently emerged in the music do- main [1–5], offering powerful general-purpose represen- 1 https://huggingface.co/ntua-slp/CultureMERT-95M 2 https://huggingface.co/ntua-slp/CultureMERT-TA-95M © A.-N. Kanatas, C. Papaioannou, and A. Potamianos. Li- censed under a Creative Commons Attribution 4.0 International License (C...

  3. [2]

    To the best of our knowledge, this is the first study to explore continual pre-training and task arith- metic for cross-cultural adaptation in MIR, demon- strating their effectiveness in music audio represen- tation learning

  4. [3]

    We propose a two-stage CPT strategy that sta- bilizes training, mitigates catastrophic forgetting, and facilitates effective adaptation under constrained computational resources

  5. [4]

    Our multi-cultural model, CultureMERT, outper- forms the original MERT-v1 by an average of 4.9% across ROC-AUC and AP on culturally diverse non- Western music tagging tasks, while exhibiting mini- mal forgetting on Western benchmarks

  6. [5]

    Our culturally adapted models surpass previous state-of-the-art results across all evaluated non- Western music tagging tasks

  7. [6]

    We analyze cross-cultural transferability, showing that single-culture adaptations exhibit varying de- grees of transfer across cultural domains. To support reproducibility and further research in cross- cultural music representation learning, we publicly release CultureMERT-95M, along with the task arithmetic vari- ant, CultureMERT-TA-95M

  8. [7]

    Western" 3 mu- sic. For

    DA TASETS For our experiments, we use a diverse set of music datasets spanning both Western and non-Western traditions. Specif- ically, we adopt the MagnaTagATune (MTAT) [39] and FMA-medium [40] datasets to represent "Western" 3 mu- sic. For "non-Western" traditions, we incorporate the Lyra corpus [41], featuring Greek traditional and folk music, along wi...

Show all 80 references
  1. [8]

    In this section, we first re- view the architecture and pre-training objective of MERT, and then present our CPT strategy for cultural adaptation

    METHOD The overall framework of our approach is illustrated in Fig- ure 1, which depicts the two-stage continual pre-training strategy for CultureMERT. In this section, we first re- view the architecture and pre-training objective of MERT, and then present our CPT strategy for...

  2. [9]

    (4) 3.2 Two-Stage Continual Pre-Training Strategy To adapt the MERT foundation model to diverse musi- cal traditions, we employ continual pre-training, which extends the training of a pre-trained model on new data, aiming to adapt it to a shifted domain or task while re- taini...

  3. [10]

    Training was conducted using the FAIRSEQ 6 framework on a single NVIDIA GeForce GTX TITAN X GPU with 12 GB of memory

    EXPERIMENTS 4.1 Implementation Details In all continual pre-training setups, we initialize our mod- els from the publicly available MERT-v1-95M 5 pre- trained checkpoint. Training was conducted using the FAIRSEQ 6 framework on a single NVIDIA GeForce GTX TITAN X GPU with 12 GB...

  4. [11]

    RESULTS AND DISCUSSION As shown in Table 2, CultureMERT, adapted via multi- cultural continual pre-training, consistently outperforms the original MERT-v1 model across all non-Western tasks and evaluation metrics, achieving an average improvement of 4.9%. It also surpasses the...

  5. [12]

    We propose a two-stage CPT strategy that incor- porates learning rate re-warming and staged adaptation for stable training

    CONCLUSIONS In this paper, we introduceCultureMERT-95M, a multi- culturally adapted music foundation model developed via continual pre-training on diverse non-Western musical tra- ditions. We propose a two-stage CPT strategy that incor- porates learning rate re-warming and sta...

  6. [13]

    Western" versus

    ETHICS STA TEMENT 7.1 Cultural Framing and Interpretive Scope We acknowledge the limitations of framing music within a "Western" versus "non-Western" dichotomy. While such terminology is commonly used in computational research for convenience, it risks oversimplifying the rich...

  7. [14]

    We also grate- fully acknowledge the Music Technology Group (MTG) at Universitat Pompeu Fabra for providing access to datasets used in this study

    ACKNOWLEDGMENTS We would like to thank the reviewers and Georgios Paraskevopoulos for their valuable and constructive feed- back, which helped us improve this work. We also grate- fully acknowledge the Music Technology Group (MTG) at Universitat Pompeu Fabra for providing acce...

  8. [15]

    MERT: acoustic music understanding model with large-scale self-supervised training,

    Y . Li, R. Yuan, G. Zhang, Y . Ma, X. Chen, H. Yin, C. Xiao, C. Lin, A. Ragni, E. Benetos et al., “MERT: acoustic music understanding model with large-scale self-supervised training,” in The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austr...

  9. [16]

    A foundation model for music informatics,

    M. Won, Y . Hung, and D. Le, “A foundation model for music informatics,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2024, Seoul, Republic of Korea, April 14-19, 2024 . IEEE, 2024, pp. 1226–1230

  10. [17]

    Jukebox: A generative model for music,

    P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever, “Jukebox: A generative model for music,” CoRR, vol. abs/2005.00341, 2020

  11. [18]

    Foundation models for music: A survey,

    Y . Ma, A. Øland, A. Ragni, B. M. D. Sette, C. Saitis, C. Donahue, C. Lin, C. Plachouras, E. Benetos, E. Quinton et al. , “Foundation models for music: A survey,”CoRR, vol. abs/2408.14340, 2024

  12. [19]

    A survey of foundation models for music understanding,

    W. Li, Y . Cai, Z. Wu, W. Zhang, Y . Chen, R. Qi, M. Dong, P. Chen, X. Dong, F. Shi et al. , “A survey of foundation models for music understanding,”CoRR, vol. abs/2409.09601, 2024

  13. [20]

    Com- putational Ethnomusicology: perspectives and chal- lenges,

    E. Gómez, P. Herrera, and F. Gómez-Martin, “Com- putational Ethnomusicology: perspectives and chal- lenges,” Journal of New Music Research, vol. 42, no. 2, June 2013, pp. 111–112

  14. [21]

    Music for all: Exploring multicultural representations in music generation mod- els,

    A. Mehta, S. Chauhan, A. Djanibekov, A. Kulkarni, G. Xia, and M. Choudhury, “Music for all: Exploring multicultural representations in music generation mod- els,” CoRR, vol. abs/2502.07328, 2025

  15. [22]

    On the suitabil- ity of state-of-the-art music information retrieval meth- ods for analyzing, categorizing and accessing non- western and ethnic music collections,

    T. Lidy, C. N. S. Jr., O. Cornelis, F. Gouyon, A. Rauber, C. A. A. Kaestner, and A. L. Koerich, “On the suitabil- ity of state-of-the-art music information retrieval meth- ods for analyzing, categorizing and accessing non- western and ethnic music collections,”Signal Process.,...

  16. [23]

    Repertoire-specific vocal pitch data gener- ation for improved melodic analysis of carnatic music,

    G. Plaja-Roglans, T. Nuttall, L. Pearson, X. Serra, and M. Miron, “Repertoire-specific vocal pitch data gener- ation for improved melodic analysis of carnatic music,” Trans. Int. Soc. Music. Inf. Retr ., vol. 6, no. 1, 2023, pp. 13–26

  17. [24]

    Com- putational approaches for the understanding of melody in carnatic music,

    G. K. Koduri, M. Miron, J. Serrà, and X. Serra, “Com- putational approaches for the understanding of melody in carnatic music,” in Proceedings of the 12th Interna- tional Society for Music Information Retrieval Confer- ence, ISMIR 2011, Miami, Florida, USA, October 24- 28, 201...

  18. [25]

    Measur- ing commonality in recommendation of cultural con- tent to strengthen cultural citizenship,

    A. Ferraro, G. Ferreira, F. Diaz, and G. Born, “Measur- ing commonality in recommendation of cultural con- tent to strengthen cultural citizenship,”Trans. Recomm. Syst., vol. 2, no. 1, 2024, pp. 10:1–10:32

  19. [26]

    Culturally aware and adapted NLP: A taxonomy and a survey of the state of the art,

    C. C. Liu, I. Gurevych, and A. Korhonen, “Culturally aware and adapted NLP: A taxonomy and a survey of the state of the art,” CoRR, vol. abs/2406.03930, 2024

  20. [27]

    Eth- ical dimensions of music information retrieval technol- ogy,

    A. Holzapfel, B. L. Sturm, and M. Coeckelbergh, “Eth- ical dimensions of music information retrieval technol- ogy,” Trans. Int. Soc. Music. Inf. Retr . , vol. 1, no. 1, 2018, pp. 44–55

  21. [28]

    Simple and scalable strategies to continually pre-train large language models,

    A. Ibrahim, B. Thérien, K. Gupta, M. L. Richter, Q. G. Anthony, E. Belilovsky, T. Lesort, and I. Rish, “Simple and scalable strategies to continually pre-train large language models,” Trans. Mach. Learn. Res., vol. 2024, 2024

  22. [29]

    Tower: An open mul- tilingual large language model for translation-related tasks,

    D. M. Alves, J. Pombal, N. M. Guerreiro, P. H. Mar- tins, J. Alves, M. A. Farajian, B. Peters, R. Rei, P. Fer- nandes, S. Agrawal et al. , “Tower: An open mul- tilingual large language model for translation-related tasks,” CoRR, vol. abs/2402.17733, 2024

  23. [30]

    Meltemi: The first open large language model for greek,

    L. V oukoutis, D. Roussis, G. Paraskevopoulos, S. Sofi- anopoulos, P. Prokopidis, V . Papavasileiou, A. Kat- samanis, S. Piperidis, and V . Katsouros, “Meltemi: The first open large language model for greek,” CoRR, vol. abs/2407.20743, 2024

  24. [31]

    Con- tinual learning under language shift,

    E. Gogoulou, T. Lesort, M. Boman, and J. Nivre, “Con- tinual learning under language shift,” in Text, Speech, and Dialogue - 27th International Conference, TSD 2024, Brno, Czech Republic, September 9-13, 2024, Proceedings, Part I , ser. Lecture Notes in Computer Science, E. Nö...

  25. [32]

    Don’t stop pretraining: Adapt language models to domains and tasks,

    S. Gururangan, A. Marasovic, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith, “Don’t stop pretraining: Adapt language models to domains and tasks,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July ...

  26. [33]

    Reuse, don’t retrain: A recipe for con- tinued pretraining of language models,

    J. Parmar, S. Satheesh, M. Patwary, M. Shoeybi, and B. Catanzaro, “Reuse, don’t retrain: A recipe for con- tinued pretraining of language models,” CoRR, vol. abs/2407.07263, 2024

  27. [34]

    Continual pre-training for cross-lingual LLM adaptation: Enhancing japanese language capa- bilities,

    K. Fujii, T. Nakamura, M. Loem, H. Iida, M. Ohi, K. Hattori, H. Shota, S. Mizuki, R. Yokota, and N. Okazaki, “Continual pre-training for cross-lingual LLM adaptation: Enhancing japanese language capa- bilities,” CoRR, vol. abs/2404.17790, 2024

  28. [35]

    Continual pre-training of large language models: How to (re)warm your model?

    K. Gupta, B. Thérien, A. Ibrahim, M. L. Richter, Q. Anthony, E. Belilovsky, I. Rish, and T. Lesort, “Continual pre-training of large language models: How to (re)warm your model?” CoRR, vol. abs/2308.04014, 2023

  29. [36]

    Continual learning of large lan- guage models: A comprehensive survey,

    H. Shi, Z. Xu, H. Wang, W. Qin, W. Wang, Y . Wang, and H. Wang, “Continual learning of large lan- guage models: A comprehensive survey,” CoRR, vol. abs/2404.16789, 2024

  30. [37]

    A practitioner’s guide to real-world continual multimodal pretraining,

    V . Udandarao, K. Roth, S. Dziadzio, A. Prabhu, M. Cherti, O. Vinyals, O. J. Hénaff, S. Albanie, Z. Akata, and M. Bethge, “A practitioner’s guide to real-world continual multimodal pretraining,” in Ad- vances in Neural Information Processing Systems 38: Annual Conference on Ne...

  31. [38]

    Adapting multilingual speech repre- sentation model for a new, underresourced language through multilingual fine-tuning and continued pre- training,

    K. Nowakowski, M. Ptaszynski, K. Murasaki, and J. Nieuwazny, “Adapting multilingual speech repre- sentation model for a new, underresourced language through multilingual fine-tuning and continued pre- training,” Inf. Process. Manag. , vol. 60, no. 2, 2023, p. 103148

  32. [39]

    Breaking language barriers: Cross-lingual continual pre-training at scale,

    W. Zheng, W. Pan, X. Xu, L. Qin, L. Yue, and M. Zhou, “Breaking language barriers: Cross-lingual continual pre-training at scale,” in Proceedings of the 2024 Con- ference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y ....

  33. [40]

    Continual pre-training mitigates forgetting in language and vision,

    A. Cossu, A. Carta, L. C. Passaro, V . Lomonaco, T. Tuytelaars, and D. Bacciu, “Continual pre-training mitigates forgetting in language and vision,” Neural Networks, vol. 179, 2024, p. 106492

  34. [41]

    Improving low-resource speech recognition with pretrained speech models: Continued pretraining vs. semi-supervised training,

    M. DeHaven and J. Billa, “Improving low-resource speech recognition with pretrained speech models: Continued pretraining vs. semi-supervised training,” CoRR, vol. abs/2207.00659, 2022

  35. [42]

    Boosting cross-domain speech recogni- tion with self-supervision,

    H. Zhu, G. Cheng, J. Wang, W. Hou, P. Zhang, and Y . Yan, “Boosting cross-domain speech recogni- tion with self-supervision,” IEEE ACM Trans. Audio Speech Lang. Process., vol. 32, 2024, pp. 471–485

  36. [43]

    Predicting positive trans- fer for improved low-resource speech recognition us- ing acoustic pseudo-tokens,

    N. San, G. Paraskevopoulos, A. Arora, X. He, P. Kaur, O. Adams, and D. Jurafsky, “Predicting positive trans- fer for improved low-resource speech recognition us- ing acoustic pseudo-tokens,” in Proceedings of the 6th Workshop on Research in Computational Linguistic Ty- pology ...

  37. [44]

    Cpt-boosted wav2vec2.0: To- wards noise robust speech recognition for classroom environments,

    A. A. Attia, D. Demszky, T. Ògúnrèmí, J. Liu, and C. Y . Espy-Wilson, “Cpt-boosted wav2vec2.0: To- wards noise robust speech recognition for classroom environments,” CoRR, vol. abs/2409.14494, 2024

  38. [45]

    Model soups: averag- ing weights of multiple fine-tuned models improves ac- curacy without increasing inference time,

    M. Wortsman, G. Ilharco, S. Y . Gadre, R. Roelofs, R. G. Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y . Carmon, S. Kornblithet al., “Model soups: averag- ing weights of multiple fine-tuned models improves ac- curacy without increasing inference time,” in Interna- tional Con...

  39. [47]

    Revisit- ing weight averaging for model merging,

    J. Choi, D. Kim, C. Lee, and S. Hong, “Revisit- ing weight averaging for model merging,” CoRR, vol. abs/2412.12153, 2024

  40. [48]

    Patching open-vocabulary models by interpolating weights,

    G. Ilharco, M. Wortsman, S. Y . Gadre, S. Song, H. Ha- jishirzi, S. Kornblith, A. Farhadi, and L. Schmidt, “Patching open-vocabulary models by interpolating weights,” in Advances in Neural Information Process- ing Systems 35: Annual Conference on Neural Infor- mation Processin...

  41. [49]

    Zipit! merging models from different tasks without training,

    G. Stoica, D. Bolya, J. Bjorner, P. Ramesh, T. Hearn, and J. Hoffman, “Zipit! merging models from different tasks without training,” in The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024

  42. [50]

    Dataless knowledge fusion by merging weights of language models,

    X. Jin, X. Ren, D. Preotiuc-Pietro, and P. Cheng, “Dataless knowledge fusion by merging weights of language models,” in The Eleventh International Con- ference on Learning Representations, ICLR 2023, Ki- gali, Rwanda, May 1-5, 2023. OpenReview.net, 2023

  43. [51]

    Editing models with task arithmetic,

    G. Ilharco, M. T. Ribeiro, M. Wortsman, L. Schmidt, H. Hajishirzi, and A. Farhadi, “Editing models with task arithmetic,” in The Eleventh International Confer- ence on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net, 2023

  44. [52]

    Mertech: Instrument playing technique detection using self-supervised pre- trained model with multi-task finetuning,

    D. Li, Y . Ma, W. Wei, Q. Kong, Y . Wu, M. Che, F. Xia, E. Benetos, and W. Li, “Mertech: Instrument playing technique detection using self-supervised pre- trained model with multi-task finetuning,” in IEEE In- ternational Conference on Acoustics, Speech and Sig- nal Processing...

  45. [53]

    Evaluation of algorithms using games: The case of music tagging,

    E. Law, K. West, M. I. Mandel, M. Bay, and J. S. Downie, “Evaluation of algorithms using games: The case of music tagging,” in Proceedings of the 10th International Society for Music Information Retrieval Conference, ISMIR 2009, Kobe International Confer- ence Center , Kobe, J...

  46. [54]

    FMA: A dataset for music analysis,

    M. Defferrard, K. Benzi, P. Vandergheynst, and X. Bresson, “FMA: A dataset for music analysis,” in Proceedings of the 18th International Society for Music Information Retrieval Conference, ISMIR 2017, Suzhou, China, October 23-27, 2017 , S. J. Cunning- ham, Z. Duan, X. Hu, and...

  47. [55]

    A dataset for greek traditional and folk music: Lyra,

    C. Papaioannou, I. Valiantzas, T. Giannakopoulos, M. A. Kaliakatsos-Papakostas, and A. Potamianos, “A dataset for greek traditional and folk music: Lyra,” in Proceedings of the 23rd International Society for Music Information Retrieval Conference, ISMIR 2022, Bengaluru, India,...

  48. [56]

    Creating research corpora for the compu- tational study of music: the case of the compmusic project,

    X. Serra, “Creating research corpora for the compu- tational study of music: the case of the compmusic project,” in AES International Conference on Seman- tic Audio 2014, London, UK, January 27-29, 2014 , C. Dittmar, G. Fazekas, and S. Ewert, Eds. Audio Engineering Society, 2014

  49. [57]

    A corpus for computational research of turk- ish makam music,

    B. Uyar, H. S. Atli, S. Sentürk, B. Bozkurt, and X. Serra, “A corpus for computational research of turk- ish makam music,” in Proceedings of the 1st Inter- national Workshop on Digital Libraries for Musicol- ogy, DLfM@JCDL 2014, London, United Kingdom, September 12, 2014 , B. ...

  50. [58]

    Computational analysis of audio record- ings and music scores for the description and discovery of ottoman-turkish makam music,

    S. Sentürk, “Computational analysis of audio record- ings and music scores for the description and discovery of ottoman-turkish makam music,” Ph.D. dissertation, Pompeu Fabra University, Spain, 2017

  51. [59]

    Corpora for music information re- search in indian art music,

    A. Srinivasamurthy, G. K. Koduri, S. Gulati, V . Ish- war, and X. Serra, “Corpora for music information re- search in indian art music,” inMusic Technology meets Philosophy - From Digital Echos to Virtual Ethos: Joint Proceedings of the 40th International Computer Music Confer...

  52. [60]

    From west to east: Who can understand the mu- sic of the others better?

    C. Papaioannou, E. Benetos, and A. Potamianos, “From west to east: Who can understand the mu- sic of the others better?” in Proceedings of the 24th International Society for Music Information Retrieval Conference, ISMIR 2023, Milan, Italy, November 5-9, 2023, A. Sarti, F. Anto...

  53. [61]

    LC- Protonets: Multi-label few-shot learning for world mu- sic audio tagging,

    C. Papaioannou, E. Benetos, and A. Potamianos, “LC- Protonets: Multi-label few-shot learning for world mu- sic audio tagging,” IEEE Open Journal of Signal Pro- cessing, vol. 6, 2025, pp. 138–146

  54. [62]

    High fidelity neural audio compression,

    A. Défossez, J. Copet, G. Synnaeve, and Y . Adi, “High fidelity neural audio compression,” Trans. Mach. Learn. Res., vol. 2023, 2023

  55. [63]

    Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

    W. Hsu, B. Bolte, Y . H. Tsai, K. Lakhotia, R. Salakhut- dinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE ACM Trans. Audio Speech Lang. Process., vol. 29, 2021, pp. 3451–3460

  56. [64]

    Overcom- ing catastrophic forgetting in neural networks,

    J. Kirkpatrick, R. Pascanu, N. C. Rabinowitz, J. Ve- ness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska et al., “Overcom- ing catastrophic forgetting in neural networks,” CoRR, vol. abs/1612.00796, 2016

  57. [65]

    Minicpm: Unveiling the potential of small language models with scalable training strategies,

    S. Hu, Y . Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y . Fang, Y . Huang, W. Zhao et al. , “Minicpm: Unveiling the potential of small language models with scalable training strategies,” CoRR, vol. abs/2404.06395, 2024

  58. [66]

    Effi- cient continual pre-training by mitigating the stability gap,

    Y . Guo, J. Fu, H. Zhang, D. Zhao, and Y . Shen, “Effi- cient continual pre-training by mitigating the stability gap,” CoRR, vol. abs/2406.14833, 2024

  59. [67]

    Continual evaluation for lifelong learning: Identifying the stability gap,

    M. D. Lange, G. M. van de Ven, and T. Tuytelaars, “Continual evaluation for lifelong learning: Identifying the stability gap,” in The Eleventh International Con- ference on Learning Representations, ICLR 2023, Ki- gali, Rwanda, May 1-5, 2023. OpenReview.net, 2023

  60. [68]

    Music4all: A new music database and its applications,

    I. A. P. Santana, F. Pinhelli, J. Donini, L. G. Catharin, R. B. Mangolin, Y . M. e Gomes da Costa, V . D. Feltrim, and M. A. Domingues, “Music4all: A new music database and its applications,” in 2020 International Conference on Systems, Signals and Image Processing, IWSSIP 202...

  61. [69]

    Continual lifelong learning with neural networks: A review,

    G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter, “Continual lifelong learning with neural networks: A review,”Neural Networks, vol. 113, 2019, pp. 54–71

  62. [70]

    On the stability-plasticity dilemma of class-incremental learning,

    D. Kim and B. Han, “On the stability-plasticity dilemma of class-incremental learning,” in IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, CVPR 2023, V ancouver , BC, Canada, June 17-24,

  63. [72]

    On layer normalization in the transformer architecture,

    R. Xiong, Y . Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y . Lan, L. Wang, and T. Liu, “On layer normalization in the transformer architecture,” in Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Vir- tual Event, ser. ...

  64. [73]

    Codified au- dio language modeling learns useful representations for music information retrieval,

    R. Castellon, C. Donahue, and P. Liang, “Codified au- dio language modeling learns useful representations for music information retrieval,” in Proceedings of the 22nd International Society for Music Information Re- trieval Conference, ISMIR 2021, Online, November 7-12, 2021 , ...

  65. [74]

    MAR- BLE: music audio representation benchmark for uni- versal evaluation,

    R. Yuan, Y . Ma, Y . Li, G. Zhang, X. Chen, H. Yin, L. Zhuo, Y . Liu, J. Huang, Z. Tian et al. , “MAR- BLE: music audio representation benchmark for uni- versal evaluation,” in Advances in Neural Informa- tion Processing Systems 36: Annual Conference on Neural Information Proc...

  66. [75]

    Mulan: A joint embedding of music audio and natural language,

    Q. Huang, A. Jansen, J. Lee, R. Ganti, J. Y . Li, and D. P. W. Ellis, “Mulan: A joint embedding of music audio and natural language,” inProceedings of the 23rd International Society for Music Information Retrieval Conference, ISMIR 2022, Bengaluru, India, December 4-8, 2022 , ...

  67. [76]

    Decoupled weight de- cay regularization,

    I. Loshchilov and F. Hutter, “Decoupled weight de- cay regularization,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019

  68. [77]

    12 maqam system and its similarity with indian raga’s (according to the indian manuscripts),

    D. Karomat, “12 maqam system and its similarity with indian raga’s (according to the indian manuscripts),” Indian Musicological Society. Journal of the Indian Musicological Society, vol. 36, 2006, p. 62

  69. [78]

    An- alyzing the effect of linguistic similarity on cross- lingual transfer: Tasks and experimental setups mat- ter,

    V . Blaschke, M. Fedzechkina, and M. ter Hoeve, “An- alyzing the effect of linguistic similarity on cross- lingual transfer: Tasks and experimental setups mat- ter,” CoRR, vol. abs/2501.14491, 2025

  70. [79]

    Adamerging: Adaptive model merging for multi-task learning,

    E. Yang, Z. Wang, L. Shen, S. Liu, G. Guo, X. Wang, and D. Tao, “Adamerging: Adaptive model merging for multi-task learning,” in The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024

  71. [80]

    Investigating the potential of task arithmetic for cross-lingual trans- fer,

    M. Parovic, I. Vulic, and A. Korhonen, “Investigating the potential of task arithmetic for cross-lingual trans- fer,” in Proceedings of the 18th Conference of the Eu- ropean Chapter of the Association for Computational Linguistics, EACL 2024 - V olume 2: Short Papers, St. Juli...

  72. [2023]

    20 196–20 204

    IEEE, 2023, pp. 20 196–20 204

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.