Pith. sign in

REVIEW 2 minor 12 references

Pretrained music embeddings outperform from-scratch models on cross-performance jazz standard recognition but remain sensitive to performer identity.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-07-02 06:11 UTC pith:ZOMXJZMZ

load-bearing objection Pretrained embeddings beat from-scratch models on jazz standard retrieval but still pick up performer identity, with a simple contrastive fix helping only partially.

arxiv 2607.00777 v1 pith:ZOMXJZMZ submitted 2026-07-01 cs.SD cs.LG

Evaluating Pretrained Music Embeddings for Cross-Performance Jazz Standard Recognition

classification cs.SD cs.LG
keywords music information retrievaljazz standardspretrained embeddingscross-performance recognitioncontrastive projectionaudio embeddingsmusic representation learning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper evaluates how well pretrained music embeddings handle recognizing the same jazz standard across different performances that vary in tempo, key, and improvisation. It shows that models trained from scratch on spectrograms overfit to the specific training performances, leading to poor generalization. In contrast, frozen pretrained embeddings achieve better retrieval results but are still influenced by the performer's identity. Applying a lightweight contrastive projection helps reduce this sensitivity to some degree. The work positions this recognition task as a challenging test for music representation models aiming to support retrieval-based identification of standards.

Core claim

From-scratch spectrogram models overfit strongly to training performances, while pretrained embeddings provide better top-k results but are sensitive to performer identity, which can be partially reduced with a lightweight contrastive projection. The findings motivate jazz standard recognition as a useful stress test for music representation models and as a step toward retrieval-based standard identification.

What carries the argument

The lightweight contrastive projection applied to frozen pretrained embeddings to reduce performer identity sensitivity during nearest-neighbor retrieval.

Load-bearing premise

The curated subset of the Jazz Trio Database adequately captures the full range of cross-performance variation without introducing dataset-specific biases that would not generalize to other jazz recordings.

What would settle it

If pretrained embeddings no longer outperform from-scratch models or the contrastive projection fails to reduce performer effects when tested on an independent set of jazz recordings with unseen performers and arrangements, the central claims would not hold.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Pretrained embeddings are more suitable than from-scratch training for cross-performance music retrieval tasks.
  • Performer identity sensitivity in embeddings can be addressed with simple projection layers.
  • Jazz standard recognition provides a practical benchmark for assessing generalization in music foundation models.
  • Retrieval methods can advance toward identifying standards without needing full supervision.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Similar performer biases may affect embedding performance in other variable music genres such as classical or folk.
  • Training future models with explicit cross-performance contrastive losses could further improve invariance to performer and arrangement.
  • The evaluation approach might extend to identifying covers or alternate versions in non-jazz music retrieval.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 2 minor

Summary. The paper evaluates pretrained music embeddings for cross-performance jazz standard recognition on a curated subset of the Jazz Trio Database. It compares a from-scratch Harmonic CNN baseline against frozen pretrained representations from music foundation models via supervised probing and nearest-neighbor retrieval, concluding that from-scratch models overfit to training performances while pretrained embeddings yield better top-k results but remain sensitive to performer identity (partially mitigated by a lightweight contrastive projection). The work positions the task as a stress test for music representations.

Significance. If the empirical results hold under scrutiny of the methods and data splits, the findings would usefully demonstrate the generalization advantages of pretrained embeddings over task-specific training on limited jazz data and provide a practical technique (contrastive projection) for reducing performer sensitivity. This contributes an application-driven benchmark that could help evaluate future music foundation models on real-world variation in tempo, key, arrangement, and improvisation.

minor comments (2)
  1. The abstract references specific models and a 'curated subset' but does not name the exact pretrained embeddings, the size of the subset, or the train/test split criteria; these details are needed to assess reproducibility and the strength of the overfitting claim.
  2. No quantitative results (e.g., top-k accuracies, statistical significance) appear in the provided abstract, making it impossible to evaluate the magnitude of the reported improvements or the effectiveness of the contrastive projection.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for their careful reading and for recognizing the potential value of jazz standard recognition as a stress test for music representations. We address the major comments below.

Circularity Check

0 steps flagged

No significant circularity

full rationale

This is an empirical comparison study that trains or probes models on a curated dataset and reports observed top-k retrieval metrics. No equations, derivations, or first-principles claims appear; results are presented as experimental observations rather than predictions derived from the model itself. The central claim (pretrained embeddings outperform from-scratch models but remain performer-sensitive) rests on direct measurement, not on any self-definitional mapping, fitted-input renaming, or self-citation chain that reduces the result to its inputs by construction.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Only the abstract is available; no methods, equations, or data details are provided to identify free parameters, axioms, or invented entities.

pith-pipeline@v0.9.1-grok · 5692 in / 1063 out tokens · 17253 ms · 2026-07-02T06:11:52.831669+00:00 · methodology

0 comments
read the original abstract

Recognizing jazz standards from audio is a challenging form of tune-level music retrieval: different performances of the same standard may vary in tempo, key, arrangement, instrumentation, improvisational content, and even whether the head melody is present. We study this problem using a curated subset of the Jazz Trio Database designed for cross-performance standard recognition. We compare a from-scratch trained Harmonic CNN baseline against frozen pretrained music representations from recent music understanding foundation models, using both supervised probing and nearest-neighbor retrieval. Our results suggest that from-scratch spectrogram models overfit strongly to training performances, while pretrained embeddings provide better top-$k$ results but are sensitive to performer identity, which can be partially reduced with a lightweight contrastive projection. Our findings motivate jazz standard recognition as a useful stress test for music representation models and as a step toward retrieval-based standard identification. Project page: https://github.com/cagries/tipofmyear.

Figures

Figures reproduced from arXiv: 2607.00777 by \c{C}a\u{g}r{\i} Eser.

Figure 1
Figure 1. Figure 1: Distribution of performances across the curated JTD subset. Windowing strategy. Each recording is converted to 24 kHz mono audio and split into 10-second windows with 5-second hop. Each window inherits the standard label of its parent performance. This creates many training examples per performance, but these windows are highly correlated within a recording. Therefore, we report performance-level metrics i… view at source ↗
Figure 2
Figure 2. Figure 2: Pipeline for the proposed standard-aware supervised contrastive retrieval approach. Frozen MERT/MuQ embeddings are projected into a retrieval space trained to pull together windows from the same standard across different performances while reducing the same-performer retrieval bias. Top-1 classification of individual 10-second windows. Per￾formance Top-1 accuracy aggregates predictions over all windows fro… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

12 extracted references · 12 canonical work pages · 1 internal anchor

  1. [1]

    Bertin-Mahieux, Thierry and Ellis, Daniel P. W. and Whitman, Brian and Lamere, Paul , booktitle =. The

  2. [2]

    Proceedings of the 4th International Society for Music Information Retrieval Conference , year =

    An Industrial-Strength Audio Search Algorithm , author =. Proceedings of the 4th International Society for Music Information Retrieval Conference , year =

  3. [3]

    Characterization and exploitation of community structure in cover song networks

    Characterization and Exploitation of Community Structure in Cover Song Networks , author =. arXiv preprint arXiv:1108.6003 , year =

  4. [4]

    2023 , url =

    Xun, Jiahao and Zhang, Shengyu and Yang, Yanting and Zhu, Jieming and Deng, Liqun and Zhao, Zhou and Dong, Zhenhua and Li, Ruiqi and Zhang, Lichao and Wu, Fei , journal =. 2023 , url =

  5. [5]

    Transactions of the International Society for Music Information Retrieval , volume =

    Jazz Trio Database: Automated Annotation of Jazz Piano Trio Recordings Processed Using Audio Source Separation , author =. Transactions of the International Society for Music Information Retrieval , volume =. 2024 , doi =

  6. [6]

    and Liu, Ruibo and Chen, Wenhu and Xia, Gus and Shi, Yemin and Huang, Wenhao and Wang, Zili and Guo, Yike and Fu, Jie , journal =

    Li, Yizhi and Yuan, Ruibin and Zhang, Ge and Ma, Yinghao and Chen, Xingran and Yin, Hanzhi and Xiao, Chenghao and Lin, Chenghua and Ragni, Anton and Benetos, Emmanouil and Gyenge, Norbert and Dannenberg, Roger B. and Liu, Ruibo and Chen, Wenhu and Xia, Gus and Shi, Yemin and Huang, Wenhao and Wang, Zili and Guo, Yike and Fu, Jie , journal =. 2024 , url =

  7. [7]

    2025 , url =

    Zhu, Haina and Zhou, Yizhi and Chen, Hangting and Yu, Jianwei and Ma, Ziyang and Gu, Rongzhi and Luo, Yi and Tan, Wei and Chen, Xie , journal =. 2025 , url =

  8. [8]

    Advances in Neural Information Processing Systems 33 , year =

    Supervised Contrastive Learning , author =. Advances in Neural Information Processing Systems 33 , year =

  9. [9]

    arXiv preprint arXiv:2506.17055 , year =

    Universal Music Representations? Evaluating Foundation Models on World Music Corpora , author =. arXiv preprint arXiv:2506.17055 , year =

  10. [10]

    arXiv preprint arXiv:2107.05677 , year =

    Codified Audio Language Modeling Learns Useful Representations for Music Information Retrieval , author =. arXiv preprint arXiv:2107.05677 , year =

  11. [11]

    arXiv preprint arXiv:2210.03799 , year =

    Supervised and Unsupervised Learning of Audio Representations for Music Understanding , author =. arXiv preprint arXiv:2210.03799 , year =

  12. [12]

    Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =

    Data-Driven Harmonic Filters for Audio Representation Learning , author =. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , pages =