{"id":"53b6bbd9-2c41-4391-bc56-809f6048dd03","arxiv_id":"2607.27909","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Kernel distances on CLaMP3/Aria embeddings score expressive MIDI performances about as well as human listeners and catch contextual corruptions invisible to attribute statistics.","lead":"New metrics compare expressive MIDI piano performances using whole-performance 'digital fingerprints' from the music models CLaMP3 and Aria instead of note-by-note statistics. In a 23-participant listening test the new scores tracked human naturalness and expression ratings about as well as traditional ones, and caught context-level distortions that older metrics miss.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Perceptual validation underpowered: MOS from ~5.6 ratings per performance yields wide, overlapping CIs; 'on par' and 'significant margin' claims in Table 3 are not statistically supported, so the core 'perceptual proxy' claim is unproven.","rationale":"The reader's weakest assumption is that chunk-level pooling washes out expressive detail. I considered this, but the perturbation experiments (Sec. 4.2) provide direct evidence that the embeddings retain some contextual sensitivity (velocity transfer, pauses), though not proof of perceptual relevance. The perceptual relevance is precisely what the listening study is supposed to establish, yet that study is statistically underpowered. A noisy MOS does not bias one metric against another, but it inflates the uncertainty to the point where the paper's comparative claims ('on par', 'significant margin') are unsupported. With ~5.6 ratings per performance, the MOS standard error is large; the reported CIs overlap for all substantive comparisons. Thus the core premise that these embeddings are perceptual proxies is not established at the reported confidence. This is a more fundamental issue than the pooling concern because even a perfect embedding cannot be validated against a noisy ground truth. The proposed check (bootstrapping the existing ratings) would directly settle whether the parity claim holds within statistical certainty. I therefore recommend keeping the CONDITIONAL verdict but emphasize that the perceptual-validity claim requires additional data or at least a statistically grounded equivalence test.","tokens_in":12195,"tokens_out":5950,"duration_ms":56946,"concrete_test":"Using the existing 810 ratings, perform a bootstrap resampling of raters (or performances) to compute a 95% confidence interval for the difference in Kendall tau-B between each deep-feature pseudo rating (KPD-CLaMP3, RMD-Aria) and the Aggregated Correlation baseline. Also compute ICC(2,k) for the MOS. If the difference CIs include zero, or ICC < 0.7, the claim that deep-feature metrics are 'on par' with attribute correlations is not statistically supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that SSL embeddings and KMD/KPD are perceptual proxies rests on the listening study (Sec. 4.3). With 23 participants and 810 ratings across 145 performances, each MOS averages ~5.6 ratings, producing a noisy ground truth. The paper does not report inter-rater reliability (e.g., ICC). The 95% CIs in Table 3 are wide and overlap for the key comparisons: KPD(CLaMP3) = 0.44±0.12 vs Aggregated Correlation = 0.51±0.09; RMD(Aria) = 0.51±0.13 vs Aggregated = 0.51±0.09; even 'CLaMP3 outperformed Aria by a significant margin' (KPD 0.44±0.12 vs 0.29±0.14) has overlapping CIs. Thus 'on par' and 'significant margin' are indistinguishable from 'unresolved by the data.' The model-level correlations in Table 4 are equally fragile: the deep metrics reverse the PianoFlow NFE ordering that humans prefer, which the authors explain as 'different levels of diversity' without quantitative support. If MOS is unreliable, the Kendall tau-B values are attenuated and unstable, and the claimed parity with attribute-scoped metrics could disappear with additional ratings. This is load-bearing because the perceptual-proxy claim is the foundation for recommending KMD/KPD as evaluation metrics; without reliable MOS, the only remaining evidence is synthetic perturbations, which show sensitivity, not perceptual validity.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the problem of objective evaluation of expressive MIDI piano performances. It argues that attribute-scoped metrics (correlation, KL divergence, reconstruction error) are limited because they treat individual expressive attributes in isolation and generally require note alignment. The authors propose two distributional, alignment-free metrics on top of self-supervised symbolic-music embeddings: Kernel Music Distance (KMD), an MMD-based distance between sets of performances, and Kernel Performance Distance (KPD), a per-score average of MMD. They also propose per-sample pseudo ratings based on Mahalanobis and Relative Mahalanobis distances in embedding space. Experiments show that KMD/KPD respond to synthetic corruptions (pauses, velocity transfer) that leave attribute correlations nearly unchanged, and a listening study with 23 participants is used to compare the pseudo ratings against human MOS. The paper releases an open-source library, Pereval. The central claims are that contextual embeddings can serve as perceptual proxies on par with traditional correlation-based metrics and that the proposed kernel metrics are alignment-free and context-aware.","tokens_in":12455,"tokens_out":3728,"duration_ms":39106,"significance":"If the claims hold, the paper makes a useful contribution: it provides a practical, reproducible evaluation toolbox for expressive MIDI performance, and it demonstrates a concrete failure mode of attribute-scoped metrics through the velocity-transfer experiment. The synthetic perturbation results in Table 2 are clean and compelling: inter-attribute correlations remain essentially unchanged while KMD/KPD values change substantially, showing that embedding-based distributional metrics detect contextual corruption that attribute statistics miss. The release of Pereval is a concrete community benefit. The main significance risk is the perceptual validation: the listening study is small, the MOS ground truth is noisy, and the headline 'on par with traditional metrics' claim rests on correlations with overlapping confidence intervals. The paper's contribution is therefore potentially valuable, but the perceptual-proxy claim is not yet established at the confidence level the text suggests.","major_comments":[{"comment":"The listening study is underpowered for the claims made. With 23 participants and 810 ratings across 145 performances, the per-performance MOS averages only about 5.6 ratings, and no inter-rater reliability (e.g., ICC) is reported. The 95% confidence intervals in Table 3 overlap for the key comparisons: KPD(CLaMP3) = 0.44±0.12 vs. Aggregated Correlation = 0.51±0.09; RMD(Aria) = 0.51±0.13 vs. Aggregated Correlation = 0.51±0.09; and KPD(CLaMP3) = 0.44±0.12 vs. KPD(Aria) = 0.29±0.14, which is nevertheless described as a 'significant margin' in Sec. 4.3. Overlap of marginal intervals is not a significance test, and with noisy MOS the Kendall tau-B values are attenuated. To support the 'perceptual proxy' claim, the authors should report the distribution of ratings per performance, ICC or a variance-component model, confidence intervals for the differences between methods, and ideally collect","section":"Sec. 4.3, Table 3"},{"comment":"The model-level comparison is internally inconsistent for the PianoFlow variants. Human naturalness ratings rank PianoFlow-2 (3.55) above PianoFlow-16 (3.24) above PianoFlow-128 (3.12), but the deep feature metrics KMD, KPD, and FMD rank PianoFlow-128 as best (e.g., KMD_CLaMP3 = 9.7 vs. 10.3 and 11.5). The text says this 'can be explained by different levels of diversity' but offers no quantitative support. Since the paper recommends these metrics for model selection, this reversal affects the central usefulness claim. The authors should either provide a concrete diversity measure that resolves the discrepancy or acknowledge that the metrics do not fully track human preference across sampling steps.","section":"Sec. 4.3, Table 4"},{"comment":"The alignment-free and context-aware properties rest on the fixed-length chunking plus pooling construction of the embeddings, but the validation of these properties is indirect. The synthetic perturbations (velocity transfer, pauses) are global transformations, and the note-shift experiment in Fig. 4 shows that the metrics are nearly insensitive to removing up to 20 notes from the beginning of each performance. This suggests that the pooled global representations may wash out finer-grained expressive structure that matters perceptually. The paper should state this limitation explicitly and, if possible, provide a perturbation that directly tests cross-chunk or phrase-level dependencies, e.g., swapping expressive timing profiles between phrase boundaries while preserving note-level marginals.","section":"Sec. 3, Sec. 4.2"}],"minor_comments":[{"comment":"Please report the exact number of ratings per performance (mean, min, max) and the procedure used to compute the 95% confidence intervals in Table 3 (bootstrap? per-score aggregation?). This is needed to interpret the '±' values.","section":"Sec. 4.3"},{"comment":"Typographical inconsistency: 'Frèchet' appears in the section heading; the standard spelling is 'Fréchet'.","section":"Sec. 3.1"},{"comment":"The table caption lists 'Kernel Perf. Distance' and 'Relative Mahalanobis', but the main text sometimes uses 'KPD' and 'RMD' without redefinition. Consider defining abbreviations in the caption.","section":"Sec. 4.3, Table 3"},{"comment":"The scatter plots are informative but the caption says 'RMD (Aria)' only in the second panel; the first panel is 'Aggregated Correlation pseudo ratings'. This is clear from the axis labels, but the caption could state both explicitly.","section":"Fig. 5"},{"comment":"The definition of inter-set correlation averages over all pairs, but it is unclear how missing notes handled by linear interpolation affect the pairing count. A brief remark would help.","section":"Sec. 2.2.2, Eq. (1)"}],"recommendation":"major_revision","confidential_remarks":"The core methodological idea is sound and the perturbation experiments are valuable. The main risk is that the listening study is too small to support the 'perceptual proxy' claim; if the authors can add data or substantially soften the claim, the paper would be acceptable. The PianoFlow ranking inconsistency should also be addressed. The paper fits ISMIR scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's real contribution is pragmatic: it adapts Kernel Audio Distance to symbolic MIDI embeddings (KMD), adds a per-score conditioned version (KPD), and releases a library called Pereval. That is a useful thing for the MIR community, and the perturbation experiments are the cleanest evidence. Velocity transfer and synthetic pauses leave attribute correlations essentially unchanged while moving KMD/KPD clearly, which nicely demonstrates that contextual embeddings catch what note-wise statistics miss. No circular fitting either: bandwidths are median heuristics, the alpha=100 rescale is arbitrary and fixed, and the covariance is estimated on the reference set. I believe that part of the paper. The systematic comparison of CLaMP3 and Aria is also valuable — CLaMP3 works out of the box, Aria needs the Mahalanobis post-processing, and that is a concrete, useful finding.\n\nWhere I would push back is the perceptual validation. Twenty-three participants and about 5.6 ratings per performance yield wide CIs, and in Table 3 the so-called 'significant margin' between KPD(CLaMP3) and KPD(Aria) is accompanied by overlapping confidence intervals. The same is true for the 'on par' claims. That does not mean the metrics are bad, but the abstract's first sentence — that these embeddings can be used as perceptual proxies — is stronger than the data support. The authors are honest enough to report CIs, but they then over-interpret them. The other soft spot is Table 4: deep metrics rank PianoFlow-128 best while humans prefer PianoFlow-2, and the 'different levels of diversity' explanation is hand-waved. If the paper claims these metrics coincide with human perception, a reversed model ranking is a problem worth real discussion, not a footnote.\n\nOne thing I would have liked is some check on whether global embeddings from fixed-length chunking plus pooling actually preserve the fine-grained expressive cues (rubato shape, articulation micro-structure) that the evaluation is supposed to measure. The perturbations are gross, and the listening study is indirect. That is a limitation, not a fatal flaw, but it should be acknowledged more explicitly.\n\nNet: the paper deserves peer review. It is a solid engineering contribution with a useful released library, and the problems are addressable — a larger or re-analyzed listening study, tempered claims, and more attention to the model-level discrepancy. I would send it to a serious referee and expect a revision. If I were working on expressive performance evaluation, I would cite it for the metrics, not for the perceptual claim.","headline":"Useful evaluation tooling for expressive MIDI, but the perceptual-proxy claim rests on a small listening study with overlapping CIs; the perturbation experiments are the strongest part.","tokens_in":13078,"tokens_out":1833,"would_cite":true,"duration_ms":21921,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Pretrained context-aware embeddings can rank expressive MIDI piano performances in line with human ratings, and kernel-based distances built on them detect contextual distortions that note-level attribute metrics miss—without requiring note","keywords":["expressive piano performance","MIDI evaluation","contextual embeddings","Kernel Music Distance","Kernel Performance Distance","maximum mean discrepancy","alignment-free metrics","human perception"],"falsifier":"A within-chunk perturbation that preserves each chunk's average embedding but changes expressive ordering or articulation (e.g., swapping velocities between adjacent notes in the same chunk); if KMD/KPD stay flat while listeners reliably hear the difference, the perceptual-proxy claim collapses.","tokens_in":11969,"feed_emoji":"🎹","tokens_out":4731,"duration_ms":46072,"temperature":0.7,"pith_summary":"This paper argues that the usual way of evaluating expressive MIDI piano performances—correlating note-level attributes like timing, velocity, and duration—misses how notes depend on each other. To fix that, it tests whether embeddings from self-supervised symbolic-music models can serve as perceptual proxies, and introduces two alignment-free distributional metrics, Kernel Music Distance and Kernel Performance Distance. In a listening study, the new metrics agree with human naturalness and expression ratings about as well as attribute correlations do, while also responding to contextual corruptions that attribute metrics ignore. If right, these deep-feature metrics give generative and rendering researchers a single scalar that captures both fidelity and diversity without requiring note-wise alignment.","feed_headline":"MIDI performance scores now match human ratings without note alignment","feed_subtitle":"Deep-feature kernel distances catch expressive shifts that per-attribute correlations miss—one scalar, no alignment.","key_machinery":"The central object is Maximum Mean Discrepancy (MMD) with a Gaussian kernel applied to global performance embeddings: Kernel Music Distance (KMD) is the rescaled squared MMD between two corpora, and Kernel Performance Distance (KPD) averages per-score MMD to account for score-performance dependence. The embeddings come from fixed-length chunking plus pooling, with CLaMP3 using BERT-like encoding followed by average pooling and Aria using the chunk's end-of-sequence token hidden state averaged across chunks. MMD's characteristic kernel ensures that two distributions are equal if and only if their mean embeddings coincide, giving an alignment-free, distributional comparison that captures both","core_discovery":"The paper's central claim is that chunk-level, pooled embeddings from pretrained symbolic-music models carry enough perceptual information to rank generated MIDI piano performances roughly as well as traditional attribute correlations. Concretely, Kernel Performance Distance with CLaMP3 embeddings reaches a Kendall tau-B of 0.44 for naturalness, within the 0.43–0.48 range of per-attribute Pearson correlations, while Aria embeddings reach 0.51 after Relative Mahalanobis post-processing. The authors also show that kernel-based metrics respond to contextual perturbations—such as transferring velocities from one performance to another—that leave attribute correlations nearly unchanged. They conc","pith_inferences":["If chunk-pooling is the bottleneck, replacing average pooling with sequence-aware or attention-pooled representations could push the metrics beyond the current ceiling of roughly 0.5 Kendall tau rather than merely matching attribute correlations.","The same KMD/KPD recipe should transfer to other instrument families and repertoires; the paper's listening study covers only Western classical solo piano, so a cross-genre replication with human ratings would test generality.","Because KPD is a distributional metric, it could serve as a training objective or early-stopping signal for generative performance models, not just as an evaluation score.","The per-sample pseudo-ratings (Mahalanobis and Relative Mahalanobis) point toward no-reference quality assessment of a single performance, a use case the paper only partially explores."],"forward_implications":["Kernel metrics on contextual embeddings can replace or complement attribute correlations when comparing human and generated MIDI performances, since they need no note alignment.","The metrics detect contextual corruptions, such as swapping velocities between interpretations, that leave per-attribute correlations unchanged, so they capture inter-note dependencies.","KPD's per-score averaging prevents popular pieces or unbalanced repertoires from dominating the evaluation.","CLaMP3 embeddings support reference-free evaluation: marginal Mahalanobis distances estimated on a training set still correlate with human ratings when the target piece is absent from the reference set.","Aria embeddings, weaker out of the box, reach human-level ranking after Mahalanobis post-processing, showing that embedding choice and post-processing matter."],"fun_headline_variants":["No alignment needed: deep embeddings match human MIDI ratings","Kernel distance on music embeddings rivals human perceptual scores","MIDI eval: contextual embeddings match human ratings without alignment","Deep-feature kernel judges expressive MIDI like humans, no alignment","Perceptual proxy: deep music embeddings score MIDI without alignment"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that averaged fixed-length chunk embeddings preserve the fine-grained expressive cues humans judge—and, secondarily, that mean human scores derived from roughly 5.6 ratings per performance are accurate enough to test that premise.","fun_headline_variants_meta":{"raw":{"variants":["No alignment needed: deep embeddings match human MIDI ratings","Kernel distance on music embeddings rivals human perceptual scores","MIDI eval: contextual embeddings match human ratings without alignment","Deep-feature kernel judges expressive MIDI like humans, no alignment","Perceptual proxy: deep music embeddings score MIDI without alignment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000645,"raw_usage":{"total_tokens":2780,"prompt_tokens":704,"completion_tokens":2076,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":448,"completion_tokens_details":{"reasoning_tokens":1993}},"tokens_in":448,"tokens_out":2076,"duration_ms":14513,"temperature":1.0,"reasoning_tokens":1993,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T23:12:23.337000+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A within-chunk perturbation that preserves each chunk's average embedding but changes expressive ordering or articulation (e.g., swapping velocities between adjacent notes in the same chunk); if KMD/KPD stay flat while listeners reliably hear the difference, the perceptual-proxy claim collapses.","supporting_citations":[],"review_version":1}