{"id":"1a0f7696-61e8-437a-93ad-74f305f20d2f","arxiv_id":"2412.11769","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The paper releases expert pairwise and label annotations for guitar timbre clips and reports cases where perceived brightness does not follow spectral centroid.","lead":"This paper introduces a dataset of 960 guitar clips rated by experienced players to map words like warm, bright, and chug onto actual guitar tones. It matters because it tests whether standard acoustic measures, such as spectral centroid for brightness, survive expert judgments, and it provides data for tone-aware music AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central counterexample to the spectral-centroid/brightness correlation may be an artifact of unnormalized loudness: the same-DI design controls pitch but not output level, and the paper provides no reliability check for the specific pairwise judgments.","rationale":"The reader's verdict is CONDITIONAL and identifies the same experimental-control weakness: applying different FX chains to identical DI clips does not hold loudness or dynamics constant, and no normalization is reported. My stress-test agrees that this is load-bearing, because the paper's main interpretive claim—that established spectral correlations fail here—depends on pairwise judgments that are neither loudness-matched nor replicated enough to rule out level artifacts. I add a second, related concern: reliability is not established for the specific contradictory pairs, since only six repeated annotations exist in the entire study and no significance testing is applied to the case study. The dataset contribution remains credible and useful: the code and data are released, the annotation interface is sensible, and the authors acknowledge small sample size in Appendix D. Therefore I do not recommend rejecting the paper; a CONDITIONAL verdict remains appropriate, with the condition being that the central counterexample be validated under loudness-matched, replicated listening tests. The proposed concrete test is inexpensive and directly targets the weakest point, so the reader's verdict should remain UNCHANGED rather than moving to ACCEPT or REJECT.","tokens_in":8049,"tokens_out":7299,"duration_ms":76261,"concrete_test":"Extract the two clips shown in the bottom row of Fig. 2 (and ideally all pairs whose brightness annotation contradicts the spectral-centroid ordering), measure their integrated loudness (EBU R128, LUFS), and compare directions: if the clip judged 'more bright' is consistently louder by more than about 2–3 LU, loudness is a plausible confound. Then re-run the same pairwise brightness questions with loudness-normalized versions (matching LUFS) using at least 10 expert listeners per pair and repeated judgments; if the contradiction is no longer significant under a two-sided binomial test at p >= 0.05, the paper's central counterexample is a level artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the conclusion that human brightness judgments contradict spectral centroid, the pairwise annotations must isolate timbre and be trustworthy. The paper's design (§3.2, §3.3.1) generates 960 clips by applying a wide range of Helix presets—including drive, EQ, and compression—to identical DI recordings. This controls pitch and musical content, but not output loudness or dynamics. Distortion/overdrive presets can be substantially louder and more compressed than clean presets; no loudness normalization or gain-matching is reported anywhere in §3 or Appendix C. The paper itself cites Melara and Marks (1990) for the claim that loudness affects timbre perception, so the 'counterexample' in Appendix C.2 (Fig. 2) could simply reflect the right clip being louder or more compressed, not brighter in a timbre-only sense. No statistical test is applied to the contradictory pair, and §4.3 reports only six repeated annotations in the entire study, so the specific comparison may be a single annotator's idiosyncratic response. Appendix D concedes the dataset is small and 'there is no objective grounding of these terms.' Thus the headline finding is not yet established; it is at best a hypothesis to test with loudness-controlled, replicated judgments.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a dataset of electric-guitar timbre descriptions, built by applying 80 Helix Native presets to 12 direct-input recordings and collecting pairwise comparisons and label annotations from expert crowdsourced guitarists. The authors propose a graph-based algorithm to unify pairwise and label annotations into per-clip-adjective scores, then use the data to revisit established correlations between acoustic features and timbral adjectives. Their headline finding is that human brightness judgments sometimes contradict spectral-centroid rankings, which they present as evidence that prevailing coarse acoustic correlates are insufficient for guitar tone semantics. The paper also releases code and data.","tokens_in":8256,"tokens_out":2594,"duration_ms":28056,"significance":"If the central claim is established, the paper delivers a reusable, carefully controlled dataset for guitar timbre semantics, with a sensible design choice (identical musical content across timbre variations) and a genuine expert-annotator pool. The graph-based unification and the honest limitations section are useful contributions, and the paper correctly names its strengths: release of dataset/code, demonstration of a crowdsourcing pipeline for niche expert communities, and a concrete case study that can be revisited. However, the headline 'contradicts prevailing theories' currently rests on a small number of anecdotal pairwise examples, with no statistical support and a potential loudness confound, so the significance as stated is not yet earned.","major_comments":[{"comment":"Section 3.3.1 states that using the same DI controls 'pitch and loudness' as confounds, but the FX chains used to generate the 960 clips include distortion, EQ, and compression, which substantially alter output level and dynamics. No loudness normalization or gain matching is reported anywhere in Section 3 or Appendix C. The bottom row of Figure 2, which is the paper's central counterexample to the spectral-centroid/brightness correlation, could therefore reflect level or compression differences rather than timbre-only brightness. The authors should either apply loudness normalization (e.g., RMS or LUFS matching) and re-run the analysis, or report per-pair output levels and show that the contradictory pair is not explained by level differences.","section":"§3.3.1 and Appendix C.2"},{"comment":"The paper reports only six repeated annotation questions in the entire study (§4.3), and the specific pairwise comparison that contradicts spectral centroid in Figure 2 appears to be based on a single annotator response. No significance test, confidence interval, or agreement measure is provided for this comparison. The claim that 'human assessments sometimes differ from previously established correlations' therefore rests on anecdotal evidence. The authors should provide replication across multiple annotators for the critical pairs, report inter-annotator agreement (e.g., proportion or kappa), and test whether the observed direction is statistically inconsistent with the spectral-centroid prediction.","section":"§4.3 and Appendix C.2"},{"comment":"The graph-based score unification in Section 3.5 introduces an arbitrary constant φ for each unit of adjective-preset correlation, and the resulting scores are released and used to produce Table 1 and the 'most relevant preset' claims. No external validation or sensitivity analysis for φ is provided, so the unified scores are not an empirically calibrated resource. The authors should either justify the choice of φ (e.g., by comparing against a hold-out set of pairwise judgments), report how Table 1 changes under reasonable variations of φ, or clearly mark these scores as provisional rather than validated.","section":"§3.5 and Table 1"}],"minor_comments":[{"comment":"The reference 'Mcadams' should be spelled 'McAdams' consistently (both in the running text and in the bibliography).","section":"References"},{"comment":"The caption of Figure 2 says 'the pairwise annotation is consistent with the spectral centroid ... whereas it is not consistent with the centroid in the bottom row' but does not state how many annotators contributed to each row; adding that information would make the anecdotal nature of the counterexample explicit.","section":"Appendix C.2"},{"comment":"The adjective table contains a line with a stray 'T o' that appears to be a typo for 'To' or a formatting artifact; it should be cleaned up for the camera-ready version.","section":"Table 2"},{"comment":"The paper says each DI segment is approximately 10 seconds long and processed with 80 presets to yield 960 samples, but does not clarify whether any audio was trimmed or faded; a short statement about clip boundaries would help reproducibility.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"The central contribution — a dataset of expert-annotated guitar timbre with controlled musical content — is potentially valuable and within scope for a data-oriented venue. The main risk is that the headline claim about contradicting spectral-centroid/brightness correlations is not yet supported by the evidence as presented. The required fixes (loudness normalization or level control, replicated pairwise judgments with statistics, and sensitivity analysis for φ) are feasible within the manuscript's scope, so I recommend major revision rather than rejection. I would also encourage the authors to frame the spectral-centroid finding as a hypothesis-generating observation unless the additional evidence supports a stronger claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dataset is the real product here: 960 clips from 12 DI recordings through 80 presets, with pairwise and label annotations from experienced guitarists, all released with code. That is a concrete, reusable resource for timbre semantics and prompt-based generation. The aligned-content design is a genuine improvement over cross-instrument or unaligned corpora, and the paper is honest about its limitations in Appendix D.\n\nThe soft spots are concentrated in the analysis section, and they matter for the headline. The claim that human brightness judgments contradict spectral centroid rests on a handful of pairwise examples, and the design does not control for loudness. Different presets—especially distortion and compression—change output level and dynamics, not just spectral shape. The paper itself cites Melara and Marks (1990) on loudness affecting timbre perception, so the Appendix C.2 counterexample could be an artifact of the right clip being louder or more compressed. No loudness normalization or gain-matching is reported. On top of that, only six repeated annotation questions exist in the whole study, so a single contradictory pair could be one annotator's idiosyncratic response. There are no significance tests or confidence intervals for the correlation analysis. The graph-based score unification in Section 3.5 uses an arbitrary phi with no external validation—workable for producing scores, but not a principled scaling.\n\nI do not think any of this sinks the dataset. The resource is still useful for retrieval, generation, and contrastive studies. But the paper's central conclusion—that human assessments sometimes differ from established acoustic correlations—is not yet established by the data as reported. It is a hypothesis that needs loudness-controlled stimuli, more repeated annotations, and some statistical grounding.\n\nThe citation pattern looks fine; the related work is appropriate. The writing is clear, and the limitations are openly stated. This deserves a serious referee: the dataset is worth reviewing, and the analysis can be fixed in revision. I would send it out with a request to address the loudness confound and quantify reliability. I would likely cite the dataset myself if I worked on timbre semantics.","headline":"A genuinely useful guitar-timbre dataset, but the headline counterexample to spectral-centroid brightness is undermined by unnormalized loudness and a thin reliability check.","tokens_in":8746,"tokens_out":2148,"would_cite":true,"duration_ms":20701,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Guitar 'brightness' resists spectral-centroid predictions","keywords":["timbre description","guitar tone","crowdsourced annotation","pairwise comparison","spectral centroid","brightness perception","audio effects","data-driven timbre"],"falsifier":"Re-run the brightness pairwise comparisons after loudness-normalizing the clips; if the cases where spectral centroid and perceived brightness disagree disappear after normalization, those cases are artifacts of level differences rather than genuine counterexamples to the centroid model.","tokens_in":7815,"feed_emoji":"🎸","tokens_out":6755,"duration_ms":55227,"temperature":0.7,"pith_summary":"This paper contends that the words guitarists use for tone—'warm,' 'chunky,' 'chuggy'—are not reliably predicted by the coarse acoustic features that timbre research usually relies on. To test this, the authors built a dataset of 960 electric-guitar clips by running 12 identical direct-input performances through 80 effect presets, then had 38 experienced guitarists compare clips pairwise and assign adjective labels. The central result is that in some pairwise comparisons, a clip with a higher spectral centroid was judged less 'bright' than a clip with a lower one, contradicting the established spectral-centroid model of brightness. The paper argues that such mismatches motivate a more data-driven, guitar-specific account of timbre semantics, and releases the dataset to support it.","feed_headline":"Guitar 'brightness' resists spectral-centroid predictions","feed_subtitle":"Expert-ranked clips reveal cases where higher spectral centroid sounds less bright, pushing timbre models beyond simple acoustics.","key_machinery":"The load-bearing object is the paired dataset: 12 direct-input guitar recordings, each processed through 80 amplifier and effects presets to yield 960 clips whose musical content is identical within each DI group. Annotators (87% with over 10 years of playing experience) gave pairwise 'more/less/equally X' judgments and multi-label adjective tags; a graph-based unification algorithm folds these into per-clip adjective scores by propagating a constant increment along pairwise relations. The brightness case study then compares the resulting rankings against the spectral centroid of each clip, computed from its spectrogram, as the acoustic feature under test.","core_discovery":"The paper's central claim is that the perception of guitar timbre adjectives is not fully captured by standard acoustic descriptors such as the spectral centroid. By holding musical content fixed—each comparison uses two different effect chains applied to the same DI recording—the experimental design isolates timbre as the variable under judgment. The authors find that while the spectral centroid correlates with 'brightness' overall, individual pairwise rankings sometimes invert the centroid ordering: a clip with a higher center of spectral mass is labeled less bright than a clip with a lower one. They take this as evidence that existing correlations, established on heterogeneous instrument recordings, do not transfer cleanly to fine-grained guitar tone, and that richer acoustic features (F0, harmonic-to-noise ratio) or machine-learned mappings are needed. The dataset itself, with 2038 annotations from expert annotators, is the enabling contribution that makes these observations possible.","pith_inferences":["If the brightness inversion survives loudness normalization, it would suggest guitar-specific timbre semantics diverge from general instrumental timbre dimensions; a cross-instrument experiment could test how far the divergence extends.","The dataset's skew toward metal and rock presets, which the authors acknowledge, means the observed adjective correlations and mismatches may not hold for other genres; sampling presets more evenly could reveal genre-dependent timbre semantics.","The graph-based scoring uses a fixed constant increment whose value is not analyzed; testing sensitivity to that constant, or comparing against a Bradley-Terry baseline, would clarify the robustness of the reported preset–adjective rankings."],"forward_implications":["The released dataset lets other researchers study guitar timbre adjectives without repeating the expert annotation collection.","Holding musical content fixed enables finer-grained timbre comparisons than cross-instrument studies, isolating timbre from pitch and content confounds.","Cross-correlations between adjectives (e.g., 'full' co-occurring with 'distorted,' 'dirty,' 'thick,' and 'dark') offer a data-driven route to defining lesser-known terms through more familiar ones.","The brightness mismatch indicates spectral centroid alone is insufficient for guitar tone, motivating models that incorporate features like F0 and harmonic-to-noise ratio.","Such resources could improve prompt-based music generation, timbre-based retrieval, and natural-language control of audio production tools."],"supporting_citations":[{"why":"Supplies the spectral-centroid account of brightness that the paper tests against in its case study; the paper's counterexamples are inversions of this correlation.","marker":"Schubert and Wolfe, 2006"},{"why":"Precedent for crowdsourcing timbre annotations of audio-effect recordings; the paper follows and adapts this methodology for guitar tones with pairwise comparisons.","marker":"Seetharaman and Pardo, 2016"},{"why":"Provides the seven-class taxonomy of timbre-descriptor origins used to argue that the dataset's adjectives span the full range of descriptor types.","marker":"Wallmark, 2019"},{"why":"Documents interactions among timbre, pitch, and loudness, justifying the paper's design choice to use identical DI content in pairwise comparisons.","marker":"Melara and Marks, 1990"},{"why":"Notes that pitch and loudness confound timbre perception, supporting the same design rationale and the interpretation of the pairwise rankings.","marker":"McAdams and Goodchild, 2017"},{"why":"Suggests F0 and harmonic-to-noise ratio as additional acoustic correlates of brightness, pointing toward the richer feature space the paper argues for.","marker":"Rosi, 2022"}],"fun_headline_variants":["Guitar brightness defies spectral centroid","Timbre adjectives: data clashes with theory","Expert ears split from spectral centroid on guitar tone","New dataset challenges spectral brightness link","Guitar timbre: when higher centroid sounds less bright"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The design assumes that running an identical guitar performance through different effect chains changes only timbre, leaving loudness and compression constant enough for pairwise judgments to be attributed purely to tone quality; the paper reports no loudness normalization or control condition that would verify this.","fun_headline_variants_meta":{"raw":{"variants":["Guitar brightness defies spectral centroid","Timbre adjectives: data clashes with theory","Expert ears split from spectral centroid on guitar tone","New dataset challenges spectral brightness link","Guitar timbre: when higher centroid sounds less bright"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000283,"raw_usage":{"total_tokens":1627,"prompt_tokens":859,"completion_tokens":768,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":475,"completion_tokens_details":{"reasoning_tokens":699}},"tokens_in":475,"tokens_out":768,"duration_ms":7729,"temperature":1.0,"reasoning_tokens":699,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:35:34.223065+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the brightness pairwise comparisons after loudness-normalizing the clips; if the cases where spectral centroid and perceived brightness disagree disappear after normalization, those cases are artifacts of level differences rather than genuine counterexamples to the centroid model.","supporting_citations":[{"cited_title":"distorted","cited_arxiv_id":null,"evidence_quote":"Supplies the spectral-centroid account of brightness that the paper tests against in its case study; the paper's counterexamples are inversions of this correlation."}],"review_version":1}