Pith. sign in

REVIEW 4 major objections 4 minor 2 cited by

An introduction to pitch strength in contemporary popular music analysis and production

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Pitch strength may be the missing low-level control for text-to-music AI.

desk verdict A transparently exploratory chapter that makes a useful negative point about HarmonicRatio, but the positive claims rest on untested listening and need validation before they carry weight. read the letter →

arxiv 2506.07473 v5 pith:KKZJLSOJ submitted 2025-06-09 cs.SD eess.AS

classification cs.SDeess.AS
keywords pitchstrengthsaliencytext-to-musicAImusicinformationretrievalHarmonicRationoisiness-inharmonicityspacepopularproductionpolyphonicdissonance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This chapter argues that pitch strength — how strongly a sound's pitch stands out against its timbre — is a low-level perceptual parameter that matters for contemporary popular music. The paper claims pitch strength varies widely both between songs and between the individual tracks (vocals, bass, drums) inside a song, and that it participates in shaping both small-scale form and large-scale song sections. The paper also argues that low pitch strength lets producers place dissonant elements together without the music sounding wrong, and that boosted upper harmonics often act as independently audible, low-pitch-strength pitches. If these claims hold, then text-to-music generative models, which work at a much higher semantic level than studio controls, would need to expose pitch strength to align with actual music production practice.

What carries the argument

The argument rests on two instruments. The first is a noisiness-inharmonicity space, built from spectral flatness and HarmonicRatio over the BEA dataset; the paper proposes that pitch strength correlates with a linear combination of these two dimensions, likely the first principal component. The second is the set of eleven reference sounds from Zwicker and Fastl (1990), numbered in decreasing pitch strength, to which musical passages are perceptually compared in the absence of formal listening tests. A spectral-peak-salience equalizer, Resonance EQ, is used to raise or lower the salience of spectral peaks, thereby audibly changing pitch strength and supporting the claim that peak salience is a causal factor.

What would settle it

Have a panel of listeners rank the pitch strength of the cited passages (for example, Primaal's vocal, bass, and drum stems, or the contrasting beats in Primaal's tracks) using pairwise comparisons against the Zwicker and Fastl reference sounds, and check whether the claimed orderings reproduce; a single failed ordinal prediction would weaken the perceptual foundation, and a measurement showing that HarmonicRatio tracks pitch strength on real popular music would weaken the paper's core distinction.

Watch

Extended reading notes

Core claim

The central claim is that pitch strength, defined by Yost as the strength of the perceived pitch of a complex sound relative to its overall timbre, is a salient and underused descriptor of contemporary popular music. Signal analysis and the author's perceptual observations indicate that pitch strength varies significantly across songs and across stems within songs, participates in both small- and large-scale musical structure by marking contrast beats and separating semiotic segments, allows low-pitch-strength elements such as roto-toms or spoken-rap vocals to coexist with otherwise dissonant tonal contexts, and is often carried by boosted upper harmonics whose individual pitches are weakly but audibly present. The paper further claims that the standard MPEG-7 feature HarmonicRatio does not measure pitch strength in the general case, despite working for iterated rippled noise, and that pitch strength may therefore be a useful addition to generative AI music models.

Load-bearing premise

The author's informal perceptual comparisons of musical passages to Zwicker and Fastl's reference sounds are assumed to rank pitch strength reliably across all the cited examples, and the paper acknowledges that most perceptual observations are the author's and should be verified with systematic listening tests.

Editorial extensions

If this is right

  • If pitch strength is a production-relevant parameter, current text-to-music models that accept only high-level captions will continue to miss a control that studio musicians routinely use.
  • Generative music models could become more useful in production by conditioning on a pitch-strength descriptor in addition to text.
  • Pitch strength provides an additional analytical axis for describing form in contemporary popular music, complementing existing semiotic segment descriptions.
  • Low-pitch-strength elements can serve as a dissonance-management resource, allowing out-of-tune or inharmonic sounds to sit inside tonal contexts without clashing.
  • Equalization and distortion that boost upper harmonics are deliberate production techniques for adding perceptual richness through weakly pitched, individually audible components.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If pitch strength were formalized as a continuous control, one could test whether professional producers can dial it independently of loudness, brightness, and harmonicity in a synthesis engine.
  • The noisiness-inharmonicity space suggests a concrete computational proxy: estimate pitch strength from the first principal component of spectral flatness and HarmonicRatio, then validate it against listening tests on studio stems.
  • The dissonance argument predicts a testable pattern in mixes: dissonant intervals will be more frequent in layers whose spectral peaks are weak or noisy, and less frequent in layers with high peak salience.
  • If text-to-music models were trained with pitch-strength annotations, their outputs might show different trade-offs between timbral clarity and harmonic ambiguity, but the paper does not demonstrate this.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces pitch strength (PS) as a low-level perceptual parameter relevant to contemporary popular music production, arguing that PS (1) varies significantly across and within songs, (2) contributes to small- and large-scale structure, (3) contributes to the handling of polyphonic dissonance, and (4) can be a feature of audible upper harmonics. It reviews psychoacoustic studies (Zwicker and Fastl, Yost, Patterson et al.), argues against the general validity of HarmonicRatio as a PS measure (Appendix B), proposes the noisiness-inharmonicity space as a proxy (Appendix D), and supports its claims with examples from songs by Vitalic, Primaal, Pink Floyd, Eminem, and others, supplemented by linked audio examples. The paper explicitly acknowledges limitations in Section 7, including that most perceptual observations are the author's and that no formal link between PS and the noisiness-inharmonicity space is established.

Significance. If the central claims hold, PS would be a production-relevant perceptual axis that generative music models should expose, making this a worthwhile contribution to the music production and MIR literature. The paper's strengths include its concrete musical examples, its honest and explicit limitation statements, and its effort to connect qualitative perceptual observations with signal features. It also reports one perceptual test in Section 6, although with insufficient detail. However, because the evidence is predominantly anecdotal and the proposed quantitative proxy is not formally validated, the contribution currently sits at the level of an informed position piece rather than a demonstrated result. The paper's own acknowledgements in Section 7 make this evident, and the revision should focus on turning the plausible hypothesis into a testable, validated framework.

major comments (4)
  1. [Section 7 (Conclusion)] Section 7 concedes that 'most perceptual observations presented are the author's; to improve their reliability, systematic listening tests should be conducted.' This concession is load-bearing because the abstract's four claims are each supported by comparative perceptual judgments without experimental data: Section 3 compares 'Eternity' and 'And it goes like' to Zwicker and Fastl's sounds; Section 4 asserts PS differences between Primaal segments; Section 5 attributes dissonance-handling to low PS in the Pink Floyd and Eminem examples. No listening-test data, inter-rater reliability, or statistical analysis is reported for any of these comparisons. As a result, the evidence cannot be distinguished from anecdote. The central claims would be considerably strengthened by a systematic listening test (e.g., pairwise PS ratings of the cited excerpts against Zwicker and Fastl's reference sounds) or by explicitly re-framing the contribution as hypothesis generation.
  2. [Section 7 / Appendix D] Section 7 states that the paper 'stops short of establishing a formal link' between PS and the noisiness-inharmonicity space. This matters because Section 3's 'Pitch strength evaluation' proposes that PS 'may correlate with a linear combination of the two dimensions,' and Figures 2, 3, and 5 use positions in this space as evidence for PS variability. Without a formal link, the figures provide no independent quantitative validation of the perceptual claims. The internal consistency is also at stake: Appendix D defines Dimension 2 using HarmonicRatio (AC1), yet Appendix B argues that HarmonicRatio does not correlate with PS in the general case. Using a feature whose validity the paper itself denies as one axis of the PS proxy requires justification. A calibration experiment mapping perceptual PS ratings onto the space would resolve both issues.
  3. [Appendix B] The claim that 'PS does not correlate with HarmonicRatio in the general case' is the motivation for seeking a new proxy, but the evidence is under-powered. Argument 1 examines eleven reference sounds without error bars or statistical analysis; Argument 2 uses only two synthetic chord pairs (minor third and fourth), with raw HR values 0.829 vs. 0.964 and normalized values 0.690 vs. 0.497, and asserts equal PS without a listening test. Two hand-picked counterexamples are insufficient to establish a general negative result. At minimum, a small formal experiment with a diverse sample of natural and synthetic tones spanning the HR range, plus a statistical test (e.g., correlation with PS ratings), is needed to support the paper's reliance on this negative result.
  4. [Section 6] The only reported perceptual data in the paper appear in Section 6, where 'perceptual tests asked respondents how many simultaneous pitches they could hear in two bass tracks and a keyboard track (Primaal, 2023c),' with mean reported values of 1.7, 1.95, and 2.45. No details are provided on the number of participants, the listening setup, the stimulus presentation, or the analysis method. Since this experiment directly supports claim (4) about audible upper harmonics, the manuscript should report the full methodology and results, or clearly label the numbers as informal demonstration.
minor comments (4)
  1. [Section 3] The heading 'T racks within a song' contains a typo; it should read 'Tracks within a song.'
  2. [Appendix B / Figure 8] The caption renders the transformation as '(1 − HarmonicRatio)0.21' without superscript formatting, and the text describes panel (b) as '10^HarmonicRatio'; this is confusing. Please render the exponent properly and explain the choice of the exponent 0.21.
  3. [Section 2] In the bullet list, 'PS of pure tones increases with duration and sound pressure' should say 'sound pressure level' for terminological consistency with the rest of the psychoacoustic literature.
  4. [Appendix C] The note that zwicker and Fastl's measures are normalized per frequency, so PS is referenced by sound number, would be clearer if it restated explicitly that all comparisons in the chapter use the 500 Hz version of the reference sounds.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the chapter is explicitly exploratory and its claims rest on perceptual examples and a proposed, unvalidated proxy, not on fitted equations or self-citation chains.

full rationale

The paper does not derive pitch strength from a fitted equation or from a parameter that is then renamed as a prediction. The noisiness-inharmonicity space is imported from the author's prior work (Deruty et al., 2024), but it is used as a proposed representation rather than as a formal derivation of pitch strength: Section 3 hypothesizes only that 'PS correlates with a linear combination of the two dimensions,' and Section 7 explicitly concedes that the chapter 'stops short of establishing a formal link between them.' The exponent 0.21 in the transformation (1-HarmonicRatio)^0.21 is described as a distribution-normalization step applied to the BEA dataset, not as a parameter fitted to measured pitch-strength values, so no quantity is fitted and then re-predicted. Similarly, the use of Zwicker and Fastl's reference sounds is an informal comparative aid, and the paper openly states that 'most perceptual observations presented are the author's' and that systematic listening tests are needed. These are evidentiary limitations, not circularity: the central claims are presented as arguments from examples and signal observations rather than as conclusions forced by definitions, fitted parameters, or a self-citation chain. The presence of self-citations for structural terminology and the noisiness-inharmonicity space is normal scholarly continuity and is not load-bearing in a circular sense, since the paper does not invoke a uniqueness theorem or an unexamined prior result to rule out alternatives. Accordingly, no specific circular step can be identified and the appropriate score is 0.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The ledger is dominated by perceptual and framework assumptions. The one explicit fitted constant is the HR normalization exponent 0.21. The main conceptual novelty, the noisiness-inharmonicity space, is imported from the author's own prior paper rather than introduced here, so it is listed as an axiom rather than an invented entity.

free parameters (2)
  • HR normalization exponent = 0.21
    The exponent in (1 - HarmonicRatio)^0.21 is chosen to normalize the HarmonicRatio distribution on the BEA dataset (Appendix B). It is a free calibration that affects the reported correlation comparison.
  • Spectral roll-off s and harmonic count for synthetic tones = s=0.8, 10 harmonics
    Hand-chosen parameters for the Mauch-Dixon model used to generate the minor-third and fourth complex tones in Appendix B, Argument 2. The conclusion that HR differs while PS is similar may depend on these values.
assumptions (5)
  • domain assumption The pitch strength measurements of Zwicker and Fastl (1990) for eleven reference sounds are valid and can be used as an ordinal scale for comparing modern music.
    Used throughout Sections 3-6 to rank the PS of songs and stems relative to sounds 1-11.
  • standard math HarmonicRatio equals the height of the first peak of the normalized autocorrelation function (AC1) and is a valid PS predictor for iterated rippled noise only.
    Adopted from Patterson et al. (2000) and ISO (2001); used as the definitional basis for the PS critique in Appendix B.
  • domain assumption The noisiness-inharmonicity space of Deruty et al. (2024), built from spectral flatness and HarmonicRatio on the BEA dataset, is a meaningful representation on which to plot PS.
    The author's own prior framework is invoked in Sections 3 and 5 and in Figures 2, 3, 5, 10 without re-derivation in this chapter.
  • standard math ISO 226:2003 equal-loudness weighting is an accurate perceptual weighting for the power spectra used here.
    Assumed in Appendix A and applied to all spectral analyses.
  • domain assumption The author's perceptual judgments of pitch strength are reliable.
    All qualitative PS comparisons are the author's; Section 7 explicitly calls for systematic listening tests.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An introduction to pitch strength in contemporary popular music analysis and production." pith.science (2026). https://pith.science/paper/KKZJLSOJ

@misc{pith2026250607473,
  author       = {Pith},
  title        = {Pith review of: An introduction to pitch strength in contemporary popular music analysis and production},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KKZJLSOJ}},
  note         = {Machine review of arXiv:2506.07473}
}
read the original abstract

Music information retrieval distinguishes between low- and high-level descriptions of music. Current generative AI models rely on text descriptions that are higher level than the controls familiar to studio musicians. Pitch strength, a low-level perceptual parameter of contemporary popular music, may be one feature that could make such AI models more suited to music production. Signal and perceptual analyses suggest that pitch strength (1) varies significantly across and inside songs; (2) contributes to both small- and large-scale structure; (3) contributes to the handling of polyphonic dissonance; and (4) may be a feature of upper harmonics made audible in a perspective of perceptual richness.

Figures

Figures reproduced from arXiv: 2506.07473 by the authors.

Figure 1
Figure 1. End guitar chord from My Bloody Valentine (1991). (a) Power spectrum; (b) weighted [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. BEA dataset, distribution of normalized noisiness and HR-inharmonicity. White arrows, [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Primaal (2023a), positions of each beat in the noisiness-inharmonicity space. White, drum [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Some implications and contrasts involving PS in (a) Primaal (2023b) and (b) Primaal [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Position of each beat in the noisiness-inharmonicity space. (a) Primaal (2023b) and (b) [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Decomposition of JID (2022) into semiotic segments. The segments in green are perceived [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Roto-tom sample ‘B-A3 2.wav’ (A3, medium velocity). (a) ISO 226-2003 weighted sample, [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: (a) HarmonicRatio for the eleven sounds used by Zwicker and Fastl (1990, p. 136). (b) [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: HarmonicRatio derived from the pairwise combination of elements in two harmonic tones. [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: BEA dataset, distribution of normalized noisiness and HR-inharmonicity. For each [PITH_FULL_IMAGE:figures/full_fig_p014_10.png]
Figure 11
Figure 11. Figure 11: Solid black line, original power spectrum. Solid gray line, power spectrum with Resonance [PITH_FULL_IMAGE:figures/full_fig_p015_11.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Methods for pitch analysis in contemporary popular music: multiple pitches from harmonic tones in Vitalic's music

    cs.SD 2025-06 unverdicted novelty 5.0 of 10

    Single quasi-harmonic tones in Vitalic's electronic music are perceived by listeners as carrying multiple simultaneous pitches, an effect the paper argues producers use deliberately.

  2. Evolving music theory for emerging musical languages

    cs.SD 2025-06 conditional novelty 3.5 of 10

    Pitch in electronic music is better understood as a listener-dependent perceptual construct than an objective property of the sound.

Reference graph

Works this paper leans on

63 extracted references · 52 canonical work pages · cited by 2 Pith papers

  1. [1]

    Agostinelli, A., Denk, T., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., et al. (2023). Music LM : Generating music from text. arXiv preprint arXiv:2301.11325 . https://arxiv.org/abs/2301.11325

  2. [2]

    Antunes, P. R. (2017). Is it possible to tune a drum? Journal of Computational Physics , 338:91--106. https://doi.org/10.1016/j.jcp.2017.02.056

  3. [3]

    Bimbot, F., Deruty, E., Sargent, G., and Vincent, E. (2012). Semiotic structure labeling of music pieces: Concepts, methods and annotation conventions. In Proceedings of the 13th International Society for Music Information Retrieval Conference , pages 235--240. ISMIR. \ ://inria.hal.science/hal-00758648/

  4. [4]

    Bimbot, F., Deruty, E., Sargent, G., and Vincent, E. (2016). System & contrast: a polymorphous model of the inner organization of structural segments within music pieces. Music Perception: An Interdisciplinary Journal , 33(5):631--661. https://inria.hal.science/hal-01188244v1

  5. [5]

    Bimbot, F., Le Blouch, O., Sargent, G., and Vincent, E. (2010). Decomposition into autonomous and comparable blocks: a structural description of music pieces. Technical report, IRISA, Rennes, France. https://inria.hal.science/inria-00553526v1

  6. [6]

    Bimbot, F., Sargent, G., Deruty, E., Guichaoua, C., and Vincent, E. (2014). Semiotic description of music structure: An introduction to the Quaero/Metiss structural annotations. In AES 53rd International Conference on Semantic Audio . Audio Engineering Society. \ ://inria.hal.science/hal-00931859v1

  7. [8]

    Bogdanov, D., Serr\' a , J., Wack, N., and Herrera, P. (2009). From low-level to high-level: Comparative study of music similarity measures. In 11th IEEE International Symposium on Multimedia, San Diego, CA, USA , pages 453--458. IEEE. https://doi.org/10.1109/ISM.2009.72

  8. [9]

    Bregman, A. S. (1994). Auditory scene analysis: The perceptual organization of sound . MIT press. https://mitpress.mit.edu/9780262521956/

Show all 63 references
  1. [10]

    Brewster, W. (2024). Gear rundown: K evin S hields of M y B loody V alentine. Mixdown . \ ://mixdownmag.com.au/features/my-bloody-valentine-kevin-shields/

  2. [11]

    Caplin, W. E. (1998). Classical form: A theory of formal functions for the instrumental music of Haydn, Mozart, and Beethoven . Oxford University Press. ISBN: 978-0195143997

  3. [13]

    Deruty, E., Bimbot, F., and Van Wymeersch, B. (2013). Methodological and musicological investigation of the system & contrast model for musical form description. Technical report, INRIA. \ ://inria.hal.science/hal-00965914/

  4. [14]

    and Grachten, M

    Deruty, E. and Grachten, M. (2022). `` M elatonin'': A case study on AI -induced musical style. In Proceedings of the 3rd Conference on AI Music Creativity . AIMC. \ ://doi.org/10.5281/zenodo.7088302

  5. [15]

    Deruty, E., Grachten, M., Lattner, S., Nistal, J., and Aouameur, C. (2022). On the development and practice of AI technology for contemporary popular music production. Transactions of the International Society for Music Information Retrieval , 5(1). \ ://doi.org/10.5334/tismir.100

  6. [16]

    Deruty, E., Meredith, D., and Lattner, S. (2024). The evolution of inharmonicity and noisiness in contemporary popular music. Journal of New Music Research , 52(5):382--409. \ ://arxiv.org/abs/2408.08127

  7. [17]

    and Pachet, F

    Deruty, E. and Pachet, F. (2015). The MIR perspective on the evolution of dynamics in mainstream music. In M \" u ller, M. and Wiering, F., editors, Proceedings of the 16th International Society for Music Information Retrieval Conference, 2015, M \' a laga, Spain, October 26-3...

  8. [18]

    No L ove [song]

    Eminem (2010). No L ove [song]. Album: Recovery. Record label: Interscope. \\(1) Original song: https://youtu.be/KV2ssT8lzj8&t=192. \\(2) Increased pitch strength: https://youtu.be/6wRKm-kUyBQ

  9. [19]

    and Munson, W

    Fletcher, H. and Munson, W. A. (1933). Loudness, its definition, measurement and calculation. Bell System Technical Journal , 12(4):377--430. \ ://doi.org/10.1002/j.1538-7305.1933.tb00403.x

  10. [20]

    Frisius, R. (2010). In search of lost harmony. Contemporary music: theoretical and philosophical perspectives , pages 77--87. ISBN 9781138251311

  11. [22]

    Holland, J. (2001). Roto-toms. In Grove Music Online . Oxford Music Online. \ ://doi.org/10.1093/gmo/9781561592630.article.23938

  12. [23]

    Y., and Ellis, D

    Huang, Q., Jansen, A., Lee, J., Ganti, R., Li, J. Y., and Ellis, D. P. W. (2022). Mulan: A joint embedding of music audio and natural language. https://arxiv.org/abs/2208.12415

  13. [24]

    Hudson, A. (2021). 'Loveless' at 30: Z oon Reflects on the Influence of M y B loody V alentine's Shoegaze Blueprint . Exlaim.ca . https://exclaim.ca/music/article/loveless_at_30_zoon_reflects_on_the_influence_of_my_bloody_valentines_shoegaze_blueprint

  14. [25]

    Hyper M usic

    Hyper Music (2024). Hyper M usic. https://www.hyper-music.com/

  15. [26]

    ISO/MPEG N4224

    ISO (2001). ISO/MPEG N4224. International Standard 15938-4 Information Technology - Multimedia Content Description Interface - Part 4 Audio . Standard, International Organization for Standardization. https://www.iso.org/standard/34231.html

  16. [27]

    Normal equal-loudness level contours-ISO 226: 2003

    ISO (2003). Normal equal-loudness level contours-ISO 226: 2003 . Standard, International Organization for Standardization, Geneva, Switzerland. https://www.iso.org/standard/34222.html

  17. [28]

    Raydar [song]

    JID (2022). Raydar [song]. Album: The Forever Story. Record label: Interscope. \ ://youtu.be/MrytARdrZZQ

  18. [29]

    Saturday C omes S low [song]

    Massive Attack (2010). Saturday C omes S low [song]. Album: Heligoland. Record label: Virgin. Partial transcription, 2'24 to 2'51: https://youtu.be/CarfITcc_r4

  19. [30]

    and Dixon, S

    Mauch, M. and Dixon, S. (2010). Approximate note transcription for the improved identification of difficult chords. In Proceedings of the 11th International Society for Music Information Retrieval Conference , pages 135--140. ISMIR. https://doi.org/10.5281/zenodo.1416598

  20. [31]

    McAdams, S. (1999). Perspectives on the contribution of timbre to musical structure. Computer music journal , 23(3):85--102. https://www.jstor.org/stable/3681242

  21. [32]

    and Fujinaga, I

    McKay, C. and Fujinaga, I. (2004). Automatic genre classification using large high-level musical feature sets. In 5th International Conference on Music Information Retrieval, Barcelona, Spain . ISMIR. https://www.music.mcgill.ca/ ich/research/ismir2004/McKay_Fujinaga_ISMIR_2004.pdf

  22. [33]

    Monelle, R. (2014). Linguistics and semiotics in music . Routledge. ISBN 978-3718652099

  23. [34]

    Moreau, N., Sikora, T., and Kim, H.-G. (2006). MPEG-7 audio and beyond: Audio content indexing and retrieval . John Wiley & Sons. https://doi.org/10.1002/0470093366

  24. [35]

    Sometimes [song]

    My Bloody Valentine (1991). Sometimes [song]. Album: Loveless. Record label: Creation. \ ://youtu.be/hSI_9P9rRt4?&t=302

  25. [36]

    Nagle, P. (2015). Spectrasonics Omnisphere 2 . Sound on Sound , Aug. 2015. \ ://www.soundonsound.com/reviews/spectrasonics-omnisphere-2

  26. [37]

    Newell, P. (2017). Recording studio design . Taylor & Francis. \ ://doi.org/10.4324/9781315675367

  27. [38]

    D., Handel, S., Yost, W

    Patterson, R. D., Handel, S., Yost, W. A., and Jaysurya Datta, A. (1996). The relative strength of the tone and noise components in iterated rippled noise. The Journal of the Acoustical Society of America , 100(5):3286--3294. https://doi.org/10.1121/1.417212

  28. [39]

    D., Yost, W

    Patterson, R. D., Yost, W. A., Handel, S., and Datta, J. A. (2000). The perceptual tone/noise ratio of merged iterated rippled noises. The Journal of the Acoustical Society of America , 107(3):1578--1588. https://doi.org/10.1121/1.428442

  29. [40]

    Peeters, G. (2004). A large set of audio features for sound description (similarity and classification) in the CUIDADO project. Technical report, IRCAM, Paris, France. http://recherche.ircam.fr/anasyn/peeters/ARTICLES/Peeters_2003_cuidadoaudiofeatures.pdf

  30. [41]

    Time [song]

    Pink Floyd (1973). Time [song]. Album: The Dark Side of the Moon. Record label: Capitol. https://youtu.be/Qr0-7Ds79zo

  31. [42]

    Boom [song]

    Primaal (2023a). Boom [song]. Publisher: Alter-K. Screen capture of ProTools session, beat, bass, and lead: https://youtu.be/vsyGQUwg80g

  32. [43]

    Cardinal [song]

    Primaal (2023b). Cardinal [song]. Publisher: Alter-K. (1) Screen capture of ProTools session, structure: https://youtu.be/o-Jw9wxRupU. (2) Screen capture of ProTools session, drum, and bass: https://youtu.be/ZHRl5-jN9ak

  33. [44]

    Elevate [song], silver [song], and whomp [song]

    Primaal (2023c). Elevate [song], silver [song], and whomp [song]. Publisher: Alter-K. Bass and keyboard tracks: https://youtu.be/Dy8Q_2UhvS4

  34. [45]

    Whomp [song]

    Primaal (2023d). Whomp [song]. Publisher: Alter-K. Screen capture of ProTools session, structure: https://youtu.be/te9NureYfeE

  35. [46]

    Yada Y ada [song]

    Primaal (2023e). Yada Y ada [song]. Publisher: Alter-K. Screen capture of ProTools session, structure: https://youtu.be/NFhDYmwz_ts

  36. [47]

    Richardson, P. G. (2010). Acoustic analysis and tuning of cylindrical membranophones . PhD thesis, Anglia Ruskin University. \ ://core.ac.uk/download/pdf/77282289.pdf

  37. [48]

    Schoenberg, A. (1967). Fundamentals of musical composition . Faber and Faber. ISBN 978-0571196586

  38. [49]

    Sethares, W. A. (2005). Tuning, timbre, spectrum, scale . Springer Science & Business Media. ISBN 978-3540761730

  39. [50]

    Snyder, B. (2000). Music and memory: An introduction . MIT press. ISBN 978-0262692373

  40. [51]

    Tenney, J. (1988). A History of `Consonance' and `Dissonance' . Excelsior Music Company. \ ://www.plainsound.org/pdfs/HCD.pdf

  41. [52]

    C., Robinson, P., and Richardson, P

    Toulson, R., Crigny, C. C., Robinson, P., and Richardson, P. (2009). The perception and importance of drum tuning in live performance and music production. The Journal on the Art of Record Production , 4. https://www.arpjournal.com/asarpwp/theme/drum-tuning/

  42. [53]

    Clouds A s W itnesses [song]

    \ uicideBoy\ (2015). Clouds A s W itnesses [song]. Album: Gray/Grey. Record label: G59 Records Inc. (1) Original song: https://youtu.be/Zi3-P4qwpGo&t=19. (2) Partial transcription, 0'19 to 0'37: https://youtu.be/c8PAPWnAato

  43. [54]

    Eternity [song]

    Vitalic (2017). Eternity [song]. Album: Voyager. Record label: Clivage Music / Universal. \ ://youtu.be/u9U_xOIdz_s

  44. [55]

    And It Goes Like [song]

    Vitalic (2023). And It Goes Like [song]. Album: Confess EP. Record label: Citizen Records / Virgin Music France. https://youtu.be/n97mOW3d1n8

  45. [56]

    Waters, R. (2020). Time [song]. Album: Us + Them. Record label: Columbia. \\(1) Original song: https://youtu.be/rePdcdLQ_mk. \\(2) Original audio, with comments: https://youtu.be/Ne3y9lrH0I0. \\(3) Processed audio, with transcription: https://youtu.be/NMrdEOMDn_c. \\(4) Transc...

  46. [57]

    Wightman, F. L. (1973). The pattern-transformation model of pitch. The Journal of the Acoustical Society of America , 54(2):407--416. https://doi.org/10.1121/1.1913592

  47. [58]

    Wu, C.-W., Dittmar, C., Southall, C., Vogl, R., Widmer, G., Hockman, J., M \"u ller, M., and Lerch, A. (2018). A review of automatic drum transcription. IEEE/ACM Transactions on Audio, Speech, and Language Processing , 26(9):1457--1483. https://doi.org/10.1109/TASLP.2018.2830113

  48. [59]

    Yost, W. A. (1996). Pitch strength of iterated rippled noise. The Journal of the Acoustical Society of America , 100(5):3329--3335. https://doi.org/10.1121/1.416973

  49. [60]

    Yost, W. A. (2009). Pitch perception. Attention, Perception, & Psychophysics , 71(8):1701--1715. https://doi.org/10.3758/APP.71.8.1701

  50. [61]

    Yost, W. A. and Hill, R. (1978). Strength of the pitches associated with ripple noise. The Journal of the Acoustical Society of America , 64(2):485--492. https://doi.org/10.1121/1.382021

  51. [62]

    A., Hill, R., and Perez-Falcon, T

    Yost, W. A., Hill, R., and Perez-Falcon, T. (1978). Pitch and pitch discrimination of broadband signals with rippled power spectra. The Journal of the Acoustical Society of America , 63(4):1166--1173. https://doi.org/10.1121/1.381824

  52. [63]

    A., Patterson, R., and Sheft, S

    Yost, W. A., Patterson, R., and Sheft, S. (1996). A time domain description for the pitch strength of iterated rippled noise. The Journal of the Acoustical Society of America , 99(2):1066--1078. https://doi.org/10.1121/1.414593

  53. [64]

    A., Sheft, S., and Patterson, R

    Yost, W. A., Sheft, S., and Patterson, R. (1994). Iterated rippled noise: Testing temporal theories of complex pitch. The Journal of the Acoustical Society of America , 95(5 Supplement):2966. https://doi.org/10.1121/1.409036

  54. [65]

    Zanoni, M., Ciminieri, D., Sarti, A., and Tubaro, S. (2023). Searching for dominant high-level features for music information retrieval. In 2012 Proceedings of the 20th European Signal Processing Conference (EUSIPCO) . IEEE. https://ieeexplore.ieee.org/document/6334107

  55. [66]

    and Fastl, H

    Zwicker, E. and Fastl, H. (1990). Pitch and pitch strength. In Psychoacoustics: Facts and Models , pages 111--149. Springer-Verlag, New York. https://doi.org/10.1007/978-3-662-09562-1_5

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.