Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Computational Analysis of Yaredawi YeZema Silt in Ethiopian Orthodox Tewahedo Church Chants

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Time-averaged pitch distributions alone distinguish the three Ethiopian Orthodox chant modes, with 98.11% within-dataset and 87.50% cross-dataset accuracy.

desk verdict First public EOTC chant dataset and a fair cross-dataset test, but the single-annotator labels and undisclosed GMM parameters leave the pitch-set revision partly unverified. read the letter →

arxiv 2412.18788 v1 pith:Y3PBKVPU submitted 2024-12-25 eess.AS cs.SDeess.SP

classification eess.AScs.SDeess.SP
keywords EthiopianOrthodoxchantYaredawiYeZemaSiltpitchdistributionmodeclassificationmusicinformationretrievalsetanalysisGe'ezEzilArarayGaussianmixturemodel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the three canonical chanting modes of the Ethiopian Orthodox Tewahedo Church—Ge'ez, Ezil, and Araray—can be told apart from the time-averaged distribution of pitch values, without relying on melodic ordering. On a newly introduced dataset of 369 single-mode clips (about 10 hours of Se'atat Zema, the Horologium chant), a small fully convolutional network trained on pitch histograms reaches 98.11% accuracy in 5-fold cross-validation and up to 87.50% on older recordings by a different chanter. The same pitch statistics, analyzed with Gaussian mixture models, lead the authors to claim that each mode has a distinctive pitch set, revising the earlier ethnomusicological statement that Ezil and Araray use the same pitch set. A sympathetic reader would care because it makes a centuries-old oral chant tradition computationally testable, and it suggests mode identity is a global, song-level property rather than a sequence of local motifs.

What carries the argument

The load-bearing object is the time-averaged pitch distribution: a histogram of frame-level pitch values (10-cent bins) extracted by pYIN from pitch contours, optionally stabilized by two methods (morphetic and masking) and calibrated for slow pitch drift by linear regression. The classifier is the M5 fully convolutional network, adapted to frequency-domain input, whose convolution makes it invariant to pitch shifting. For the musicological comparison, the paper aligns each recording's pitch distribution to an anchor by maximizing cross-correlation, averages the aligned distributions per mode, and fits a Gaussian mixture model; components with variance under 100 cents are taken as representative pitches.

What would settle it

Have an independent expert chanter re-label a random subset of the source recordings into Ge'ez, Ezil, and Araray with clip boundaries; if the network trained on the original labels performs at chance on the independently labeled subset, or if the GMM pitch sets do not re-emerge, the classification and pitch-set claims would collapse.

Watch

Extended reading notes

Core claim

The central discovery, as the paper states it, is that Yaredawi YeZema Silt is identifiable from the distribution of stabilized pitch contours: the mode is a song-level property that a pitch-shift-invariant classifier can read from time-averaged pitch statistics. Under 5-fold cross-validation the best calibrated, non-stabilized pitch distribution reaches 98.11% accuracy on the new dataset; in the cross-dataset test, the non-calibrated raw and masking-stabilized distributions reach 87.50%. Comparing the GMM-estimated pitch sets, the paper finds Ge'ez uses a flexible chain of thirds (g1–g2–g3 with g2 as a stable anchor), while Ezil has five fairly evenly spaced pitches (intervals 200–300 cents) and Araray has five unevenly spaced pitches (intervals 172–347 cents), so Ezil and Araray do not share a pitch set as previously reported.

Load-bearing premise

The load-bearing assumption is that the single annotator's manual mode labels and clip boundaries are correct; there is no second annotator or inter-annotator agreement reported, and mislabeled or mis-segmented clips would propagate into every accuracy and pitch-set result.

Editorial extensions

If this is right

  • Yaredawi YeZema Silt is a song-level property that can be captured by time-averaged pitch statistics; classification accuracy improves as the input grows from 5 seconds to full length.
  • Pitch-shift invariance is sufficient for mode identity, since a convolutional network that treats transposed pitch distributions as the same class performs well both within and across datasets.
  • The three modes have distinctive pitch sets and pitch variances, revising the earlier claim that Ezil and Araray share a pitch set.
  • The public dataset enables reproducible benchmarks for EOTC chant music information retrieval, including future temporal, lyrical, and music-generation tasks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: If mode identity is truly readable from pitch histograms, automatic indexing of the large unlabeled corpus of EOTC recordings could become feasible, helping preservation and education efforts.
  • Editorial extension: The stable-region and calibration pipeline originally developed for Georgian vocal music transfers cleanly to this monophonic liturgical tradition, suggesting it may work for other oral chant traditions with flexible pitch sets.
  • Editorial extension: The claimed Ezil/Araray distinction invites a perceptual test: synthesized phrases using the GMM pitch sets could be played to expert chanters to see whether the computational boundary corresponds to an audibly recognized mode change.
  • Editorial extension: Because a time-averaged distribution suffices, the modes may be characterized by tonal color rather than melodic order, which predicts that pitch-order-shuffled clips within a mode should remain classifiable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper presents a computational study of the three chanting modes (Ge’ez, Ezil, Araray) of Ethiopian Orthodox Tewahedo Church chant. The authors introduce a new dataset of 369 manually segmented and labeled single-mode clips from Se’atat Zema, extract time-averaged features (pitch-distribution histograms from pYIN with optional stable-region extraction and linear pitch calibration, plus mel-spectrogram, MFCC, and chromagram baselines), and train an adapted M5 CNN classifier. Within-dataset 5-fold CV accuracy reaches 98.11% for the calibrated non-stabilized pitch distribution on full-length audio; cross-dataset testing on Shelemay and Jeffery’s independently recorded 24-item anthology reaches up to 87.50%. The paper then aligns per-recording pitch distributions, fits GMMs, and uses Table 4 to argue that the three modes have distinct pitch sets, revising Shelemay et al.’s claim that Ezil and Araray share the same pitch set.

Significance. If the results hold, the paper would provide the first computational evidence that Yaredawi YeZema Silt is recognizable from time-averaged pitch distributions at the song level, and it would offer a concrete revision to the ethnomusicological record on EOTC chant pitch sets. The strengths include a newly described dataset with mode annotations made publicly available on GitHub, an external validation set from Shelemay and Jeffery’s anthology, use of established stable-region extraction methods, and a clearly presented classification benchmark. These are valuable contributions to the under-resourced MIR area of EOTC chant research and to comparative ethnomusicology. However, the significance is currently tempered by unresolved issues in annotation reliability, performer-confounding in the within-dataset experiment, under-specification of the GMM analysis, and the absence of uncertainty quantification on the headline accuracies.

major comments (4)
  1. [Section 3 (Dataset)] The dataset annotations are the linchpin of every downstream claim, but Section 3 reports that all 369 clips were segmented and labeled manually by a single annotator, with no inter-annotator agreement statistic, no second-opinion protocol, and no audit of segmentation boundaries. The 5-fold CV results in Table 2, the cross-dataset transfer to [4], and the per-mode averaged pitch distributions in Figure 3 all use these labels as ground truth. A systematic labeling bias or a segmentation error at a mode boundary would propagate directly into the classification accuracies and the GMM pitch-set estimates. The cross-dataset experiment provides independent evidence that the labels are consistent with Shelemay’s annotations, but it uses only 24 clips and does not validate the segmentation of the 369 training clips. Please report an annotation reliability protocol (e.g., a second annotator on a subset, boundary agreement measures, or a stability analysis under relabeling of ambiguous cases) and discuss how remaining ambiguity affects the reported numbers.
  2. [Section 3 (Dataset) and Section 4.3 (Results)] All 369 within-dataset clips are recordings by a single scholar (Section 3). The high within-dataset accuracies in Table 2 may therefore partly reflect speaker or recording characteristics rather than the chanting mode itself, especially for timbral features such as MFCC and mel-spectrogram. The cross-dataset result is the appropriate check, but only the full-length full-audio conditions are reported there, and the 24-item test set is too small to establish equal performance across feature classes. Please quantify the performer confound (if possible, via additional recordings from other chanters) or at minimum temper the within-dataset comparisons to account for this single-performer design.
  3. [Section 5, Table 4] The GMM fitting procedure that grounds the central musicological claim is under-specified. The text says the fitting is 'initialized by user-specified mean values' and that only components with variance below 100 cents are retained as representative pitches, but it does not disclose the number of GMM components per mode, the initial mean values, or the algorithm used to select the 'maximal weights within one octave.' These choices directly determine the number and positions of the representative pitches in Table 4, and hence the conclusion that Ezil and Araray use distinct pitch sets. If the component count or initialization is drawn from [4]’s reported pitch sets, the comparison is partly circular; if it differs between modes, the comparison is confounded. Please disclose the full initialization, report all fitted components (not only those passing the threshold), and provide sensitivity analyses over component counts and variance thresholds.
  4. [Section 4.3, Table 2] All classification accuracies are reported as point estimates without error bars or significance tests. The within-dataset 5-fold CV numbers have no per-fold standard deviation, and the cross-dataset column is based on only 24 instances, so the difference between 87.50% and 75.00% corresponds to three clips and can hardly be distinguished from noise. As a result, qualitative comparisons in Section 4.3 (pitch distributions 'greatly outperform' other features, stabilization 'helps for most of the cases,' masking reduces the cross-dataset gap better than morphetic) are not statistically supported. Please report per-fold variability, confidence intervals, and tests (e.g., paired tests across folds or binomial CIs for the cross-dataset results) for at least the headline comparisons.
minor comments (5)
  1. [Section 4.1] The pYIN time resolution is stated as '128 samples (5.8 ms)', but 128 samples at 44.1 kHz corresponds to approximately 2.9 ms; if 5.8 ms was intended, the sample count should be 256.
  2. [Section 4.2] The text says Shelemay’s recordings have '0.33 and 4.05 seconds' as shortest and longest durations, but Table 1 lists minute-scale durations; these units should be corrected to minutes.
  3. [Section 5] The sentence 'g1 and g2 form approximately a minor third g2' contains a typo; it should be 'a minor third between g1 and g2'.
  4. [Table 4] The note-name notation (G3, ˙g1, E4, e1, etc.) mixed with a reference pitch of 82.4 Hz is confusing; please define the octave convention and spell out how each note name relates to the cents scale.
  5. [Figure 3] The red background used to mark Shelemay recordings is not legible in grayscale; please use an additional visible marker or a colorbar and axis labels to make the figure self-contained.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: classification benchmarks use independently assigned mode labels and external cross-dataset validation, and the pitch-set analysis is descriptive.

full rationale

The paper's derivation chain is self-contained rather than circular. The central classification experiment (Section 4) trains a supervised neural classifier on time-averaged pitch-distribution features to predict YeZema Silt labels that were produced by manual segmentation and listening (Section 3), not by the feature-extraction pipeline itself, so the accuracy figures in Table 2 are not true by construction. The cross-dataset experiment uses recordings and labels from Shelemay and Jeffery [4], an independent source, which provides external validation rather than a self-citation loop. The pitch-set analysis in Section 5 is a descriptive GMM summary of the same pitch-distribution features; it does not feed predictions back into the labels or into the classifier, and the claim that Ezil and Araray differ in pitch sets is a new empirical comparison against [3,4], not a renaming of those authors' results. Two concerns are noted but do not rise to demonstrable circularity under the stated standard: the GMM is 'initialized by user-specified mean values' (Section 5) without disclosing those values, and the manual mode labels have no inter-annotator reliability check. However, without evidence that the initialization means equal the reported pitch centers, or that the labels were generated from the same pitch distributions used as features, no specific reduction of the output to the input can be exhibited. These are reproducibility and validation risks, not definitional or self-citation circularity. Accordingly, the paper receives a score of 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new theoretical entities. The main uncharged inputs are the user-specified GMM initialization and thresholds, plus domain assumptions about label accuracy, pitch-tracking reliability, and the sufficiency of pitch-distribution features.

free parameters (3)
  • GMM initial mean values = Not reported
    Section 5 says 'The GMM fitting process is initialized by user-specified mean values to enhance convergence [10]'. These hand-chosen means can shape which peaks are found as representative pitches.
  • Number of GMM components per mode = Not reported
    The GMM complexity is not stated; representative pitches are later selected from components with variance < 100 cents and maximal weight within an octave, so the component count affects the reported pitch sets.
  • Variance threshold for representative pitches = 100 cents
    Components with variance larger than 100 cents are excluded; this cutoff is hand-chosen and determines which pitches are listed in Table 4.
assumptions (4)
  • domain assumption The manual mode labels for all 369 audio segments are correct, including the boundaries where a long recording switches modes.
    Single annotator segmented and labeled the dataset (Section 3); no inter-annotator agreement is reported, and all classification and pitch-set results depend on these labels.
  • domain assumption pYIN pitch tracking reliably estimates the fundamental frequency of this monophonic vocal style with no accompaniment.
    The entire feature pipeline depends on pYIN (Section 4.1); if pitch tracking fails on the vocal timbre or ornamentation, the pitch distributions and GMM pitch sets are invalid.
  • ad hoc to paper Time-averaged pitch distributions are a sufficient representation for identifying YeZema Silt at the song level.
    The paper states that the classification results support this assumption (Section 5) and then uses the same feature type for the musicological analysis; the assumption is central to the analysis approach.
  • domain assumption The Shelemay and Jeffery anthology recordings [4] are comparable to the proposed dataset and were produced in the same chant tradition.
    Cross-dataset results (Table 2) are treated as out-of-domain validation; if the anthology recordings differ in recording conditions or performance practice, the transfer accuracy and pitch-distribution comparisons would be affected.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Computational Analysis of Yaredawi YeZema Silt in Ethiopian Orthodox Tewahedo Church Chants." pith.science (2026). https://pith.science/paper/Y3PBKVPU

@misc{pith2026241218788,
  author       = {Pith},
  title        = {Pith review of: Computational Analysis of Yaredawi YeZema Silt in Ethiopian Orthodox Tewahedo Church Chants},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Y3PBKVPU}},
  note         = {Machine review of arXiv:2412.18788}
}
read the original abstract

Despite its musicological, cultural, and religious significance, the Ethiopian Orthodox Tewahedo Church (EOTC) chant is relatively underrepresented in music research. Historical records, including manuscripts, research papers, and oral traditions, confirm Saint Yared's establishment of three canonical EOTC chanting modes during the 6th century. This paper attempts to investigate the EOTC chants using music information retrieval (MIR) techniques. Among the research questions regarding the analysis and understanding of EOTC chants, Yaredawi YeZema Silt, namely the mode of chanting adhering to Saint Yared's standards, is of primary importance. Therefore, we consider the task of Yaredawi YeZema Silt classification in EOTC chants by introducing a new dataset and showcasing a series of classification experiments for this task. Results show that using the distribution of stabilized pitch contours as the feature representation on a simple neural network-based classifier becomes an effective solution. The musicological implications and insights of such results are further discussed through a comparative study with the previous ethnomusicology literature on EOTC chants. By making this dataset publicly accessible, we aim to promote future exploration and analysis of EOTC chants and highlight potential directions for further research, thereby fostering a deeper understanding and preservation of this unique spiritual and cultural heritage.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Frame-Level Pansori Mode Classification with Complementary Audio Representations

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A new 46-hour frame-level pansori mode corpus with multi-representation classifiers shows mode-relevant generalization and representation disagreement patterns consistent with musicological theory.

Reference graph

Works this paper leans on

29 extracted references · 26 canonical work pages · cited by 1 Pith paper

  1. [4]

    Focusing on such features also supports our subsequent discussion on the pitch distribu- tions of different chanting modes (see Section 5)

    YAREDA WI YEZEMA SILT CLASSIFICATION As a preliminary study, we only consider using time- averaged audio features (i.e., the features ignoring the in- formation lying in the temporal dimension) for Yaredawi YeZema Silt classification. Focusing on such features also supports our subsequent discussion on the pitch distribu- tions of different chanting modes...

  2. [1]

    Computa- tional Analysis of Yaredawi YeZema Silt in Ethiopian Orthodox Tewa- hedo Church Chants

    INTRODUCTION The Ethiopian Orthodox Tewahedo Church (EOTC) chants hold immense cultural and religious significance in Ethiopia, yet they are largely overlooked [1].1 The EOTC chant is believed to have originated with Saint Yared (505– 571), who composed the three EOTC chanting modes 1 The Eritrean Orthodox Tewahedo Church, which separated from the EOTC ad...

  3. [2]

    chromatic auxiliary notes around the outer fifth

    BACKGROUND OF EOTC CHANTS 2.1 Features and Performance Traditions The spiritual schools of the EOTC have several depart- ments, locally known as Guba’e bet . These departments include Nibab-bet (reading practice), Zema-bet (introduc- tory to advanced level offices chanting), Qidase-bet (or Kidase-bet, liturgical chants), Qine-bet 5 (or Kine-bet, po- etry)...

  4. [3]

    From the available audio books, we se- lected the Se’atat Zema (Horologium chant), which is part of Qidase-bet department

    DATASET The dataset was manually collected from the Eat the Book website, 7 a hub of numerous audio books for most of the teachings in the EOTC school departments, with full and partial coverage. From the available audio books, we se- lected the Se’atat Zema (Horologium chant), which is part of Qidase-bet department. All the audios selected for our datase...

  5. [5]

    Lastly, there is a clear trend that a longer input audio leads to a better per- formance

    using stabilization tends to reduce the performance gap between within-dataset and cross-dataset scenarios, and 3) the masking method can reduce this gap better than the morphetic method does, though the morphetic method typ- ically has better classification accuracy. Lastly, there is a clear trend that a longer input audio leads to a better per- formance...

  6. [6]

    the tonic of the Ge’ez mode

    ANALYSIS OF YAREDA WI YEZEMA SILT The goal of our analysis of YeZema Silt is using com- putational tools to individually identify the pitches uti- lized in the three chanting modes. Any attempt to this relies on some music theoretical assumptions. The clas- sification results presented in Section 4.3 supports two as- sumptions that facilitate the analysis...

  7. [7]

    CONCLUSION In this paper, we presented a research on a relatively under- explored music genre, the Ethiopian Orthodox Tewahedo Church (EOTC) chant, from three computational perspec- tives. First, through a rigorous data cleaning and annota- tion process, we presented a new and high-quality EOTC chant dataset, which can be extended for various music in- fo...

  8. [8]

    The oral traditions and beliefs that mainly preserve these chants for more than 1500 years should be credited properly

    ETHICS STATEMENT The chants we are working on belongs to the Ethiopian Orthodox Tewahedo Church (EOTC). The oral traditions and beliefs that mainly preserve these chants for more than 1500 years should be credited properly. This work aims to contribute on promoting the chants and finding solutions for their easy understanding. There is no intention to mod...

Show all 29 references
  1. [9]

    The sacred chant of Ethiopian monotheis- tic churches: Music in black Jewish and Christian com- munities,

    A. Kebede, “The sacred chant of Ethiopian monotheis- tic churches: Music in black Jewish and Christian com- munities,” in The Black Perspective in Music, J. South- ern, Ed. Brandeis University: JSTOR, 1980, pp. 20– 34

  2. [10]

    The musician and transmission of re- ligious tradition: The multiple roles of the Ethiopian Debtera,

    K. Shelemay, “The musician and transmission of re- ligious tradition: The multiple roles of the Ethiopian Debtera,” Journal of Religion in Africa, pp. 242-260 , vol. 22, 1992

  3. [11]

    Oral and written transmission in Ethiopian Christian chant,

    K. K. Shelemay, P. Jeffery, and I. Monson, “Oral and written transmission in Ethiopian Christian chant,” Early Music History, vol. 12, pp. 55–117, 1993

  4. [12]

    K. K. Shelemay and P. Jeffery, Ethiopian Christian Liturgical Chant: An Anthology, Part 1 to Part 3. AR Editions, Inc., 1993, vol. 1

  5. [13]

    De- tecting stable regions in frequency trajectories for tonal analysis of traditional Georgian vocal music,

    S. Rosenzweig, F. Scherbaum, and M. Müller, “De- tecting stable regions in frequency trajectories for tonal analysis of traditional Georgian vocal music,” in Inter- national Society for Music Information Retrieval Con- ference (ISMIR) , Delft, The Netherlands, 2019, pp. 352–359

  6. [14]

    Finding tori: Self-supervised learning for analyzing korean folk song,

    D. Han, R. C. Repetto, and D. Jeong, “Finding tori: Self-supervised learning for analyzing korean folk song,” in International Society of Music Information Retrieval Conference (ISMIR) , Milan, Italy, 2023, p. 440–447

  7. [15]

    KDC: an open corpus for computational research of dastg¯ahi music,

    B. Nikzat and R. Caro Repetto, “KDC: an open corpus for computational research of dastg¯ahi music,” inInter- national Society of Music Information Retrieval Con- ference (ISMIR), Bengaluru, India, 2022, pp. 321–328

  8. [16]

    Exploring the correspondence of melodic contour with gesture in raga alap singing,

    S. Nadkarni, S. Roychowdhury, P. Rao, and M. Clay- ton, “Exploring the correspondence of melodic contour with gesture in raga alap singing,” in International So- ciety of Music Information Retrieval Conference (IS- MIR), Milan, Italy, 2023, pp. 21–28

  9. [17]

    Creating a corpus of jingju (beijing opera) music and possibilities for melodic analysis,

    R. Caro Repetto and X. Serra, “Creating a corpus of jingju (beijing opera) music and possibilities for melodic analysis,” inInternational Society of Music In- formation Retrieval Conference (ISMIR) , Taipei, Tai- wan, 2014, pp. 313–318

  10. [18]

    Tuning systems of traditional georgian singing determined from a new corpus of field record- ings,

    F. Scherbaum, N. Mzhavanadze, S. Rosenzweig, and M. Müller, “Tuning systems of traditional georgian singing determined from a new corpus of field record- ings,” Musicologist, vol. 6, no. 2, pp. 142–168, 2022

  11. [19]

    Classifying cul- tural music using melodic features,

    A. Vidwans, P. Verma, and P. Rao, “Classifying cul- tural music using melodic features,” in 2020 Interna- tional Conference on Signal Processing and Commu- nications (SPCOM), 2020, pp. 1–5

  12. [20]

    Traditional education of the ethiopian orthodox church and its potential for tourism develop- ment (1975-present),

    M. Tsegaye, “Traditional education of the ethiopian orthodox church and its potential for tourism develop- ment (1975-present),” Master’s Thesis, Addia Ababa University, Addis Ababa, Ethiopia, 2011. [Online]. Available: http://etd.aau.edu.et/handle/123456789/248

  13. [21]

    Ethiopian Orthodox Tewahido Church Aquaquam Zema classification model using deep learning,

    B. T. Dagnew, “Ethiopian Orthodox Tewahido Church Aquaquam Zema classification model using deep learning,” Master’s Thesis, Bahir Dar University, Bahir Dar, Ethiopia, 2023

  14. [22]

    “Ethiopian Statistical Agency: Population and Housing Census. (2007). 2007 Census Results. retrieved from https://www.statsethiopia.gov.et/wp- content/uploads/2019/06/population-and-housing- census-2007-national_statistical.pdf,” Online, ac- cessed: April 11, 2024

  15. [23]

    Concatenative hymn syn- thesis from Yared notations,

    G. Zemedu and Y . Assabie, “Concatenative hymn syn- thesis from Yared notations,” in Advances in Nat- ural Language Processing , A. Przepiórkowski and M. Ogrodniczuk, Eds. Cham: Springer International Publishing, 2014, pp. 400–411

  16. [24]

    Computer-assisted analysis of field recordings: A case study of Georgian funeral songs,

    S. Rosenzweig, F. Scherbaum, and M. Müller, “Computer-assisted analysis of field recordings: A case study of Georgian funeral songs,” ACM Journal on Computing and Cultural Heritage , vol. 16, no. 1, dec 2022. [Online]. Available: https://doi.org/10.1145/ 3551645

  17. [25]

    Pyin: A fundamental fre- quency estimator using probabilistic threshold distri- butions,

    M. Mauch and S. Dixon, “Pyin: A fundamental fre- quency estimator using probabilistic threshold distri- butions,” in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 659–663

  18. [26]

    Torchaudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for pytorch,

    J. Hwang, M. Hira, C. Chen, X. Zhang, Z. Ni, G. Sun, P. Ma, R. Huang, V . Pratap, Y . Zhang, A. Kumar, C.- Y . Yu, C. Zhu, C. Liu, J. Kahn, M. Ravanelli, P. Sun, S. Watanabe, Y . Shi, Y . Tao, R. Scheibler, S. Cornell, S. Kim, and S. Petridis, “Torchaudio 2.1: Advancing speech...

  19. [27]

    librosa: Audio and mu- sic signal analysis in python,

    B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and mu- sic signal analysis in python,” in In Proceedings of the 14th python in science conference , vol. 8, 2015, pp. 18–25

  20. [28]

    Very deep convolutional neural networks for raw waveforms,

    W. Dai, C. Dai, S. Qu, J. Li, and S. Das, “Very deep convolutional neural networks for raw waveforms,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 421–425

  21. [29]

    Invariances and data augmentation for super- vised music transcription,

    J. Thickstun, Z. Harchaoui, D. P. Foster, and S. M. Kakade, “Invariances and data augmentation for super- vised music transcription,” in IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2018, pp. 2241–2245

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.