REVIEW 4 major objections 5 minor 1 cited by
Computational Analysis of Yaredawi YeZema Silt in Ethiopian Orthodox Tewahedo Church Chants
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Time-averaged pitch distributions alone distinguish the three Ethiopian Orthodox chant modes, with 98.11% within-dataset and 87.50% cross-dataset accuracy.
desk verdict First public EOTC chant dataset and a fair cross-dataset test, but the single-annotator labels and undisclosed GMM parameters leave the pitch-set revision partly unverified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the time-averaged pitch distribution: a histogram of frame-level pitch values (10-cent bins) extracted by pYIN from pitch contours, optionally stabilized by two methods (morphetic and masking) and calibrated for slow pitch drift by linear regression. The classifier is the M5 fully convolutional network, adapted to frequency-domain input, whose convolution makes it invariant to pitch shifting. For the musicological comparison, the paper aligns each recording's pitch distribution to an anchor by maximizing cross-correlation, averages the aligned distributions per mode, and fits a Gaussian mixture model; components with variance under 100 cents are taken as representative pitches.
What would settle it
Have an independent expert chanter re-label a random subset of the source recordings into Ge'ez, Ezil, and Araray with clip boundaries; if the network trained on the original labels performs at chance on the independently labeled subset, or if the GMM pitch sets do not re-emerge, the classification and pitch-set claims would collapse.
Extended reading notes
Core claim
The central discovery, as the paper states it, is that Yaredawi YeZema Silt is identifiable from the distribution of stabilized pitch contours: the mode is a song-level property that a pitch-shift-invariant classifier can read from time-averaged pitch statistics. Under 5-fold cross-validation the best calibrated, non-stabilized pitch distribution reaches 98.11% accuracy on the new dataset; in the cross-dataset test, the non-calibrated raw and masking-stabilized distributions reach 87.50%. Comparing the GMM-estimated pitch sets, the paper finds Ge'ez uses a flexible chain of thirds (g1–g2–g3 with g2 as a stable anchor), while Ezil has five fairly evenly spaced pitches (intervals 200–300 cents) and Araray has five unevenly spaced pitches (intervals 172–347 cents), so Ezil and Araray do not share a pitch set as previously reported.
Load-bearing premise
The load-bearing assumption is that the single annotator's manual mode labels and clip boundaries are correct; there is no second annotator or inter-annotator agreement reported, and mislabeled or mis-segmented clips would propagate into every accuracy and pitch-set result.
Editorial extensions
If this is right
- Yaredawi YeZema Silt is a song-level property that can be captured by time-averaged pitch statistics; classification accuracy improves as the input grows from 5 seconds to full length.
- Pitch-shift invariance is sufficient for mode identity, since a convolutional network that treats transposed pitch distributions as the same class performs well both within and across datasets.
- The three modes have distinctive pitch sets and pitch variances, revising the earlier claim that Ezil and Araray share a pitch set.
- The public dataset enables reproducible benchmarks for EOTC chant music information retrieval, including future temporal, lyrical, and music-generation tasks.
Reading between the lines
- Editorial extension: If mode identity is truly readable from pitch histograms, automatic indexing of the large unlabeled corpus of EOTC recordings could become feasible, helping preservation and education efforts.
- Editorial extension: The stable-region and calibration pipeline originally developed for Georgian vocal music transfers cleanly to this monophonic liturgical tradition, suggesting it may work for other oral chant traditions with flexible pitch sets.
- Editorial extension: The claimed Ezil/Araray distinction invites a perceptual test: synthesized phrases using the GMM pitch sets could be played to expert chanters to see whether the computational boundary corresponds to an audibly recognized mode change.
- Editorial extension: Because a time-averaged distribution suffices, the modes may be characterized by tonal color rather than melodic order, which predicts that pitch-order-shuffled clips within a mode should remain classifiable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a computational study of the three chanting modes (Ge’ez, Ezil, Araray) of Ethiopian Orthodox Tewahedo Church chant. The authors introduce a new dataset of 369 manually segmented and labeled single-mode clips from Se’atat Zema, extract time-averaged features (pitch-distribution histograms from pYIN with optional stable-region extraction and linear pitch calibration, plus mel-spectrogram, MFCC, and chromagram baselines), and train an adapted M5 CNN classifier. Within-dataset 5-fold CV accuracy reaches 98.11% for the calibrated non-stabilized pitch distribution on full-length audio; cross-dataset testing on Shelemay and Jeffery’s independently recorded 24-item anthology reaches up to 87.50%. The paper then aligns per-recording pitch distributions, fits GMMs, and uses Table 4 to argue that the three modes have distinct pitch sets, revising Shelemay et al.’s claim that Ezil and Araray share the same pitch set.
Significance. If the results hold, the paper would provide the first computational evidence that Yaredawi YeZema Silt is recognizable from time-averaged pitch distributions at the song level, and it would offer a concrete revision to the ethnomusicological record on EOTC chant pitch sets. The strengths include a newly described dataset with mode annotations made publicly available on GitHub, an external validation set from Shelemay and Jeffery’s anthology, use of established stable-region extraction methods, and a clearly presented classification benchmark. These are valuable contributions to the under-resourced MIR area of EOTC chant research and to comparative ethnomusicology. However, the significance is currently tempered by unresolved issues in annotation reliability, performer-confounding in the within-dataset experiment, under-specification of the GMM analysis, and the absence of uncertainty quantification on the headline accuracies.
major comments (4)
- [Section 3 (Dataset)] The dataset annotations are the linchpin of every downstream claim, but Section 3 reports that all 369 clips were segmented and labeled manually by a single annotator, with no inter-annotator agreement statistic, no second-opinion protocol, and no audit of segmentation boundaries. The 5-fold CV results in Table 2, the cross-dataset transfer to [4], and the per-mode averaged pitch distributions in Figure 3 all use these labels as ground truth. A systematic labeling bias or a segmentation error at a mode boundary would propagate directly into the classification accuracies and the GMM pitch-set estimates. The cross-dataset experiment provides independent evidence that the labels are consistent with Shelemay’s annotations, but it uses only 24 clips and does not validate the segmentation of the 369 training clips. Please report an annotation reliability protocol (e.g., a second annotator on a subset, boundary agreement measures, or a stability analysis under relabeling of ambiguous cases) and discuss how remaining ambiguity affects the reported numbers.
- [Section 3 (Dataset) and Section 4.3 (Results)] All 369 within-dataset clips are recordings by a single scholar (Section 3). The high within-dataset accuracies in Table 2 may therefore partly reflect speaker or recording characteristics rather than the chanting mode itself, especially for timbral features such as MFCC and mel-spectrogram. The cross-dataset result is the appropriate check, but only the full-length full-audio conditions are reported there, and the 24-item test set is too small to establish equal performance across feature classes. Please quantify the performer confound (if possible, via additional recordings from other chanters) or at minimum temper the within-dataset comparisons to account for this single-performer design.
- [Section 5, Table 4] The GMM fitting procedure that grounds the central musicological claim is under-specified. The text says the fitting is 'initialized by user-specified mean values' and that only components with variance below 100 cents are retained as representative pitches, but it does not disclose the number of GMM components per mode, the initial mean values, or the algorithm used to select the 'maximal weights within one octave.' These choices directly determine the number and positions of the representative pitches in Table 4, and hence the conclusion that Ezil and Araray use distinct pitch sets. If the component count or initialization is drawn from [4]’s reported pitch sets, the comparison is partly circular; if it differs between modes, the comparison is confounded. Please disclose the full initialization, report all fitted components (not only those passing the threshold), and provide sensitivity analyses over component counts and variance thresholds.
- [Section 4.3, Table 2] All classification accuracies are reported as point estimates without error bars or significance tests. The within-dataset 5-fold CV numbers have no per-fold standard deviation, and the cross-dataset column is based on only 24 instances, so the difference between 87.50% and 75.00% corresponds to three clips and can hardly be distinguished from noise. As a result, qualitative comparisons in Section 4.3 (pitch distributions 'greatly outperform' other features, stabilization 'helps for most of the cases,' masking reduces the cross-dataset gap better than morphetic) are not statistically supported. Please report per-fold variability, confidence intervals, and tests (e.g., paired tests across folds or binomial CIs for the cross-dataset results) for at least the headline comparisons.
minor comments (5)
- [Section 4.1] The pYIN time resolution is stated as '128 samples (5.8 ms)', but 128 samples at 44.1 kHz corresponds to approximately 2.9 ms; if 5.8 ms was intended, the sample count should be 256.
- [Section 4.2] The text says Shelemay’s recordings have '0.33 and 4.05 seconds' as shortest and longest durations, but Table 1 lists minute-scale durations; these units should be corrected to minutes.
- [Section 5] The sentence 'g1 and g2 form approximately a minor third g2' contains a typo; it should be 'a minor third between g1 and g2'.
- [Table 4] The note-name notation (G3, ˙g1, E4, e1, etc.) mixed with a reference pitch of 82.4 Hz is confusing; please define the octave convention and spell out how each note name relates to the cents scale.
- [Figure 3] The red background used to mark Shelemay recordings is not legible in grayscale; please use an additional visible marker or a colorbar and axis labels to make the figure self-contained.
Circularity Check
No significant circularity: classification benchmarks use independently assigned mode labels and external cross-dataset validation, and the pitch-set analysis is descriptive.
full rationale
The paper's derivation chain is self-contained rather than circular. The central classification experiment (Section 4) trains a supervised neural classifier on time-averaged pitch-distribution features to predict YeZema Silt labels that were produced by manual segmentation and listening (Section 3), not by the feature-extraction pipeline itself, so the accuracy figures in Table 2 are not true by construction. The cross-dataset experiment uses recordings and labels from Shelemay and Jeffery [4], an independent source, which provides external validation rather than a self-citation loop. The pitch-set analysis in Section 5 is a descriptive GMM summary of the same pitch-distribution features; it does not feed predictions back into the labels or into the classifier, and the claim that Ezil and Araray differ in pitch sets is a new empirical comparison against [3,4], not a renaming of those authors' results. Two concerns are noted but do not rise to demonstrable circularity under the stated standard: the GMM is 'initialized by user-specified mean values' (Section 5) without disclosing those values, and the manual mode labels have no inter-annotator reliability check. However, without evidence that the initialization means equal the reported pitch centers, or that the labels were generated from the same pitch distributions used as features, no specific reduction of the output to the input can be exhibited. These are reproducibility and validation risks, not definitional or self-citation circularity. Accordingly, the paper receives a score of 0.
Assumptions & free parameters
free parameters (3)
- GMM initial mean values =
Not reported
- Number of GMM components per mode =
Not reported
- Variance threshold for representative pitches =
100 cents
assumptions (4)
- domain assumption The manual mode labels for all 369 audio segments are correct, including the boundaries where a long recording switches modes.
- domain assumption pYIN pitch tracking reliably estimates the fundamental frequency of this monophonic vocal style with no accompaniment.
- ad hoc to paper Time-averaged pitch distributions are a sufficient representation for identifying YeZema Silt at the song level.
- domain assumption The Shelemay and Jeffery anthology recordings [4] are comparable to the proposed dataset and were produced in the same chant tradition.
Cite this review
Pith. "Pith review of Computational Analysis of Yaredawi YeZema Silt in Ethiopian Orthodox Tewahedo Church Chants." pith.science (2026). https://pith.science/paper/Y3PBKVPU
@misc{pith2026241218788,
author = {Pith},
title = {Pith review of: Computational Analysis of Yaredawi YeZema Silt in Ethiopian Orthodox Tewahedo Church Chants},
year = {2026},
howpublished = {\url{https://pith.science/paper/Y3PBKVPU}},
note = {Machine review of arXiv:2412.18788}
}
read the original abstract
Despite its musicological, cultural, and religious significance, the Ethiopian Orthodox Tewahedo Church (EOTC) chant is relatively underrepresented in music research. Historical records, including manuscripts, research papers, and oral traditions, confirm Saint Yared's establishment of three canonical EOTC chanting modes during the 6th century. This paper attempts to investigate the EOTC chants using music information retrieval (MIR) techniques. Among the research questions regarding the analysis and understanding of EOTC chants, Yaredawi YeZema Silt, namely the mode of chanting adhering to Saint Yared's standards, is of primary importance. Therefore, we consider the task of Yaredawi YeZema Silt classification in EOTC chants by introducing a new dataset and showcasing a series of classification experiments for this task. Results show that using the distribution of stabilized pitch contours as the feature representation on a simple neural network-based classifier becomes an effective solution. The musicological implications and insights of such results are further discussed through a comparative study with the previous ethnomusicology literature on EOTC chants. By making this dataset publicly accessible, we aim to promote future exploration and analysis of EOTC chants and highlight potential directions for further research, thereby fostering a deeper understanding and preservation of this unique spiritual and cultural heritage.
Forward citations
Cited by 1 Pith paper
-
Frame-Level Pansori Mode Classification with Complementary Audio Representations
A new 46-hour frame-level pansori mode corpus with multi-representation classifiers shows mode-relevant generalization and representation disagreement patterns consistent with musicological theory.
Reference graph
Works this paper leans on
-
[4]
YAREDA WI YEZEMA SILT CLASSIFICATION As a preliminary study, we only consider using time- averaged audio features (i.e., the features ignoring the in- formation lying in the temporal dimension) for Yaredawi YeZema Silt classification. Focusing on such features also supports our subsequent discussion on the pitch distribu- tions of different chanting modes...
-
[1]
Computa- tional Analysis of Yaredawi YeZema Silt in Ethiopian Orthodox Tewa- hedo Church Chants
INTRODUCTION The Ethiopian Orthodox Tewahedo Church (EOTC) chants hold immense cultural and religious significance in Ethiopia, yet they are largely overlooked [1].1 The EOTC chant is believed to have originated with Saint Yared (505– 571), who composed the three EOTC chanting modes 1 The Eritrean Orthodox Tewahedo Church, which separated from the EOTC ad...
arXiv 2024
-
[2]
chromatic auxiliary notes around the outer fifth
BACKGROUND OF EOTC CHANTS 2.1 Features and Performance Traditions The spiritual schools of the EOTC have several depart- ments, locally known as Guba’e bet . These departments include Nibab-bet (reading practice), Zema-bet (introduc- tory to advanced level offices chanting), Qidase-bet (or Kidase-bet, liturgical chants), Qine-bet 5 (or Kine-bet, po- etry)...
work page 2007
-
[3]
DATASET The dataset was manually collected from the Eat the Book website, 7 a hub of numerous audio books for most of the teachings in the EOTC school departments, with full and partial coverage. From the available audio books, we se- lected the Se’atat Zema (Horologium chant), which is part of Qidase-bet department. All the audios selected for our datase...
-
[5]
Lastly, there is a clear trend that a longer input audio leads to a better per- formance
using stabilization tends to reduce the performance gap between within-dataset and cross-dataset scenarios, and 3) the masking method can reduce this gap better than the morphetic method does, though the morphetic method typ- ically has better classification accuracy. Lastly, there is a clear trend that a longer input audio leads to a better per- formance...
-
[6]
ANALYSIS OF YAREDA WI YEZEMA SILT The goal of our analysis of YeZema Silt is using com- putational tools to individually identify the pitches uti- lized in the three chanting modes. Any attempt to this relies on some music theoretical assumptions. The clas- sification results presented in Section 4.3 supports two as- sumptions that facilitate the analysis...
-
[7]
CONCLUSION In this paper, we presented a research on a relatively under- explored music genre, the Ethiopian Orthodox Tewahedo Church (EOTC) chant, from three computational perspec- tives. First, through a rigorous data cleaning and annota- tion process, we presented a new and high-quality EOTC chant dataset, which can be extended for various music in- fo...
-
[8]
ETHICS STATEMENT The chants we are working on belongs to the Ethiopian Orthodox Tewahedo Church (EOTC). The oral traditions and beliefs that mainly preserve these chants for more than 1500 years should be credited properly. This work aims to contribute on promoting the chants and finding solutions for their easy understanding. There is no intention to mod...
Show all 29 references
-
[9]
The sacred chant of Ethiopian monotheis- tic churches: Music in black Jewish and Christian com- munities,
A. Kebede, “The sacred chant of Ethiopian monotheis- tic churches: Music in black Jewish and Christian com- munities,” in The Black Perspective in Music, J. South- ern, Ed. Brandeis University: JSTOR, 1980, pp. 20– 34
1980
-
[10]
The musician and transmission of re- ligious tradition: The multiple roles of the Ethiopian Debtera,
K. Shelemay, “The musician and transmission of re- ligious tradition: The multiple roles of the Ethiopian Debtera,” Journal of Religion in Africa, pp. 242-260 , vol. 22, 1992
1992
-
[11]
Oral and written transmission in Ethiopian Christian chant,
K. K. Shelemay, P. Jeffery, and I. Monson, “Oral and written transmission in Ethiopian Christian chant,” Early Music History, vol. 12, pp. 55–117, 1993
1993
-
[12]
K. K. Shelemay and P. Jeffery, Ethiopian Christian Liturgical Chant: An Anthology, Part 1 to Part 3. AR Editions, Inc., 1993, vol. 1
1993
-
[13]
De- tecting stable regions in frequency trajectories for tonal analysis of traditional Georgian vocal music,
S. Rosenzweig, F. Scherbaum, and M. Müller, “De- tecting stable regions in frequency trajectories for tonal analysis of traditional Georgian vocal music,” in Inter- national Society for Music Information Retrieval Con- ference (ISMIR) , Delft, The Netherlands, 2019, pp. 352–359
2019
-
[14]
Finding tori: Self-supervised learning for analyzing korean folk song,
D. Han, R. C. Repetto, and D. Jeong, “Finding tori: Self-supervised learning for analyzing korean folk song,” in International Society of Music Information Retrieval Conference (ISMIR) , Milan, Italy, 2023, p. 440–447
2023
-
[15]
KDC: an open corpus for computational research of dastg¯ahi music,
B. Nikzat and R. Caro Repetto, “KDC: an open corpus for computational research of dastg¯ahi music,” inInter- national Society of Music Information Retrieval Con- ference (ISMIR), Bengaluru, India, 2022, pp. 321–328
2022
-
[16]
Exploring the correspondence of melodic contour with gesture in raga alap singing,
S. Nadkarni, S. Roychowdhury, P. Rao, and M. Clay- ton, “Exploring the correspondence of melodic contour with gesture in raga alap singing,” in International So- ciety of Music Information Retrieval Conference (IS- MIR), Milan, Italy, 2023, pp. 21–28
2023
-
[17]
Creating a corpus of jingju (beijing opera) music and possibilities for melodic analysis,
R. Caro Repetto and X. Serra, “Creating a corpus of jingju (beijing opera) music and possibilities for melodic analysis,” inInternational Society of Music In- formation Retrieval Conference (ISMIR) , Taipei, Tai- wan, 2014, pp. 313–318
2014
-
[18]
Tuning systems of traditional georgian singing determined from a new corpus of field record- ings,
F. Scherbaum, N. Mzhavanadze, S. Rosenzweig, and M. Müller, “Tuning systems of traditional georgian singing determined from a new corpus of field record- ings,” Musicologist, vol. 6, no. 2, pp. 142–168, 2022
2022
-
[19]
Classifying cul- tural music using melodic features,
A. Vidwans, P. Verma, and P. Rao, “Classifying cul- tural music using melodic features,” in 2020 Interna- tional Conference on Signal Processing and Commu- nications (SPCOM), 2020, pp. 1–5
2020
-
[20]
Traditional education of the ethiopian orthodox church and its potential for tourism develop- ment (1975-present),
M. Tsegaye, “Traditional education of the ethiopian orthodox church and its potential for tourism develop- ment (1975-present),” Master’s Thesis, Addia Ababa University, Addis Ababa, Ethiopia, 2011. [Online]. Available: http://etd.aau.edu.et/handle/123456789/248
1975
-
[21]
Ethiopian Orthodox Tewahido Church Aquaquam Zema classification model using deep learning,
B. T. Dagnew, “Ethiopian Orthodox Tewahido Church Aquaquam Zema classification model using deep learning,” Master’s Thesis, Bahir Dar University, Bahir Dar, Ethiopia, 2023
2023
-
[22]
“Ethiopian Statistical Agency: Population and Housing Census. (2007). 2007 Census Results. retrieved from https://www.statsethiopia.gov.et/wp- content/uploads/2019/06/population-and-housing- census-2007-national_statistical.pdf,” Online, ac- cessed: April 11, 2024
2007
-
[23]
Concatenative hymn syn- thesis from Yared notations,
G. Zemedu and Y . Assabie, “Concatenative hymn syn- thesis from Yared notations,” in Advances in Nat- ural Language Processing , A. Przepiórkowski and M. Ogrodniczuk, Eds. Cham: Springer International Publishing, 2014, pp. 400–411
2014
-
[24]
Computer-assisted analysis of field recordings: A case study of Georgian funeral songs,
S. Rosenzweig, F. Scherbaum, and M. Müller, “Computer-assisted analysis of field recordings: A case study of Georgian funeral songs,” ACM Journal on Computing and Cultural Heritage , vol. 16, no. 1, dec 2022. [Online]. Available: https://doi.org/10.1145/ 3551645
2022
-
[25]
Pyin: A fundamental fre- quency estimator using probabilistic threshold distri- butions,
M. Mauch and S. Dixon, “Pyin: A fundamental fre- quency estimator using probabilistic threshold distri- butions,” in 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2014, pp. 659–663
2014
-
[26]
Torchaudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for pytorch,
J. Hwang, M. Hira, C. Chen, X. Zhang, Z. Ni, G. Sun, P. Ma, R. Huang, V . Pratap, Y . Zhang, A. Kumar, C.- Y . Yu, C. Zhu, C. Liu, J. Kahn, M. Ravanelli, P. Sun, S. Watanabe, Y . Shi, Y . Tao, R. Scheibler, S. Cornell, S. Kim, and S. Petridis, “Torchaudio 2.1: Advancing speech...
2023
-
[27]
librosa: Audio and mu- sic signal analysis in python,
B. McFee, C. Raffel, D. Liang, D. P. Ellis, M. McVicar, E. Battenberg, and O. Nieto, “librosa: Audio and mu- sic signal analysis in python,” in In Proceedings of the 14th python in science conference , vol. 8, 2015, pp. 18–25
2015
-
[28]
Very deep convolutional neural networks for raw waveforms,
W. Dai, C. Dai, S. Qu, J. Li, and S. Das, “Very deep convolutional neural networks for raw waveforms,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 421–425
2017
-
[29]
Invariances and data augmentation for super- vised music transcription,
J. Thickstun, Z. Harchaoui, D. P. Foster, and S. M. Kakade, “Invariances and data augmentation for super- vised music transcription,” in IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2018, pp. 2241–2245
2018
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.