Pith. sign in

REVIEW 4 major objections 8 minor 35 references

Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models

T0 review · 4 major / 8 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Self-supervised speech models pre-trained on high-resource languages can recognize lexical tones in Angami, Ao, and Mizo, with the best accuracy coming from the middle transformer layers.

desk verdict A useful first probing benchmark for three under-resourced Tibeto-Burman tone languages, but the speaker-independent split may not isolate tone from word identity. read the letter →

arxiv 2506.03606 v1 pith:LXGFPVAN submitted 2025-06-04 eess.AS cs.AIcs.CLeess.SP

classification eess.AScs.AIcs.CLeess.SP
keywords lexicaltonerecognitionself-supervisedspeechmodelsWav2Vec2.0low-resourcelanguagesAngamiAoMizolayer-wiseprobing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that embeddings from self-supervised speech models built for high-resource languages can be repurposed for lexical tone recognition in three low-resource Tibeto-Burman languages of North-East India: Angami, Ao, and Mizo, without any fine-tuning. Using four base Wav2vec2.0 models pre-trained on English, Mandarin, and Vietnamese, and a linear classifier on per-layer averaged embeddings, the authors find that tone recognition is best for Mizo and worst for Angami, with the strongest results for the model pre-trained on the largest Mandarin dataset. The central claim is that the middle transformer layers, roughly layers 4 to 7, carry the most tone information in all three languages, regardless of the pre-training language. The paper also argues that tone inventory shape and dialectal variation, not just model choice, set the ceiling: Mizo's mix of level and contour tones helps, Angami's four level tones and multiple dialects hurt, and dialect-independent evaluation drops accuracy markedly. If right, this would offer a low-cost route to tone-aware speech technology for under-resourced tonal languages without hand-designed pitch features.

What carries the argument

The operative mechanism is layer-wise probing of Wav2vec2.0 base models: for each tone-bearing unit, the frame-level hidden representations from each transformer layer are averaged into a single feature vector, and a linear SVM classifies that vector into tone categories under speaker-independent cross-validation. This design isolates the information contained in the embeddings themselves, because the classifier is deliberately simple. The central finding is that a single middle layer, typically the fourth through seventh, makes tones linearly separable in all three languages, and that this layer-wise pattern is stable across models pre-trained on tonal and non-tonal languages.

What would settle it

Take the same embeddings and the same linear probe, replace the manual labels with a second independent annotation pass for a subset of tokens, and measure middle-layer F1 on the new labels; if label agreement is low or the F1 drops to chance, the claim that middle layers encode tone is an artifact of annotation noise.

Watch

Extended reading notes

Core claim

On the paper's own terms, tone recognition can be read directly out of frozen SSL embeddings. With a linear SVM trained on per-layer averaged Wav2vec2.0 embeddings, the chinese-wav2vec2-base model reaches macro F1-scores of about 61.2 percent for Angami, 73.9 percent for Ao, and 88.03 percent for Mizo. Performance rises from lower layers, peaks in the middle layers, and declines in higher layers for every model and language, indicating that the middle transformer layers specialize in tone-discriminative information. Pre-training on a tonal language with a large corpus helps transfer to other tonal languages, though the English-pretrained wav2vec2-base outperforms the smaller Mandarin model, so data scale and model quality matter too. Language-internal factors also matter: Mizo's mixture of level and contour tones is recognized best, Angami's four-level system with dialectal variation is recognized worst, and removing dialect information from training drops Angami F1 from about 61 to 38 percent and Ao F1 from about 75 to 48 percent.

Load-bearing premise

The results stand on the hand-marked tone labels for Angami, Ao, and Mizo and on the published tone inventories being accurate, and the paper reports no check of whether different annotators agree; if labels are wrong or inconsistent across dialects, every reported accuracy and the language ranking would change.

Editorial extensions

If this is right

  • Tone labels in these languages can be predicted from frozen SSL embeddings with substantial accuracy, so tone annotation or tone-based downstream tasks do not necessarily require fine-tuning on labelled target-language data.
  • Systems built on these models should read tone information from layers 4 through 7 rather than from the final layer, because the final layers are less tone-specific.
  • Pre-training on a tonal language with a large corpus helps transfer to other tonal languages, though model quality and data size matter as much as the tonal status of the pre-training data.
  • Linguistic factors set an upper bound: a four-level tone system with dialectal variation is harder than a three-level system, which is harder than a mixed level-and-contour system with more homogeneous data.
  • Dialect-independent evaluation drops by roughly 20 to 27 F1 points, showing that dialect variation, not only tone category, is encoded in the embeddings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if the middle-layer result generalizes beyond these three languages, layer selection alone could improve tone-related probing in any related low-resource language with unlabeled audio, before any model training or fine-tuning is attempted.
  • Editorial inference: the large dialect drop suggests the embeddings encode dialect-specific tone realizations, so a practical extension would be to train a small dialect classifier first and use its predictions to condition or normalize tone classification.
  • Editorial inference: the reported confusion between Mizo low and falling tones points to a concrete experiment: time-normalizing f0 contours before embedding extraction should make those two tones more separable if the model's difficulty is range rather than shape.
  • Editorial inference: the paper leaves open whether a nonlinear probe or light fine-tuning on a few labelled tones would push Angami past its observed ceiling, which would separate a limitation of linear classifiers from a limitation of the embeddings themselves.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 8 minor

Summary. This paper reports layer-wise tone-recognition experiments using four wav2vec2.0 base models pre-trained on Mandarin Chinese, Vietnamese, and English, evaluated on Angami, Ao, and Mizo. Embeddings from each transformer layer are averaged over tone-bearing units and classified with a linear SVM under speaker-independent and dialect-independent cross-validation. The paper's main claims are that Mizo tones are recognized best and Angami worst, that middle transformer layers (about layers 4-8) are the most important for tone recognition across all pre-training languages, and that tone inventory and dialectal variation modulate the results.

Significance. The study targets an important and under-resourced problem: transfer of SSL representations to lexical-tone recognition in low-resource Tibeto-Burman languages. Its strengths include a systematic layer-wise comparison across four models and three languages, per-tone analyses, consideration of dialect variation, and public code. If the central empirical claims survive, the paper would be a useful reference for practitioners and for probing studies of tonal information in SSL models. However, the interpretation hinges on whether the classifiers generalize to unseen words, an issue the current experimental design does not address, and several summary-level statements are inconsistent with the reported layer-wise numbers.

major comments (4)
  1. [§2.4, §2.2] The speaker-independent CV splits only by speakers, not by lexical items. The corpora are built from closed word lists (Ao: 40 target words per dialect; Mizo: 64 trisyllabic tone combinations with five phrases; Angami: a closed set of monosyllabic/disyllabic words), so every test token's word or phrase has also been seen in training. Because tone is lexically determined, a linear SVM can decode the lexical item from the embedding and map it to its tone without learning a general tone-to-acoustic mapping. This is a real risk here, particularly because the paper itself cites [34] to the effect that middle/deeper layers encode lexical information; the same mechanism would produce the observed middle-layer peak. The dialect-independent experiments do not remove the confound, since they still train and test on the same closed word lists. To support the claim that SSL embeddings encode tone, the authors should add word-disjoint or phrase-disjoint splits (e.g., leave-target-words-out, or at least a per-word generalization analysis).
  2. [§3.1, §3.2, §4] Several summary-level claims about model ranking are contradicted by the reported layer-wise numbers. Section 3.1 states that 'For all the languages, chinese-wav2vec2-base yields the best performance' and reports an Ao maximum of 73.9%, but Section 3.2 reports that for Ao, wav2vec2-base-vi achieves 74.93% at layer 8, exceeding the chinese-wav2vec2-base peak of 73.91% at layer 6. Similarly, the conclusion that 'Models pre-trained on tonal languages consistently outperform non-tonal counterparts' is contradicted by the Mizo result in Section 3.2, where the English wav2vec2-base reaches 80.61% at layer 4, higher than the mandarin-wav2vec2 model. These inconsistencies must be reconciled before the model-ranking and pre-training-language conclusions can be accepted.
  3. [§1, §3.2] The introduction states that tone recognition is 'more effective in the middle layers, specifically between the 4th and the 6th layers', but the results in Section 3.2 include peaks at layer 7 for Angami (wav2vec2-base-vi, 59.80%), layer 8 for Ao (wav2vec2-base-vi, 74.93%), and layer 7 for Mizo (wav2vec2-base-vi). The '4th to 6th' wording is too specific and should be replaced by a broad middle-band statement, or the 4-6 window should be justified with a significance test or a consistent definition of 'middle' across models and languages.
  4. [§3.1-§3.4] The F1 scores are presented without any chance-level baseline or comparison with the authors' earlier f0-based tone recognition systems [1-3]. For four-way tone classification (Angami, Mizo) chance accuracy is 25%, and for Ao's three tones it is 33%. Providing these baselines is necessary to support statements such as 'considerable accuracy' and to interpret the layer-wise curves, and the earlier f0-based systems offer a natural reference point for assessing the claimed advantage of SSL embeddings.
minor comments (8)
  1. [§2.3] The first model is called 'wav2vec2-base2' in the bullet list, but the rest of the paper and the HuggingFace link refer to 'wav2vec2-base'; please make the naming consistent.
  2. [§2.3] The sentence 'After computing the embeddings using the milliseconds duration threshold' should read 'using the 50 ms duration threshold'.
  3. [§2.3] The feature vector is described as 746-dimensional, whereas wav2vec2 base has a hidden size of 768; please clarify what the reduced dimensionality corresponds to (for example, removal of special tokens or padding frames).
  4. [§2.2] The Ao corpus description gives 2,880 utterances per dialect, which over three dialects is 8,640, while the text reports a total of 25,920 tokens; the arithmetic needs to be reconciled.
  5. [§2.2] The sentence about the Mizo speakers, 'Their average is 22 years', should read 'Their average age is 22 years'.
  6. [Figure 2] The caption says 'classification accuracies (in percentage)', but the color bars are on a 0-1 scale; please make the units consistent.
  7. [References] References [29] and [35] are the same article (Tzudir, Sarmah, and Prasanna 2021, JASA); the duplicate should be merged.
  8. [§2.2] The tone labels are described as 'meticulously annotated using Praat', but no inter-annotator agreement or validation against an external source is reported; please add at least a summary of label reliability, since all conclusions depend on these labels.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper reports direct measurements and does not derive any headline result from a fitted quantity or self-citation chain.

full rationale

This is an empirical measurement study rather than a derivation. Every headline result — per-language F1 scores, the layer-wise peak for each SSL model, tone-wise heatmaps, and the SI/DI cross-validation comparisons — is read directly off a fixed evaluation protocol (linear SVM on averaged layer embeddings, Section 2.4) and is not produced from any fitted parameter that is then renamed as a prediction. The tone inventories and dialect groupings cited from the authors' prior work ([28]-[32]) serve as external ground-truth definitions for the labels; the paper does not use the recognition scores to infer those inventories, so importing them is not circular. The SI-CV design may share lexical items across speaker folds, and the Praat tone labels lack reported inter-annotator agreement, but those are data-validation or evaluation-confound concerns external to the claim chain, not cases where an output equals an input by construction. No self-citation is load-bearing for the layer-wise or language-ranking claims, which are independently reproducible from the reported protocol and data splits.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The paper is a measurement study: no constants are fitted to produce the result, and no new entities are introduced. The free parameters are limited to the hand-chosen 0.05 s token duration threshold and the unreported SVM regularization value, both of which shift the numeric results. The load-bearing assumptions are the accuracy of the manual tone labels, the correctness of the tone inventories taken from the authors' prior work, and the methodological choices of mean-pooling frame embeddings and using a linear probe. These are all domain assumptions rather than ad hoc inventions, and none of them alone would invalidate the qualitative mid-layer pattern if mildly violated.

free parameters (2)
  • minimum tone token duration threshold = 0.05 s
    Hand-chosen cutoff to keep temporal context; removes 6% of Angami, 33% of Ao and 28% of Mizo tokens (Table 1), so all reported accuracies depend on it; no sensitivity analysis is given.
  • linear SVM regularization C = not reported
    Section 2.4 specifies a linear SVM but no C value; the default may be used, but the choice is invisible and could shift the F1 numbers.
assumptions (4)
  • domain assumption Manual Praat tone annotations and tone-boundary labels for Angami, Ao, and Mizo are accurate.
    Section 2.2 states boundaries were annotated using Praat; no inter-annotator reliability or validation is provided, and every F1 score depends on these labels.
  • domain assumption The tonal inventories and dialect groupings of the three languages, drawn from refs [28-32] by the same research group, are correct.
    Section 2.1 describes the tones using these references; the interpretation of Mizo's falling/low confusion and Angami's difficulty relies on these inventories.
  • domain assumption Averaging frame-level embeddings over a tone segment preserves the tone-discriminative information needed for linear classification.
    Section 2.3 averages corresponding frames within a segment into a single vector; the paper does not test max-pooling, attention, or frame-level classification, so the entire pipeline rests on this pooling choice.
  • domain assumption Linear separability of tone classes in the embedding space is adequate to measure how well tone information is encoded.
    Section 2.4 uses a linear SVM deliberately to attribute performance to embeddings; this conflates encoding quality with linear readout capacity, but it is a common probing assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models." pith.science (2026). https://pith.science/paper/LXGFPVAN

@misc{pith2026250603606,
  author       = {Pith},
  title        = {Pith review of: Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LXGFPVAN}},
  note         = {Machine review of arXiv:2506.03606}
}
read the original abstract

This study explores the use of self-supervised learning (SSL) models for tone recognition in three low-resource languages from North Eastern India: Angami, Ao, and Mizo. We evaluate four Wav2vec2.0 base models that were pre-trained on both tonal and non-tonal languages. We analyze tone-wise performance across the layers for all three languages and compare the different models. Our results show that tone recognition works best for Mizo and worst for Angami. The middle layers of the SSL models are the most important for tone recognition, regardless of the pre-training language, i.e. tonal or non-tonal. We have also found that the tone inventory, tone types, and dialectal variations affect tone recognition. These findings provide useful insights into the strengths and weaknesses of SSL-based embeddings for tonal languages and highlight the potential for improving tone recognition in low-resource settings. The source code is available at GitHub 1 .

Figures

Figures reproduced from arXiv: 2506.03606 by the authors.

Figure 1
Figure 1. Layer-wise comparison of F1-score for three languages (Angami, Ao, Mizo) across various models. The [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Heatmaps illustrating tone classification accuracies (in percentage) across layers for Angami, Ao, and Mizo. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

35 extracted references · 34 canonical work pages

  1. [34]

    Layer-wise analysis of a self-supervised speech representation model

    Ankita Pasad, Ju-Chieh Chou, and Karen Livescu. Layer-wise analysis of a self-supervised speech representation model. In2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 914–921. IEEE, 2021

  2. [1]

    Lexical Tone Recognition in Mizo Using Acoustic-Prosodic Features

    Parismita Gogoi, Abhishek Dey, Wendy Lalhminghlui, Priyankoo Sarmah, and S R Mahadeva Prasanna. Lexical Tone Recognition in Mizo Using Acoustic-Prosodic Features. InProceedings of The 12th Language Resources and Evaluation Conference, pages 6460–6463, Marseille, France, May 2020. European Language Resources Association

  3. [2]

    Parismita Gogoi, Moakala Tzudir, Priyankoo Sarmah, and S. R. M. Prasanna. Automatic tone recognition of Ao language. InProceedings of the 10th International Conference on Speech Prosody, pages 1005–1008, 2020

  4. [3]

    Gogoi, S

    P. Gogoi, S. Kalita, W. Lalhminghlui, P. Sarmah, and S. R. M. Prasanna. Learning Mizo Tones from F0 Contours Using 1D-CNN. InSpeech and Computer: 23rd International Conference, SPECOM 2021, pages 214–225, St. Petersburg, Russia, 2021

  5. [4]

    Detection of Mizo Tones

    Biswajit Dev Sarma, Priyankoo Sarmah, Wendy Lalhminghlui, and S R Mahadeva Prasanna. Detection of Mizo Tones. InInterspeech, pages 934–937, 2015

  6. [5]

    Disambiguation of Isolated Manipuri Tonal Contrast Word Pairs Using Acoustic Features.ACM Transactions on Asian and Low-Resource Language Information Processing, 23(3):1–18, 2024

    Thiyam Susma Devi and Pradip K Das. Disambiguation of Isolated Manipuri Tonal Contrast Word Pairs Using Acoustic Features.ACM Transactions on Asian and Low-Resource Language Information Processing, 23(3):1–18, 2024

  7. [6]

    A hidden Markov model applied to Chinese four- tone recognition

    Xi-Xian Chen, Chang-Nian Cai, Peng Guo, and Ying Sun. A hidden Markov model applied to Chinese four- tone recognition. InICASSP’87. IEEE International Conference on Acoustics, Speech, and Signal Processing, volume 12, pages 797–800. IEEE, 1987. 6 Running Title for Header

  8. [7]

    Hidden markov model for Mandarin lexical tone recognition

    W-J Yang, J-C Lee, Y-C Chang, and H-C Wang. Hidden markov model for Mandarin lexical tone recognition. IEEE Transactions on Acoustics, Speech, and Signal Processing, 36(7):988–992, 1988

Show all 35 references
  1. [8]

    Tone recognition of polysyllabic words in Mandarin speech.Computer Speech & Language, 3(3):253–264, 1989

    Lih-Cherng Liu, Wu-Ji Yang, Hsiao-Chuan Wang, and Yueh-Chin Chang. Tone recognition of polysyllabic words in Mandarin speech.Computer Speech & Language, 3(3):253–264, 1989

  2. [9]

    Using tone information in Cantonese continuous speech recognition.ACM Transactions on Asian Language Information Processing (TALIP), 1(1):83–102, 2002

    Tan Lee, Wai Lau, Yiu Wing Wong, and PC Ching. Using tone information in Cantonese continuous speech recognition.ACM Transactions on Asian Language Information Processing (TALIP), 1(1):83–102, 2002

  3. [10]

    Tone recognition of continuous cantonese speech based on support vector machines.Speech Communication, 45(1):49–62, 2005

    Gang Peng and William S-Y Wang. Tone recognition of continuous cantonese speech based on support vector machines.Speech Communication, 45(1):49–62, 2005

  4. [11]

    Mandarin tone recognition by multi-layer perceptron

    P-C Chang, S-W Sun, and S-H Chen. Mandarin tone recognition by multi-layer perceptron. InInternational Conference on Acoustics, Speech, and Signal Processing, pages 517–520. IEEE, 1990

  5. [12]

    Highly accurate Mandarin tone classification in the absence of pitch information

    Neville Ryant, Malcolm Slaney, Mark Liberman, Elizabeth Shriberg, and Jiahong Yuan. Highly accurate Mandarin tone classification in the absence of pitch information. InProceedings of Speech Prosody, 2014

  6. [13]

    Mandarin tone classification without pitch tracking

    Neville Ryant, Jiahong Yuan, and Mark Liberman. Mandarin tone classification without pitch tracking. In2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014

  7. [14]

    Tone classification in Mandarin Chinese using convolu- tional neural networks

    Charles Chen, Razvan Bunescu, Li Xu, and Chang Liu. Tone classification in Mandarin Chinese using convolu- tional neural networks. InInterspeech 2016, pages 2150–2154, 2016

  8. [15]

    Q. Gao, S. Sun, and Y . Yang. Tonenet: A CNN model of tone Classification of Mandarin Chinese. InInterspeech 2019, pages 3367–3371, Graz, Austria, 2019. ISCA

  9. [16]

    Tang and M

    J. Tang and M. Li. End-to-End Mandarin Tone Classification with Short Term Context information. In2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), pages 878–883, Tokyo, Japan, 2021. IEEE

  10. [17]

    Huang, K

    H. Huang, K. Wang, Y . Hu, and S. Li. Encoder-Decoder Based Pitch Tracking and Joint Model Training for Mandarin Tone Classification. In2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6943–6947, Toronto, Canada, 2021. IEEE

  11. [18]

    Y . F. Liu and X. L. Lu. Learning and consolidating the contextualized contour representations of tones from f0 sequences and durational variations via transformers.The Journal of the Acoustical Society of America, 156(5):3353–3372, 2024

  12. [19]

    Liu and X.-L

    Y .-F. Liu and X.-L. Lu. J-tonenet: A transformer-based encoding network for improving tone classification in continuous speech via f0 sequences. InInterspeech 2023, pages 2138–2142, Dublin, Ireland, 2023. ISCA

  13. [20]

    Wav2vec 2.0: A framework for self-supervised learning of speech representations.Advances in neural information processing systems, 33:12449– 12460, 2020

    Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. Wav2vec 2.0: A framework for self-supervised learning of speech representations.Advances in neural information processing systems, 33:12449– 12460, 2020

  14. [21]

    Hubert: Self-supervised speech representation learning by masked prediction of hidden units

    Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrah- man Mohamed. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM transactions on audio, speech, and language processing, 2...

  15. [22]

    A layer-wise analysis of Mandarin and English suprasegmentals in SSL speech models

    Anton de la Fuente and Dan Jurafsky. A layer-wise analysis of Mandarin and English suprasegmentals in SSL speech models. InProc. Interspeech 2024, pages 1290–1294, 2024

  16. [23]

    What can an accent identifier learn? probing phonetic and prosodic information in a wav2vec2-based accent identification model

    Mu Yang, Ram C Shekar, Okim Kang, and John H Hansen. What can an accent identifier learn? probing phonetic and prosodic information in a wav2vec2-based accent identification model. InProc. Interspeech 2024, 2023

  17. [24]

    Self-supervised speech representation learning: A review.IEEE Journal of Selected Topics in Signal Processing, 16(6):1179–1210, 2022

    Abdelrahman Mohamed, Hung-yi Lee, Lasse Borgholt, Jakob D Havtorn, Joakim Edin, Christian Igel, Katrin Kirchhoff, Shang-Wen Li, Karen Livescu, Lars Maaløe, et al. Self-supervised speech representation learning: A review.IEEE Journal of Selected Topics in Signal Processing, 16(...

  18. [25]

    Jing Zhao and Wei-Qiang Zhang. Improving automatic speech recognition performance for low-resource languages with self-supervised models.IEEE Journal of Selected Topics in Signal Processing, 16(6):1227–1241, 2022

  19. [26]

    Performance analysis of speech encoders for low-resource slu and asr in tunisian dialect

    Salima Mdhaffar, Haroun Elleuch, Fethi Bougares, and Yannick Estève. Performance analysis of speech encoders for low-resource slu and asr in tunisian dialect. InProceedings of The Second Arabic Natural Language Processing Conference, pages 130–139, 2024

  20. [27]

    Encoding of lexical tone in self-supervised models of spoken language

    Gaofei Shen, Michaela Watkins, Afra Alishahi, Arianna Bisazza, and Grzegorz Chrupała. Encoding of lexical tone in self-supervised models of spoken language. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Hum...

  21. [28]

    V oice onset time in Angami.Linguistics of the Tibeto-Burman Area, 47(1):32–56, 2024

    Viyazonuo Terhiija and Priyankoo Sarmah. V oice onset time in Angami.Linguistics of the Tibeto-Burman Area, 47(1):32–56, 2024

  22. [29]

    Analysis and modeling of dialect information in Ao, a low resource language.The Journal of the Acoustical Society of America, 149(5):2976–2987, 2021

    Moakala Tzudir, Priyankoo Sarmah, and SR Prasanna. Analysis and modeling of dialect information in Ao, a low resource language.The Journal of the Acoustical Society of America, 149(5):2976–2987, 2021

  23. [30]

    PhD thesis, Indian Institute of Technology Guwahati, 2023

    Moakala Tzudir.Automatic Dialect Identification in Ao, a Low Resource Language. PhD thesis, Indian Institute of Technology Guwahati, 2023

  24. [31]

    A Preliminary Acoustic Study of Mizo Vowels and Tones.Journal of the Acoustical Society of India, 37(3):121–129, 2010

    Priyankoo Sarmah and Caroline R Wiltshire. A Preliminary Acoustic Study of Mizo Vowels and Tones.Journal of the Acoustical Society of India, 37(3):121–129, 2010

  25. [32]

    PhD thesis, Indian Institute of Technology Guwahati, 2023

    Wendy Lalhminghlui.The Phonetics and Phonology of Tones in Mizo. PhD thesis, Indian Institute of Technology Guwahati, 2023

  26. [33]

    Praat: doing phonetics by computer [computer program].Version, 6:1, November 2020

    Boersma, Paul and Weenink, David. Praat: doing phonetics by computer [computer program].Version, 6:1, November 2020

  27. [35]

    Analysis and modeling of dialect information in Ao, a low resource language.The Journal of the Acoustical Society of America, 149(5):2976–2987, 2021

    Moakala Tzudir, Priyankoo Sarmah, and SR Mahadeva Prasanna. Analysis and modeling of dialect information in Ao, a low resource language.The Journal of the Acoustical Society of America, 149(5):2976–2987, 2021. 8

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.