REVIEW 4 major objections 5 minor 40 references
V(is)owel: An Interactive Vowel Chart to Understand What Makes Visual Pronunciation Effective in Second Language Learning
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read An interactive vowel chart that plots a learner's tongue position against a target speaker's line gives untrained learners a tangible goal and more practice than audio-only feedback.
desk verdict A genuinely novel interactive vowel chart with a think-aloud study, but the feedback accuracy errors are load-bearing and the abstract overstates the engagement result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrier of the argument is V(is)owel, an interactive vowel chart that maps the first two formant frequencies onto a trapezoidal chart labeled with tongue height and front-back position. A projective (homography) transformation, calibrated on four English corner vowels, maps each speaker's idiosyncratic vowel space onto a fixed chart; an autocorrelation-based boundary detector finds vowel onsets and offsets; a formant extractor, selected by a smoothness metric, supplies the first and second formants; and the resulting time-varying line is paired with audio playback so learners can hear and see the same token. This machinery matters because it is what turns an untrained learner's pronunciation into a visible, comparable target line, and the think-aloud protocol reveals how learners use that line while adjusting their speech.
What would settle it
Run the same think-aloud comparison with a corrected extraction pipeline that keeps vowel-offset mean absolute error under 10 ms and second-formant mean absolute error under 100 Hz; if the practice advantage and the goal-oriented comments disappear, the reported effect depended on inaccurate visuals. A more direct test would pair the chart with ultrasound tongue imaging to check whether the plotted line matches the measured tongue position for the same recordings.
Extended reading notes
Core claim
On its own terms, V(is)owel is presented as an interactive vowel chart that accepts vowels in any word context, shows diphthong trajectories over time, and is used, as far as the authors know, in the first visual pronunciation-training study to capture learners' real-time interpretations as they adjust their pronunciation. The paper's central finding is that the plotted line acted as an external goal: all eight participants referenced the Spanish speaker's line when deciding whether to re-record, whereas with audio-only they relied on their own judgment, and participants recorded more per word with V(is)owel than with audio-only (averages of 2.27 versus 1.64 recordings per word, with a Friedman test p=0.0339). The authors conclude that the visualization provides actionable feedback because it maps tongue movement onto a physical chart, and they recommend that future pronunciation tools include explicit anatomical feedback tied to physical movement, while acknowledging that algorithmic inaccuracies, especially in the second formant and in vowel offset times, sometimes showed users incorrect lines.
Load-bearing premise
The study assumes the plotted vowel lines are accurate enough that learners' interpretations and extra practice were responses to their actual tongue position rather than to extraction artifacts; the paper's own measurements of second-formant and vowel-offset errors leave room for that assumption to fail.
Editorial extensions
If this is right
- Designers of pronunciation feedback for phonetically untrained learners should include explicit anatomical visual feedback that maps directly onto physical movement, rather than relying on correctness colors or percentages.
- A visible target-speaker line can increase practice volume even when learners distrust the accuracy of the visualization, because it gives them a concrete thing to work toward.
- Vowel-focused visualization did not show a clear blinder effect in this study: participants continued to mention consonants, pitch, and overall pronunciation while practicing with V(is)owel.
- Extracting vowels in any word context, including rhotics and diphthongs, is feasible enough for a real-time practice tool, though second-formant and vowel-offset accuracy need improvement before the feedback can be fully trusted.
- Learners tended to trust the visualization over their own hearing, which means the accuracy of the plotted line is not just a technical detail but a central design constraint.
Reading between the lines
- A corrected extraction pipeline, with vowel-offset mean absolute error under 10 ms and second-formant mean absolute error under 100 Hz, would allow a cleaner test of whether the visual-goal effect is caused by accurate articulation feedback or by the mere presence of a target line.
- The same goal-tracking mechanism could plausibly generalize to other articulatory dimensions, such as lip rounding or nasalization, by adding equivalent visual dimensions to a chart; the paper gestures toward this in its discussion of extending V(is)owel to other languages.
- Because learners trusted the visualization over their own ears, small algorithmic errors may be more harmful in a visual tool than in audio-only training, so accuracy requirements should be treated as a design constraint rather than a backend afterthought.
- A longitudinal study would be needed to separate the motivating effect of novelty from the motivating effect of the visual goal itself, a limitation the authors themselves note.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces V(is)owel, an interactive vowel chart that plots a learner's first and second formants over time and plays back the associated audio, with the goal of providing actionable pronunciation feedback to phonetically untrained second-language learners. The authors report a within-subject study with eight American English speakers practicing Spanish minimal pairs under two conditions, V(is)owel and an audio-only control, using think-aloud protocols, practice counts, NASA-TLX, SUS, and exit interviews. They also report a technical evaluation of their vowel-boundary and formant-extraction algorithms against manual Praat annotations. The main claimed results are that participants practiced more with V(is)owel, that all participants used the visualized target as a goal, and that the tool provided actionable feedback that audio alone did not.
Significance. If the central claims were fully supported, the paper would make a useful contribution to computer-assisted pronunciation training by showing which aspects of articulatory visualizations help untrained learners and by demonstrating a system that works beyond isolated hVd contexts. The study has notable strengths: a counterbalanced within-subject design, think-aloud data that capture real-time interpretation, an iterative design process with pretesting, and a technical evaluation that openly benchmarks extraction accuracy against an external manual annotation standard. The paper also reports accuracy metrics transparently, which is commendable. However, the significance is undercut by a factual inconsistency in the engagement data and by accuracy levels that are acknowledged as faulty in the Discussion; both issues bear directly on the claim that V(is)owel provides actionable feedback. The absence of any pronunciation improvement measure further limits the scope of the conclusions.
major comments (4)
- [Abstract, §5.1.3, Table 3] The abstract states that "all participants practicing words with V(is)owel more than with audio-only," but Table 3 shows that P7 recorded fewer times per word with V(is)owel (1.875) than with audio-only (2.250). This contradiction invalidates the "all participants" claim and should be corrected in the abstract and in Section 5.1.3. The Friedman test result may still hold, but the summary of the finding must not claim unanimity where the data show an exception.
- [§5.3.1, §5.3.2, §5.4, §5.5] The technical evaluation reports F2 mean absolute error of 116.3 Hz against a 100 Hz tolerance, vowel-offset MAEs up to 265 ms (Table 4), and large formant-error rates of 5–26% per participant (Table 6). Section 5.4 attributes the higher practice counts to participants having "guidance on what to change," and Section 5.5 states that participants "put greater emphasis on the visualization than on their perception of pronunciation." If a substantial fraction of the plotted lines are misleading, then the behavioral and think-aloud evidence is partly a reaction to incorrect visuals, and the claim that V(is)owel provides "actionable feedback" is not supported in its current form. The paper needs either an accuracy-conditioned analysis (e.g., comparing engagement and interpretations on trials with small versus large formant errors) or a substantially weakened claim that separates perceived guidance from verified guidance.
- [§7, Abstract] The study measures engagement, self-reported confidence, and think-aloud interpretations, but it does not measure pronunciation improvement before and after practice. Nevertheless, the conclusion in Section 7 states that "personalized feedback can have a significant impact on a learners' pronunciation in a second language," and the abstract opens with a general claim that "visual feedback speeds up learners' improvement." These causal learning claims are not supported by the reported data. The paper should either add a learning-outcome measure or explicitly reframe the conclusions as being about engagement and perceived actionability rather than demonstrated pronunciation gains.
- [Discussion] The Discussion states that "based on offset times, V(is)owel currently gives faulty feedback," which directly conflicts with the abstract's claim that "V(is)owel is effective at providing actionable feedback." This tension is not merely stylistic: faulty feedback can be actionable in the narrow sense that users can act on it, but the paper does not establish that acting on it moves learners toward the target. The authors should reconcile these statements, either by defining the limited sense in which feedback is actionable or by removing the effectiveness claim.
minor comments (5)
- [§3.2.2] The sentence "We use Praat's format extraction library" should read "formant extraction library."
- [§5.1.2, Familiarity] The text says "Pt expressed how this made it more difficult to reach their goals," where "Pt" appears to be a placeholder for a participant identifier; please replace it with the intended label.
- [§5.1.2, Using the Spanish Speaker's Line as a Goal] The phrase "far away from the spa line" appears to be a typo; it should presumably be "Spanish speaker's line."
- [Table 5 caption] The caption says "Bold indicates an offset that is outside allowed tolerance," but the table contains formant frequency MAE values and no bold entries; the caption appears to have been copied from Table 4 and should be corrected.
- [Appendix A.1] The Friedman test result would be more informative if the test statistic and effect size were reported alongside the p-value, given the small sample size of eight participants.
Circularity Check
No significant circularity: the study's claims are empirical, measured against external benchmarks, and no prediction reduces to a fitted input or self-citation.
full rationale
The paper contains no derivation chain that reduces to its own inputs. Its central claims are empirical observations from an eight-participant within-subject study: participants recorded more words with V(is)owel than with audio-only (Friedman p=0.0339), and think-aloud comments indicated that the visual target served as a goal. These are measurements, not values predicted from a fitted model. The only fitted component is the speaker-calibration homography, which maps four corner vowels to a fixed trapezoid; it is a speaker normalization, and the plotted positions of practice vowels are measured formant values transformed by that map, not predictions of the outcome claim. The technical evaluation is benchmarked externally against manual Praat annotations (Tables 4-6), and the paper explicitly reports large errors, including F2 MAE above tolerance and offsets up to 265 ms, rather than hiding them. The discussion even concedes that V(is)owel 'currently gives faulty feedback.' The conclusion that feedback is 'actionable' is inferred from participant statements and practice counts; whether that inference is valid despite the accuracy faults is a correctness and construct-validity concern, not a circularity concern. There are no load-bearing self-citations, no imported uniqueness theorems, and no renamed known result. Accordingly, no circular step can be exhibited, and the appropriate score is 0.
Assumptions & free parameters
free parameters (3)
- Autocorrelation threshold for vowel onset and offset =
not reported (tuned by trial and error)
- Vowel chart trapezoid ratio =
height 1, top width 1.33, bottom width 0.833
- Formant-count smoothness metric coefficient =
0.8^(i-2)
assumptions (4)
- domain assumption Formant frequencies F1 and F2 map monotonically to tongue height and tongue frontness or backness.
- domain assumption Manual Praat formant annotations are an acceptable ground truth for evaluating extraction accuracy.
- domain assumption Think-aloud utterances and exit interviews reflect participants' underlying cognitive processes.
- ad hoc to paper The autocorrelation threshold identifies vowel boundaries correctly across speakers and words.
Cite this review
Pith. "Pith review of V(is)owel: An Interactive Vowel Chart to Understand What Makes Visual Pronunciation Effective in Second Language Learning." pith.science (2026). https://pith.science/paper/EPAYPG7V
@misc{pith2026250706202,
author = {Pith},
title = {Pith review of: V(is)owel: An Interactive Vowel Chart to Understand What Makes Visual Pronunciation Effective in Second Language Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/EPAYPG7V}},
note = {Machine review of arXiv:2507.06202}
}
read the original abstract
Visual feedback speeds up learners' improvement of pronunciation in a second language. The visual combined with audio allows speakers to see sounds and differences in pronunciation that they are unable to hear. Prior studies have tested different visual methods for improving pronunciation, however, we do not have conclusive understanding of what aspects of the visualizations contributed to improvements. Based on previous work, we created V(is)owel, an interactive vowel chart. Vowel charts provide actionable feedback by directly mapping physical tongue movement onto a chart. We compared V(is)owel with an auditory-only method to explore how learners parse visual and auditory feedback to understand how and why visual feedback is effective for pronunciation improvement. The findings suggest that designers should include explicit anatomical feedback that directly maps onto physical movement for phonetically untrained learners. Furthermore, visual feedback has the potential to motivate more practice since all eight of the participants cited using the visuals as a goal with V(is)owel versus relying on their own judgment with audio alone. Their statements are backed up by all participants practicing words with V(is)owel more than with audio-only. Our results indicate that V(is)owel is effective at providing actionable feedback, demonstrating the potential of visual feedback methods in second language learning.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Jessica Barlow and Judith Gierut. 2002. Minimal Pair Approaches to Phonological Remediation. Seminars in speech and language 23 (03 2002), 57–68. https: //doi.org/10.1055/s-2002-24969
-
[2]
R. Bellman and R. Kalaba. 1959. On adaptive control processes. IRE Transactions on Automatic Control 4, 2 (1959), 1–9. https://doi.org/10.1109/TAC.1959.1104847
-
[3]
Cindy Blanco and Ari Moline. 2025. Covering all the bases: Duolingo’s approach to listening skills. https://blog.duolingo.com/covering-all-the-bases-duolingos- approach-to-listening-skills/
work page 2025
-
[4]
Heather Bliss, Jennifer Abel, and Bryan Gick. 2018. Computer-Assisted Visual Articulation Feedback in L2 Pronunciation Instruction: A Review. Journal of Second Language Pronunciation 4 (05 2018), 129–153. https://doi.org/10.1075/ jslp.00006.bli
work page 2018
- [5]
-
[6]
Paul A. Booth. 1991. Errors and theory in human-computer interaction. Acta Psychologica 78, 1 (1991), 69–96. https://doi.org/10.1016/0001-6918(91)90005-K
-
[7]
Adam Brown. 1995. Minimal pairs: minimal importance? ELT Journal 49, 2 (04 1995), 169–175. https://doi.org/10.1093/elt/49.2.169 arXiv:https://academic.oup.com/eltj/article-pdf/49/2/169/9791152/169.pdf
-
[8]
Yaohua Bu, Tianyi Ma, Weijun Li, Hang Zhou, Jia Jia, Shengqi Chen, Kaiyuan Xu, Dachuan Shi, Haozhe Wu, Zhihan Yang, et al. 2021. PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–14
work page 2021
Show all 40 references
-
[9]
Michael Carey. 2004. CALL Visual Feedback for Pronunciation of Vowels: Kay Sona-Match. CALICO Journal 21, 3 (2004), 571–601. https://www.jstor.org/ stable/24149798 Publisher: Equinox Publishing Ltd
2004
-
[10]
Qiaoyi Chen, Siyu Liu, Kaihui Huang, Xingbo Wang, Xiaojuan Ma, Junkai Zhu, and Zhenhui Peng. 2024. RetAssist: Facilitating Vocabulary Learners with Generative Images in Story Retelling Practices. In Proceedings of the 2024 ACM Designing Interactive Systems Conference (Copenhag...
2024
-
[11]
Isabelle Darcy, Brian Rocca, and Zoie Hancock. 2021. A Window into the Class- room: How Teachers Integrate Pronunciation Instruction. RELC Journal 52, 1 (2021), 110–127. https://doi.org/doi/10.1177/0033688220964269
2021 doi
-
[12]
Keya Rani Das and AHMR Imon. 2016. A brief review of tests for normality. American Journal of Theoretical and Applied Statistics 5, 1 (2016), 5–12
2016
-
[13]
Tracey Derwing. 2003. What Do ESL Students say about Their Accents?Canadian Modern Language Review-revue Canadienne Des Langues Vivantes - CAN MOD LANG REV 59 (06 2003), 547–567. https://doi.org/10.3138/cmlr.59.4.547
2003 doi
-
[14]
Greene, David B
Beth G. Greene, David B. Pisoni, and Thomas D. Carrell. 1984. Recognition of speech spectrograms. The Journal of the Acoustical Society of America 76, 1 (July 1984), 32–43. https://doi.org/10.1121/1.391035
1984 doi
-
[15]
Mohd Hilmi Hamzah and Abdullah Bawodood. 2019. TEACHING ENGLISH SOUNDS VIA MINIMAL PAIRS: THE CASE OF YEMENI EFL LEARNERS.Journal of English Language and Literature 6 (10 2019), 97–102. https://doi.org/10.33329/ joell.63.97
2019
-
[16]
Feifei Han. 2016. Integrating pronunciation instruction with passage-level read- ing instruction. Pronunciation in the classroom: The overlooked essential (2016), 143–151
2016
-
[17]
Sandra G Hart. 1986. NASA task load index (TLX). (1986)
1986
-
[18]
Soon-Hin Hew and Mitsuru Ohki. 2004. Effect of Animated Graphic Annotations and Immediate Visual Feedback in Aiding Japanese Pronunciation Learning: A Comparative Study. CALICO Journal 21, 2 (2004), 397–419. https://www.jstor. org/stable/24149403 Publisher: Equinox Publishing Ltd
2004
-
[19]
Martin Hinton. 2013. An Aptitude for Speech: The Importance of Mimicry Ability in Foreign Language Pronunciation. In Teaching and Researching English Accents in Native and Non-native Speakers , Ewa Waniek-Klimczak and Linda R. Shockey (Eds.). Springer, 103–111. https://doi.org...
2013 doi
-
[20]
Murat Hismanoglu and Sibel Hismanoglu. 2010. Language teachers’ preferences of pronunciation teaching techniques: traditional or modern? Procedia-Social and Behavioral Sciences 2, 2 (2010), 983–989
2010
-
[21]
David Huggins-Daines, Mohit Kumar, Arthur Chan, Alan W Black, Mosur Ravis- hankar, and Alexander I Rudnicky. 2006. Pocketsphinx: A free, real-time contin- uous speech recognition system for hand-held devices. In 2006 IEEE international conference on acoustics speech and signal...
2006
-
[22]
Yannick Jadoul, Bill Thompson, and Bart De Boer. 2018. Introducing parselmouth: A python interface to praat. Journal of Phonetics 71 (2018), 1–15
2018
-
[23]
Kelly Moser and Tianlan Wei. 2024. COVID-19 and the pre-existing language teacher supply crisis. Language Teaching Research 28, 5 (2024), 1940–1975
2024
-
[24]
Murray J Munro, Tracey M Derwing, and Ron I Thomson. 2015. Setting segmental priorities for English learners: Evidence from a longitudinal study. International Review of Applied Linguistics in Language Teaching 53, 1 (2015), 39–60
2015
-
[25]
Offerman and Daniel J
Heather M. Offerman and Daniel J. Olson. 2016. Visual feedback and second language segmental production: The generalizability of pronunciation gains. System 59 (July 2016), 45–60. https://doi.org/10.1016/j.system.2016.03.003
2016 doi
-
[26]
Daniel J Olson. 2014. Benefits of visual feedback on segmental production in the L2 classroom. (2014)
2014
-
[27]
Annu Paganus, Vesa-Petteri Mikkonen, Tomi Mäntylä, Sami Nuuttila, Jouni Isoaho, Olli Aaltonen, and Tapio Salakoski. 2006. The vowel game: continuous real-time visualization for pronunciation learning with vowel charts. InAdvances in Natural Language Processing: 5th Internation...
2006
-
[28]
Indah Rahman and Isna Nur. 2017. THE USE OF MINIMAL PAIR TECHNIQUE IN TEACHING PRONUNCIATION AT THE SECOND YEAR STUDENTS OF SMAN 4 BANTIMURUNG. 4 (2017), 276. https://doi.org/10.24252/Eternal.V42.2018.A11
2017 doi
-
[29]
Rajiv Rao. 2019. Key issues in the teaching of Spanish pronunciation . Routledge London, United Kingdom
2019
-
[30]
Ivana Rehman. 2021. Real-time formant extraction for second language vowel production training. Ph. D. Dissertation. Iowa State University
2021
-
[31]
Schafer and Lawrence R
Ronald W. Schafer and Lawrence R. Rabiner. 1970. System for Au- tomatic Formant Analysis of Voiced Speech. The Journal of the Acoustical Society of America 47, 2B (02 1970), 634–648. https: //doi.org/10.1121/1.1911939 arXiv:https://pubs.aip.org/asa/jasa/article- pdf/47/2B/634/...
1970 doi
-
[32]
Michael R Sheldon, Michael J Fillyaw, and W Douglas Thompson. 1996. The use and interpretation of the Friedman test in the analysis of ordinal-scale data in repeated measures designs. Physiotherapy Research International 1, 4 (1996), 221–228
1996
-
[33]
Laura Sicola and Isabelle Darcy. 2015. Integrating pronunciation into the lan- guage classroom. The handbook of English pronunciation (2015), 471–487
2015
-
[34]
Lars St, Svante Wold, et al. 1989. Analysis of variance (ANOVA). Chemometrics and intelligent laboratory systems 6, 4 (1989), 259–272
1989
-
[35]
Magdalena Szyszka. 2015. Good English Pronunciation Users and Their Pro- nunciation Learning Strategies. Research in Language 13 (03 2015), 93–106. https://doi.org/10.1515/rela-2015-0017
2015 doi
-
[36]
Jeff Tennant and RC Gardner. 2004. The computerized mini-AMTB. Calico Journal (2004), 245–263
2004
-
[37]
R. I. Thomson and T. M. Derwing. 2014. The Effectiveness of L2 Pronunciation Instruction: A Narrative Review. 36, 3 (12 2014), 326–344. https://doi.org/10. 1093/applin/amu076
2014
-
[38]
Michael Wei. 2006. A Literature Review on Strategies for Teaching Pronunciation. Online submission (2006)
2006
-
[39]
Martijn Wieling, Eliza Margaretha, and John Nerbonne. 2011. Inducing phonetic distances from dialect variation. Computational Linguistics in the Netherlands Journal 1 (2011), 109–118
2011
-
[40]
weak ... strong
Ka-Wa Yuen, Wai-Kim Leung, Peng-fei Liu, Ka-Ho Wong, Xiao-jun Qian, Wai-Kit Lo, and Helen Meng. 2011. Enunciate: An internet-accessible computer-aided pronunciation training system and related user evaluations. In 2011 International Conference on Speech Database and Assessment...
2011
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.