{"id":"7e322082-f19e-482d-80e3-fcea71dd90d0","arxiv_id":"2507.06202","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An interactive vowel chart that plots learners' vowels against a native target increased practice recordings over audio-only for most of eight L2 Spanish learners, but formant errors and the lack of a learning outcome temper the claim.","lead":"V(is)owel is an interactive vowel chart that shows a learner's extracted vowels alongside a native speaker's target, pairing each plotted vowel with playable audio. It was tested with eight phonetically untrained L2 Spanish learners against an audio-only condition, and the main finding is that the visual target motivated more recorded practice and gave learners something concrete to adjust.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Feedback accuracy errors (F2 MAE 116 Hz, offsets up to 265 ms, 5–26% large formant errors) mean participant-reported 'actionable feedback' may have been based on incorrect visuals; the central claim needs an accuracy-conditioned analysis.","rationale":"The reader's weakest assumption identifies the same load-bearing issue: the study's design conclusions are inferred from what participants said and did in response to an on-screen visualization whose accuracy is internally reported as marginal or worse for many recordings. My stress-test pass sharpens this: the offset errors are not just a background technical limitation; they directly corrupt the line-length information that participants explicitly tried to use, and the formant errors can flip the qualitative position of the line relative to the target. Because participants reported trusting the visualization over their own hearing, even a minority of corrupted displays can materially shape the think-aloud themes and the decision to re-record. The paper's own Discussion acknowledges faulty feedback, yet the abstract's 'actionable feedback' claim and the conclusions do not carry that caveat. A concrete accuracy-conditioned re-analysis would settle whether the 'actionable feedback' and engagement findings survive when only qualitatively correct visuals are considered. Until then, CONDITIONAL is the appropriate verdict: the paper should not be rejected because the interface and qualitative method are promising and the authors are transparent about errors, but it should not be accepted as demonstrating effective feedback without this check or a softened claim. I found no other concern that is more load-bearing than this one; sample size, novelty, and the absence of a learning outcome are secondary limitations already noted by the reader and the authors.","tokens_in":19869,"tokens_out":3746,"duration_ms":43155,"concrete_test":"On the recorded study audio, recompute formant trajectories with manual Praat correction and apply the same calibration transform; for each displayed recording, label whether the line's qualitative affordance (closer/farther from target, higher/lower, longer/shorter, and which side of the target it lands on) matches the ground-truth-corrected line. Then re-run the qualitative coding (Sections 5.1.2, 5.4, 5.5) and the engagement comparison (Table 3) on only the subset of interactions in which the displayed line was qualitatively correct. If the 'actionable feedback' themes and greater practice counts survive the accurate-subset analysis, the concern is largely answered; if the themes disappear or flip, the paper should be revised to claim perceived actionability and to condition design conclusions on feedback accuracy.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that V(is)owel provides actionable feedback and that participants practiced more because the visual gave them a basis to adjust pronunciation—depends on the plotted lines being sufficiently faithful to the learners' vowels. Section 5.3 reports mean F2 MAE of 116 Hz against a 100 Hz tolerance (Table 5), vowel-offset MAE up to 265 ms in Table 4 (P3 'canoa'), and 5–26% large formant errors per participant (Table 6). Section 5.3.1 states that offset errors were 'completely out of tolerance,' increasing the chance that learners saw accurate beginnings but inaccurate endings—exactly the line-length and endpoint information that §5.5 and §5.1.2 say participants tried to interpret ('I wasn't able to make a clear connection between the length of the line drawn and what I was saying,' P2). Because participants trusted the visualization over their own hearing (§5.5), a substantial fraction of the 'guidance on what to change' they cited may have been guidance toward an incorrect target or a misplotted current line. The engagement difference (Friedman p=0.0339) is also partly driven by participants recording again to resolve visual-auditory conflict (§5.4); if the visual is wrong, that conflict is a system artifact rather than evidence about effective feedback. The Discussion concedes 'V(is)owel currently gives faulty feedback,' but the abstract and conclusions still assert that it is 'effective at providing actionable feedback.' No learning outcome is measured, so the claim rests entirely on self-report and practice counts in the presence of known feedback errors.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces V(is)owel, an interactive vowel chart that plots a learner's first and second formants over time and plays back the associated audio, with the goal of providing actionable pronunciation feedback to phonetically untrained second-language learners. The authors report a within-subject study with eight American English speakers practicing Spanish minimal pairs under two conditions, V(is)owel and an audio-only control, using think-aloud protocols, practice counts, NASA-TLX, SUS, and exit interviews. They also report a technical evaluation of their vowel-boundary and formant-extraction algorithms against manual Praat annotations. The main claimed results are that participants practiced more with V(is)owel, that all participants used the visualized target as a goal, and that the tool provided actionable feedback that audio alone did not.","tokens_in":20177,"tokens_out":3179,"duration_ms":37116,"significance":"If the central claims were fully supported, the paper would make a useful contribution to computer-assisted pronunciation training by showing which aspects of articulatory visualizations help untrained learners and by demonstrating a system that works beyond isolated hVd contexts. The study has notable strengths: a counterbalanced within-subject design, think-aloud data that capture real-time interpretation, an iterative design process with pretesting, and a technical evaluation that openly benchmarks extraction accuracy against an external manual annotation standard. The paper also reports accuracy metrics transparently, which is commendable. However, the significance is undercut by a factual inconsistency in the engagement data and by accuracy levels that are acknowledged as faulty in the Discussion; both issues bear directly on the claim that V(is)owel provides actionable feedback. The absence of any pronunciation improvement measure further limits the scope of the conclusions.","major_comments":[{"comment":"The abstract states that \"all participants practicing words with V(is)owel more than with audio-only,\" but Table 3 shows that P7 recorded fewer times per word with V(is)owel (1.875) than with audio-only (2.250). This contradiction invalidates the \"all participants\" claim and should be corrected in the abstract and in Section 5.1.3. The Friedman test result may still hold, but the summary of the finding must not claim unanimity where the data show an exception.","section":"Abstract, §5.1.3, Table 3"},{"comment":"The technical evaluation reports F2 mean absolute error of 116.3 Hz against a 100 Hz tolerance, vowel-offset MAEs up to 265 ms (Table 4), and large formant-error rates of 5–26% per participant (Table 6). Section 5.4 attributes the higher practice counts to participants having \"guidance on what to change,\" and Section 5.5 states that participants \"put greater emphasis on the visualization than on their perception of pronunciation.\" If a substantial fraction of the plotted lines are misleading, then the behavioral and think-aloud evidence is partly a reaction to incorrect visuals, and the claim that V(is)owel provides \"actionable feedback\" is not supported in its current form. The paper needs either an accuracy-conditioned analysis (e.g., comparing engagement and interpretations on trials with small versus large formant errors) or a substantially weakened claim that separates perceived guidance from verified guidance.","section":"§5.3.1, §5.3.2, §5.4, §5.5"},{"comment":"The study measures engagement, self-reported confidence, and think-aloud interpretations, but it does not measure pronunciation improvement before and after practice. Nevertheless, the conclusion in Section 7 states that \"personalized feedback can have a significant impact on a learners' pronunciation in a second language,\" and the abstract opens with a general claim that \"visual feedback speeds up learners' improvement.\" These causal learning claims are not supported by the reported data. The paper should either add a learning-outcome measure or explicitly reframe the conclusions as being about engagement and perceived actionability rather than demonstrated pronunciation gains.","section":"§7, Abstract"},{"comment":"The Discussion states that \"based on offset times, V(is)owel currently gives faulty feedback,\" which directly conflicts with the abstract's claim that \"V(is)owel is effective at providing actionable feedback.\" This tension is not merely stylistic: faulty feedback can be actionable in the narrow sense that users can act on it, but the paper does not establish that acting on it moves learners toward the target. The authors should reconcile these statements, either by defining the limited sense in which feedback is actionable or by removing the effectiveness claim.","section":"Discussion"}],"minor_comments":[{"comment":"The sentence \"We use Praat's format extraction library\" should read \"formant extraction library.\"","section":"§3.2.2"},{"comment":"The text says \"Pt expressed how this made it more difficult to reach their goals,\" where \"Pt\" appears to be a placeholder for a participant identifier; please replace it with the intended label.","section":"§5.1.2, Familiarity"},{"comment":"The phrase \"far away from the spa line\" appears to be a typo; it should presumably be \"Spanish speaker's line.\"","section":"§5.1.2, Using the Spanish Speaker's Line as a Goal"},{"comment":"The caption says \"Bold indicates an offset that is outside allowed tolerance,\" but the table contains formant frequency MAE values and no bold entries; the caption appears to have been copied from Table 4 and should be corrected.","section":"Table 5 caption"},{"comment":"The Friedman test result would be more informative if the test statistic and effect size were reported alongside the p-value, given the small sample size of eight participants.","section":"Appendix A.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is an interesting exploratory study, and the transparent reporting of technical accuracy is a strength. The main concern is that the central 'actionable feedback' claim is undermined by the paper's own accuracy data, and the engagement summary contains a direct factual inconsistency with Table 3. These issues are fixable with reanalysis, additional conditioning on accuracy, and substantial rewriting of the abstract and conclusions, so major revision rather than rejection seems appropriate. I would also urge the editor to ask the authors to clarify the scope of their claims: without a learning-outcome measure, the paper should not be framed as demonstrating pronunciation improvement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read of arXiv:2507.06202. The real contribution: a working interactive vowel chart that handles vowels in arbitrary word contexts, shows diphthong trajectories, and integrates audio playback, plus a think-aloud study of how phonetically untrained learners interpret the chart. That last part is genuinely new—prior vowel chart studies used reflexive interviews and limited contexts. The technical evaluation against manual Praat annotations is a mark in its favor; they report boundary and formant errors rather than just screenshots.\n\nWhat the paper does well: the design process is transparent, the tutorial iterations are thoughtful, and the qualitative analysis gives a credible picture of learners treating the Spanish speaker's line as a target and struggling with line length. The engagement difference (Friedman p=0.034) is real but borderline, and with N=8 it is suggestive, not conclusive.\n\nThe soft spots are real and some are load-bearing. The abstract says 'all eight' participants practiced more with V(is)owel, but Table 3 shows P7 recorded less (1.875 vs 2.25). That overstatement should be fixed. More importantly, the stress-test concern lands: the computed feedback is too inaccurate to support the central claim. F2 MAE is 116 Hz against their own 100 Hz tolerance, offsets miss by up to 265 ms, and large formant errors run 5–26% across participants. The Discussion concedes 'V(is)owel currently gives faulty feedback,' yet the abstract and conclusions call it 'effective at providing actionable feedback.' Since participants trusted the visual over their own ears, some of the 'guidance on what to change' they reported was likely guidance based on misplotted lines. The engagement effect may in part reflect participants re-recording to resolve a visual-auditory conflict that is a system artifact, not a learning affordance. And no pronunciation learning outcome was measured, so the claim rests on self-report and practice counts.\n\nThis is fixable: release the data, report an accuracy-conditioned analysis (e.g., how many of the displayed lines were within tolerance), soften the claims to 'perceived actionable feedback,' and add a small longitudinal or transfer measure. As is, the paper deserves a serious referee but should not be accepted without major revision. I'd send it to review, and tell the authors to address the accuracy-conditioned analysis and the P7 discrepancy.","headline":"A genuinely novel interactive vowel chart with a think-aloud study, but the feedback accuracy errors are load-bearing and the abstract overstates the engagement result.","tokens_in":20736,"tokens_out":2089,"would_cite":false,"duration_ms":22383,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An interactive vowel chart that plots a learner's tongue position against a target speaker's line gives untrained learners a tangible goal and more practice than audio-only feedback.","keywords":["visual feedback","computer-assisted pronunciation training","vowel chart","formant extraction","second language learning","user evaluation","phonetics","think-aloud study"],"falsifier":"Run the same think-aloud comparison with a corrected extraction pipeline that keeps vowel-offset mean absolute error under 10 ms and second-formant mean absolute error under 100 Hz; if the practice advantage and the goal-oriented comments disappear, the reported effect depended on inaccurate visuals. A more direct test would pair the chart with ultrasound tongue imaging to check whether the plotted line matches the measured tongue position for the same recordings.","tokens_in":19646,"feed_emoji":"🗣️","tokens_out":6548,"duration_ms":70324,"temperature":0.7,"pith_summary":"The paper claims that an interactive vowel chart, which plots a learner's first and second formants as a moving line on a tongue-position chart beside a target speaker's line, gives phonetically untrained second-language learners feedback they can act on, and that this visual guidance motivates more practice than audio-only feedback. In a within-subject think-aloud study with eight American-English speakers practicing Spanish, all eight participants described using the Spanish speaker's plotted line as a goal, and the group recorded significantly more words with the chart than with audio alone. The authors interpret this as evidence that the visualization turns pronunciation practice from relying on one's own ears into working toward a visible target. They also report that the chart's formant extraction is not yet precise enough for full trust, and they use the participants' interpretations to recommend explicit anatomical feedback that maps directly onto tongue movement.","feed_headline":"Vowel chart beats audio-only at keeping L2 learners practicing","feed_subtitle":"By plotting tongue position in real time, the chart gave all eight learners a visible target and more re-recordings.","key_machinery":"The carrier of the argument is V(is)owel, an interactive vowel chart that maps the first two formant frequencies onto a trapezoidal chart labeled with tongue height and front-back position. A projective (homography) transformation, calibrated on four English corner vowels, maps each speaker's idiosyncratic vowel space onto a fixed chart; an autocorrelation-based boundary detector finds vowel onsets and offsets; a formant extractor, selected by a smoothness metric, supplies the first and second formants; and the resulting time-varying line is paired with audio playback so learners can hear and see the same token. This machinery matters because it is what turns an untrained learner's pronunciation into a visible, comparable target line, and the think-aloud protocol reveals how learners use that line while adjusting their speech.","core_discovery":"On its own terms, V(is)owel is presented as an interactive vowel chart that accepts vowels in any word context, shows diphthong trajectories over time, and is used, as far as the authors know, in the first visual pronunciation-training study to capture learners' real-time interpretations as they adjust their pronunciation. The paper's central finding is that the plotted line acted as an external goal: all eight participants referenced the Spanish speaker's line when deciding whether to re-record, whereas with audio-only they relied on their own judgment, and participants recorded more per word with V(is)owel than with audio-only (averages of 2.27 versus 1.64 recordings per word, with a Friedman test p=0.0339). The authors conclude that the visualization provides actionable feedback because it maps tongue movement onto a physical chart, and they recommend that future pronunciation tools include explicit anatomical feedback tied to physical movement, while acknowledging that algorithmic inaccuracies, especially in the second formant and in vowel offset times, sometimes showed users incorrect lines.","pith_inferences":["A corrected extraction pipeline, with vowel-offset mean absolute error under 10 ms and second-formant mean absolute error under 100 Hz, would allow a cleaner test of whether the visual-goal effect is caused by accurate articulation feedback or by the mere presence of a target line.","The same goal-tracking mechanism could plausibly generalize to other articulatory dimensions, such as lip rounding or nasalization, by adding equivalent visual dimensions to a chart; the paper gestures toward this in its discussion of extending V(is)owel to other languages.","Because learners trusted the visualization over their own ears, small algorithmic errors may be more harmful in a visual tool than in audio-only training, so accuracy requirements should be treated as a design constraint rather than a backend afterthought.","A longitudinal study would be needed to separate the motivating effect of novelty from the motivating effect of the visual goal itself, a limitation the authors themselves note."],"forward_implications":["Designers of pronunciation feedback for phonetically untrained learners should include explicit anatomical visual feedback that maps directly onto physical movement, rather than relying on correctness colors or percentages.","A visible target-speaker line can increase practice volume even when learners distrust the accuracy of the visualization, because it gives them a concrete thing to work toward.","Vowel-focused visualization did not show a clear blinder effect in this study: participants continued to mention consonants, pitch, and overall pronunciation while practicing with V(is)owel.","Extracting vowels in any word context, including rhotics and diphthongs, is feasible enough for a real-time practice tool, though second-formant and vowel-offset accuracy need improvement before the feedback can be fully trusted.","Learners tended to trust the visualization over their own hearing, which means the accuracy of the plotted line is not just a technical detail but a central design constraint."],"supporting_citations":[{"why":"Supplies the criteria for effective feedback (natural and logical, able to facilitate comparison) used to justify choosing an articulatory-based visualization over correctness displays.","marker":"[4]"},{"why":"Prior evidence that vowel-chart visual feedback improves vowel pronunciation, providing the baseline the paper extends.","marker":"[9]"},{"why":"Prior real-time vowel-chart visualization, limited to single vowels, whose scope V(is)owel extends to any word context.","marker":"[27]"},{"why":"Prior real-time formant-based vowel training restricted to limited word contexts, supplying a comparison point and formant-plotting approach.","marker":"[30]"},{"why":"Prior comparison showing that visual-plus-audio feedback outperforms audio-only feedback, motivating the study's two-condition design.","marker":"[18]"},{"why":"Provides the formant-extraction library used in the system's backend to compute the first two formants.","marker":"[5]"},{"why":"Dynamic Time Warping algorithm tested and rejected for vowel boundary extraction because it marked vowel onsets too early.","marker":"[2]"},{"why":"PocketSphinx aligner tested and rejected for vowel boundary extraction, motivating the autocorrelation-based extractor.","marker":"[21]"}],"fun_headline_variants":["Interactive vowel chart turns practice into a visible target","Vowel chart gives L2 learners a visual goal that boosts reps","Visual tongue map keeps learners recording more than audio alone","Chart's plotted line becomes the goal, driving more practice","Vowel chart motivates more re-recordings by showing a goal line"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes the plotted vowel lines are accurate enough that learners' interpretations and extra practice were responses to their actual tongue position rather than to extraction artifacts; the paper's own measurements of second-formant and vowel-offset errors leave room for that assumption to fail.","fun_headline_variants_meta":{"raw":{"variants":["Interactive vowel chart turns practice into a visible target","Vowel chart gives L2 learners a visual goal that boosts reps","Visual tongue map keeps learners recording more than audio alone","Chart's plotted line becomes the goal, driving more practice","Vowel chart motivates more re-recordings by showing a goal line"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000207,"raw_usage":{"total_tokens":1413,"prompt_tokens":972,"completion_tokens":441,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":588,"completion_tokens_details":{"reasoning_tokens":359}},"tokens_in":588,"tokens_out":441,"duration_ms":5014,"temperature":1.0,"reasoning_tokens":359,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:08:53.605663+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same think-aloud comparison with a corrected extraction pipeline that keeps vowel-offset mean absolute error under 10 ms and second-formant mean absolute error under 100 Hz; if the practice advantage and the goal-oriented comments disappear, the reported effect depended on inaccurate visuals. A more direct test would pair the chart with ultrasound tongue imaging to check whether the plotted line matches the measured tongue position for the same recordings.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the criteria for effective feedback (natural and logical, able to facilitate comparison) used to justify choosing an articulatory-based visualization over correctness displays."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior evidence that vowel-chart visual feedback improves vowel pronunciation, providing the baseline the paper extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior real-time vowel-chart visualization, limited to single vowels, whose scope V(is)owel extends to any word context."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior real-time formant-based vowel training restricted to limited word contexts, supplying a comparison point and formant-plotting approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior comparison showing that visual-plus-audio feedback outperforms audio-only feedback, motivating the study's two-condition design."},{"cited_title":"1992-2022","cited_arxiv_id":null,"evidence_quote":"Provides the formant-extraction library used in the system's backend to compute the first two formants."},{"cited_title":"Bellman and R","cited_arxiv_id":null,"evidence_quote":"Dynamic Time Warping algorithm tested and rejected for vowel boundary extraction because it marked vowel onsets too early."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"PocketSphinx aligner tested and rejected for vowel boundary extraction, motivating the autocorrelation-based extractor."}],"review_version":1}