{"id":"71ce912c-be89-4548-85ca-24262d2bc243","arxiv_id":"2507.19026","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An interactive dubbing system with visual stress-timing aids improved ESL learners' self-reported rhythm perception in a 12-participant within-subjects study.","lead":"RhythmTA is an interactive dubbing system that shows ESL learners visual cues for stress and timing while they practice English rhythm. In a 12-participant study, learners reported better rhythm perception and comparison with the visual system than with a text-only baseline.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central effectiveness claim rests only on self-reported Likert ratings; no objective pre/post measure of rhythm perception or production is reported, so 'effectively enhances rhythm perception' is not established.","rationale":"I read the paper as an HCI systems contribution whose central empirical claim is that RhythmTA improves ESL learners' rhythm perception. That claim requires evidence that perceptual ability changed, not merely that participants believed it changed. The user study's Likert scales measure beliefs and subjective ease; in a within-subjects design where the experimental interface is visually rich and the baseline intentionally minimal, positive ratings are expected and do not establish transferable learning. I agree with the reader that the stress-detection pipeline trained only on British English and the empirically set thresholds (Sec. 4.2 M2/M3) are genuine threats to correctness; if the visual labels are wrong, the system can mis-teach. But that concern applies to the accuracy of the content being taught, whereas the self-report issue invalidates the central effectiveness claim regardless of pipeline accuracy. Both issues are addressable, and both support the reader's CONDITIONAL verdict rather than changing it. I therefore mark agreement as partial: the reader's stated weakest assumption is the pipeline, while I locate the load-bearing weakness in the outcome measure, though the reader's rationale also notes the absence of objective production metrics.","tokens_in":21832,"tokens_out":5422,"duration_ms":60446,"concrete_test":"Add a pre/post objective rhythm-perception task using novel, untrained English sentences (e.g., forced-choice same/different stress-pattern discrimination or stressed-syllable identification) and collect pre/post dubbing recordings scored by blinded expert raters on stress placement and stress-timing accuracy. If perception scores or expert production ratings do not improve significantly for RhythmTA over baseline, the headline must be weakened to 'self-reported perception and confidence gains.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing gap is in the outcome measurement, not the stress-detection pipeline. The Abstract's claim that RhythmTA 'effectively enhances learners' rhythm perception' is supported in Sec. 5 only by questionnaire items Q11-Q12 and Q16-Q19, which ask whether participants feel more confident, feel improved, or find rhythm easier to perceive. No objective rhythm-perception test (e.g., stress identification or same/different discrimination on held-out stimuli) was administered before versus after use, and no acoustic or expert-rated production metric is reported. The only behavioral measures in Sec. 5.2.1 are time per attempt and number of attempts, which do not measure rhythmic ability. A within-subjects comparison of a visually rich system against a transcript-only baseline is especially susceptible to demand characteristics and preference effects: participants know which condition is novel and can report greater perceived ease without any actual change in perceptual skill. The paper's own limitation statement (Sec. 6.3) concedes the evaluation captured 'initial impressions and short-term learning outcomes,' but it does not acknowledge that the learning-improvement measures are entirely self-reported. If the headline claim is about actual learning, the evidence is at present insufficient; this is more load-bearing than the British-English stress-model domain shift because even a perfectly accurate pipeline cannot validate the claim without an objective outcome measure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents RhythmTA, an interactive dubbing-based system for ESL rhythm training. The system extracts word-level stress from speech using VOSK ASR, a Conformer-based stress classifier trained on Aix-MARSEC, and an nPVI-based rhythm-group segmentation algorithm; it visualizes stress timing through rhythm notes, rhythm groups, and rhythm waterfalls, and generates tolerant corrective feedback. A formative study with nine instructors yields six design requirements, and a within-subjects user study with twelve ESL learners compares RhythmTA with a transcript-only baseline. The study reports significantly higher self-reported ratings for perception, comparison, correction, learning improvement, and learning experience, along with SUS scores around 83, and the authors conclude that RhythmTA effectively enhances rhythm perception and shows significant potential for improving rhythm production.","tokens_in":22068,"tokens_out":5177,"duration_ms":53003,"significance":"The system design is thoughtful and addresses a real gap: independent, feedback-rich rhythm practice without an instructor. The formative study is credible, the three-stage dubbing workflow is well motivated, and the visualizations are a useful contribution to computer-assisted pronunciation training. The within-subjects comparison is appropriately counterbalanced, the baseline reasonably represents existing transcript-only dubbing applications, and the reported stress-detection accuracy of 85.44% on the Aix-MARSEC test set indicates a functional pipeline. If the learning-effectiveness claim were backed by objective outcome measures, this would be a strong contribution. As it stands, the evidence supports a well-received, usable system with promising qualitative results, but not a demonstrated improvement in rhythm perception or production.","major_comments":[{"comment":"The claim that RhythmTA 'effectively enhances learners' rhythm perception' rests entirely on self-report items Q16-Q19 (and to a lesser extent Q11-Q15); no objective pre/post measure of rhythm perception or production is reported. The only behavioral metrics in §5.2.1, duration per attempt and number of attempts, do not measure rhythmic ability. Because the comparison is between a visually rich novel system and a transcript-only baseline, demand characteristics are a serious concern: participants can report greater perceived ease and improvement without any actual change in perceptual skill. The limitation statement in §6.3 concedes that the evaluation captured 'initial impressions and short-term learning outcomes,' but it does not acknowledge that the learning-improvement measures are entirely self-reported. The authors should either add an objective outcome measure (e.g., a pre/post stress-identification or rhythm-discrimination test, or expert ratings of recorded production) or substantially soften the abstract and conclusion wording.","section":"§5.2.2, §6.3, Abstract"},{"comment":"The stress detector is trained on Aix-MARSEC, which contains British English BBC broadcasts, yet the user study uses American English video clips and non-native ESL speech from Mandarin, Cantonese, Korean, Japanese, and German L1 speakers. No validation accuracy is reported for these inputs. Because the visual aids, rhythm groups, and corrective feedback are all generated from predicted stress labels and intervals, systematic misdetection on accented or conversational speech could teach incorrect rhythmic patterns, which would undermine the pedagogical claims. The authors should validate the stress detector on the actual target and user speech materials used in the study, or explicitly scope the system's claims to varieties and conditions where the detector has been shown to work.","section":"§4.2 M2, §6.3"},{"comment":"The rhythm-group segmentation threshold tau=18 and the fuzzy word-match threshold of 0.62 are set empirically, but no sensitivity analysis or selection procedure is reported. The grouping, waterfalls, and corrective feedback that users see depend directly on these thresholds; a small perturbation could change the number of rhythm groups and the feedback content. Because the design requirements DR4-DR6 are implemented through these parameters, the authors should report how the segmentation and feedback output vary over a reasonable range of tau and fuzzy-match thresholds, and should justify the chosen values beyond an empirical hand-wave.","section":"§4.2 M3, §4.4"},{"comment":"The statistical analysis reports a large number of Wilcoxon signed-rank tests (Q1, Q5, Q11-Q23, and individual SUS items) without correction for multiple comparisons and without effect sizes. With N=12 and more than a dozen tests, uncorrected p-values overstate the strength of the evidence. The authors should report exact p-values, effect sizes (e.g., matched rank-biserial correlation), and either apply a multiple-comparison correction or explicitly frame the analysis as exploratory. This matters because these tests are the only quantitative support for the central claims of perceived learning improvement.","section":"§5.2.2"}],"minor_comments":[{"comment":"The text states 'We then used a paired t-test for significance analysis,' but the reported statistics are Z values (Z = 1.27, Z = 2.79); the authors should clarify which test was used and report the corresponding statistic consistently.","section":"§5.2.1"},{"comment":"There is a typo: 'local relayer' should be 'local replayer' in the paragraph discussing the local replayer.","section":"§5.3.3"},{"comment":"The caption uses a placeholder '? = 0.06' for Q22; the actual p-value should be reported, and 'flipped' should be defined as 'reverse-coded' for clarity.","section":"Figure 5 caption and §5.2.2"},{"comment":"The materials description says V2 and V3 were chosen to maintain similar speech characteristics, but the reported speech rates differ notably (147.4 vs. 126.9 words per minute); the authors should either justify that this difference is acceptable or report a matching procedure that accounts for it.","section":"§5.1 Materials"},{"comment":"The sentence 'the current rhythm extraction models is trained' has a subject-verb agreement error and should read 'the current rhythm extraction model is trained.'","section":"§6.3"}],"recommendation":"major_revision","confidential_remarks":"The system contribution is solid and the qualitative findings are useful, but the abstract's learning-effectiveness claim outruns the evidence. If the authors add an objective perception or production measure, or carefully rewrite the claims as perceived benefit and initial impressions, the paper would be publishable. The stress-detection domain shift also deserves validation because it affects the correctness of the pedagogical feedback."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRhythmTA is a real system contribution, not a repackaging. The novel bits are the rhythm-group segmentation (sliding-window nPVI on inter-stress intervals) and the rhythm-waterfall comparison view. Those are concrete, visual, and tied to a specific pedagogical problem—learners can't hear stress timing—and the paper develops them carefully through a formative study with nine instructors. The user study is honest within-subjects work: counterbalanced, think-aloud, transcript-only baseline, and the effect sizes on the facilitation items (Q14, Q15) are large enough to take seriously as perceived support.\n\nThe soft spot is exactly where the abstract makes its strongest claim. \"Effectively enhances rhythm perception\" is supported only by questionnaire items asking people whether they feel more confident or found rhythm easier to perceive. There is no objective stress-identification or discrimination test before/after, and no acoustic or expert-rated production measure. The behavioral data (time per attempt, attempts per clip) don't measure rhythmic ability. So the learning-improvement claim is weaker than the wording suggests. The paper's own limitation section says the evaluation captured initial impressions and short-term outcomes, but doesn't flag that the improvement metrics are entirely self-reported. I'd want that acknowledged.\n\nThe pipeline concerns are addressable but real: stress detection is trained on Aix-MARSEC British English and not validated on accented ESL speech; tau=18 and the fuzzy-match threshold 0.62 are empirical; no sensitivity analysis; multiple Wilcoxon tests uncorrected. These don't overturn the perception-facilitation result, but they do limit how far you can generalize the system's feedback accuracy. Also no code or data released, which matters for a system paper.\n\nThe citation pattern looks fine, and the writing is clear. It's a UIST-style contribution: the value is in the artifact and the study of its perceived usefulness. For that, it's solid. For a claim of actual learning, the evidence isn't there yet. A longitudinal study with an objective perception measure would settle it.\n\nWho should read it: anyone building prosody visualization or dubbing-based language tools. I'd send it to review; a good reviewer can push for the objective measure without asking them to redo the system.","headline":"A solid systems paper with a genuinely useful rhythm visualization; the main weakness is that the headline learning claim rests on self-reported perception scores, not objective outcomes.","tokens_in":22598,"tokens_out":1705,"would_cite":true,"duration_ms":18268,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a visual-aided dubbing system lets ESL learners train English speech rhythm on their own, with a twelve-user study showing improved rhythm perception and promising gains in production.","keywords":["speech rhythm training","visual aids","dubbing practice","audio and speech interfaces","English language learning","stress detection","rhythm visualization","ESL learners"],"falsifier":"Run the pipeline on speech with expert-annotated stress from a different corpus, especially accented ESL speech, and compare labels: if accuracy falls well below the reported 85.44%, the visual guidance is unreliable. Alternatively, a pre/post test in which learners who trained with RhythmTA identify stress or tap beats in unfamiliar audio with no visuals would settle whether the perception gain transfers beyond the interface.","tokens_in":21617,"feed_emoji":"🎤","tokens_out":7127,"duration_ms":68340,"temperature":0.7,"pith_summary":"RhythmTA is presented as a way for English-as-a-second-language learners to practice speech rhythm on their own, without a teacher: the system automatically marks which words are stressed and when they occur in any English speech, and renders that information visually on a timeline during dubbing. The design is grounded in interviews with nine spoken-English instructors, who reported that learners struggle to hear rhythmic patterns, cannot easily compare their own speech with a native model, and fall back on first-language rhythm habits. In a within-subjects study with twelve ESL learners, the system was compared against a transcript-only dubbing baseline that resembles existing pronunciation apps. Participants rated RhythmTA significantly higher on perceiving rhythm, identifying deviations, correcting mistakes, and self-reported learning gains, and they made more practice attempts per clip. The authors interpret this as evidence that visual rhythm notation can substitute for instructor feedback for perception, with production improvement as a promising but not yet directly demonstrated outcome.","feed_headline":"Visual dubbing cues train English rhythm without a teacher","feed_subtitle":"A twelve-person study reports clearer rhythm perception and self-correction than transcript-only dubbing apps.","key_machinery":"The central object is the rhythm notation: a dual-track visual layout with the transcript on top and a horizontal timeline of circles beneath, one circle per word, horizontally positioned by the word's utterance time, filled if the word carries stress and hollow if not. Consecutive stressed words whose intervals form a steady beat are grouped and color-coded by average interval length, and a rhythm waterfall links each target group to the corresponding range of the user's notation in the reflection stage. The notation is produced by a three-module pipeline—VOSK transcription and word alignment, wav2vec 2.0 plus Conformer stress classification (trained on Aix-MARSEC, 85.44% test accuracy), and a sliding-window nPVI segmentation with an empirically set threshold tau=18—and it carries the system's three functions: making rhythm perceptible during listening, guiding repetition in real time, and making deviations visible for self-correction.","core_discovery":"On its own terms, the paper claims that English speech rhythm can be taught as a visible, inspectable object rather than only an auditory one. RhythmTA's pipeline transcribes speech, classifies each word as stressed or unstressed with a wav2vec 2.0/Conformer model trained on the Aix-MARSEC British-English corpus (85.44% test accuracy), and segments consecutive stressed words into 'rhythm groups' when their intervals stay regular by a normalized Pairwise Variability Index threshold. The interface then shows a rhythm notation: filled versus hollow dots on a timeline, color-coded groups for steady beats, and a parallel 'rhythm waterfall' connecting target groups to the user's utterance. In the evaluation, twelve ESL learners used RhythmTA and a simplified baseline in counterbalanced order; on seven-point scales, RhythmTA significantly outperformed the baseline on perceiving rhythm in target and self speech, reproducing rhythm, identifying deviations, correcting errors, and four self-reported learning-improvement items. The paper frames the result as effective enhancement of rhythm perception and significant potential for improving rhythm production.","pith_inferences":["A natural next test the authors did not run is whether perception gains transfer to unaided listening; if rhythm notation acts as a crutch, learners might improve inside the interface yet show no gain when the visuals disappear.","Because the stress model is trained only on British broadcast English, accent robustness is the main generalization risk; fine-tuning or evaluation on accented and conversational speech would be the direct extension.","The rhythm-group segmentation could double as a material index, so learners could search for clips by tempo, beat regularity, or density of rhythm groups, matching practice content to their level.","The same visual-notation idea may apply beyond ESL, for instance to first-language prosody training, accent coaching, or speech therapy, though the paper only studies ESL learners."],"forward_implications":["Learners can practice rhythm on any English video, because the pipeline transcribes, detects stress, and segments rhythm groups automatically without hand-prepared materials.","The parallel comparison view shows where the user's stress timing deviates from the target, reducing the memory load of switching between two audio recordings.","Lenient rule-based feedback on stress accuracy, beat stability, and pace similarity gives a concrete goal for the next dubbing attempt; participants made more attempts per clip with RhythmTA than with the baseline.","In the twelve-participant study, RhythmTA significantly outperformed a transcript-only dubbing baseline on perceived rhythm, deviation identification, error correction, and self-reported learning improvement, with no significant increase in time per attempt.","Production improvement is reported as promising rather than proven; the study's significant results are perceptual and self-reported."],"supporting_citations":[{"why":"Supplies the Aix-MARSEC British-English broadcast corpus with expert stress annotations used to train and test the stress detector.","marker":"[3]"},{"why":"Provides the wav2vec 2.0 self-supervised speech representations fed into the stress classifier.","marker":"[4]"},{"why":"Provides the Conformer architecture used to build the binary word-stress detection model.","marker":"[18]"},{"why":"Defines the normalized Pairwise Variability Index used by the sliding-window algorithm to segment stress intervals into rhythm groups.","marker":"[17]"},{"why":"Supplies evidence that wav2vec 2.0 representations are sensitive to stress, supporting the feature choice.","marker":"[7]"},{"why":"Provides the System Usability Scale used to measure and compare usability across the two systems.","marker":"[5]"},{"why":"A commercial dubbing application whose feature set defines the baseline that RhythmTA is compared against.","marker":"[38]"},{"why":"Another commercial dubbing application lacking rhythm-specific feedback, used as context for the comparison baseline.","marker":"[56]"}],"fun_headline_variants":["See the beat: visual dubbing trains ESL rhythm","Visual dubbing cues help ESL learners hear and fix rhythm","RhythmTA turns English stress into visible dubbing cues","Visual rhythm feedback in dubbing boosts ESL self-correction","Dubbing with visual cues improves rhythm perception and production"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the stress detector, trained on about six hours of British English BBC radio speech, labels stressed words correctly in any target clip and in the accented speech of ESL learners; if those labels are wrong, the visual aids and feedback teach the wrong rhythm.","fun_headline_variants_meta":{"raw":{"variants":["See the beat: visual dubbing trains ESL rhythm","Visual dubbing cues help ESL learners hear and fix rhythm","RhythmTA turns English stress into visible dubbing cues","Visual rhythm feedback in dubbing boosts ESL self-correction","Dubbing with visual cues improves rhythm perception and production"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1361,"prompt_tokens":939,"completion_tokens":422,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":343}},"tokens_in":555,"tokens_out":422,"duration_ms":4456,"temperature":1.0,"reasoning_tokens":343,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:02:09.752257+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the pipeline on speech with expert-annotated stress from a different corpus, especially accented ESL speech, and compare labels: if accuracy falls well below the reported 85.44%, the visual guidance is unreliable. Alternatively, a pre/post test in which learners who trained with RhythmTA identify stress or tap beats in unfamiliar audio with no visuals would settle whether the perception gain transfers beyond the interface.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Aix-MARSEC British-English broadcast corpus with expert stress annotations used to train and test the stress detector."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the wav2vec 2.0 self-supervised speech representations fed into the stress classifier."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Conformer architecture used to build the binary word-stress detection model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the normalized Pairwise Variability Index used by the sliding-window algorithm to segment stress intervals into rhythm groups."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies evidence that wav2vec 2.0 representations are sensitive to stress, supporting the feature choice."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the System Usability Scale used to measure and compare usability across the two systems."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A commercial dubbing application whose feature set defines the baseline that RhythmTA is compared against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Another commercial dubbing application lacking rhythm-specific feedback, used as context for the comparison baseline."}],"review_version":2}