{"id":"b7b3ef75-a758-40f4-a709-470e2db1bcb6","arxiv_id":"2502.05758","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An AV2vec lipreading system transferred from English to Chinese, adapted to target speakers, and ensembled across face and lip inputs achieves 77.3% CER on ChatCLR.","lead":"This paper applies cross-lingual transfer, speaker adaptation with KL-divergence regularization, and an ensemble of face-based and lip-based models to Chinese lipreading, building on the authors' earlier AV2vec self-supervised audio-visual model. The combined system reaches a 77.3% character error rate on the ChatCLR evaluation set, a number below the top 2024 challenge team but not a statistically significant difference.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline benchmark claim lacks statistical and protocol support: 77.3% vs leaderboard 77.6% is within the reported ±1.5 CI and Section 5.3.4 concedes no significant difference, yet the abstract presents 'lower than the top result' as a decisive new state of the art.","rationale":"The paper is internally consistent and mostly honest about its own uncertainty: Section 5.3.4 explicitly says the difference from the top-performing teams is not significant at the 95% level, and Section 6 acknowledges that the 77.3% CER is still high. Those disclosures are credit to the authors. However, the abstract and conclusion drop the caveat and claim a lower CER than the top result. That is precisely the load-bearing assertion, because the paper's contribution is framed as setting a new benchmark on ChatCLR. The reader identified the same weakness: the post-competition, organizer-assisted evaluation is not shown to be protocol-equivalent to the leaderboard runs, and the difference is within statistical noise. No code or weights are released, so external verification is impossible. My concern is therefore not about the method's internal logic, which appears sound, but about the validity of the headline comparison as stated. The proposed test would settle it directly: organizer certification of identical scoring plus a paired bootstrap over utterances would show whether the 0.3-point gap is a real, reproducible effect or an artifact of protocol differences and sampling noise. Given the authors' own statistical admission, the verdict should remain CONDITIONAL until the abstract is corrected or the evidence is strengthened.","tokens_in":19473,"tokens_out":5299,"duration_ms":57884,"concrete_test":"Obtain from the ChatCLR organizers documented confirmation that the post-hoc scoring of the authors' final models used the exact official evaluation script and text normalization applied to leaderboard submissions, ideally with identical reference transcriptions and CER computation. In parallel, from the authors' per-utterance predictions on the official evaluation set, run a paired bootstrap test of the CER difference between their 77.3% ensemble and the leaderboard's 77.6%; if the organizer confirmation is unavailable or the 95% confidence interval for the difference includes zero, revise the abstract and conclusion to say 'comparable to the top result' rather than 'lower than'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the abstract's final sentence: a 77.3% CER on ChatCLR, 'lower than the top result' of 77.6% (Table 5). Two conditions are needed for that claim to carry the weight the paper places on it. First, the post-competition evaluation must be protocol-comparable to the leaderboard. Section 5.3.4 says only that organizers assisted; no scoring script, text normalization, decoding configuration, or confirmation that the same official evaluation pipeline was used is given. M4 was trained after the deadline, so the post-hoc system is not the one that obtained 4th place and may reflect extra development cycles and validation-set model selection. Second, the 0.3 percentage-point gap is far smaller than the 95% confidence interval of ±1.5 reported for the 77.3% result, and the paper itself states there is no significant difference at the 95% level. A 77.3 vs 77.6 difference is not a meaningful benchmark improvement under the paper's own statistics. The remaining experimental findings (transfer helps under data scarcity, ensemble helps on the internal validation set) are conditionally supported, but the headline 'new benchmark' assertion is not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a target-speaker lipreading system for Chinese built on the authors' AV2vec self-supervised audio-visual encoder. The contributions are: (i) cross-lingual transfer of an English-pretrained AV2vec model to Chinese lipreading, (ii) speaker-dependent adaptation of a speaker-independent model using a KLD-regularized loss, and (iii) test-time ensembling of models trained on full-face and lip-ROI inputs. Experiments on the 2024 ChatCLR Challenge task 2 report that the full ensemble achieves 77.3% CER on the official evaluation set, which is numerically lower than the top challenge result of 77.6%; the authors also note that this difference is not statistically significant at the 95% level. Additional validation-set results examine transfer under reduced target-language data, per-speaker adaptation, and ensemble combinations.","tokens_in":19748,"tokens_out":9652,"duration_ms":92331,"significance":"The paper addresses a practical and under-explored problem: lipreading for a target speaker in a low-resource language, using cross-lingual self-supervised pretraining and speaker adaptation. The proposal is concrete and the validation set includes useful ablations. If the benchmark claim were fully supported, the system would be a meaningful advance for Chinese visual speech recognition. The strengths include the use of a self-distillation pretraining approach that avoids multi-iteration clustering, and the explicit comparison of face vs lip-ROI inputs with a complementary ensemble. However, the headline comparison is not statistically established, and the cross-lingual transfer result is inconsistent across input types, which tempers the significance. The paper does not provide code, but the experimental protocol is mostly transparent.","major_comments":[{"comment":"The abstract's final sentence presents the 77.3% CER as 'lower than the top result' from the challenge, but the paper's own statistical analysis (Section 5.3.4) states that the difference from the top-performing teams is not significant at the 95% confidence level. The gap (0.3 percentage points) is well inside the reported bootstrap 95% confidence interval of ±1.5 for the 77.3% result, and Table 5 does not provide confidence intervals for the leaderboard systems. The benchmark claim therefore rests on a point estimate comparison that the authors themselves do not support. Please either replace this claim with a properly tested statement (e.g., matched-pair significance test on the same evaluation transcripts) or remove it from the abstract and conclusion.","section":"Abstract; Section 5.3.4, Table 5"},{"comment":"The cross-lingual transfer advantage under data scarcity is not consistent across input types. With 10 hours of Chinese data, AV2vec-tf-lip (88.9% CER) outperforms AV2vec-lip (92.9%), but AV2vec-tf-face (94.5%) is worse than AV2vec-face (88.1%). The text states that 'a similar trend emerged' for the face-based models, which is contradicted by the reported numbers. Since the first highlight claims that cross-lingual transfer enhances lipreading with limited target-language data, the claim must be qualified to the lip-ROI setting or an explanation must be given for why face-based transfer fails. This is load-bearing for the contribution on transfer learning.","section":"Section 5.3.1, Figure 4"},{"comment":"The paper overstates the speaker-adaptation result. The text in Section 5.3.2 reports significant average improvements only for AV2vec-tf-lip and AV2vec-tf-face, while the non-transfer models show no significant changes; per-speaker analysis is 'mostly non-significant.' Nevertheless, Section 5.3.2 begins by stating that speaker adaptation 'effectively reduced CER,' and the highlights claim that speaker adaptation boosts specific-speaker accuracy. These statements should be tied to the model types and speakers for which the effect is significant, and the per-speaker analysis should include a summary test (e.g., sign test over 12 speakers) rather than relying on individual confidence-interval overlap.","section":"Section 5.3.2, Figure 5, Table 4"},{"comment":"The comparison to the challenge leaderboard is not demonstrated to be protocol-comparable. The final ensemble includes M4, which was trained after the challenge deadline, and the evaluation was performed post-hoc 'with the assistance of the Challenge organizers.' The paper does not specify the scoring script, text normalization, decoding configuration, or whether the same official evaluation pipeline was used for the reported 77.3% result as for the leaderboard runs. Without this information, a reader cannot exclude the possibility that differences in decoding or scoring contribute to the 0.3-point gap. Please document the post-hoc evaluation protocol and, if possible, obtain an official score from the organizers under the original protocol.","section":"Section 5.3.4, Table 5"}],"minor_comments":[{"comment":"The paragraph says 'The results are summarized in Table 6,' but the leaderboard results appear in Table 5; Table 6 is the model architecture table. Please correct the cross-reference.","section":"Section 5.3.4"},{"comment":"The baseline name 'REVAn' should be 'RAVEn' for consistency with the reference list and the text.","section":"Table 2"},{"comment":"The sentence 'The last 8 layers were averaged representations of the for the teacher' is missing a word; presumably 'outputs of the teacher.'","section":"Section 5.2"},{"comment":"The phrase 'also leaded to better lipreading results' should be 'also led to better lipreading results.'","section":"Conclusion"},{"comment":"The displayed abstract contains a typo 'di?erent speakers' (likely an OCR artifact); ensure the final version is printed correctly.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The benchmark claim is the paper's main promotional device, and it is currently not statistically or protocolically established. I would require the authors to remove or substantially qualify that claim, and to add the missing evaluation-protocol details, before accepting. The cross-lingual transfer inconsistency is also concerning; the authors need to present the negative face-transfer result honestly and discuss why it occurs. The novelty over their own AV2vec work is incremental, but the application to Chinese target-speaker lipreading is useful. I would not recommend rejection if the authors are willing to revise the claims and provide the requested evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline number is the problem. The abstract's \"lower than the top result\" is technically true but meaningless under the paper's own statistics: 77.3 versus 77.6 is 0.3 points, inside the reported ±1.5 confidence interval, and Section 5.3.4 openly admits no significant difference at 95% confidence. The stress-test concern about protocol comparability also lands. The post-competition evaluation was done \"with the assistance of the Challenge organizers,\" but the paper never shows that the scoring pipeline, text normalization, or decoding setup matches the official leaderboard. Without that, the comparison to other teams is not a controlled benchmark result.\n\nWhat is actually new and useful: the English-to-Chinese transfer evaluation for AV2vec, the KLD-regularized speaker adaptation on ChatCLR, and the lip-ROI plus face ensemble. The experiments are standard but solidly run, with bootstrap confidence intervals. The data-scarcity result for the lip model is the strongest part: with 10 hours of Chinese data, transfer clearly beats training from scratch (88.9 vs 92.9 CER). That is a practical, reproducible finding.\n\nSoft spots, in proportion. First, the benchmark overclaim needs to be fixed. Second, there is an internal contradiction in the data-scarcity story: the text says \"a similar trend emerged\" for the face models, but the numbers show the opposite — AV2vec-tf-face at 10 hours is 94.5, worse than the non-transfer AV2vec-face at 88.1. That needs a correction or a hedged conclusion. Third, speaker adaptation is significant only for the two transfer models; per-speaker gains are mostly non-significant, so \"boosts\" in the abstract should be softened to \"can improve.\" Fourth, Figure 4 lacks error bars, which weakens the data-scarcity comparisons. Fifth, no code or weights are released, so the claims are not independently checkable.\n\nOverall, the method is defensible and the paper is honestly reported in its statistical caveats — the problem is the framing, not the core pipeline. I would send this to peer review because the data-scarcity finding is a genuinely useful empirical result, and a competent referee will catch the face-transfer contradiction and the protocol issue. With those fixed, it becomes a reasonable contribution to the lipreading community.","headline":"The 77.3% benchmark claim is a non-significant gap with an unverified protocol match, but the cross-lingual data-scarcity experiment is the real contribution worth a second look.","tokens_in":20300,"tokens_out":2506,"would_cite":false,"duration_ms":26237,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A transferred self-supervised lipreading model, tuned per speaker and ensembled across lip and face views, reaches a 77.3% character error rate on the ChatCLR benchmark, below the top 2024 challenge result.","keywords":["lipreading","visual speech recognition","self-supervised learning","speaker adaptation","cross-lingual transfer","model ensemble","audio-visual self-distillation","ChatCLR"],"falsifier":"Run the authors' four-model ensemble on the official ChatCLR evaluation set under the challenge's own decoding and scoring script. If the CER comes out at or above the top leaderboard result of 77.6%, or if re-scoring under the two protocols gives non-identical numbers for the same model, the claim of beating the top result is falsified.","tokens_in":19271,"feed_emoji":"👄","tokens_out":8583,"duration_ms":82302,"temperature":0.7,"pith_summary":"Lipreading usually assumes a generic speaker, but casual conversation in a home setting has large per-speaker variation, and building a system from scratch for Chinese is expensive in both data and compute. The paper argues that a pipeline of three moves fixes both problems: an audio-visual self-distillation model pretrained on English transfers to Chinese lipreading; the model is then fine-tuned per target speaker with a Kullback-Leibler divergence regularizer that limits overfitting to that speaker's small video set; and at test time the system averages predictions from lip-only and full-face models, which fail on different inputs. The reported result is a character error rate of 77.3% on the official ChatCLR evaluation set, below the top result of the 2024 Chat-scenario Chinese Lipreading Challenge. The paper's own confidence-interval analysis, however, shows that this gap is not statistically significant.","feed_headline":"Lipreading pipeline beats top ChatCLR result with 77.3% CER","feed_subtitle":"Cross-language transfer, per-speaker tuning, and lip+face ensembling push character error below the winner.","key_machinery":"The load-bearing machinery is fourfold. AV2vec pretraining is a student-teacher self-distillation: the student receives masked and modality-dropped audio/video and regresses, with a mean-squared error, to targets produced online by a teacher whose weights are the student's exponential moving average, giving a label-free visual encoder. The lipreading head is a hybrid CTC/attention encoder-decoder, trained with cross-entropy plus CTC and decoded by joint CTC/attention beam search. Speaker adaptation fine-tunes that model on the target speaker's videos with a weighted sum of the cross-entropy loss, a KLD term, and the CTC loss, where the KLD term pins the adapted model's output distribution to the speaker-independent model's, preventing overfitting on the small adaptation set. Test-time ensemble averaging over lip-ROI and full-face models exploits the observation that the two input views make different, complementary mistakes.","core_discovery":"The paper's central claim is that these three techniques are complementary and can be combined on top of a single pretrained visual encoder. Concretely, an audio-visual self-distillation model (AV2vec) pretrained on English LRS3, when fine-tuned on Chinese labeled video with a hybrid CTC/attention decoder, outperforms supervised-from-scratch training as target-language data shrinks. Fine-tuning that speaker-independent model on a target speaker's own videos with a KLD term ($\\rho=0.1$) yields per-speaker gains for most of the twelve ChatCLR speakers. Equal-weight ensembling of lip-ROI and full-face models gives further gains, and the four-model ensemble reaches 77.3% CER on the official evaluation set. The paper is explicit that this is below the top leaderboard entry but that the difference is not significant at the 95% confidence level.","pith_inferences":["The same transfer-then-adapt recipe should carry the English-pretrained encoder into other low-resource target languages, since the paper's rationale is that visemes are broadly shared across languages; a cheap test would be transferring to another language with a small lipreading corpus.","The equal-weight lip/face ensemble is a fixed fusion; a learned or input-adaptive fusion that suppresses the face stream when occluding hands are detected could beat the fixed average in the microphone-in-hand cases the paper shows.","Re-scoring under the official challenge protocol is the decisive check on the benchmark claim; until that is done, the practical takeaway is the combination's internal gains, not the leaderboard comparison."],"forward_implications":["Cross-lingual transfer makes an English-pretrained audio-visual encoder usable for Chinese lipreading, with the advantage growing as the amount of target-language data decreases.","Per-speaker adaptation with KLD regularization improves most of the twelve target speakers' lipreading accuracy relative to the speaker-independent model.","Ensembling lip-ROI and full-face models reduces CER by roughly 3-4% relative to the average of the two single models, outperforming ensembling the same input type with different random seeds.","On the official ChatCLR evaluation set, the full four-model ensemble achieves 77.3% CER, numerically below the top challenge result, though the 95% confidence intervals overlap."],"supporting_citations":[{"why":"Supplies AV2vec, the self-distillation encoder that every model in the paper is built from.","marker":"Zhang et al., 2023"},{"why":"Supplies the LRS3 English dataset used for AV2vec pretraining and for the source-language transfer.","marker":"Afouras et al., 2018"},{"why":"Supplies the CN-CVS Mandarin corpus used as target-language pretraining and finetuning data.","marker":"Chen et al., 2023"},{"why":"Describes the MISP challenge dataset whose far-field videos form the ChatCLR training split.","marker":"Chen et al., 2022"},{"why":"Supplies the hybrid CTC/attention objective and joint decoding used for finetuning and testing.","marker":"Watanabe et al., 2017"},{"why":"Supplies the KL-divergence regularization idea for speaker adaptation.","marker":"Yu et al., 2013"},{"why":"Provides the AV-HuBERT baseline that AV2vec is compared against on LRS3 and motivates the pretraining setup.","marker":"Shi et al., 2022"},{"why":"Provides the prior study of speaker adaptation in lipreading that this paper extends.","marker":"Gimeno-Gómez and Martínez-Hinarejos, 2023"}],"fun_headline_variants":["Lipreading hits 77.3% CER with speaker adaptation","Cross-lingual lipreading: pretrain English, adapt Chinese","Speaker-aware lipreading improves on LRS3 to ChatCLR","Lipread: Self-distillation, speaker tuning, lip+face ensemble","AV2vec adapts to Chinese lipreading, 77.3% CER"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline benchmark claim assumes the authors' post-competition evaluation of their final model on the official ChatCLR set followed the same protocol as the challenge leaderboard runs; the paper does not demonstrate that its decoding, text normalization, and scoring are identical.","fun_headline_variants_meta":{"raw":{"variants":["Lipreading hits 77.3% CER with speaker adaptation","Cross-lingual lipreading: pretrain English, adapt Chinese","Speaker-aware lipreading improves on LRS3 to ChatCLR","Lipread: Self-distillation, speaker tuning, lip+face ensemble","AV2vec adapts to Chinese lipreading, 77.3% CER"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000746,"raw_usage":{"total_tokens":3349,"prompt_tokens":990,"completion_tokens":2359,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":2265}},"tokens_in":606,"tokens_out":2359,"duration_ms":17732,"temperature":1.0,"reasoning_tokens":2265,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T18:05:37.288983+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the authors' four-model ensemble on the official ChatCLR evaluation set under the challenge's own decoding and scoring script. If the CER comes out at or above the top leaderboard result of 77.6%, or if re-scoring under the two protocols gives non-identical numbers for the same model, the claim of beating the top result is falsified.","supporting_citations":[{"cited_title":", author Wan, G","cited_arxiv_id":null,"evidence_quote":"Supplies AV2vec, the self-distillation encoder that every model in the paper is built from."},{"cited_title":", author Wang, D","cited_arxiv_id":null,"evidence_quote":"Supplies the CN-CVS Mandarin corpus used as target-language pretraining and finetuning data."},{"cited_title":", author Zhou, H","cited_arxiv_id":null,"evidence_quote":"Describes the MISP challenge dataset whose far-field videos form the ChatCLR training split."},{"cited_title":", author Hori, T","cited_arxiv_id":null,"evidence_quote":"Supplies the hybrid CTC/attention objective and joint decoding used for finetuning and testing."},{"cited_title":", author Hsu, W.N","cited_arxiv_id":null,"evidence_quote":"Provides the AV-HuBERT baseline that AV2vec is compared against on LRS3 and motivates the pretraining setup."},{"cited_title":", author Mart \\' nez-Hinarejos, C.D","cited_arxiv_id":null,"evidence_quote":"Provides the prior study of speaker adaptation in lipreading that this paper extends."}],"review_version":1}