REVIEW 3 major objections 4 minor 38 references
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Adding a SyncNet-based lip-synchrony loss to the duration predictor of a direct audio-visual speech-to-speech translation model improves dubbed lip-sync by 9.2% while leaving translation quality and naturalness unchanged.
desk verdict A practical lip-sync improvement for AVS2S, but the missing duration-loss-only ablation and shared SyncNet train/eval metric leave the sync-loss attribution unproven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing component is the duration predictor, a network that decides how long each discrete target speech unit should be held before the vocoder turns it into audio. The paper fine-tunes only this module with a combined loss made of a sync loss and a duration loss. The sync loss uses SyncNet, a pretrained audio-visual synchrony scorer, to penalize generated audio whose temporal alignment with the original video's lip movements is poor; the duration loss keeps predicted unit durations close to the source speech duration. Since the unit-to-unit decoder is frozen, the sync loss can only act through timing, stretching or compressing units rather than changing the content of the translation. That frozen-decoder constraint is exactly what makes the method lightweight, and also what sets its ceiling.
What would settle it
Run the same fine-tuning on a language pair whose target speech units cannot reproduce the source language's visible mouth-shape distinctions, and measure LSE-D on held-out videos; if the score does not improve over the untuned baseline, the timing-only assumption is the limiting factor and the reported gains depend on content-level choices.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that lip-synchrony in a direct AVS2S system can be improved by fine-tuning only the duration predictor with a combined sync and duration loss, while leaving the unit-to-unit decoder frozen. The predictor learns to stretch or compress each target speech unit so that the generated audio lands on moments where the source video's mouth shape is compatible, and the synchrony scorer supplies the gradient that points toward those moments. The method reaches an average lip-sync distance (LSE-D, lower is better) of 10.67, a 9.2% reduction from the speech-overlay baseline, and the improvement is consistent in every language direction. Automated measures of audio quality (PESQ) and translation quality (BLASER-2.0 and ASR-BLEU) stay essentially flat, so the sync gain is not bought by degrading the translation. Ablations show both losses are needed and that starting from the pretrained checkpoint matters.
Load-bearing premise
The method assumes that getting the timing of the target speech units right is enough to make the lips look synchronized, because the unit-to-unit decoder is frozen and cannot choose different sounds that would better match the speaker's visible mouth movements.
Editorial extensions
If this is right
- Dubbed videos that keep the original footage can become more realistic without any face synthesis, avoiding visual artifacts and the identity/likeness risks of regenerated mouths.
- The sync gain is consistent across four target languages, so the timing-based loss transfers without language-specific tuning.
- Because only the duration predictor is fine-tuned, the method can be applied on top of an existing frozen unit-based translation pipeline with a modest training step.
- The reported metrics show the lip-sync improvement does not come at the expense of speech naturalness or translation fidelity.
Reading between the lines
- An implication beyond the paper's claims is that the method's ceiling is set by the frozen decoder: if the available target speech units cannot approximate the source speaker's visible mouth shapes, retiming alone will saturate, and further sync gains would require content-level changes.
- The paper's paraphrase experiment hints that content-level choice is a stronger lever on lip-sync than timing, but the authors find that unconstrained paraphrase generation lowers translation quality; constraining paraphrase generation to preserve meaning while matching mouth shapes is a natural next step.
- The same duration-predictor fine-tuning recipe could transfer to any unit-based speech synthesizer that must align with a reference video, independent of the translation component.
- A human perceptual study of dubbing would be a stronger test of practical benefit than the geometric LSE-D metric, which measures embedding distance rather than perceived naturalness.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper addresses lip-synchrony in direct audio-visual speech-to-speech translation (AVS2S) by fine-tuning the duration predictor of an existing AV2AV model with two auxiliary losses: a standard duration loss (Eq. 1) and a SyncNet-based lip-synchrony loss (Eq. 2). The unit-to-unit translation decoder and vocoder remain frozen, and the generated speech is overlaid on the original video without modifying the visual content. Experiments on LRS3 across English-to-Spanish, Portuguese, Italian, and French show that the proposed method reduces the average LSE-D from 11.75 to 10.67 relative to the AV2AV-Speech baseline, with roughly unchanged PESQ, BLASER, and ASR-BLEU. Ablations and an additional paraphrase-based translation-model experiment are reported.
Significance. If the reported results are robust, the paper offers a practical, lightweight way to improve lip-sync in dubbing without face synthesis, thereby avoiding deepfake-related ethical concerns. The work is clearly motivated, builds on an open-source strong baseline, and evaluates on four language pairs with consistent trends. The main caveats are that the attribution of the improvement to the sync loss is not yet established (missing D-only control) and that the evaluation metric and training loss share the same SyncNet expert. The paper also provides a useful negative result in Section 6.2 showing that content-level changes improve synchrony further at the cost of translation quality.
major comments (3)
- [Section 6.1, Table 4] The ablation table compares LS+D FT, LS+D from scratch, and LS FT, but it never includes a duration-loss-only (D-only) condition. Since the duration loss in Eq. 1 already forces the generated speech to match source durations, it may by itself reduce LSE-D by improving temporal alignment. Without a D-only row, the improvement attributed to the sync loss (LS) in the headline result (Table 2) is not identified, because the combined LS+D FT result could be driven entirely by Ldur. Please add a D-only fine-tuning run from the same pretrained checkpoint and report its LSE-C/LSE-D.
- [Section 4.2 and Eq. (2)] The evaluation metrics LSE-C and LSE-D are computed with SyncNet [27], which is exactly the network that provides the training loss Lsync in Eq. (2). Because Lsync is defined in terms of SyncNet confidence, the fine-tuning procedure directly optimizes the network that produces the headline evaluation metric. The reported gains may therefore partly reflect train/test overlap rather than a genuinely better audio-visual alignment. The authors should either add an evaluation with a loss-agnostic lip-sync metric (e.g., a different pretrained sync detector or human ratings) or show that the improvement persists when Lsync is replaced by a different objective; at minimum, a model trained with Ldur only should be compared to quantify the overlap.
- [Section 5.2] The statement "Ours significantly outperform both Synthetic and AV2AV-Speech approaches in terms of lip-synchrony scores (p-value<0.05)" is not accompanied by any details of the statistical test: number of samples (videos or utterances), variance of LSE-D across samples, test type (paired or unpaired), or the value of the test statistic. Given that Table 2 shows per-language LSE-D differences of about 0.7-1.9 and that the baseline itself varies by language, a paired test with confidence intervals on the mean difference is needed to substantiate the claim. Please provide the full statistical details.
minor comments (4)
- [Tables 5 and 6] The arrows in the headers of Tables 5 and 6 are inconsistent with the definitions in Table 2: LSE-D is better when lower (should be ↓) and LSE-C is better when higher (should be ↑). The current notation ('LSE-D (↑)' and 'LSE-C (↓)' in Table 5) is reversed and will confuse readers.
- [Eq. (1)] Please define the units of d and d_p (e.g., predicted log-duration from the length predictor) and clarify whether the 'target duration' d is the source speech duration or a derived alignment target. This will make the loss and the fine-tuning procedure reproducible.
- [References] References [7] and [21] are the same work (Wav2Lip) and should be consolidated; the duplication is visible in the reference list.
- [Table 3] The claim of 'no degradation' in PESQ, BLASER, and ASR-BLEU is based on very small numeric differences (e.g., BLASER differences of 0.001-0.003). Reporting these without variance or significance testing is acceptable as a descriptive statement, but please soften the claim to 'no substantial degradation' or add supporting intervals.
Circularity Check
No circular derivation: the reported LSE-D gain is an empirical, held-out result; the SyncNet overlap between Eq. 2 and LSE-D is a metric-independence caveat, and the only author self-citation is non-load-bearing.
full rationale
This paper does not derive a result from first principles; it fine-tunes a duration predictor with Ltotal = Lsync + 10*Ldur (Eq. 3) and measures LSE-C/LSE-D, PESQ, BLASER, and ASR-BLEU on a held-out LRS3 test split (Section 4.1 states there is no overlap between Trainval and Test). Lsync (Eq. 2) is a SyncNet-based synchronization score and LSE-D is the SyncNet embedding distance used by Wav2Lip [7], so the training objective and headline metric share the same expert network. That is a genuine methodological caveat about independent measurement, but it is not an equation-level circularity: optimizing the training loss does not by construction force the held-out LSE-D value, and the paper reports independent checks (PESQ, BLASER, ASR-BLEU) that are not derived from SyncNet. The framework builds on AV2AV [6] by different authors, so there is no load-bearing self-citation chain. The only author self-citation, [16], motivates isometric length matching in Section 6.2, and the paper's own length-match experiment rejects that principle (LSE-D 10.35 vs. 10.18 in Table 6), so the citation is not load-bearing. The missing D-only ablation in Table 4 and the unsupported 'p-value<0.05' claim in Section 5.2 weaken causal attribution of the sync-loss contribution, but under-determination of an effect is an experimental-design issue, not circularity under the definitions here. The central claim is therefore self-contained empirical evidence rather than a circular derivation.
Assumptions & free parameters
free parameters (1)
- lambda (sync loss weight) =
10
assumptions (5)
- domain assumption SyncNet provides a valid measure of lip-synchrony, and optimizing its score improves real perceived synchrony.
- domain assumption Lip-synchrony can be improved by adjusting unit durations alone while the unit-to-unit decoder is frozen.
- domain assumption PESQ is an acceptable proxy for naturalness of overlaid translated speech.
- domain assumption The pre-trained AV2AV checkpoint from Choi et al. [6] is a valid starting point and a strong baseline.
- domain assumption The LRS3 test set has no overlap with training data and is representative for dubbing evaluation.
Cite this review
Pith. "Pith review of Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation." pith.science (2026). https://pith.science/paper/SBGM3ZVV
@misc{pith2026241216530,
author = {Pith},
title = {Pith review of: Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation},
year = {2026},
howpublished = {\url{https://pith.science/paper/SBGM3ZVV}},
note = {Machine review of arXiv:2412.16530}
}
read the original abstract
Audio-Visual Speech-to-Speech Translation typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony-ensuring that the movements of the lips match the spoken content-essential for maintaining realism in dubbed videos. Despite its importance, the inclusion of lip-synchrony constraints in AVS2S models has been largely overlooked. This study addresses this gap by integrating a lip-synchrony loss into the training process of AVS2S models. Our proposed method significantly enhances lip-synchrony in direct audio-visual speech-to-speech translation, achieving an average LSE-D score of 10.67, representing a 9.2% reduction in LSE-D over a strong baseline across four language pairs. Additionally, it maintains the naturalness and high quality of the translated speech when overlaid onto the original video, without any degradation in translation quality.
Reference graph
Works this paper leans on
-
[27]
Jointly optimizing translations and speech timing to improve isochrony in automatic dubbing,
Alexandra Chronopoulou, Brian Thompson, Prashant Mathur, Yogesh Virkar, Surafel M. Lakew, and Marcello Federico, “Jointly optimizing translations and speech timing to improve isochrony in automatic dubbing,” 2023
work page 2023
-
[1]
However, improving lip synchrony should not compromise translation quality and natural- ness [4, 5]
INTRODUCTION Traditionally, research has emphasized the importance of lip syn- chrony – the alignment of translated audio with the visible mouth movements of the original actors—as a key factor in maintaining the quality and realism of dubbed content [1, 2, 3]. However, improving lip synchrony should not compromise translation quality and natural- ness [4...
-
[2]
RELA TED WORK Lip synchronization has emerged as a crucial research area with wide-ranging applications, particularly in automatic dubbing for translation [1]. Numerous studies have focused on improving syn- chrony in dubbing through various approaches, such as isometric translation, where translations are generated to match the length of the source text ...
work page Pith review arXiv 2024
-
[3]
METHODOLOGY Our overall framework of Audio-Visual Speech-to-Speech Transla- tion (A VS2S) system in depicted in Fig. 1 and is based off Choi et al. [6]. The framework consists of the following: visuals and speech content from the original video are fed to an Audio-Visual (A V) En- coder which processes lip region and speech content and convert them into d...
-
[4]
EXPERIMENTAL SETTINGS 4.1. Dataset We leverage LRS3 [28] which is a large-scale video data consisting of thousands of spoken sentences collected from TED talks. The dataset statistics are provided in Table 1. Importantly, there is no overlap between the videos used in the test set and those used in the Trainval sets. In our experiments, we use the Trainva...
-
[5]
Baselines We use the latest work of A V2A V [6] as a strong baseline for this work
EXPERIMENTAL RESULTS 5.1. Baselines We use the latest work of A V2A V [6] as a strong baseline for this work. This is the only open-source A V translation model that has research permissive license. In our experiments with A V2A V , we did not evaluate their full audio-visual outputs. Instead, we use the generated speech by A V renderer and overlay it on ...
-
[6]
not trading off lip-synchrony improvements over speech translation quality and naturalness
ABLA TIONS 6.1. Duration Prediction This ablation study assesses the impact of lip-sync loss (LS loss), duration loss (D. loss), and model initialization on lip-sync perfor- mance for English to Spanish translation, measured using LSE-C (higher is better) and LSE-D (lower is better). As shown in Table 4, our model ( LS+D FT ), which incorporates both lip-...
-
[7]
CONCLUSIONS This study focuses on generating speech that aligns seamlessly with the original video, avoiding the need for facial synthesis and en- suring high-quality dubbing. Our A VS2S framework incorporates lip-synchrony and duration loss to enhance the alignment between speech and lip movements in audio-visual translation models. By concentrating sole...
Show all 38 references
-
[8]
Frederic Chaume, Audiovisual Translation: Dubbing, Transla- tion Practices Explained. St. Jerome Pub, Manchester, UK, 1st edition, 2012
2012
-
[9]
Neural dubber: Dubbing for videos according to scripts,
Chenxu Hu, Qiao Tian, Tingle Li, Wang Yuping, Yuxuan Wang, and Hang Zhao, “Neural dubber: Dubbing for videos according to scripts,” in Advances in Neural Information Pro- cessing Systems. 2021, vol. 34, pp. 16582–16595, Curran As- sociates, Inc
2021
-
[10]
Neural style-preserving visual dubbing,
Hyeongwoo Kim, Mohamed Elgharib, Michael Zollh ¨ofer, Hans-Peter Seidel, Thabo Beeler, Christian Richardt, and Christian Theobalt, “Neural style-preserving visual dubbing,” ACM Transactions on Graphics, vol. 38, no. 6, pp. 1–13, 2019
2019
-
[11]
An empirical take on the dubbing vs. subtitling debate: An eye movement study,
Elisa Perego, David Orrego-Carmona, and Sara Bottiroli, “An empirical take on the dubbing vs. subtitling debate: An eye movement study,” Lingue e Linguaggi , vol. 19, pp. 255–274, 2016
2016
-
[12]
Dub- bing in practice: A large scale study of human localization with insights for automatic dubbing,
William Brannon, Yogesh Virkar, and Brian Thompson, “Dub- bing in practice: A large scale study of human localization with insights for automatic dubbing,” Transactions of the Associa- tion for Computational Linguistics, vol. 11, pp. 419–435, 2022
2022
-
[13]
Av2av: Direct audio-visual speech to audio-visual speech translation with unified audio-visual speech representation,
Jeong Yun Choi, Se Jin Park, Minsu Kim, and Yong Man Ro, “Av2av: Direct audio-visual speech to audio-visual speech translation with unified audio-visual speech representation,” in CVPR 2024, 2024
2024
-
[14]
A lip sync expert is all you need for speech to lip generation in the wild,
K R Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, and C. V . Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in Proceedings of the 28th ACM International Conference on Multimedia, 2020
2020
-
[15]
Exposing lip-syncing deepfakes from mouth inconsistencies,
Soumyya Kanti Datta, Shan Jia, and Siwei Lyu, “Exposing lip-syncing deepfakes from mouth inconsistencies,” 2024
2024
-
[16]
The deepfake detection challenge dataset,
Brian Dolhansky, Joanna Bitton, Ben Pflaum, Jikuo Lu, Russ Howes, Menglin Wang, and Cristian Canton-Ferrer, “The deepfake detection challenge dataset,” CoRR, vol. abs/2006.07397, 2020
2006 arXiv
-
[17]
Faces of the future: How generative ai is redefining likeness and identity in the age of artificial intelligence,
Ben Bariach, Bernie Hogan, and Keegan McBride, “Faces of the future: How generative ai is redefining likeness and identity in the age of artificial intelligence,” Feb 2024
2024
-
[18]
Face/off: Changing the face of movies with deepfakes,
G Murphy, D Ching, J Twomey, and C Linehan, “Face/off: Changing the face of movies with deepfakes,” PLoS One, vol. 18, no. 7, pp. e0287503, 2023
2023
-
[19]
Regu- lating deep fakes: legal and ethical considerations,
E Meskys, J Kalpokiene, P Jurcys, and A Liaudanskas, “Regu- lating deep fakes: legal and ethical considerations,” Journal of Intellectual Property Law & Practice, vol. 15, no. 1, pp. 24–31, 2020
2020
-
[20]
Vasa-1: Lifelike audio-driven talking faces gener- ated in real time,
Sicheng Xu, Guojun Chen, Yu-Xiao Guo, Jiaolong Yang, Chong Li, Zhenyu Zang, Yizhong Zhang, Xin Tong, and Bain- ing Guo, “Vasa-1: Lifelike audio-driven talking faces gener- ated in real time,” ArXiv abs/2404.10667, 2024
2024 arXiv
-
[21]
Av-transpeech: Audio-visual robust speech-to- speech translation,
Rongjie Huang, Huadai Liu, Xize Cheng, Yi Ren, Lin Li, Zhe Ye, Jinzheng He, Lichao Zhang, Jinglin Liu, Xiaoyue Yin, and Zhou Zhao, “Av-transpeech: Audio-visual robust speech-to- speech translation,” ArXiv abs/2305.15403, 2023
2023 arXiv
-
[22]
Mixspeech: Cross-modality self-learning with audio- visual stream mixup for visual speech translation and recogni- tion,
Xize Cheng, Lin Li, Tao Jin, Rongjie Huang, Wang Lin, Ze- han Wang, Huangdai Liu, Yejin Wang, Aoxiong Yin, and Zhou Zhao, “Mixspeech: Cross-modality self-learning with audio- visual stream mixup for visual speech translation and recogni- tion,” in 2023 IEEE/CVF International C...
2023
-
[23]
Isometric mt: Neural machine translation for automatic dubbing,
Surafel M Lakew, Yogesh Virkar, Prashant Mathur, and Mar- cello Federico, “Isometric mt: Neural machine translation for automatic dubbing,” arXiv preprint arXiv:2112.08682, 2021
2021 arXiv
-
[24]
Duration modeling of neural tts for au- tomatic dubbing,
Johanes Effendi, Yogesh Virkar, Roberto Barra-Chicote, and Marcello Federico, “Duration modeling of neural tts for au- tomatic dubbing,” in ICASSP 2022 - 2022 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 8037–8041
2022
-
[25]
Prosodic alignment for off-screen automatic dubbing,
Yogesh Virkar, Marcello Federico, Robert Enyedi, and Roberto Barra-Chicote, “Prosodic alignment for off-screen automatic dubbing,” 2022
2022
-
[26]
Improving isochronous machine translation with target factors and auxil- iary counters,
Proyag Pal, Brian Thompson, Yogesh Virkar, Prashant Mathur, Alexandra Chronopoulou, and Marcello Federico, “Improving isochronous machine translation with target factors and auxil- iary counters,” 2023
2023
-
[28]
A lip sync expert is all you need for speech to lip generation in the wild,
K R Prajwal, Rudrabha Mukhopadhyay, Vinay P. Nambood- iri, and C.V . Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” inProceedings of the 28th ACM International Conference on Multimedia, New York, NY , USA, 2020, MM ’20, p. 484–492, Assoc...
2020
-
[29]
Towards realistic visual dubbing with heterogeneous sources,
Tianyi Xie, Liucheng Liao, Cheng Bi, Benlai Tang, Xiang Yin, Jianfei Yang, Mingjie Wang, Jiali Yao, Yang Zhang, and Ze- jun Ma, “Towards realistic visual dubbing with heterogeneous sources,” in Proceedings of the 29th ACM International Con- ference on Multimedia , New York, NY...
2021
-
[30]
Learning audio-visual speech representation by masked multimodal cluster prediction,
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrah- man Mohamed, “Learning audio-visual speech representation by masked multimodal cluster prediction,” in International Conference on Learning Representations, 2021, pp. 2–15
2021
-
[31]
Many-to-many spoken language translation via unified speech and text representation learning with unit-to-unit translation,
Minsu Kim, Jeongsoo Choi, Dahun Kim, and Yong Man Ro, “Many-to-many spoken language translation via unified speech and text representation learning with unit-to-unit translation,” arXiv preprint arXiv:2308.01831, pp. 1–15, 2023
2023 arXiv
-
[32]
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Advances in Neural Information Process- ing Systems, 2020, vol. 33, pp. 17022–17033
2020
-
[33]
Direct speech-to-speech translation with discrete units,
Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Sravya Popuri, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, et al., “Direct speech-to-speech translation with discrete units,” arXiv preprint arXiv:2107.05604, pp. 3–5, 2021
2021 arXiv
-
[34]
Out of time: automated lip sync in the wild,
J.S. Chung and A. Zisserman, “Out of time: automated lip sync in the wild,” in Workshop on Multi-view Lip-reading, ACCV , 2016
2016
-
[35]
Deep audio-visual speech recognition,
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman, “Deep audio-visual speech recognition,” IEEE transactions on pattern analysis and ma- chine intelligence, vol. 44, no. 12, pp. 8717–8727, 2018
2018
-
[36]
Seamlessm4t: Massively multilin- gual and multimodal machine translation,
Lo ¨ıc Barrault and Others, “Seamlessm4t: Massively multilin- gual and multimodal machine translation,” 2023
2023
-
[37]
Bleu: a method for automatic evaluation of machine translation,
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 2002, pp. 311–318
2002
-
[38]
Decoupled weight decay regularization,
Ilya Loshchilov and Frank Hutter, “Decoupled weight decay regularization,” 2019
2019
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.