REVIEW 4 major objections 5 minor 26 references
ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The VLSP 2022–2023 MT shared tasks produced a public Vietnamese–Chinese and Vietnamese–Lao benchmark, with human post-editing scores naming SDS (2022) and Bluesky (2023) as winners.
desk verdict The released VLSP 2022-2023 MT benchmark is worth having, but this report's own tables contradict its announced winner, so it cannot be trusted as the official record. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the post-editing human evaluation framework: for each translation direction, outputs from all participating systems are randomly assigned in equal numbers to five professional translators, who post-edit them; the resulting edited texts become new reference translations, and systems are ranked by a human evaluation error score computed against all collected post-edits. This protocol is what turns subjective judgments into a ranked table, and it is the basis for the paper's official winner announcements.
What would settle it
Recompute the final Lao–Vietnamese standings from the Human/FinalScore column of Table 9: MTA AI's two-direction mean (61.31 + 51.03)/2 = 56.17 exceeds Bluesky's (54.28 + 51.37)/2 = 52.83, which contradicts the paper's announcement of Bluesky as champion; additionally, any release of individual translator scores that showed large disagreements would undermine the reliability of the human rankings.
Extended reading notes
Core claim
On its own terms, the paper establishes that the VLSP 2022–2023 machine translation tasks can be meaningfully evaluated by combining automatic metrics (BLEU, SacreBLEU) with a post-editing protocol in which five translators edit each system's output and the edited versions serve as additional references. Based on that protocol, the official final standings are: for Chinese–Vietnamese, SDS first, VBD-MT second, JNLP third, VC-Datamining fourth; for Lao–Vietnamese, Bluesky first, MTA AI second, BGSV AI third. The paper also contributes the ViBidirectionMT-Eval dataset, including human post-edits, as a reusable resource for these language pairs.
Load-bearing premise
The entire ranking rests on the assumption that the human post-editing scores—averaged across five translators with no reported inter-annotator agreement or significance testing—accurately reflect translation quality.
Editorial extensions
If this is right
- Future MT systems for Vietnamese–Chinese and Vietnamese–Lao can be compared directly against the published scores and dataset.
- The post-edited outputs provide additional reference translations that can be used for training and evaluation beyond the original test set.
- The demonstrated success of data synthesis, back-translation, and mBART fine-tuning offers a template for other low-resource language pairs in the region.
Reading between the lines
- The announced 2023 winner does not follow from the paper's own numbers: averaging the two direction scores in Table 9 gives MTA AI an arithmetic mean of 56.17 versus Bluesky's 52.83, so the stated ranking appears inconsistent with the presented data.
- Because no inter-annotator agreement or significance test is reported, the human scores should be treated as a descriptive summary rather than a statistically grounded comparison.
- The same post-editing protocol could be extended to other under-resourced Southeast Asian language pairs, such as Vietnamese–Khmer, using the released dataset as a template.
- Readers should verify the reproducibility of the human scoring by checking whether the five translators' individual post-edit scores are released alongside the aggregate.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports on the VLSP 2022 and VLSP 2023 machine translation shared tasks for Vietnamese-Chinese and Vietnamese-Lao, respectively. It describes the released ViBidirectionMT-Eval dataset, the participating systems, and the automatic and human evaluations. The paper's main claims are the official rankings: SDS wins the 2022 Chinese-Vietnamese task and Bluesky wins the 2023 Lao-Vietnamese task, with MTA AI second and BGSV AI third. The paper also states that human post-editing evaluation was the decisive factor for the official rankings.
Significance. If the results and rankings were reliable, the paper would be a useful reference for low-resource machine translation between Vietnamese-Chinese and Vietnamese-Lao, and the released dataset on HuggingFace would be a community resource. The paper also contains strengths that should be acknowledged: the dataset covers four translation directions, uses 1,000 test sentences per direction, includes both automatic and human evaluation, and makes a distinction between constrained and unconstrained systems. However, the central ranking claims are undermined by internal inconsistencies in the reported scores and by a lack of detail in the human evaluation methodology, so the paper cannot currently serve as an authoritative record of the shared tasks.
major comments (4)
- [§5, Tables 5, 6, and 9] The announced 2023 winner contradicts the data in the paper's own tables. The text states 'The Final Score column of Table 9 shows that the winning team is Bluesky, second is MTA AI, third is BGSV AI', but Table 9 contains no Final Score column. Averaging the Human columns of Tables 5 and 6 gives MTA AI (51.03+61.31)/2 = 56.17 versus Bluesky (54.28+51.37)/2 = 52.83; under the stated averaging rule, MTA AI would be the 2023 champion. The paper must provide the exact scoring formula and a corrected Table 9, or the central ranking claim is unsupported.
- [§6] The official rankings depend entirely on human post-editing scores, but the paper never defines the scoring function referred to as 'human evaluation error', reports no inter-annotator agreement, and provides no significance testing. Section 6 also states that there were five submissions per task and five translators, yet Section 4 reports 5 official teams in 2022 and 7 in 2023, and Tables 5 and 6 list 7 and 5 teams respectively. These inconsistencies make it impossible to verify the final rankings.
- [§5, VLSP 2023 paragraph] The sentence 'Due to various reasons, the S-NLP team could not complete the technical report, so we have removed the S-NLP team from the final standings' is duplicated from the 2022 discussion; S-NLP appears only in the 2022 tables. This duplication calls into question the reliability of the surrounding final-standings text and must be corrected.
- [§5, Tables 5 and 6] The tables list SacreBLEU and 'Human' scores but do not define what the Human score represents, how it was computed, or whether higher values are better. The Description column in Table 9 labels Bluesky as 'SacreBLEU highest,' which refers to a different criterion from the human-based ranking claimed in the text; the relationship between the automatic and human metrics needs to be stated explicitly.
minor comments (5)
- [Abstract and throughout] The manuscript contains numerous grammatical errors ('an results', 'submission were evaluated', 'this approarch') that should be corrected in a revision.
- [References] The citation markers are malformed ('[ Papineni]', '[ Matt:2018]', '[ Cho2014LearningPR]'), and the reference list includes many entries that are not cited in the text (e.g., LAMBADA, TruthfulQA, BIG-Bench) while missing standard, correctly formatted references for BLEU and SacreBLEU.
- [§4.1.1 and §4.2.3] Large blocks of text appear to be pasted from the participating teams' reports, including repeated figures and equations (e.g., Equations (1)-(3) in Section 4.2.3), which makes it difficult to distinguish the organizers' own evaluation description from the teams' internal reports.
- [Tables 5 and 6] Team names are inconsistent across the paper ('F AIZ AIO' vs. 'Faiz AIO', 'BLUESKY' vs. 'Bluesky', 'TESTLA V100' vs. 'TeslaV100'), and 'ScareBLEU' appears instead of 'SacreBLEU' in the text.
- [Tables 7 and 8] The labels 'Result' and 'Final Score' are used without defining the unit or the direction of the scale; the text should explicitly state that these are percentages with higher values indicating better quality.
Circularity Check
No circularity: the paper is an empirical evaluation report whose rankings rest on external human post-editing scores and SacreBLEU, with no fitted parameter or derivation that reduces to its inputs.
full rationale
This paper contains no derivation chain, predictive model, fitted parameter, or first-principles claim, so the standard circularity patterns do not apply. The load-bearing assertions are the reported VLSP 2022 and 2023 rankings, which are presented as the outcomes of external evaluation: SacreBLEU scores against reference translations and human post-editing scores collected from professional translators. The organizers do evaluate their own released corpus with their own protocol, but that is normal shared-task practice and is not circular: the scores are produced by independent translators and an external metric implementation, not derived from the paper's own claims. The self-referential element is limited to the organizers releasing and describing the dataset they organized, which does not make the ranking claim equivalent to its input by construction. There are serious internal consistency and reporting problems in Table 9, notably that the announced 2023 winner (Bluesky) is contradicted by the average of the human scores shown for Bluesky and MTA AI, and the promised Final Score column is absent. However, an internally inconsistent ranking is an evaluation-validity or reporting error, not circularity: the ranking does not reduce by definition to the data, it merely conflicts with the data as printed. The manuscript also contains extraneous or misplaced material (e.g., the S-NLP removal passage appearing in the 2023 discussion), but these do not create a circular derivation. Under the instructed standard, no specific equation or defined quantity is shown to be equivalent to its own input, so the appropriate verdict is no significant circularity with score 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Five professional translators' post-edits yield a stable and valid ordering of MT systems.
- domain assumption The 1,000-sentence private test sets are representative of the target news and general domains.
- domain assumption The technical reports submitted by teams accurately describe the systems they ran.
Cite this review
Pith. "Pith review of ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair." pith.science (2026). https://pith.science/paper/23DLYLL2
@misc{pith2026250108621,
author = {Pith},
title = {Pith review of: ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair},
year = {2026},
howpublished = {\url{https://pith.science/paper/23DLYLL2}},
note = {Machine review of arXiv:2501.08621}
}
read the original abstract
This paper presents an results of the VLSP 2022-2023 Machine Translation Shared Tasks, focusing on Vietnamese-Chinese and Vietnamese-Lao machine translation. The tasks were organized as part of the 9th, 10th annual workshop on Vietnamese Language and Speech Processing (VLSP 2022, VLSP 2023). The objective of the shared task was to build machine translation systems, specifically targeting Vietnamese-Chinese and Vietnamese-Lao translation (corresponding to 4 translation directions). The submission were evaluated on 1,000 pairs for testing (news and general domains) using established metrics like BLEU [11] and SacreBLEU [12]. Additionally, system outputs also were evaluated with human judgment provided by experts in Chinese and Lao languages. These human assessments played a crucial role in ranking the performance of the machine translation models, ensuring a more comprehensive evaluation.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus
Julien Abadji et al. “Towards a Cleaner Document-Oriented Multilingual Crawled Corpus”. In: arXiv e-prints, arXiv:2201.06642 (Jan. 2022), arXiv:2201.06642. arXiv: 2201.06642 [cs.CL]
arXiv 2022
-
[2]
Proceedings of the Ninth Workshop on Statistical Machine Translation
Ondˇ rej Bojar et al., eds. Proceedings of the Ninth Workshop on Statistical Machine Translation. Baltimore, Maryland, USA: Association for Computational Linguistics, June 2014. doi: 10 . 3115/v1/W14-33. url: https://aclanthology.org/W14-3300
work page 2014
-
[3]
Tom B. Brown et al. Language Models are Few-Shot Learners. 2020. arXiv: 2005.14165 [cs.CL]
arXiv 2020
-
[4]
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Peter Clark et al. “Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge”. In: arXiv:1803.05457v1 (2018)
arXiv 2018
-
[5]
BERT: Pre-training of Deep Bidirectional Transformers for Language Un- derstanding
Jacob Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Un- derstanding. 2019. arXiv: 1810.04805 [cs.CL]
arXiv 2019
-
[6]
Sentence Extraction-Based Machine Reading Comprehension for Vietnamese
Phong Nguyen-Thuan Do et al. Sentence Extraction-Based Machine Reading Comprehension for Vietnamese. 2021. arXiv: 2105.09043 [cs.CL]
work page Pith review arXiv 2021
-
[7]
KenLM: Faster and Smaller Language Model Queries
Kenneth Heafield. “KenLM: Faster and Smaller Language Model Queries”. In: Proceedings of the Sixth Workshop on Statistical Machine Translation. Ed. by Chris Callison-Burch et al. Edinburgh, Scotland: Association for Computational Linguistics, July 2011, pp. 187–197. url: https:// aclanthology.org/W11-2123
work page 2011
-
[8]
Measuring Massive Multitask Language Understanding
Dan Hendrycks et al. Measuring Massive Multitask Language Understanding. 2021. arXiv: 2009. 03300 [cs.CY]
work page 2021
Show all 26 references
-
[9]
Long Short-Term Memory
Sepp Hochreiter and J¨ urgen Schmidhuber. “Long Short-Term Memory”. In: Neural Comput. 9.8 (Nov. 1997), pp. 1735–1780. issn: 0899-7667. doi: 10.1162/neco.1997.9.8.1735 . url: https://doi.org/10.1162/neco.1997.9.8.1735
1997 doi
-
[10]
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. TruthfulQA: Measuring How Models Mimic Human Falsehoods. 2022. arXiv: 2109.07958 [cs.CL]
2022 arXiv
-
[11]
Don’t Give Me the Details, Just the Sum- mary! Topic-Aware Convolutional Neural Networks for Extreme Summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. “Don’t Give Me the Details, Just the Sum- mary! Topic-Aware Convolutional Neural Networks for Extreme Summarization”. In:Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Ed. by Ellen Rilo...
2018 doi
-
[12]
Vistral-7B-Chat - Towards a State-of-the-Art Large Language Model for Vietnamese
Chien Van Nguyen et al. “Vistral-7B-Chat - Towards a State-of-the-Art Large Language Model for Vietnamese”. In: (2023)
2023
-
[13]
PhoGPT: Generative Pre-training for Vietnamese
Dat Quoc Nguyen et al. “PhoGPT: Generative Pre-training for Vietnamese”. In: arXiv preprint arXiv:2311.02945 (2023). 14
2023 arXiv
-
[14]
A Vietnamese Dataset for Evaluating Machine Reading Comprehension
Kiet Van Nguyen et al. A Vietnamese Dataset for Evaluating Machine Reading Comprehension
-
[15]
New Vietnamese Corpus for Machine Reading Comprehension of Health News Articles
Kiet Van Nguyen et al. New Vietnamese Corpus for Machine Reading Comprehension of Health News Articles. 2021. arXiv: 2006.11138 [cs.CL]
2021 arXiv
-
[16]
VinaLLaMA: LLaMA-based Vietnamese Foundation Model
Quan Nguyen, Huy Pham, and Dung Dao. VinaLLaMA: LLaMA-based Vietnamese Foundation Model. 2023. arXiv: 2312.11011 [cs.CL]
2023 arXiv
-
[17]
SeaLLMs - Large Language Models for Southeast Asia
Xuan-Phi Nguyen* et al. “SeaLLMs - Large Language Models for Southeast Asia”. In: (2023). eprint: arXiv:2312.00738
2023 arXiv
-
[18]
Introducing ChatGPT
OpenAI. Introducing ChatGPT. 2022. url: https://openai.com/blog/chatgpt
2022
-
[19]
The LAMBADA dataset: Word prediction requiring a broad discourse con- text
Denis Paperno et al. The LAMBADA dataset: Word prediction requiring a broad discourse con- text. 2016. arXiv: 1606.06031 [cs.CL]
2016 arXiv
-
[20]
Language Models are Unsupervised Multitask Learners
Alec Radford et al. “Language Models are Unsupervised Multitask Learners”. In: 2019. url: https://api.semanticscholar.org/CorpusID:160025533
2019
-
[21]
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Mirac Suzgun et al. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. 2022. arXiv: 2210.09261 [cs.CL]
2022 arXiv
-
[22]
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron et al. Llama 2: Open Foundation and Fine-Tuned Chat Models. 2023. arXiv: 2307.09288 [cs.CL]
2023 arXiv
-
[23]
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Alex Wang et al. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. 2019. arXiv: 1804.07461 [cs.CL]
2019 arXiv
-
[24]
A Vietnamese Multitask Language Understanding Benchmark Suite for Large Lan- guage Models
ZaloAI-Jaist. A Vietnamese Multitask Language Understanding Benchmark Suite for Large Lan- guage Models. https://github.com/ZaloAI-Jaist/VMLU. 2023
2023
-
[25]
HellaSwag: Can a Machine Really Finish Your Sentence?
Rowan Zellers et al. “HellaSwag: Can a Machine Really Finish Your Sentence?” In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. 15
2019
-
[2020]
arXiv: 2009.14725 [cs.CL]
2009 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.