Pith. sign in

REVIEW 4 major objections 5 minor 26 references

ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The VLSP 2022–2023 MT shared tasks produced a public Vietnamese–Chinese and Vietnamese–Lao benchmark, with human post-editing scores naming SDS (2022) and Bluesky (2023) as winners.

desk verdict The released VLSP 2022-2023 MT benchmark is worth having, but this report's own tables contradict its announced winner, so it cannot be trusted as the official record. read the letter →

arxiv 2501.08621 v1 pith:23DLYLL2 submitted 2025-01-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords low-resourcemachinetranslationVietnamese-ChineseVietnamese-LaoVLSPsharedtaskhumanpost-editingevaluationBLEUSacrebenchmarkdataset
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports the organization and results of two machine translation shared tasks run under the VLSP evaluation campaign: Vietnamese–Chinese in 2022 and Vietnamese–Lao in 2023. It introduces the released evaluation dataset ViBidirectionMT-Eval, describes the participating systems, and presents both automatic scores and human post-editing scores for four translation directions. The paper's central claim is that the human post-editing evaluation, conducted by five professional translators per task, provides a reliable basis for the official system rankings, which name SDS as the 2022 champion and Bluesky as the 2023 champion. A sympathetic reader would care because the dataset and rankings are a reusable public benchmark for low-resource Southeast Asian language pairs.

What carries the argument

The load-bearing mechanism is the post-editing human evaluation framework: for each translation direction, outputs from all participating systems are randomly assigned in equal numbers to five professional translators, who post-edit them; the resulting edited texts become new reference translations, and systems are ranked by a human evaluation error score computed against all collected post-edits. This protocol is what turns subjective judgments into a ranked table, and it is the basis for the paper's official winner announcements.

What would settle it

Recompute the final Lao–Vietnamese standings from the Human/FinalScore column of Table 9: MTA AI's two-direction mean (61.31 + 51.03)/2 = 56.17 exceeds Bluesky's (54.28 + 51.37)/2 = 52.83, which contradicts the paper's announcement of Bluesky as champion; additionally, any release of individual translator scores that showed large disagreements would undermine the reliability of the human rankings.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that the VLSP 2022–2023 machine translation tasks can be meaningfully evaluated by combining automatic metrics (BLEU, SacreBLEU) with a post-editing protocol in which five translators edit each system's output and the edited versions serve as additional references. Based on that protocol, the official final standings are: for Chinese–Vietnamese, SDS first, VBD-MT second, JNLP third, VC-Datamining fourth; for Lao–Vietnamese, Bluesky first, MTA AI second, BGSV AI third. The paper also contributes the ViBidirectionMT-Eval dataset, including human post-edits, as a reusable resource for these language pairs.

Load-bearing premise

The entire ranking rests on the assumption that the human post-editing scores—averaged across five translators with no reported inter-annotator agreement or significance testing—accurately reflect translation quality.

Editorial extensions

If this is right

  • Future MT systems for Vietnamese–Chinese and Vietnamese–Lao can be compared directly against the published scores and dataset.
  • The post-edited outputs provide additional reference translations that can be used for training and evaluation beyond the original test set.
  • The demonstrated success of data synthesis, back-translation, and mBART fine-tuning offers a template for other low-resource language pairs in the region.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The announced 2023 winner does not follow from the paper's own numbers: averaging the two direction scores in Table 9 gives MTA AI an arithmetic mean of 56.17 versus Bluesky's 52.83, so the stated ranking appears inconsistent with the presented data.
  • Because no inter-annotator agreement or significance test is reported, the human scores should be treated as a descriptive summary rather than a statistically grounded comparison.
  • The same post-editing protocol could be extended to other under-resourced Southeast Asian language pairs, such as Vietnamese–Khmer, using the released dataset as a template.
  • Readers should verify the reproducibility of the human scoring by checking whether the five translators' individual post-edit scores are released alongside the aggregate.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper reports on the VLSP 2022 and VLSP 2023 machine translation shared tasks for Vietnamese-Chinese and Vietnamese-Lao, respectively. It describes the released ViBidirectionMT-Eval dataset, the participating systems, and the automatic and human evaluations. The paper's main claims are the official rankings: SDS wins the 2022 Chinese-Vietnamese task and Bluesky wins the 2023 Lao-Vietnamese task, with MTA AI second and BGSV AI third. The paper also states that human post-editing evaluation was the decisive factor for the official rankings.

Significance. If the results and rankings were reliable, the paper would be a useful reference for low-resource machine translation between Vietnamese-Chinese and Vietnamese-Lao, and the released dataset on HuggingFace would be a community resource. The paper also contains strengths that should be acknowledged: the dataset covers four translation directions, uses 1,000 test sentences per direction, includes both automatic and human evaluation, and makes a distinction between constrained and unconstrained systems. However, the central ranking claims are undermined by internal inconsistencies in the reported scores and by a lack of detail in the human evaluation methodology, so the paper cannot currently serve as an authoritative record of the shared tasks.

major comments (4)
  1. [§5, Tables 5, 6, and 9] The announced 2023 winner contradicts the data in the paper's own tables. The text states 'The Final Score column of Table 9 shows that the winning team is Bluesky, second is MTA AI, third is BGSV AI', but Table 9 contains no Final Score column. Averaging the Human columns of Tables 5 and 6 gives MTA AI (51.03+61.31)/2 = 56.17 versus Bluesky (54.28+51.37)/2 = 52.83; under the stated averaging rule, MTA AI would be the 2023 champion. The paper must provide the exact scoring formula and a corrected Table 9, or the central ranking claim is unsupported.
  2. [§6] The official rankings depend entirely on human post-editing scores, but the paper never defines the scoring function referred to as 'human evaluation error', reports no inter-annotator agreement, and provides no significance testing. Section 6 also states that there were five submissions per task and five translators, yet Section 4 reports 5 official teams in 2022 and 7 in 2023, and Tables 5 and 6 list 7 and 5 teams respectively. These inconsistencies make it impossible to verify the final rankings.
  3. [§5, VLSP 2023 paragraph] The sentence 'Due to various reasons, the S-NLP team could not complete the technical report, so we have removed the S-NLP team from the final standings' is duplicated from the 2022 discussion; S-NLP appears only in the 2022 tables. This duplication calls into question the reliability of the surrounding final-standings text and must be corrected.
  4. [§5, Tables 5 and 6] The tables list SacreBLEU and 'Human' scores but do not define what the Human score represents, how it was computed, or whether higher values are better. The Description column in Table 9 labels Bluesky as 'SacreBLEU highest,' which refers to a different criterion from the human-based ranking claimed in the text; the relationship between the automatic and human metrics needs to be stated explicitly.
minor comments (5)
  1. [Abstract and throughout] The manuscript contains numerous grammatical errors ('an results', 'submission were evaluated', 'this approarch') that should be corrected in a revision.
  2. [References] The citation markers are malformed ('[ Papineni]', '[ Matt:2018]', '[ Cho2014LearningPR]'), and the reference list includes many entries that are not cited in the text (e.g., LAMBADA, TruthfulQA, BIG-Bench) while missing standard, correctly formatted references for BLEU and SacreBLEU.
  3. [§4.1.1 and §4.2.3] Large blocks of text appear to be pasted from the participating teams' reports, including repeated figures and equations (e.g., Equations (1)-(3) in Section 4.2.3), which makes it difficult to distinguish the organizers' own evaluation description from the teams' internal reports.
  4. [Tables 5 and 6] Team names are inconsistent across the paper ('F AIZ AIO' vs. 'Faiz AIO', 'BLUESKY' vs. 'Bluesky', 'TESTLA V100' vs. 'TeslaV100'), and 'ScareBLEU' appears instead of 'SacreBLEU' in the text.
  5. [Tables 7 and 8] The labels 'Result' and 'Final Score' are used without defining the unit or the direction of the scale; the text should explicitly state that these are percentages with higher values indicating better quality.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is an empirical evaluation report whose rankings rest on external human post-editing scores and SacreBLEU, with no fitted parameter or derivation that reduces to its inputs.

full rationale

This paper contains no derivation chain, predictive model, fitted parameter, or first-principles claim, so the standard circularity patterns do not apply. The load-bearing assertions are the reported VLSP 2022 and 2023 rankings, which are presented as the outcomes of external evaluation: SacreBLEU scores against reference translations and human post-editing scores collected from professional translators. The organizers do evaluate their own released corpus with their own protocol, but that is normal shared-task practice and is not circular: the scores are produced by independent translators and an external metric implementation, not derived from the paper's own claims. The self-referential element is limited to the organizers releasing and describing the dataset they organized, which does not make the ranking claim equivalent to its input by construction. There are serious internal consistency and reporting problems in Table 9, notably that the announced 2023 winner (Bluesky) is contradicted by the average of the human scores shown for Bluesky and MTA AI, and the promised Final Score column is absent. However, an internally inconsistent ranking is an evaluation-validity or reporting error, not circularity: the ranking does not reduce by definition to the data, it merely conflicts with the data as printed. The manuscript also contains extraneous or misplaced material (e.g., the S-NLP removal passage appearing in the 2023 discussion), but these do not create a circular derivation. Under the instructed standard, no specific equation or defined quantity is shown to be equivalent to its own input, so the appropriate verdict is no significant circularity with score 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters are fit in this paper. The central claims rest on domain assumptions about test-set representativeness and the validity of human post-editing as a ranking signal; these assumptions are not backed by agreement statistics or sampling details in the text.

assumptions (3)
  • domain assumption Five professional translators' post-edits yield a stable and valid ordering of MT systems.
    Section 6 assigns outputs to five translators and uses post-edit scores for official rankings, but gives no inter-annotator agreement, scoring-function details, or significance tests.
  • domain assumption The 1,000-sentence private test sets are representative of the target news and general domains.
    Section 2 asserts consistency of domain, while Section 6 reveals the evaluation set is a consecutive block of sentences, which is not necessarily a random or balanced sample.
  • domain assumption The technical reports submitted by teams accurately describe the systems they ran.
    Section 4 summarizes methods from team reports, but several ranked teams, including S-NLP, Faiz AIO, and Humble Bees, are noted as having no technical report, so some summaries may be missing or unverifiable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair." pith.science (2026). https://pith.science/paper/23DLYLL2

@misc{pith2026250108621,
  author       = {Pith},
  title        = {Pith review of: ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/23DLYLL2}},
  note         = {Machine review of arXiv:2501.08621}
}
read the original abstract

This paper presents an results of the VLSP 2022-2023 Machine Translation Shared Tasks, focusing on Vietnamese-Chinese and Vietnamese-Lao machine translation. The tasks were organized as part of the 9th, 10th annual workshop on Vietnamese Language and Speech Processing (VLSP 2022, VLSP 2023). The objective of the shared task was to build machine translation systems, specifically targeting Vietnamese-Chinese and Vietnamese-Lao translation (corresponding to 4 translation directions). The submission were evaluated on 1,000 pairs for testing (news and general domains) using established metrics like BLEU [11] and SacreBLEU [12]. Additionally, system outputs also were evaluated with human judgment provided by experts in Chinese and Lao languages. These human assessments played a crucial role in ranking the performance of the machine translation models, ensuring a more comprehensive evaluation.

Figures

Figures reproduced from arXiv: 2501.08621 by the authors.

Figure 1
Figure 1. Flow of data processing and model training Figure 1: Flow of data processing and model training [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. System flow machine translation After evaluating various models, the team selected Fairseq [ott-etal-2019-fairseq] for the baseline system, as it demonstrated superior performance on the public test set. For Chinese-Vietnamese translation, the model achieved a BLEU score of 38.0, which improved to 38.8 with the inclusion of back-translation; for Vietnamese-Chinese translation, the BLEU score increased from 37.8 to 3… view at source ↗
Figure 3
Figure 3. Overview of PhraseTransformer (CrossH) using [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Training phases of mBART and Transformer WMT models. [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Training phases of mT5 small and m2m 100-418M models. The [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 1
Figure 1. Figure 1: Overview of our proposed machine translation framework, which includes Lao-to-Vietnamese and Oif thid hitltifkhih ild [PITH_FULL_IMAGE:figures/full_fig_p009_1.png]
Figure 7
Figure 7. Figure 7: An example evaluation by human for Lao - Vietnamese machine translation task [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

26 extracted references · 12 canonical work pages

  1. [1]

    Towards a Cleaner Document-Oriented Multilingual Crawled Corpus

    Julien Abadji et al. “Towards a Cleaner Document-Oriented Multilingual Crawled Corpus”. In: arXiv e-prints, arXiv:2201.06642 (Jan. 2022), arXiv:2201.06642. arXiv: 2201.06642 [cs.CL]

  2. [2]

    Proceedings of the Ninth Workshop on Statistical Machine Translation

    Ondˇ rej Bojar et al., eds. Proceedings of the Ninth Workshop on Statistical Machine Translation. Baltimore, Maryland, USA: Association for Computational Linguistics, June 2014. doi: 10 . 3115/v1/W14-33. url: https://aclanthology.org/W14-3300

  3. [3]

    Brown et al

    Tom B. Brown et al. Language Models are Few-Shot Learners. 2020. arXiv: 2005.14165 [cs.CL]

  4. [4]

    Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

    Peter Clark et al. “Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge”. In: arXiv:1803.05457v1 (2018)

  5. [5]

    BERT: Pre-training of Deep Bidirectional Transformers for Language Un- derstanding

    Jacob Devlin et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Un- derstanding. 2019. arXiv: 1810.04805 [cs.CL]

  6. [6]

    Sentence Extraction-Based Machine Reading Comprehension for Vietnamese

    Phong Nguyen-Thuan Do et al. Sentence Extraction-Based Machine Reading Comprehension for Vietnamese. 2021. arXiv: 2105.09043 [cs.CL]

  7. [7]

    KenLM: Faster and Smaller Language Model Queries

    Kenneth Heafield. “KenLM: Faster and Smaller Language Model Queries”. In: Proceedings of the Sixth Workshop on Statistical Machine Translation. Ed. by Chris Callison-Burch et al. Edinburgh, Scotland: Association for Computational Linguistics, July 2011, pp. 187–197. url: https:// aclanthology.org/W11-2123

  8. [8]

    Measuring Massive Multitask Language Understanding

    Dan Hendrycks et al. Measuring Massive Multitask Language Understanding. 2021. arXiv: 2009. 03300 [cs.CY]

Show all 26 references
  1. [9]

    Long Short-Term Memory

    Sepp Hochreiter and J¨ urgen Schmidhuber. “Long Short-Term Memory”. In: Neural Comput. 9.8 (Nov. 1997), pp. 1735–1780. issn: 0899-7667. doi: 10.1162/neco.1997.9.8.1735 . url: https://doi.org/10.1162/neco.1997.9.8.1735

  2. [10]

    TruthfulQA: Measuring How Models Mimic Human Falsehoods

    Stephanie Lin, Jacob Hilton, and Owain Evans. TruthfulQA: Measuring How Models Mimic Human Falsehoods. 2022. arXiv: 2109.07958 [cs.CL]

  3. [11]

    Don’t Give Me the Details, Just the Sum- mary! Topic-Aware Convolutional Neural Networks for Extreme Summarization

    Shashi Narayan, Shay B. Cohen, and Mirella Lapata. “Don’t Give Me the Details, Just the Sum- mary! Topic-Aware Convolutional Neural Networks for Extreme Summarization”. In:Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Ed. by Ellen Rilo...

  4. [12]

    Vistral-7B-Chat - Towards a State-of-the-Art Large Language Model for Vietnamese

    Chien Van Nguyen et al. “Vistral-7B-Chat - Towards a State-of-the-Art Large Language Model for Vietnamese”. In: (2023)

  5. [13]

    PhoGPT: Generative Pre-training for Vietnamese

    Dat Quoc Nguyen et al. “PhoGPT: Generative Pre-training for Vietnamese”. In: arXiv preprint arXiv:2311.02945 (2023). 14

  6. [14]

    A Vietnamese Dataset for Evaluating Machine Reading Comprehension

    Kiet Van Nguyen et al. A Vietnamese Dataset for Evaluating Machine Reading Comprehension

  7. [15]

    New Vietnamese Corpus for Machine Reading Comprehension of Health News Articles

    Kiet Van Nguyen et al. New Vietnamese Corpus for Machine Reading Comprehension of Health News Articles. 2021. arXiv: 2006.11138 [cs.CL]

  8. [16]

    VinaLLaMA: LLaMA-based Vietnamese Foundation Model

    Quan Nguyen, Huy Pham, and Dung Dao. VinaLLaMA: LLaMA-based Vietnamese Foundation Model. 2023. arXiv: 2312.11011 [cs.CL]

  9. [17]

    SeaLLMs - Large Language Models for Southeast Asia

    Xuan-Phi Nguyen* et al. “SeaLLMs - Large Language Models for Southeast Asia”. In: (2023). eprint: arXiv:2312.00738

  10. [18]

    Introducing ChatGPT

    OpenAI. Introducing ChatGPT. 2022. url: https://openai.com/blog/chatgpt

  11. [19]

    The LAMBADA dataset: Word prediction requiring a broad discourse con- text

    Denis Paperno et al. The LAMBADA dataset: Word prediction requiring a broad discourse con- text. 2016. arXiv: 1606.06031 [cs.CL]

  12. [20]

    Language Models are Unsupervised Multitask Learners

    Alec Radford et al. “Language Models are Unsupervised Multitask Learners”. In: 2019. url: https://api.semanticscholar.org/CorpusID:160025533

  13. [21]

    Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

    Mirac Suzgun et al. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. 2022. arXiv: 2210.09261 [cs.CL]

  14. [22]

    Llama 2: Open Foundation and Fine-Tuned Chat Models

    Hugo Touvron et al. Llama 2: Open Foundation and Fine-Tuned Chat Models. 2023. arXiv: 2307.09288 [cs.CL]

  15. [23]

    GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

    Alex Wang et al. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. 2019. arXiv: 1804.07461 [cs.CL]

  16. [24]

    A Vietnamese Multitask Language Understanding Benchmark Suite for Large Lan- guage Models

    ZaloAI-Jaist. A Vietnamese Multitask Language Understanding Benchmark Suite for Large Lan- guage Models. https://github.com/ZaloAI-Jaist/VMLU. 2023

  17. [25]

    HellaSwag: Can a Machine Really Finish Your Sentence?

    Rowan Zellers et al. “HellaSwag: Can a Machine Really Finish Your Sentence?” In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. 15

  18. [2020]

    arXiv: 2009.14725 [cs.CL]

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.