REVIEW 4 major objections 5 minor 32 references
An Attempt to Develop a Neural Parser based on Simplified Head-Driven Phrase Structure Grammar on Vietnamese
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A simplified HPSG neural parser, adapted to Vietnamese with PhoBERT, is claimed to reach 82.34% constituency F1 on VietTreebank and 89.04% on the VLSP 2023 private test, with lower LAS attributed to label-preserving random permutations…
desk verdict Honest transfer attempt whose headline SOTA numbers are confounded by unconstrained random permutation of training data; fixable with an ablation and data release. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the joint-span simplified HPSG tree: a constituent-style tree in which every span carries both a syntactic category and a head word, so dependency arcs are embedded in the same span structure. The parser scores spans and head-arcs with a biaffine attention model over self-attentive token representations built from PhoBERT or XLM-RoBERTa, character, word, and POS embeddings, and decodes with dynamic programming. The paper's adaptation also adds a pre-processing repair: roughly one thousand non-compliant constituency/dependency tree pairs in the training and development splits are randomly permuted so that every pair satisfies the simplified HPSG constraints. That repair, plus the pretrained Vietnamese encoders, is what carries the claimed transfer from Penn Treebank to Vietnamese.
What would settle it
Retrain the same parser using linguist-corrected versions of the roughly one thousand tree pairs that were randomly permuted: if LAS does not rise above the reported 78.42 or F1 drops below 82.34, the paper's account of the repair's effect is falsified.
Extended reading notes
Core claim
The paper claims that a joint-span simplified HPSG neural parser, originally built for the Penn Treebank, can be transferred to Vietnamese by swapping in PhoBERT or XLM-RoBERTa encoders and by repairing the ~15% of VietTreebank/VnDT tree pairs that violate simplified HPSG rules through random permutation of training and development items. In the authors' experiments this adaptation reaches an 82.34% constituency F-score on the VietTreebank test set with predicted POS tags, above the self-attentive PhoBERT-large baseline's 80.55%, and an 89.04% F-score on the VLSP 2023 private test, slightly above Stanza's 88.73%. On dependency parsing it reports a UAS of 85.73 on VnDT, higher than the PhoNLP baseline, while LAS stays lower, which the authors attribute to keeping original dependency labels untouched during permutation. The paper's conclusion is that simplified HPSG deserves more linguistic-expert attention when building Vietnamese treebanks.
Load-bearing premise
The load-bearing premise is that randomly permuting the roughly one thousand non-compliant tree pairs in the training and development sets, without linguistic review, produces valid training trees rather than noise that quietly trains the parser to accept ungrammatical structures.
Editorial extensions
If this is right
- With PhoBERT-large and predicted POS tags, the joint-span HPSG parser reaches 82.34% constituency F1 on the VietTreebank test set, above the self-attentive PhoBERT-large baseline's 80.55%.
- On the VLSP 2023 private test the same model scores 89.04% F1, beating the Stanza baseline's 88.73%, while the two are nearly tied on the public test (86.05% vs 85.87%).
- On VnDT dependency parsing the parser reports UAS 85.73 with PhoBERT-large, above the PhoNLP baseline's 84.95, but lower LAS, which the authors attribute to keeping original dependency labels while randomly permuting arcs.
- Tuning the constituency/dependency loss weight to 0.9 on the development set indicates that the random-permutation repair did not disrupt the balance between the two parsing tasks.
Reading between the lines
- If the random-permutation repair is as benign as the development curves suggest, the same preprocessing could be tried on other languages whose treebanks pair constituency and dependency annotations, with the 15% violation rate as a natural diagnostic.
- The paper's LAS result points to a direct follow-up: re-run the same architecture on the permuted spans but with linguist-corrected head and dependency labels; if LAS rises without hurting UAS, the label-preserving choice is the cause.
- Because the VLSP 2023 head rules were written by non-linguists, an expert-written version of those rules is a testable way to see whether the HPSG parser's margin over Stanza widens further on out-of-domain data.
- A more skeptical probe would compare per-tree error patterns on the original versus permuted training trees; if the parser performs worse on sentences whose trees were permuted, the repair is introducing a systematic bias rather than a neutral fix.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper adapts the joint-span HPSG parser of Zhou and Zhao (ACL 2019) to Vietnamese by replacing its word encoders with PhoBERT and XLM-RoBERTa and by modifying the training data. Because roughly 15% of the constituent/dependency tree pairs in VietTreebank (VTB) and VnDT do not satisfy the simplified HPSG constraints, the authors randomly permute the non-compliant pairs in the training and development sets without linguistic constraints. They report a constituency F1 of 82.34 on the VTB test set, a dependency UAS of 85.73 on VnDT, and an F1 of 89.04 on the VLSP 2023 private test, claiming state-of-the-art results. For the VLSP 2023 treebank, which has no dependency annotation, the authors design head-percolation rules and use ClearNLP to convert constituency trees into dependency trees. The paper's central claim is that this simplified-HPSG parser outperforms prior Vietnamese parsers when trained on the modified corpora.
Significance. If the reported results were obtained under controlled conditions, the paper would provide a useful data point that the Zhou-Zhao joint-span HPSG architecture transfers to Vietnamese when combined with PhoBERT-large, and it would support the authors' suggestion that treebank construction should pay more attention to simplified HPSG constraints. The paper has several strengths: the test sets are left unaltered, the models are run five times, baseline systems are replicated in some cases, and the authors are transparent about the non-linguistic nature of their data modifications and head rules. Code and toolkit links are provided. The weakness is that the headline comparisons are not controlled: the HPSG parser is trained on randomly permuted data while the baselines are trained on the original treebanks, and no ablation separates the effect of the data modification from the effect of the architecture. The VLSP 2023 claim also rests on an unvalidated constituency-to-dependency conversion. These issues are fixable with additional experiments, but they are load-bearing for the stated state-of-the-art conclusions.
major comments (4)
- [§5.2, Tables 1–2] The HPSG parser is trained on data in which roughly 1,000 non-compliant tree pairs were "randomly permuted ... not bound by linguistic constraints," while the Self-Attentive, Biaffine, and PhoNLP baselines were trained on the original treebanks. This changes the training distribution, so the reported F1 and UAS gains cannot currently be attributed to the simplified HPSG architecture rather than to the permutation itself. The λ sweep in Figure 3 only varies the loss weight and does not compare permuted versus unpermuted training data, so the statement in §5.2 that the permutation's impact was "minimal" is not supported. I request an ablation training the same HPSG parser on (a) the original, unpermuted data and (b) data with non-compliant pairs removed, reporting F1, UAS, and LAS for both conditions.
- [§5.2, Figure 4, Table 3] The VLSP 2023 result depends on head-percolation rules written "by individuals with a non-linguistic engineering background" and on a ClearNLP conversion, but no validation of the resulting dependency trees is reported. Because the HPSG parser requires dependency supervision and the 89.04 private-test F1 exceeds the Stanza parser by only 0.31 points, the quality of this conversion could be decisive for the claimed advantage. Please compare the converted dependencies against an existing rule set or gold-standard dependencies, or provide a manual error analysis of the converted trees.
- [Tables 1 and 2] The headline numbers are presented without variance, even though the models are run five times. Additionally, Table 2 marks a result as "significantly different" when the lowest of five HPSG runs is compared with the replicated PhoNLP score using an unspecified paired t-test; selecting the minimum run before testing is not a valid significance procedure. Please report mean ± standard deviation over all five seeds for every row, and use a proper paired test (e.g., bootstrap or matched-seed comparison) for claims of statistical significance.
- [§3.1 vs. §5.2] The paper is inconsistent about which data were permuted: §3.1 says "samples from the training and development sets were adjusted" for the VTB and VnDT corpora, while §5.2 says "our modifications were limited to the training and development subsets of the VnDT dataset." This distinction matters because it determines whether the VTB constituency F1 comparison is also confounded by the modified training data. Please clarify whether the constituency trees in the VTB training set were modified or left untouched.
minor comments (5)
- [§5.4, Figure 5(a)] The text says the system "scored zero on categories like Nby and ADJb," but Figure 5(a) shows Nby at 8.00% and ADJb at 38.79%; please correct the mismatch.
- [§3.1, reference [25]] The sentence "Nguyễn et al. [25] introduce our project on developing a Vietnamese lexicon" is confusing, because reference [25] is a lexicon paper and not the current project; please rephrase.
- [Footnote 6] The URL path contains a duplicated suffix "run_constituency .py .py"; please fix the typo.
- [Figure 4] The head-percolation rules are printed as one long, low-resolution rule string; please provide a readable table of the rules or move them to an appendix.
- [Table 1, caption and §5.2] The caption says predicted POS tags were generated with the VnCoreNLP toolkit, while §5.2 describes a fine-tuned Stanza tagger with PhoBERT-large; please clarify which tagger produced the POS tags used in Table 1 and Table 2.
Circularity Check
No significant circularity: the reported F1/UAS/LAS scores are empirical benchmarks on unmodified test sets, not results forced by construction or by self-citation.
full rationale
The paper contains no derivation chain in which a predicted quantity is defined in terms of the fitted quantity or vice versa. The closest step is Section 5.2, where roughly 1,000 VTB/VnDT training and development tree pairs are randomly permuted, 'not bound by linguistic constraints,' to satisfy the simplified HPSG criteria of Zhou and Zhao [13]. This is a training-data modification, and it may raise concerns about comparability with baselines trained on the original corpora, but it is not circular: the VnDT test set is explicitly left unaltered, and the VLSP 2023 private test is an external evaluation set. The reported constituency F1 of 82.34, UAS of 85.73, and VLSP F1 of 89.04 are computed against original gold annotations, so they are not identities or fitted values renamed as predictions. The tuning of the loss weight λ on development data (Section 5.2, Figure 3) is standard hyperparameter selection, not a fitted input called a prediction. Self-citations such as [29] are used as empirical baselines and are not load-bearing. No specific equation-level reduction or self-citation chain forces the paper's conclusions, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- loss weight lambda (λ) =
0.9
assumptions (3)
- domain assumption Simplified HPSG tree constraints from [13] apply to Vietnamese syntax as a valid grammar formalism.
- ad hoc to paper Randomly permuting non-compliant tree pairs yields valid training data.
- ad hoc to paper Head-percolation rules in Figure 4 adequately convert VLSP 2023 constituency trees to dependency trees.
Cite this review
Pith. "Pith review of An Attempt to Develop a Neural Parser based on Simplified Head-Driven Phrase Structure Grammar on Vietnamese." pith.science (2026). https://pith.science/paper/5KMRARF2
@misc{pith2026241117270,
author = {Pith},
title = {Pith review of: An Attempt to Develop a Neural Parser based on Simplified Head-Driven Phrase Structure Grammar on Vietnamese},
year = {2026},
howpublished = {\url{https://pith.science/paper/5KMRARF2}},
note = {Machine review of arXiv:2411.17270}
}
read the original abstract
In this paper, we aimed to develop a neural parser for Vietnamese based on simplified Head-Driven Phrase Structure Grammar (HPSG). The existing corpora, VietTreebank and VnDT, had around 15% of constituency and dependency tree pairs that did not adhere to simplified HPSG rules. To attempt to address the issue of the corpora not adhering to simplified HPSG rules, we randomly permuted samples from the training and development sets to make them compliant with simplified HPSG. We then modified the first simplified HPSG Neural Parser for the Penn Treebank by replacing it with the PhoBERT or XLM-RoBERTa models, which can encode Vietnamese texts. We conducted experiments on our modified VietTreebank and VnDT corpora. Our extensive experiments showed that the simplified HPSG Neural Parser achieved a new state-of-the-art F-score of 82% for constituency parsing when using the same predicted part-of-speech (POS) tags as the self-attentive constituency parser. Additionally, it outperformed previous studies in dependency parsing with a higher Unlabeled Attachment Score (UAS). However, our parser obtained lower Labeled Attachment Score (LAS) scores likely due to our focus on arc permutation without changing the original labels, as we did not consult with a linguistic expert. Lastly, the research findings of this paper suggest that simplified HPSG should be given more attention to linguistic expert when developing treebanks for Vietnamese natural language processing.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
From Treebank Conversion to Automatic Dependency Parsing for Vietnamese
Dat Quoc Nguyen, Dai Quoc Nguyen, Son Bao Pham, Phuong-Thai Nguyen, and Minh Le Nguyen. “From Treebank Conversion to Automatic Dependency Parsing for Vietnamese”. In: Natural Language Processing and Information Sys- tems. Ed. by Elisabeth Métais, Mathieu Roche, and Maguelonne Teisseire. Cham: Springer International Publishing, 2014, pp. 196–207. ISBN : 97...
work page 2014
-
[2]
Ensuring annotation consistency and accuracy for Vietnamese treebank
Quy T. Nguyen, Yusuke Miyao, Ha T. T. Le, and Nhung T. H. Nguyen. “Ensuring annotation consistency and accuracy for Vietnamese treebank”. In: Language Resources and Evaluation 52.1 (Mar. 2018), pp. 269–315. ISSN : 1574-0218
work page 2018
-
[3]
Carl Pollard and Ivan A. Sag. Head-Driven Phrase Structure Grammar. Chicago: The University of Chicago Press, 1994
work page 1994
-
[4]
Implementing a Vietnamese syntactic parser using HPSG
Ba Lam Do and Thanh Huong Le. “Implementing a Vietnamese syntactic parser using HPSG”. In: The International Conference on Asian Language Processing (IALP). 2008
work page 2008
-
[5]
BKTreebank: Building a Vietnamese Dependency Tree- bank
Kiem-Hieu Nguyen. “BKTreebank: Building a Vietnamese Dependency Tree- bank”. In: Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) . Miyazaki, Japan: European Language Resources Association (ELRA), May 2018
work page 2018
-
[6]
Building a treebank for Vietnamese dependency pars- ing
Luong Nguyen Thi, Linh Ha My, Hung Nguyen Viet, Huyen Nguyen Thi Minh, and Phuong Le Hong. “Building a treebank for Vietnamese dependency pars- ing”. In: The 2013 RIVF International Conference on Computing & Communi- cation Technologies - Research, Innovation, and Vision for Future (RIVF). 2013, pp. 147–151
work page 2013
-
[7]
PhoBERT: Pre-trained language models for Vietnamese
Dat Quoc Nguyen and Anh Tuan Nguyen. “PhoBERT: Pre-trained language models for Vietnamese”. In: Findings of the Association for Computational Lin- guistics: EMNLP 2020. Online: Association for Computational Linguistics, Nov. 2020, pp. 1037–1042
work page 2020
-
[8]
Unsupervised Cross-lingual Representation Learning at Scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guil- laume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. “Unsupervised Cross-lingual Representation Learning at Scale”. In: Proceedings of the 58th Annual Meeting of the Association for Com- putational Linguistics . Online: Association for Co...
work page 2020
Show all 32 references
-
[9]
Deep Biaffine Attention for Neu- ral Dependency Parsing
Timothy Dozat and Christopher D. Manning. “Deep Biaffine Attention for Neu- ral Dependency Parsing”. In: 5th International Conference on Learning Rep- resentations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017
2017
-
[10]
The Pisa Lectures
Noam Chomsky. The Pisa Lectures . Berlin, New Y ork: De Gruyter Mouton,
-
[11]
Generating T yped Dependency Parses from Phrase Structure Parses
Marie-Catherine de Marneffe, Bill MacCartney, and Christopher D. Manning. “Generating T yped Dependency Parses from Phrase Structure Parses”. In: Pro- ceedings of the Fifth International Conference on Language Resources and Eval- uation (LREC’06) . Genoa, Italy: European Langu...
2006
-
[12]
Dependency Parser for Chinese Constituent Parsing
Xuezhe Ma, Xiaotian Zhang, Hai Zhao, and Bao-Liang Lu. “Dependency Parser for Chinese Constituent Parsing”. In: CIPS-SIGHAN Joint Conference on Chi- nese Language Processing. 2010
2010
-
[13]
Head-Driven Phrase Structure Grammar Parsing on Penn Treebank
Junru Zhou and Hai Zhao. “Head-Driven Phrase Structure Grammar Parsing on Penn Treebank”. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Florence, Italy: Association for Computational Linguistics, July 2019, pp. 2396–2408
2019
-
[14]
Stanza: A Python Natural Language Processing Toolkit for Many Hu- man Languages
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Man- ning. “Stanza: A Python Natural Language Processing Toolkit for Many Hu- man Languages”. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations....
2020
-
[15]
VLSP 2023 chal- lenge on Vietnamese Constituency Parsing
Thi-Minh-Huyen Nguyen, Xuan-Luong Vu, and My-Linh Ha. “VLSP 2023 chal- lenge on Vietnamese Constituency Parsing”. In: 2023
2023
-
[16]
Building a Large Syntactically-Annotated Cor- pus of Vietnamese
Phuong-Thai Nguyen, Xuan-Luong Vu, Thi-Minh-Huyen Nguyen, Van-Hiep Nguyen, and Hong-Phuong Le. “Building a Large Syntactically-Annotated Cor- pus of Vietnamese”. In: Proceedings of the Third Linguistic Annotation Work- shop (LAW III) . Suntec, Singapore: Association for Comput...
2009
-
[17]
A Minimal Span-Based Neural Constituency Parser
Mitchell Stern, Jacob Andreas, and Dan Klein. “A Minimal Span-Based Neural Constituency Parser”. In: Proceedings of the 55th Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers) . Vancouver, Canada: Association for Computational Linguistics...
2017
-
[18]
What’s Going On in Neural Con- stituency Parsers? An Analysis
David Gaddy, Mitchell Stern, and Dan Klein. “What’s Going On in Neural Con- stituency Parsers? An Analysis”. In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Hu- man Language Technologies, Volume 1 (Long Pap...
2018
-
[19]
Variants of Long Short-Term Memory for Sentiment Analysis on Vietnamese Students’ Feedback Corpus
Vu Duc Nguyen, Kiet Van Nguyen, and Ngan Luu-Thuy Nguyen. “Variants of Long Short-Term Memory for Sentiment Analysis on Vietnamese Students’ Feedback Corpus”. In: 2018 10th International Conference on Knowledge and Systems Engineering (KSE). 2018, pp. 306–311
2018
-
[20]
Prosodic Boundary Prediction Model for Vietnamese Text-To- Speech
Nguyen Thi Thu Trang, Nguyen Hoang Ky, Albert Rilliard, and Christophe d’ Alessandro. “Prosodic Boundary Prediction Model for Vietnamese Text-To- Speech”. In: Interspeech 2021 . Brno, Czech Republic: ISCA, Aug. 2021, pp. 3885–3889
2021
-
[21]
VLSP 2020 Shared Task: Universal Depen- dency Parsing for Vietnamese
Ha My Linh, Nguyen Thi Minh Huyen, Vu Xuan Luong, Nguyen Thi Luong, Phan Thi Hue, and Le Van Cuong. “VLSP 2020 Shared Task: Universal Depen- dency Parsing for Vietnamese”. In: Proceedings of the 7th International Work- shop on Vietnamese Language and Speech Processing . Hanoi,...
2020
-
[22]
Vietnamese transition-based de- pendency parsing with supertag features
Kiet V . Nguyen and Ngan Luu-Thuy Nguyen. “Vietnamese transition-based de- pendency parsing with supertag features”. In: 2016 Eighth International Confer- ence on Knowledge and Systems Engineering (KSE) . 2016, pp. 175–180. An Attempt to Develop a Neural Parser based on Simpli...
2016
-
[23]
LSTM Easy-first Dependency Parsing with Pre-trained Word Embeddings and Character-level Word Embeddings in Vietnamese
Binh Duc Nguyen, Kiet Van Nguyen, and Ngan Luu-Thuy Nguyen. “LSTM Easy-first Dependency Parsing with Pre-trained Word Embeddings and Character-level Word Embeddings in Vietnamese”. In: 2018 10th International Conference on Knowledge and Systems Engineering (KSE). 2018, pp. 187–192
2018
-
[24]
A neural joint model for Vietnamese word segmentation, POS tagging and dependency parsing
Dat Quoc Nguyen. “A neural joint model for Vietnamese word segmentation, POS tagging and dependency parsing”. In: Proceedings of the The 17th Annual Workshop of the Australasian Language Technology Association . Sydney, Aus- tralia: Australasian Language Technology Association...
2019
-
[25]
A lexicon for Vietnamese language processing
Thị Minh Huyền Nguyễn, Laurent Romary, Mathias Rossignol, and Xuân Lương Vũ. “A lexicon for Vietnamese language processing”. In: Language Resources and Evaluation 40 (2006), pp. 291–309
2006
-
[26]
Attention Is All Y ou Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. “Attention Is All Y ou Need”. In: Proceedings of the 31st International Conference on Neural Informa- tion Processing Systems. Long Beach, California, ...
2017
-
[27]
In-order Transition- based Parsing for Vietnamese
John Bauer, Hung Bui, Vy Thai, and Christopher Manning. “In-order Transition- based Parsing for Vietnamese”. In:Journal of Computer Science and Cybernetics 39.3 (Sept. 2023), pp. 207–221
2023
-
[28]
VnCoreNLP: A Vietnamese Natural Language Processing Toolkit
Thanh Vu, Dat Quoc Nguyen, Dai Quoc Nguyen, Mark Dras, and Mark Johnson. “VnCoreNLP: A Vietnamese Natural Language Processing Toolkit”. In: Pro- ceedings of the 2018 Conference of the North American Chapter of the Associ- ation for Computational Linguistics: Demonstrations . N...
2018
-
[29]
An Empirical Study for Vietnamese Constituency Parsing with Pre-training
Tuan-Vi Tran, Xuan-Thien Pham, Duc-Vu Nguyen, Kiet Van Nguyen, and Ngan Luu-Thuy Nguyen. “An Empirical Study for Vietnamese Constituency Parsing with Pre-training”. In: 2021 RIVF International Conference on Computing and Communication Technologies (RIVF). 2021, pp. 1–6
2021
-
[30]
PhoNLP: A joint multi-task learning model for Vietnamese part-of-speech tagging, named entity recognition and de- pendency parsing
Linh The Nguyen and Dat Quoc Nguyen. “PhoNLP: A joint multi-task learning model for Vietnamese part-of-speech tagging, named entity recognition and de- pendency parsing”. In: Proceedings of the 2021 Conference of the North Ameri- can Chapter of the Association for Computationa...
2021
-
[31]
Strongly incremental constituency parsing with graph neural networks
Kaiyu Y ang and Jia Deng. “Strongly incremental constituency parsing with graph neural networks”. In: Proceedings of the 34th International Conference on Neu- ral Information Processing Systems . NIPS ’20. Vancouver, BC, Canada: Curran Associates Inc., 2020. ISBN : 9781713829546
2020
-
[1993]
ISBN : 9783110884166
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.