{"id":"479444d2-ab3a-4860-b1d9-1945da5d64d4","arxiv_id":"2507.19995","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"VLQA is a new expert-annotated Vietnamese legal question-answering dataset with 3,129 citizen questions, long-form answers, and a 59,636-article statutory corpus.","lead":"The authors build VLQA, a dataset of over 3,000 real Vietnamese legal questions with expert-verified answers and links to roughly 60,000 law articles. It is meant to give researchers a benchmark for legal information retrieval and question answering in Vietnamese, a low-resource language.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The gold labels depend on a single expert reviewer's unmeasured judgment: Table 3 shows 25.8% of annotations were revised, and with no IAA the final labels lack independent verification.","rationale":"I agree with the reader's conditional verdict. The dataset construction pipeline is clearly described, and the paper's internal reviewer statistics are useful evidence, especially the 74.18% exact-match rate, which shows that many student annotations were consistent with the expert reviewer. However, the final labels are effectively a one-person gold standard: the students' double annotation is filtered through a single legal expert, and no inter-annotator agreement is reported. Because 25.8% of samples were modified during review, the robustness of the final labels cannot be inferred from the students' initial agreement. The EQUALS size comparison in Table 1 also deserves scrutiny, but it affects the 'largest' superlative rather than the usability of the benchmark, so I do not treat it as the primary blocker. The proposed independent re-annotation test would settle whether the final labels are reliable; if it cannot be run because the data are not released, the dataset claim remains conditional.","tokens_in":22275,"tokens_out":7222,"duration_ms":91641,"concrete_test":"Sample 200 test questions (or the full test set if feasible). Have two independent legal experts, not involved in construction, annotate from scratch: for each question, return the set of relevant articles from the full 59,636-article corpus and a correctness verdict on the gold answer. Compute exact article-set agreement and Cohen's or Fleiss' kappa between the two experts and between each expert and the released VLQA labels. Also examine the 14.83% 'article-modified' cases specifically. If the independent experts agree with VLQA article sets at >=80% exact match and kappa >=0.7, the single-reviewer concern is resolved; otherwise the benchmark's ground truth should be treated as unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that VLQA is an expert-annotated benchmark whose gold article/answer labels support reliable evaluation. That claim rests on the correctness of the final labels, which are produced by five senior law students double-annotating and then one legal expert reviewing (Section 3.1.3). The paper reports reviewer statistics (Table 3) but no inter-annotator agreement at either the student or reviewer level. Table 3 is double-edged: it shows the review process caught errors (74.18% exact matches, 14.83% article-only changes, 3.80% answer-only changes, 7.19% both changed), but precisely because 25.8% of student outputs were changed, the decisive quality signal is the expert reviewer's own judgment, and there is no evidence about its reliability or consistency. One expert's systematic blind spots, such as a particular legal subdomain or difficulty tracking temporal amendments, would propagate directly into the gold labels and hence into every retrieval/QA result on the test set. The absence of a release link also prevents independent inspection of the data, so the current paper provides no way to distinguish a high-quality benchmark from a single-reviewer artifact.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces VLQA, a Vietnamese legal question-answering dataset consisting of 3,129 question-answer-article triplets collected from public legal consultation forums and verified by a team of five senior law students under the supervision of one legal expert. The dataset is built on a corpus of 59,636 statutory articles from 2,162 Vietnamese legal documents, and is split into training, development, and test subsets. The authors conduct a statistical analysis of the dataset and benchmark multiple legal article retrieval models (BM25, fastText, SBERT, BGE-m3, mBERT, BGE-reranker) and question-answering models (extractive transformers, generative Vietnamese models, and several LLMs) using automatic metrics and a human evaluation of 100 sampled outputs. The paper claims that VLQA is the first comprehensive, large-scale, expert-annotated Vietnamese benchmark for legal QA and the largest real-world expert-verified LQA dataset covering any statutory domain.","tokens_in":22491,"tokens_out":3824,"duration_ms":44454,"significance":"If the dataset is released and its quality is validated, VLQA would be a valuable resource for legal NLP in Vietnamese, a low-resource language, and would complement existing benchmarks such as LLeQA and EQUALS. The paper's strengths include a detailed and transparent description of the construction pipeline, the presentation of reviewer-modification statistics in Table 3, concrete examples of corrections in Table 4, and an honest qualitative analysis of LLM hallucination behavior in Tables 14 and 15. The experimental section covers a wide range of baseline models and both automatic and human evaluation. The main caveats are that the dataset is not yet publicly available, no inter-annotator agreement is reported, and the 'largest' claim appears to conflict with the same table's listing of EQUALS, which has more questions. These issues are central to the paper's contribution as a benchmark dataset, but they are addressable in revision.","major_comments":[{"comment":"The claim that VLQA is 'the largest real-world expert-verified LQA dataset covering any statutory domain' (Section 1) is not supported by the paper's own Table 1, which lists EQUALS with 6,914 questions from a web-forum source. Since EQUALS is also a real-world dataset, the 'largest' claim requires an explicit justification of why EQUALS does not count (e.g., lack of expert verification) and evidence for that distinction. As written, the claim is internally inconsistent with the presented comparison, and this affects a central contribution of the paper.","section":"Section 1 and Section 3.2.1, Table 1"},{"comment":"The gold labels rest on a single legal expert's review, yet no inter-annotator agreement is reported at either the student-annotation stage or the reviewer stage. Table 3 shows that the reviewer modified 25.8% of the student annotations, so the final labels depend substantially on the reviewer's individual judgment. Without a measured agreement statistic (e.g., Cohen's kappa between the two student annotators, or between two independent expert reviews on a subset), the 'high-quality' and 'reliable' claims for the dataset are not established. A systematic blind spot of the reviewer would propagate directly into all retrieval and QA evaluations. I recommend reporting IAA metrics and, if feasible, having a second legal expert independently verify a random subset.","section":"Section 3.1.3, Table 3"},{"comment":"The paper states that the dataset and source code 'will be publicly released soon,' but no release link, data availability statement, or repository identifier is provided in the current manuscript. For a dataset paper, the central artifact must be accessible for independent verification and for the benchmark to be usable by the community. The manuscript should include a URL or a clear availability plan (e.g., an anonymous link for review) before acceptance.","section":"Section 1 and Section 7"},{"comment":"The question-answering evaluation does not specify which retriever and which value of k are used to supply the 'top-k relevant articles' to the QA models. This makes the QA results in Tables 10 and 11 non-reproducible and complicates interpretation, since QA performance depends critically on the quality of the retrieved context. Please state the retriever, k, and whether the same retrieved articles are used for all QA models.","section":"Section 5.2.2, Tables 10 and 11"}],"minor_comments":[{"comment":"The manuscript contains placeholder-like metadata: 'Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009' and a 2018 copyright notice. These should be updated or removed.","section":"Section 1 and first page"},{"comment":"Table 1 reports an average gold-answer length of 216.74 words, whereas Section 6.3 states that gold answers average 198 words. Please clarify whether the latter figure refers only to the 100-sample human-evaluation subset or to the test set, and align the numbers.","section":"Table 1 vs. Section 6.3"},{"comment":"The corpus collection step says documents with identical titles but different subjects are excluded. This may remove legally distinct instruments; please clarify the criterion and its potential effect on corpus coverage.","section":"Section 3.1.1"},{"comment":"The reasoning-type percentages sum to more than 100%, which the caption notes is because one question may require multiple reasoning abilities. The table itself does not show individual question counts; adding counts for each reasoning type would help interpret the percentages.","section":"Table 7"}],"recommendation":"major_revision","confidential_remarks":"The paper's central contribution is a dataset, and the main risks are the unverified single-reviewer gold labels and the unresolved comparison with EQUALS. The authors should be encouraged to release a data sample for review and report IAA; if these are addressed, the paper could become a solid benchmark contribution. The overlap with ALQAC (shared co-author) is minor and does not appear to affect the novelty of the dataset."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"VLQA is a real contribution to low-resource legal NLP: 3,129 real Vietnamese citizen questions, expert-verified article references and long-form answers, over a 59,636-article corpus. That combination did not exist for Vietnamese, and the construction pipeline is described in enough detail that you can see where the data come from and how they were filtered. The statistical analyses and baseline experiments (BM25, dense retrieval, extractive and generative QA, plus LLM evaluation) are standard but appropriate, and the human evaluation of 100 samples gives useful qualitative signal beyond ROUGE/BERTScore.\n\nThe soft spots are two, and they are the ones you would expect. First, there is no release link; the paper says 'will be publicly released soon,' so nothing can be independently verified. Second, no inter-annotator agreement is reported. Five law students double-annotated, and one legal expert reviewed. Table 3 shows the expert changed 25.8% of the samples. That is double-edged: it suggests the review was real, but it also means the final gold labels are effectively one reviewer's judgment. Without IAA, we cannot tell how reliable the construction process is. These issues are fixable—release the data, report agreement—but until then the 'high-quality' claim is unsubstantiated.\n\nThe 'largest' claim also overreaches. Table 1 lists EQUALS with 6,914 questions, so VLQA is not the largest by question count. If the claim is about the statutory corpus (59,636 articles), say that. As written it looks like an error.\n\nNone of this is load-bearing to the resource's value. The data are not circular; they come from external forums, not model outputs. The design is sound. I would want the data and IAA before using it, but this is a serious dataset paper, not a desk reject.\n\nWho it is for: Vietnamese legal NLP and low-resource legal QA/IR researchers. They will get real value if the dataset is released. Recommendation: send to peer review and require major revision—data release, an IAA section, and a corrected 'largest' statement. If those land, this becomes a useful benchmark.","headline":"A genuinely needed Vietnamese legal QA dataset, built carefully enough to be useful, but the missing release and single-reviewer ground truth make the 'largest' and 'high-quality' claims untestable as submitted.","tokens_in":23000,"tokens_out":3493,"would_cite":false,"duration_ms":43090,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces VLQA, a new benchmark of 3,129 real-world Vietnamese legal questions with expert-verified answers and statute citations, and uses it to show that current LLMs score well on automatic metrics but often hallucinate in…","keywords":["Vietnamese legal question answering","legal information retrieval","statutory article retrieval","expert-annotated dataset","low-resource NLP","large language models","legal hallucination","benchmark"],"falsifier":"Take a random sample of, say, 100 VLQA test questions, have a separate group of licensed Vietnamese lawyers independently identify the relevant statutory articles without seeing the gold labels, and measure agreement; if the independent lawyers match the dataset's article sets on substantially fewer than 90% of questions, the gold standard and the benchmark rankings built on it are not stable.","tokens_in":22114,"feed_emoji":"⚖️","tokens_out":4914,"duration_ms":50996,"temperature":0.7,"pith_summary":"The paper introduces VLQA, a dataset of 3,129 legal questions asked by Vietnamese citizens, each paired with an expert-verified answer and citations to relevant articles from a corpus of roughly 59,636 statutory provisions. The authors claim this is the largest expert-verified, long-form legal QA dataset for any statutory domain, and they position it as a benchmark for two tasks: legal article retrieval and legal question answering. They evaluate sparse and dense retrievers, extractive and generative QA models, and several large language models, finding that while LLMs achieve high scores on automated metrics, human evaluation shows their outputs often contain factual errors or hallucinated content. The paper's core contribution is the dataset itself, together with baseline measurements that reveal how far current systems are from reliable legal assistance.","feed_headline":"3,129 Vietnamese legal questions now a public benchmark","feed_subtitle":"Expert-verified statute citations ground each of 3,129 citizen questions, pitting LLMs against retrieval baselines.","key_machinery":"The central object is the VLQA benchmark: 3,129 question-answer-article triplets built from real citizen questions, paired with a hierarchical corpus of 59,636 Vietnamese law articles. The supporting mechanism is the construction pipeline: regular-expression parsing of roughly 430,000 raw forum posts to extract question-answer-article references, independent annotation by five senior law students, and final verification by a supervising legal expert, whose revisions changed either answers or cited articles on about 26% of samples.","core_discovery":"The central claim is that VLQA is the first comprehensive, large-scale, expert-annotated legal QA dataset for Vietnamese, built from 3,129 questions sourced from public legal consultation platforms, each validated by legal professionals and linked to relevant articles in a 59,636-article corpus spanning 27 legal domains. The dataset supports two tasks: statutory article retrieval (given a question, find the relevant articles) and legal question answering (produce a long-form answer). The paper reports that a fine-tuned multilingual BERT retriever outperforms zero-shot dense and sparse baselines, and that GPT-4o-mini achieves the highest automatic QA scores; however, qualitative analysis of 100 test samples shows that LLM outputs frequently contain incomplete, logically incorrect, or hallucinated elements even when the text is fluent.","pith_inferences":["Because the questions come from public forums and cover everyday concerns, VLQA could also be used to study how LLMs explain legal rights to laypeople, a use beyond model ranking that the paper does not explore.","A natural stress test for VLQA would be to measure whether a legal QA system trained on it actually helps non-expert users apply the law to new situations, which would test practical utility rather than benchmark scores.","The absence of reported inter-annotator agreement means that a re-annotation study with independent legal experts would be the immediate next check on the gold labels; if agreement is low, the comparative ranking of retrieval and QA models could shift.","If VLQA is updated as Vietnamese statutes change, it could serve as a longitudinal test of whether models track legal amendments, since the paper already shows that repealed provisions must be caught during annotation."],"forward_implications":["VLQA gives Vietnamese legal AI a common testbed for article retrieval and answer generation, enabling fair comparison of models on realistic layperson questions.","The baseline results establish that domain-adapted dense retrieval beats zero-shot models, but even the best retriever identifies fewer than half of the relevant articles in its top two, leaving substantial room for improvement.","Automated metrics such as ROUGE and BERTScore overestimate LLM quality in legal QA, since human review shows fluent, well-structured answers can still contain incorrect figures or fabricated legal references.","The finding that smaller LLMs degrade with few-shot examples while larger ones improve suggests that context length and memory capacity limit legal QA performance on this benchmark."],"supporting_citations":[{"why":"Provides the closest bar-exam-style legal QA comparison, one of the datasets VLQA claims to surpass in scale and real-world grounding.","marker":"[11]"},{"why":"The prior large-scale legal QA dataset (multiple-choice from the Chinese bar exam) that anchors the comparison for dataset size and answer format.","marker":"[47]"},{"why":"The prior long-form legal QA dataset from Belgian consulting platforms whose design VLQA extends to Vietnamese statutory law.","marker":"[23]"},{"why":"A real-world Chinese legal QA dataset from web forums that VLQA compares against in source and answer type.","marker":"[5]"},{"why":"The statutory article retrieval dataset in French that motivates the retrieval task and the corpus-scale comparison.","marker":"[22]"},{"why":"The prior Vietnamese legal QA competition dataset that VLQA contrasts with, being narrow and jurist-posed rather than real-world.","marker":"[26]"}],"fun_headline_variants":["First large Vietnamese legal QA dataset: 3,129 expert-checked questions","Vietnamese legal QA benchmark with 3,129 questions and 59k statutes","New Vietnamese legal QA dataset exposes LLM hallucination gaps","VLQA: 3,129 expert-validated questions for Vietnamese legal QA"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the article citations and answers produced by five senior law students and one supervising expert are correct and complete for every one of the 3,129 questions; if those gold labels are wrong or miss relevant provisions, every retrieval and QA score in the paper is comparing models against a flawed ground truth.","fun_headline_variants_meta":{"raw":{"variants":["First large Vietnamese legal QA dataset: 3,129 expert-checked questions","Vietnamese legal QA benchmark with 3,129 questions and 59k statutes","New Vietnamese legal QA dataset exposes LLM hallucination gaps","VLQA: 3,129 expert-validated questions for Vietnamese legal QA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000615,"raw_usage":{"total_tokens":2842,"prompt_tokens":913,"completion_tokens":1929,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":1849}},"tokens_in":529,"tokens_out":1929,"duration_ms":16890,"temperature":1.0,"reasoning_tokens":1849,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T13:50:36.378835+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of, say, 100 VLQA test questions, have a separate group of licensed Vietnamese lawyers independently identify the relevant statutory articles without seeing the gold labels, and measure agreement; if the independent lawyers match the dataset's article sets on substantially fewer than 90% of questions, the gold standard and the benchmark rankings built on it are not stable.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the closest bar-exam-style legal QA comparison, one of the datasets VLQA claims to surpass in scale and real-world grounding."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The statutory article retrieval dataset in French that motivates the retrieval task and the corpus-scale comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The prior Vietnamese legal QA competition dataset that VLQA contrasts with, being narrow and jurist-posed rather than real-world."}],"review_version":1}