{"id":"92788490-0f7d-4b13-bfe5-43d9b74d809a","arxiv_id":"2412.10960","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"With only a dictionary and parallel sentences, GPT-4o-mini generated XLE grammar rules and mostly accurate lexical entries for Moklen, though the evaluation is not independently verified.","lead":"Researchers gave an AI language model a dictionary and sample sentences for Moklen, an endangered language in Thailand, and asked it to write a formal grammar. The model produced plausible rules and dictionary entries, but the paper's evaluation is too self-referential and small to confirm the method works.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never runs XLE on its generated grammar; without a parser check, 'coherent XLE grammar' and the 86/100 accuracy estimate remain unverified, and BERTScore alone cannot carry the central claim.","rationale":"The reader's verdict and my read align on rejection, but for a slightly different emphasis. The reader's weakest_assumption is the unpublished Spencer (2022) gold standard; that is real and serious. I would put the most load-bearing condition one step earlier: a grammar in XLE is coherent only if it can be compiled and parse sentences, and the paper never attempts this. The grammar-rule accuracy evaluation in Section 6.2 is a manual qualitative comparison, the lexical-entry evaluation in Section 6.3 is manual, and the Section 6.1 translation scores are reported after selecting among 48 context combinations on the same 40-sentence test set. In addition, the condition that produces 0.7110 in Section 6.1 explicitly says the 'XLE gold standard grammar' was incorporated, so it is not clear the generated grammar produced the gain. All three evaluation strands would be superseded by running the generated grammar in XLE on held-out sentences. If that parse test succeeds, the central claim has direct support; if it fails, the abstract's claim of coherent generated XLE grammar is false. One concrete test is enough to arbitrate, so I recommend keeping the reader's REJECT verdict rather than softening to CONDITIONAL until the parser result is available.","tokens_in":10076,"tokens_out":5760,"duration_ms":53652,"concrete_test":"Compile the generated grammar rules and the 100 generated lexical entries (at least for the TD+C+S condition) into an XLE grammar file, load it in the XLE parser, and parse a held-out set of Moklen sentences that were not used to select the best context (for example, 10 of the 40 test sentences or newly collected field-note sentences). Report parse coverage and compare the resulting f-structures against the gold-standard translations. If parse coverage is low or the f-structures conflict with the gold standard on the held-out set, the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim asserts that GPT-4o-mini, given a bilingual dictionary and few parallel sentences, generates coherent XLE grammar and lexical entries for Moklen. The load-bearing condition is that the output is actually a coherent, parseable LFG/XLE grammar. The paper never checks this condition. Section 6.2 scores grammar rules by manual comparison to a gold standard 'attempt[ed]' from Spencer (2022), an unpublished manuscript by the first author, and Section 6.3 reports 86/100 lexical entries as accurate by the authors' manual assessment. Section 6.1's headline BERTScore improvement (0.7110 for TD+C+S) is obtained after choosing among 48 context combinations on the same 40-sentence evaluation set, and the text says the 'XLE gold standard grammar' was incorporated in that condition, not the generated grammar. No experiment feeds the generated grammar and lexical entries to the XLE parser. If XLE rejects even the bitext sentences from which the grammar was induced, then the output is not a coherent XLE grammar, and the manual accuracy and BERTScore results cannot establish the central claim. The unpublished gold-standard issue compounds this: without an independent grammar source or a parser check, every accuracy judgment is either circular or unverifiable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using GPT-4o-mini with in-context learning to generate formal XLE (LFG) grammar rules and lexical entries for Moklen, an endangered Austronesian language, using only a bilingual dictionary and a small parallel bitext. The method consists of dictionary-based tokenisation, sense mapping, concatenation of per-word translations, and a prompt that varies several context types (bitext, tokenisation, dictionary, concatenated sentence, XLE example, self-explanation). Translation BERTScore is used to select among 48 context combinations; grammar rules are evaluated qualitatively against a gold standard constructed from Spencer (2022); lexical entries for 100 words not in the bitext are manually assessed against dictionary definitions. The paper reports that the TD+C+S context combination achieves BERTScore F1 of 0.7110 for Moklen-to-English translation when an XLE gold-standard grammar is added to the prompt, and that 86 of 100 generated lexical entries are accurate.","tokens_in":10274,"tokens_out":6885,"duration_ms":60932,"significance":"If properly validated, the approach would be a valuable contribution to language documentation, where manual grammar engineering is a major bottleneck. The paper deserves credit for working with a real endangered language, deliberately using a small model, and for stating that data, code, and model generations are in the supplementary material. However, the current evidence does not establish the central claim that the model generates a coherent XLE grammar: no parser is run on the generated grammar, the grammar gold standard is an unpublished manuscript by the first author, and the headline translation number appears to be selected on the test set. These are not merely presentation issues; they affect the validity of the main results. The idea is promising, but the paper needs substantially stronger validation before publication.","major_comments":[{"comment":"The reported peak BERTScore F1 of 0.7110 for TD+C+S is obtained after exploring 48 context combinations, and the evaluation appears to use the same 40-sentence parallel set described in Section 5.2 as the basis for selecting the best context. If so, the headline number is an in-sample selection optimum, and the comparison among contexts does not support the claim that TD+C+S is the best setting. Please clarify the exact split between the ablation set and the evaluation set, and ideally re-run context selection on a development set with a held-out test set.","section":"Section 6.1"},{"comment":"The grammar-rule evaluation relies on a gold standard that the authors 'attempt to create one based on Spencer (2022)', an unpublished manuscript by the first author. That source is neither available to readers nor independently verified, so the accuracy judgements for grammar rules are non-reproducible and potentially circular. More importantly, the paper never runs the XLE parser on the generated grammar and lexical entries, even though 'coherent XLE grammar' is the central claim. The translation improvements in Section 6.1 are obtained by adding the 'XLE gold standard grammar' to the prompt, not the grammar produced by the model, so those results do not validate the generated grammar. Section 6.2 also reports no quantitative accuracy or coverage score for the grammar rules.","section":"Sections 5.2 and 6.2"},{"comment":"The 86/100 lexical-entry accuracy is assessed by comparing generated entries to dictionary definitions that were also provided as context in the prompt (Sections 4.3 and 5.2). This makes the evaluation partly circular: high agreement with the input dictionary does not demonstrate that the model has learned Moklen grammar or generalised beyond the provided lexicon. In addition, the manual assessment is reported without a blind protocol or inter-annotator agreement, so the 86% figure is difficult to interpret.","section":"Section 6.3"}],"minor_comments":[{"comment":"The abstract contains typographical and grammatical errors, including 'Yes!' at the start and 'We takes Moklen as a case study'; please copyedit the entire manuscript.","section":"Abstract"},{"comment":"There are unresolved cross-references 'See ?? for full details' in Sections 5.1 and 6.1, and several figure captions are malformed or incomplete (for example, Figure 3 lacks a clear legend and readable axis labels).","section":"Sections 5.1 and 6.1"},{"comment":"The paper lists BLEU, ROUGE, METEOR, chrF, and BERTScore as evaluation metrics, but Section 6.1 reports only BERTScore; please report the other metrics or explicitly justify their omission.","section":"Section 5.2"},{"comment":"The tokeniser depends on the assumption that every word in the Moklen bitext appears in the dictionary; please verify this against the actual data and report the number of out-of-dictionary tokens, if any.","section":"Section 4.1"},{"comment":"The section title is 'Grammar Rules: Accuracy and Completeness', but no accuracy numbers are given; either add a quantitative evaluation or rename the section to reflect the qualitative discussion.","section":"Section 6.2"},{"comment":"The limitation section addresses typological generality, but it does not mention the lack of parser validation or the reliance on an unpublished gold standard; both should be acknowledged.","section":"Limitation"}],"recommendation":"major_revision","confidential_remarks":"The central problem is that the paper's main claim, the generation of a coherent XLE grammar, is never validated by running the XLE parser. The evaluation of grammar rules uses an unpublished manuscript by the first author as the gold standard, which creates a verifiability and independence concern. If the authors can add an XLE parser check, a held-out evaluation split, and an independent or at least blinded grammar assessment, a future revision could be viable; in its current form, the evidence is not sufficient for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something fresh: instead of using in-context learning to translate, it asks an LLM to write formal XLE grammar rules and lexical entries for an undocumented language from a dictionary and a few parallel sentences. That is a real task, worth trying, and the pipeline they build is easy to follow. The choice of Moklen is sensible, the prompting design is systematic, and the limitation section is honest about the single-language and English-bias concerns. I also credit them for using a cheap model; if this works, it lowers the skill and cost barrier for grammar engineering in documentation work.\n\nThe soft spots are not minor, though. The biggest one: the paper never runs the generated grammar through the XLE parser. They claim the LLM produces a coherent XLE grammar, but the only check is manual comparison of rules to a gold standard that was 'attempted' from the first author's unpublished manuscript, plus a manual count of lexical entries. A grammar that XLE rejects is not a coherent XLE grammar, regardless of what a human judges. BERTScore doesn't rescue this, especially because the headline number (0.7110) is from the condition where the gold standard grammar was put into the prompt, not the generated one. On top of that, the best context combination was selected from 48 options using the same 40-sentence test set that produces the reported scores, so the peak is in-sample. And the lexical-entry accuracy is checked against the very dictionary entries that were fed into the model, which is a weak form of verification. The stress-test note is right: these issues compound, and the central claim as stated is not established.\n\nThat said, the paper is not sloppy in its thinking. The pipeline is described clearly, the related work is relevant, and the authors acknowledge important limitations. The problems are in the evaluation design, not in the core idea. This is the kind of paper that could become solid after major revision: run XLE on the generated grammar, use a held-out test set for any context selection, get an independent grammar sketch or at least a released artifact for reproducibility, and grade the rules with a parser check rather than a handwritten gold standard.\n\nWho is this for? Computational linguists working on grammar engineering, endangered-language documentation, and ICL for low-resource languages. It deserves a serious referee, because the task is new and the failure mode is instructive. My recommendation: engage with it, but make the revision demands concrete and strict. If the parser check fails, the authors need to say so and adjust the claim.","headline":"A genuinely new task with a clean pipeline, but the evaluation never checks whether the generated grammar actually parses, so the central claim is not yet supported.","tokens_in":10826,"tokens_out":1350,"would_cite":false,"duration_ms":14973,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An LLM with in-context learning can generate coherent grammar for endangered Moklen from just a dictionary and parallel sentences, with 86 of 100 lexical entries judged accurate.","keywords":["endangered languages","Moklen","in-context learning","grammar induction","Lexical-Functional Grammar","XLE","low-resource NLP","language documentation"],"falsifier":"Have a second linguist with no access to Spencer (2022) independently write a reference grammar of Moklen from primary fieldwork, then compare it rule-by-rule with the gold standard and with the generated XLE grammar; if the independent grammar disagrees on core word order, parts of speech, or lexical senses, the evaluation baseline is unreliable. A more computational check is to feed the generated grammar to the XLE parser on the 40 held-out sentences and require a successful parse and f-structure for sentences the gold grammar parses.","tokens_in":9853,"feed_emoji":"🗣️","tokens_out":6520,"duration_ms":53876,"temperature":0.7,"pith_summary":"This paper tries to show that a large language model, prompted with only a Moklen-to-English dictionary and a handful of parallel sentences, can write a formal grammar for Moklen in the XLE formalism, plus lexical entries for words it never saw in sentences. The authors use Moklen, an endangered Austronesian language of southern Thailand with fewer than 1,000 speakers, as a test case because it is isolating and lacks inflectional morphology. They report that the best prompt combination, which adds tokenised dictionary senses and the model's own self-explanation, raises Moklen-to-English BERTScore F1 to 0.7110 when the generated grammar is included, and that 86 of 100 generated lexical entries are accurate against dictionary definitions. If the claim holds, grammar writing for endangered languages could start from the kinds of data linguists already collect, without pretraining a model or hiring a grammar engineer. The paper itself notes that English grammatical bias and the isolating typology of Moklen may have made the task easier.","feed_headline":"LLM drafts usable grammar for endangered Moklen from scant data","feed_subtitle":"With only a bilingual dictionary and parallel sentences, 86 of 100 generated lexical entries were judged accurate.","key_machinery":"The mechanism is an in-context learning pipeline with five stages: a dictionary-based tokeniser using longest-match segmentation; sense mapping that attaches English glosses to each Moklen token; concatenation of word-by-word glosses into a string that mirrors Moklen word order; prompting the LLM with these materials plus XLE documentation and optional self-explanation; and finally using the generated grammar rules to prompt lexical entries for unseen dictionary words. The load-bearing formal object is the XLE grammar, a computational implementation of Lexical-Functional Grammar whose rules and lexical entries pair constituent structure with functional structure. The pipeline's effect is to give the LLM enough aligned local evidence to infer word order, parts of speech, and lexical schemata by analogy with English, without any Moklen text in pretraining.","core_discovery":"The central discovery is that GPT-4o-mini, a small API-based LLM, can induce a coherent XLE grammar and useful lexical entries for a language it has not been trained on, provided its prompt contains a bilingual dictionary, tokenised and sense-mapped parallel sentences, and XLE documentation. The authors find that the most effective context for translation is tokenised dictionary senses plus concatenated word-by-word meanings plus self-explanation (TD+C+S), and that adding grammar to the prompt improves translation scores across most settings. In direct evaluation, 86 out of 100 lexical entries generated for Moklen words absent from the bitext were judged accurate and coherent, while the main observed error type was missing secondary senses of polysemous words and an over-application of English parts of speech such as determiners. This is framed as evidence that LLMs can assist language documentation, not replace it.","pith_inferences":["A testable extension is to run the identical pipeline on a morphologically rich endangered language; the paper's limitation note suggests the isolating typology of Moklen may be doing much of the work, so success on agreement-heavy languages would be a stronger test.","The BERTScore gain after adding grammar could partly reflect the evaluator's sensitivity to content-word overlap rather than true syntactic understanding, so a parse-based evaluation of the generated XLE grammar against independently collected sentences would separate the two.","Hallucinated categories that the model inserted into Moklen grammar, such as determiners, might function as hypotheses for field linguists to check, turning model errors into a discovery aid.","Dictionary completeness is the quiet precondition: the tokeniser fails if the bitext contains words absent from the dictionary, so field applications would need a morphological guesser or a mechanism to flag out-of-vocabulary items."],"forward_implications":["If correct, grammar construction for endangered languages could begin from a bilingual wordlist and a few dozen recorded sentences, greatly lowering the cost of formal grammar writing.","Generated lexical entries can extend a dictionary to words not present in any sentence corpus, since 86 of 100 entries in the study were judged accurate.","Adding the generated grammar to the prompt improves translation quality, suggesting grammar rules and lexical resources can feed back into machine translation.","The same prompting strategy may transfer to other formal grammar frameworks beyond XLE, such as dependency grammars or HPSG, as the paper suggests."],"supporting_citations":[{"why":"Supplies the benchmark for learning to translate a new language from a grammar book, which the translation evaluation follows.","marker":"Tanzer et al. (2024)"},{"why":"The work this paper builds on; contributes the in-context linguistic description methodology and multi-task evaluation for endangered languages.","marker":"Zhang et al. (2024)"},{"why":"XLE Documentation provides the formal grammar and lexical entry syntax that the model is prompted to reproduce.","marker":"Crouch et al. (2011)"},{"why":"The Moklen-Thai-English pilot dictionary is the source of the roughly 1,000 vocabulary items used for tokenisation, sense mapping, and lexical entry generation.","marker":"Pittayaporn et al. (2022)"},{"why":"The unpublished grammatical sketch of Moklen is used to build the gold standard grammar against which generated rules are judged.","marker":"Spencer (2022)"},{"why":"Provides the Thai-based orthography system that underlies the written Moklen data in the dictionary and bitext.","marker":"Pittayaporn and Choemprayong (Forthcoming)"}],"fun_headline_variants":["LLM induces Moklen grammar from dictionary and bitext alone","86/100 lexical entries accurate: LLM aids Moklen grammar","In-context learning: LLM writes grammar for unknown tongue","Small LLM crafts XLE grammar for endangered Moklen","LLM grammar induction rescues endangered Moklen with 86% accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The accuracy judgements presuppose that the unpublished reference grammar of Moklen used as the gold standard is correct and complete, and that the dictionary covers every word appearing in the bitext; if either premise fails, the reported accuracy numbers stop supporting the conclusions.","fun_headline_variants_meta":{"raw":{"variants":["LLM induces Moklen grammar from dictionary and bitext alone","86/100 lexical entries accurate: LLM aids Moklen grammar","In-context learning: LLM writes grammar for unknown tongue","Small LLM crafts XLE grammar for endangered Moklen","LLM grammar induction rescues endangered Moklen with 86% accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001024,"raw_usage":{"total_tokens":4283,"prompt_tokens":877,"completion_tokens":3406,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":493,"completion_tokens_details":{"reasoning_tokens":3318}},"tokens_in":493,"tokens_out":3406,"duration_ms":21269,"temperature":1.0,"reasoning_tokens":3318,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:26:32.966374+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a second linguist with no access to Spencer (2022) independently write a reference grammar of Moklen from primary fieldwork, then compare it rule-by-rule with the gold standard and with the generated XLE grammar; if the independent grammar disagrees on core word order, parts of speech, or lexical senses, the evaluation baseline is unreliable. A more computational check is to feed the generated grammar to the XLE parser on the 40 held-out sentences and require a successful parse and f-structure for sentences the gold grammar parses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the benchmark for learning to translate a new language from a grammar book, which the translation evaluation follows."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The work this paper builds on; contributes the in-context linguistic description methodology and multi-task evaluation for endangered languages."},{"cited_title":"Kaplan, Tracy Holloway King, John T","cited_arxiv_id":null,"evidence_quote":"XLE Documentation provides the formal grammar and lexical entry syntax that the model is prompted to reproduce."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Moklen-Thai-English pilot dictionary is the source of the roughly 1,000 vocabulary items used for tokenisation, sense mapping, and lexical entry generation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The unpublished grammatical sketch of Moklen is used to build the gold standard grammar against which generated rules are judged."}],"review_version":1}