{"id":"4862880c-b75e-4abe-a0e0-35abcb98180f","arxiv_id":"2606.09767","paper_version":1,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Synthetic dictionary data with LoRA fine-tuning teaches structure to Q'eqchi' NMT but fails to transfer semantics to organic inputs, showing overfitting to template constraints.","lead":"This paper creates synthetic parallel text from community dictionaries to train an NMT model for Q'eqchi' Mayan without web scraping, then applies LoRA fine-tuning to an mT5 model. It reports strong grammar learning on synthetic tests but a large drop on real glossary text, plus negative results from multi-task learning.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"High in-domain BLEU (42.02) may reflect template memorization rather than acquisition of morphology/VOS rules, since test data likely follows the same generative templates as training.","rationale":"The reader's weakest_assumption directly targets the same inference point (whether constrained synthetic variance supports generalization). The in-domain vs. organic contrast is the paper's key evidence, yet the in-domain result itself is the least secured link for the 'teach' claim. Full text does not appear to alter this based on the abstract's description of the pipeline.","tokens_in":1798,"tokens_out":363,"duration_ms":19741,"concrete_test":"Create a held-out synthetic test set by applying the same dictionary-to-sentence pipeline but with held-out word stems and at least one additional morphological combination or minor word-order perturbation not present in the training templates; recompute BLEU. If the score drops below ~15 while the original in-domain score remains high, the 'teaching' claim is supported; if it stays comparable, the in-domain result is consistent with template matching.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that synthetic constraints 'effectively teach complex agglutinative morphology and VOS word order' rests on the in-domain BLEU score. However, if the synthetic corpus is produced by a fixed set of dictionary-derived templates (as described), an in-domain test split drawn from the same distribution would be expected to yield high overlap even with limited rule learning. The structural-semantic gap on the organic glossary is then ambiguous: it could stem from lexical mismatch, domain shift, or true failure to internalize productive rules. No evidence is given that the test set contains novel morphological combinations or syntactic variants absent from training templates.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces a data synthesis methodology for low-resource NMT on Q'eqchi' Mayan by transforming community-sourced dictionaries into a synthetic corpus and applying LoRA-based PEFT on mT5-base. It claims that this approach achieves high structural acquisition with an in-domain BLEU of 42.02, demonstrating effective teaching of agglutinative morphology and VOS word order, but reveals a structural-semantic gap with BLEU 0.59 on organic glossary data due to overfitting to synthetic template constraints. An ablation with multi-task learning shows negative transfer.","tokens_in":1945,"tokens_out":539,"duration_ms":21234,"significance":"Should the central claims hold after addressing evaluation concerns, this work would contribute to data-sovereign methods for Indigenous language NMT by showing synthetic data's utility as a structural primer. The contrast between in-domain and out-of-domain performance underscores the limitations of template-based synthesis for semantic grounding. The efficient use of LoRA adapters is noted as a practical strength for low-resource settings.","major_comments":[{"comment":"The claim that 'synthetic constraints effectively teach complex agglutinative morphology and VOS word order' based on the in-domain BLEU score of 42.02 is undermined by the lack of evidence that the test set includes novel morphological combinations or syntactic variants absent from the training templates; high scores may result from memorization of fixed dictionary-derived templates rather than rule acquisition.","section":"Abstract"},{"comment":"The structural-semantic gap interpretation of the BLEU 0.59 on the organic glossary lacks controls or details on template generation rules, data splits, baseline comparisons, or statistical significance, making it unclear whether the gap stems from lexical mismatch, domain shift, or true failure to internalize productive rules.","section":"Abstract"},{"comment":"The ablation study reporting negative transfer in the Multi-Task Learning architecture does not provide specifics on the auxiliary tasks, how they lead to competition for LoRA capacity, or quantitative metrics supporting the over-optimization conclusion.","section":"Ablation study"}],"minor_comments":[{"comment":"The abstract refers to a 'massive synthetic corpus' without reporting its size, the number of templates, or the exact rules used for generation, which would aid in assessing the constrained structural variance.","section":null},{"comment":"More information on the organic glossary evaluation setup, including how it differs from the synthetic data in terms of vocabulary and syntax, would clarify the nature of the observed gap.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our manuscript. We address each major comment point-by-point below, agreeing where revisions are needed to strengthen the claims and providing the strongest honest defense based on the work presented.","responses":[{"response":"We acknowledge that the in-domain BLEU of 42.02 alone does not conclusively demonstrate acquisition of productive rules versus memorization, as the test set composition relative to training templates was not explicitly analyzed for novelty. The synthetic data generation applies rule-based transformations from the dictionary that combinatorially vary morphology and syntax, but without a breakdown of unseen combinations, the interpretation remains suggestive rather than definitive. We will revise the abstract and add a dedicated analysis quantifying novel morphological and syntactic variants in the test set, reporting performance stratified by novelty to better support the claim.","revision_made":"yes","referee_comment":"[Abstract] The claim that 'synthetic constraints effectively teach complex agglutinative morphology and VOS word order' based on the in-domain BLEU score of 42.02 is undermined by the lack of evidence that the test set includes novel morphological combinations or syntactic variants absent from the training templates; high scores may result from memorization of fixed dictionary-derived templates rather than rule acquisition."},{"response":"The manuscript interprets the low out-of-domain BLEU as evidence of a structural-semantic gap due to template overfitting, consistent with the observed maintenance of grammatical integrity but poor lexical grounding. However, we agree that without the requested controls, alternative explanations cannot be ruled out. We will revise to add explicit details on template generation rules, data split methodology, baseline model comparisons (e.g., untuned mT5), and statistical significance testing of the BLEU difference to clarify the gap's source.","revision_made":"yes","referee_comment":"[Abstract] The structural-semantic gap interpretation of the BLEU 0.59 on the organic glossary lacks controls or details on template generation rules, data splits, baseline comparisons, or statistical significance, making it unclear whether the gap stems from lexical mismatch, domain shift, or true failure to internalize productive rules."},{"response":"The ablation demonstrates negative transfer, which we attribute to auxiliary tasks competing for limited LoRA capacity and leading to over-optimization on synthetic markers. We agree the section lacks the requested specifics. We will expand it to detail the auxiliary tasks, describe the capacity competition mechanism (e.g., via shared adapter updates), and report quantitative metrics such as per-task BLEU deltas and indicators of overfitting to support the conclusion.","revision_made":"yes","referee_comment":"[Ablation study] The ablation study reporting negative transfer in the Multi-Task Learning architecture does not provide specifics on the auxiliary tasks, how they lead to competition for LoRA capacity, or quantitative metrics supporting the over-optimization conclusion."}],"tokens_in":1482,"tokens_out":612,"duration_ms":20272,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The one or two things to know: this case study on Q'eqchi' finds that synthetic data from dictionaries plus LoRA fine-tuning produces high in-domain BLEU but almost none on organic text, and multi-task learning leads to negative transfer instead of improvement.\n\nThey convert community dictionaries into synthetic parallel sentences via templates, then apply parameter-efficient tuning on mT5. The in-domain result of 42 BLEU is taken as evidence that the constraints teach agglutinative morphology and VOS order. The organic glossary evaluation at 0.59 BLEU shows the model sticks to grammatical patterns but doesn't pick up natural lexical grounding. They also ran an ablation with multi-task learning that showed negative transfer, which they link to the adapters over-optimizing for synthetic markers.\n\nWhat stands out is the focus on data sovereignty and avoiding extractive scraping for an Indigenous language. Reporting the specific gap between structural and semantic performance, along with the negative MTL outcome, adds a concrete observation to the literature on synthetic data for low-resource NMT.\n\nThe soft spots are around the evaluation design. The abstract gives no information on how the templates were generated or what the data splits look like, so it's possible the in-domain test set overlaps heavily with training patterns. That makes the claim of learning morphology rest on shaky ground, as the stress-test note suggests. If the test data doesn't include novel combinations, the high score could be memorization rather than generalization. The paper would benefit from more controls there.\n\nThis is aimed at people building translation tools for similar low-resource Indigenous languages. A reader working on synthetic data strategies or PEFT for NMT would find the ablation and the reported gap useful to think about. It deserves serious referee time because the application is timely and the negative result on MTL is worth documenting, even if the methods need fleshing out.\n\nI'd recommend sending it for peer review, with the expectation that reviewers will ask for details on template construction and additional tests for productive rule learning.","headline":"Synthetic dictionary data plus LoRA gets high in-domain BLEU on Q'eqchi' but near-zero on organic text, with the in-domain gains likely reflecting template overlap rather than productive morphology learning.","tokens_in":2436,"tokens_out":492,"would_cite":false,"duration_ms":28501,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Synthetic data from dictionaries teaches Q'eqchi' morphology and VOS order to an NMT model yet leaves a lexical gap on real sentences.","keywords":["Q'eqchi' Mayan","low-resource NMT","data synthesis","LoRA fine-tuning","agglutinative morphology","synthetic corpus","structural-semantic gap"],"falsifier":"Evaluating the trained model on a set of naturally occurring Q'eqchi' sentences with independent English translations; if BLEU scores stay near 0.59 even after further increases in template diversity, the structural-semantic gap cannot be closed by synthesis alone.","tokens_in":2694,"feed_emoji":"","tokens_out":696,"duration_ms":22137,"temperature":0.7,"pith_summary":"The paper establishes that generating parallel sentence pairs from community dictionaries and fine-tuning mT5 with LoRA adapters produces strong in-domain performance on an agglutinative language. This shows the synthetic templates successfully encode complex morphology and verb-object-subject syntax. When the same model is tested on an organic glossary, however, performance collapses because the outputs remain grammatically well-formed but fail to match natural lexical choices. The authors therefore conclude that synthetic bootstrapping supplies an effective structural foundation but cannot supply the semantic flexibility of authentic language without additional real data.","feed_headline":"Dictionary synthesis trains Mayan NMT to 42 BLEU but fails on real text","feed_subtitle":"Model acquires morphology and VOS order from templates yet scores 0.59 BLEU on organic glossary due to lexical mismatch.","key_machinery":"Dictionary-derived synthetic corpus generated via constrained templates, paired with LoRA adapters on mT5-base for parameter-efficient adaptation.","core_discovery":"Transforming community-sourced dictionaries into a massive synthetic corpus and applying LoRA adapters to mT5-base yields BLEU 42.02 on in-domain synthetic test data, confirming that constrained templates teach agglutinative morphology and VOS word order. Evaluation on an organic glossary drops to BLEU 0.59, exposing a structural-semantic gap in which the model preserves grammatical integrity but lacks lexical grounding. An ablation with multi-task learning produces negative transfer, indicating that auxiliary tasks over-optimize the adapters for synthetic markers at the expense of flexibility on natural inputs.","pith_inferences":["The pipeline supports data-sovereignty goals by avoiding any web scraping of target-language text.","The same dictionary-to-synthetic approach could be tested on other Mayan or agglutinative languages to determine whether the observed gap is language-specific.","Increasing the diversity of generation templates might reduce the rigidity that forces organic inputs into learned patterns."],"forward_implications":["Synthetic templates alone suffice to teach complex morphology and VOS order when test data matches the training distribution.","Performance measured only on synthetic data overestimates readiness for real-language use.","Multi-task learning on top of LoRA adapters causes negative transfer by competing for limited adapter capacity.","Authentic parallel data is required after synthetic priming to refine semantic mappings via curriculum learning."],"fun_headline_variants":["Synthetic dictionaries train Q'eqchi' NMT to 42 BLEU but score 0.59 on glossary","Q'eqchi' NMT from synthetic data scores 42 BLEU but 0.59 on real glossary","Mayan NMT from dictionary synthesis scores 42 BLEU synthetic 0.59 on organic glossary","LoRA adapters with synthetic Q'eqchi' data score 42 BLEU in domain 0.59 on glossary"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The constrained structural patterns in the synthetic templates are broad enough to let the model generalize to the syntactic fluidity and lexical variety of natural Q'eqchi' sentences.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic dictionaries train Q'eqchi' NMT to 42 BLEU but score 0.59 on glossary","Q'eqchi' NMT from synthetic data scores 42 BLEU but 0.59 on real glossary","Mayan NMT from dictionary synthesis scores 42 BLEU synthetic 0.59 on organic glossary","LoRA adapters with synthetic Q'eqchi' data score 42 BLEU in domain 0.59 on glossary"]},"model":"grok-4.3","cost_usd":0.017721,"raw_usage":{"total_tokens":7575,"prompt_tokens":758,"num_sources_used":0,"completion_tokens":111,"cost_in_usd_ticks":177212000,"prompt_tokens_details":{"text_tokens":758,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":6706,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":758,"tokens_out":111,"duration_ms":41723,"temperature":1.0,"reasoning_tokens":6706,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T16:18:01.434579+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Evaluating the trained model on a set of naturally occurring Q'eqchi' sentences with independent English translations; if BLEU scores stay near 0.59 even after further increases in template diversity, the structural-semantic gap cannot be closed by synthesis alone.","supporting_citations":[],"review_version":1}