{"id":"50f062ad-c36a-4834-be20-372116346fd3","arxiv_id":"2502.09778","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"Word-by-word retrieval-based prompting with GPT-4 improves morpheme-level glossing over the SIGMORPHON 2023 baseline, and a 3-best oracle beats the tuned challenge winner in five of seven languages.","lead":"This paper tests whether giving GPT-4 word-by-word examples and instructions helps it create interlinear glosses for seven endangered or underdocumented languages. The system beats the official baseline on morpheme-level scores in all languages, and a three-best oracle outperforms the previous challenge winner in five, suggesting LLMs could assist human annotators.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The strongest claim assumes GPT-4o has not memorized the SIGMORPHON 2023 test glosses; the Appendix A canary (Tsez only, verbatim completion, safety refusals) does not rule this out, so the reported gains may be inflated.","rationale":"The most defensible reading of the paper is as an empirical demonstration that retrieval-based word-by-word prompting can produce useful gloss suggestions. The headline numbers depend on GPT-4o's test-set novelty. Appendix A is the only support for novelty, and it is not probative for gloss memorization: safety refusals on verbatim completion are compatible with memorization, only Tsez is tested, and recall of glosses is a different capability from recall of surface strings. This is load-bearing because all score comparisons—including the small margins in Gitksan and Uspanteko—would be inflated if the model had seen gold labels. The reader's conditional verdict already rests on this same assumption, so the stress-test agrees. I do not see an internal inconsistency in the oracle claim: it is explicitly an oracle and the paper does not overstate it as human performance. The main missing piece is a direct leakage probe. If the proposed probe shows no memorization, the central claim stands as a useful conditional result; if it shows memorization, the comparisons should be recomputed on a truly held-out set. Because the paper already acknowledges the need for more evidence in its Limitations section, no change to the conditional verdict is needed.","tokens_in":993,"tokens_out":895,"duration_ms":79937,"concrete_test":"Run a memorization probe that asks for glosses, not sentence completions, and covers all seven languages. For each language, randomly sample 100 test sentences and prompt GPT-4o with the Section 4.1 scaffold but with all retrieved examples, tag distributions, and reverse-lookup material removed—only the sentence, its translation, and the JSON instruction remain. In a matched control condition, replace the target word with a nonce string that does not occur in the training data while keeping the rest of the sentence and the translation identical. If top-1 gloss accuracy on the original test tokens substantially exceeds accuracy on nonce controls, or if retrieval-ablated accuracy remains close to the reported retrieval-augmented scores, the model is exploiting memorized test-set glosses.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparisons—beating the BERT baseline at morpheme level in all seven languages and the 3-best oracle surpassing the challenge winner in five—require that the test-set gold glosses were not seen by GPT-4o during pretraining. The only evidence offered, Appendix A, asks GPT-4o to complete Tsez test sentences verbatim and reports refusals. This does not test whether the model can recall test-sentence glosses: (i) it is run for Tsez only, not the other six languages; (ii) it tests surface-form completion, not gloss-label retrieval; and (iii) \"I'm sorry, I can't provide verbatim text...\" is a safety-policy response that can occur even when the underlying text is perfectly memorized. A model could memorize the gold gloss for a test token while still refusing to emit the full sentence. Because the reported margins are sometimes tiny (Gitksan morpheme-level 8.68 vs 8.54; Uspanteko 57.59 vs 57.24), even a small amount of memorization could flip the \"beats baseline in every language\" claim. The paper's Limitations section concedes single runs without significance testing, but that is a variance issue; undetected test-set memorization would invalidate the comparison regardless of variance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a word-by-word retrieval-based prompting approach to interlinear glossing, applied to the seven languages of the SIGMORPHON 2023 shared task. The system retrieves exact and approximate matches, reverse-indexed translation words, and tag-frequency summaries for each target word, and asks GPT-4o to return a 3-best list of glosses. The main claims are that the system beats the BERT baseline at morpheme-level accuracy for all seven languages, that a 3-best oracle exceeds the challenge winner's word-level scores in five languages, and that in a Tsez case study, automatically generated linguistic instructions reduce confusions between syncretic tags such as PFV.CVB and PST.UNW. The paper reports improved word-level accuracy on Tsez from 75.28 to 75.86 after adding these instructions, and a reduction in CVB-related confusions.","tokens_in":17693,"tokens_out":4292,"duration_ms":40238,"significance":"If the results hold, the paper provides a useful empirical data point for interactive glossing: word-by-word prompting is competitive with whole-sentence prompting, the k-best oracle quantifies a plausible upper bound for human-in-the-loop annotation, and the Tsez instruction-following case study is one of the few demonstrations that LLMs can apply abstract linguistic rules to concrete glossing decisions. The authors deserve credit for honestly stating their limitations: single runs, no significance testing, a reduced Arapaho test set, and the use of a proprietary model. The paper also makes its code and results available, which strengthens reproducibility. However, the significance is tempered by the small margins in several comparisons, the lack of any variance analysis, and an unconvincing canary test for test-set memorization; these issues directly affect the central empirical claims.","major_comments":[{"comment":"The canary test in Appendix A is too weak to support the assumption that GPT-4o has not memorized the SIGMORPHON 2023 test glosses. It is run only for Tsez, it probes verbatim sentence completion rather than gloss recall, and the observed refusals ('I'm sorry...') are consistent with a safety policy that would apply even to perfectly memorized text. Because the claim that the system beats the BERT baseline in all seven languages at morpheme level (Section 4.2, Table 4) and the oracle comparisons in Section 5 depend on test-set novelty, this is a load-bearing gap. Please provide stronger evidence, for example a gloss-recall probe on held-out training versus test sentences, or a comparison on a newly collected evaluation set.","section":"Appendix A"},{"comment":"All reported scores come from single runs, as the Limitations section states, but the key margins are often smaller than any plausible run-to-run variation. For example, Gitksan morpheme accuracy is 8.68 versus the baseline's 8.54 (Table 4), and Uspanteko morpheme accuracy is 57.59 versus 57.24; the Tsez instruction improvement is 75.28 to 75.86 (Section 6.2). Without variance estimates or repeated runs, the headline claim of beating the baseline in every language is not established. I ask for at least a few repeated runs at different temperatures or seeds for the languages with the smallest margins, or an explicit error-bar analysis.","section":"Section 4.2, Tables 1 and 4; Limitations"},{"comment":"The Arapaho results in all tables are based on the first 100 test sentences only (arp*), yet they are placed next to the full-test baseline and challenge-winner scores from the shared task. A 100-sentence subset is not guaranteed to be representative, so the comparisons for arp* (e.g., word-level 66.19 vs. 71.14 and morpheme-level 52.57 vs. 44.19) are not controlled and should not be used to infer relative system quality. The paper should either run the full Arapaho test set or report the baseline and winner scores computed on the same 100-sentence subset.","section":"Tables 1, 2, 4, 5; Section 3"},{"comment":"The conclusion states that the system 'surpass[es] the word-level test scores of the Track 1 challenge winner in Gitksan, Lezgi, Nyangbo, Tsez, and Uspanteko' without immediately qualifying that this is the 3-best oracle, not the system's 1-best output. The distinction is important because the oracle uses gold tags at test time to choose among candidates; it is a ceiling for human-in-the-loop annotation, not a deployable predictor. The abstract and Table 2 caption make the oracle status clear, but the conclusion should too.","section":"Section 7, first paragraph"}],"minor_comments":[{"comment":"The caption contains a broken phrase: 'Table 2 in the shows word-level scores' should read 'Table 2 shows word-level scores'.","section":"Table 2 caption"},{"comment":"There are two typos: 'subection 4.2' and 'subection 5' should be 'subsection 4.2' and 'subsection 5'.","section":"Section 4.1"},{"comment":"In Table 4, the rows for ddo (Tsez) and usp (Uspanteko) both report 57.59 under 'Ours'; please verify that this is not a copy-paste error.","section":"Table 4"},{"comment":"The JSON example in Appendix D is not valid JSON because the 'glosses' field uses triple underscores inside quotes; the formatting should be cleaned up to match the intended 'best three glosses' output.","section":"Appendix D"},{"comment":"The notation 'arp*' is used in table captions but not defined at first use; please state explicitly that it refers to the first 100 sentences of the Arapaho test set.","section":"Tables 1, 2, 4, 5"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the memorization concern: the canary test in Appendix A is not convincing, but it is addressable with additional experiments. If the authors can provide a stronger leakage check or evaluate on a fresh dataset, the paper could become publishable. The single-run and Arapaho-subset issues are also fixable. I recommend major revision rather than rejection because the central approach is plausible, the limitations are disclosed, and the requested changes are within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi colleague,\n\nMy read of Elsner and Liu's glossing paper: the genuinely new piece is the Tsez instruction-generation experiment. They build contrastive minimal pairs from the training data, ask GPT-4 to write disambiguation rules, then inject those rules into the prompt. The error analysis is careful—they show both successes and failures, including a rule that is factually wrong about the grammar, and they trace how the chain-of-thought sometimes over-relies on the English translation. That is a real contribution worth building on.\n\nThe rest is more incremental. Word-by-word retrieval prompting is a variation of Ginn et al. 2024a, and the 3-best oracle is a sensible evaluation idea rather than a new model. The empirical claims are mostly in the expected direction, but weaker than the abstract suggests. The margins over the baseline are tiny in several languages—Gitksan morpheme accuracy 8.68 vs 8.54, Uspanteko 57.59 vs 57.24—and there are no significance tests, just single runs. Arapaho is evaluated on only the first 100 test sentences and compared to full-test published scores, which is not a clean comparison.\n\nThe bigger soft spot is the data-leakage check. The Appendix A canary asks GPT-4 to complete Tsez test sentences verbatim; the model refuses, and the authors take that as evidence it hasn't memorized the test data. That doesn't rule out memorization of glosses or labels, and the refusals could just be safety policy. Given how thin the margins are, even partial memorization could flip the \"beats the baseline everywhere\" claim. This deserves a proper membership test or at least a direct probe of gloss recall.\n\nWhat the paper does well is honesty. The Limitations section openly reports the cost-driven choices, the lack of significance testing, and the single-language focus of the instruction experiment. The oracle discussion is clearly labeled as a measure of potential, not actual human performance.\n\nBottom line: this is a serious empirical paper with one genuinely novel component and an exemplary limitations section. It deserves a real referee, but the referee should ask for variance estimates, a full Arapaho evaluation, and a stronger memorization check before the central claims are accepted.\n\nI'd bring it to the reading group—the instruction-generation part will spark discussion—and I'd cite it for that idea. Send it to review.","headline":"Genuinely novel instruction-generation experiment wrapped in an honest but statistically fragile empirical package; the data-leakage check is too weak to support the headline margins.","tokens_in":18252,"tokens_out":2817,"would_cite":true,"duration_ms":26152,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Retrieval-based word-by-word prompting of GPT-4 beats the BERT baseline at interlinear glossing in all seven tested languages.","keywords":["interlinear glossing","low-resource languages","LLM prompting","retrieval-augmented generation","morphological syncretism","Tsez","SIGMORPHON","human-in-the-loop annotation"],"falsifier":"Ask GPT-4 to produce the gold gloss for a random sample of Tsez (and other language) test sentences with no retrieval and no candidate tags; if it reliably outputs the gold glosses, then the retrieval scores are inflated by memorization. A cheaper check is to run the canary test in all seven languages asking for verbatim gloss lines rather than sentence completions.","tokens_in":17147,"feed_emoji":"📖","tokens_out":4771,"duration_ms":38307,"temperature":0.7,"pith_summary":"The paper argues that a word-by-word retrieval prompting strategy can make large language models useful assistants for interlinear glossing, the morpheme-by-morpheme annotation used in language documentation. It shows that this approach beats the official BERT baseline in all seven SIGMORPHON 2023 languages on morpheme-level accuracy, and that when a human picks among three generated glosses, the resulting oracle beats the tuned challenge winner in five languages. It also demonstrates that automatically generated linguistic instructions can reduce confusion between syncretic Tsez tags, such as perfective converb versus past unwitnessed, by about ten percent. The larger claim is that LLMs' ability to follow natural-language instructions makes them a practical interactive tool for linguists, not just a batch predictor.","feed_headline":"Word-by-word GPT-4 prompts beat BERT at interlinear glossing","feed_subtitle":"With three candidate glosses per word, an oracle beats the tuned winner in 5 of 7 languages.","key_machinery":"The load-bearing mechanism is word-by-word retrieval prompting with k-best elicitation. Each word gets up to three exact-match sentences, up to three approximate matches sharing a four-character substring, reverse-indexed words from the metalanguage translation, and a frequency summary of the word's tags over the training corpus; the LLM must return three ordered glosses in JSON. For the syncretism case, a second machine generates contrastive instances for a confused tag pair, asks the LLM for concise syntactic rules (with a hardcoded good/bad example pattern), and then inserts those rules plus a chain-of-thought step into the glossing prompt.","core_discovery":"The central discovery is that the model's top guess is often wrong while one of its three candidates is correct, and that this near-miss behavior is what makes LLM prompting valuable in human-in-the-loop annotation. For each target word the system retrieves exact and approximate matches, reverse-matches words from the translation, and supplies a distribution of the word's common tags; the LLM then ranks three candidate glosses. A Jaccard-based oracle over those three candidates yields word-level scores above the tuned sequence-model winner in Gitksan, Lezgi, Nyangbo, Tsez, and Uspanteko, and the same oracle beats the baseline everywhere. The paper also shows that asking the LLM to write disambiguation rules from contrastive examples, then injecting those rules into the prompt, reduces the dominant Tsez error class—confusions between PFV.CVB and PST.UNW—from 102 to 74 test-set errors, raising word accuracy from 75.28 to 75.86.","pith_inferences":["The oracle result may understate how well a real annotator would do, because the Jaccard oracle assumes perfect selection among the three options; a realistic annotator makes mistakes, so the practical benefit would likely be smaller but could still be positive.","If the method is applied to a truly unseen language with no training examples at all, the reverse retrieval and frequency summaries would disappear, leaving only the translation and the LLM's priors; the paper does not test this zero-example regime, but the Gitksan result (only 31 training sentences) suggests substantial degradation is likely.","A direct test of the memorization concern would be to run the same retrieval prompts on newly collected IGT from the same languages; if the gains vanish, the scores depend on test-set leakage rather than the method."],"forward_implications":["A human annotator working with a 3-best oracle could accept one of the top three glosses at rates exceeding the tuned sequence model in most languages, suggesting LLM suggestions can speed manual glossing.","Because word-by-word prompting uses only six retrieved examples per word, it can be cheaper and more interpretable than sentence-level prompting that retrieves up to 100 sentences.","The Tsez result indicates that LLMs can apply abstract grammatical instructions to concrete data when instructions are generated from contrastive examples, contrasting with repeated failures in low-resource translation.","The instruction-generation pipeline could be applied to other syncretic or confusable tag pairs in languages beyond Tsez, and to morphological disambiguation more broadly."],"supporting_citations":[{"why":"Supplies the seven-language dataset and the official baseline scores that the paper compares against.","marker":"Ginn et al. (2023)"},{"why":"Supplies the tuned challenge winner (hard-attention encoder-decoder) that the oracle and 1-best system are measured against.","marker":"Girrbach (2023)"},{"why":"Provides the sentence-level LLM prompt baseline that the paper compares to and extends to word-by-word retrieval.","marker":"Ginn et al. (2024a)"},{"why":"Shows that adding translation embeddings improves glossing, motivating the use of translations in prompts.","marker":"Yang et al. (2024a)"},{"why":"Descriptive grammar used to validate the Tsez syncretism analysis and the accuracy of generated instructions.","marker":"Polinsky (2014)"},{"why":"Motivates the canary test used to argue against memorization.","marker":"Carlini et al. (2019)"},{"why":"Documents LLM failure to benefit from grammar text in Kalamang, which the Tsez instruction result contrasts with.","marker":"Aycock et al. (2024)"}],"fun_headline_variants":["LLM near-miss glossing beats tuned winner in 5 of 7 languages","Word-by-word LLM prompts outgloss BERT and auto-write grammar rules","A 3-gloss oracle defeats the tuned champion for low-resource glossing","LLM-written rules cut a Tsez error class by 28%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"GPT-4 has not memorized the test-set glosses, so the reported scores reflect the prompting method rather than recall of the answer key.","fun_headline_variants_meta":{"raw":{"variants":["LLM near-miss glossing beats tuned winner in 5 of 7 languages","Word-by-word LLM prompts outgloss BERT and auto-write grammar rules","A 3-gloss oracle defeats the tuned champion for low-resource glossing","LLM-written rules cut a Tsez error class by 28%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001165,"raw_usage":{"total_tokens":4813,"prompt_tokens":925,"completion_tokens":3888,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":3800}},"tokens_in":541,"tokens_out":3888,"duration_ms":26688,"temperature":1.0,"reasoning_tokens":3800,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T20:31:39.268072+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask GPT-4 to produce the gold gloss for a random sample of Tsez (and other language) test sentences with no retrieval and no candidate tags; if it reliably outputs the gold glosses, then the retrieval scores are inflated by memorization. A cheaper check is to run the canary test in all seven languages asking for verbatim gloss lines rather than sentence completions.","supporting_citations":[],"review_version":1}