REVIEW 2 major objections 5 minor 34 references
Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Russenorsk, a dead Russian-Norwegian trade pidgin, can be analyzed and reconstructed by an LLM pipeline built on a new structured dictionary.
desk verdict Useful artifacts and an honest ablation, but the central 'hypotheses align' claim lacks a defensible mapping procedure. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the agentic discovery pipeline: a large language model is first prompted with the cleaned 691-entry Russenorsk lexicon to propose lexical-origin hypotheses; it is then given the same lexicon to infer phonetic and morphological transformations; and finally it receives dictionary entries and example sentences to infer grammatical rules. The reconstruction stage reuses all three outputs as context and instructs the model to prefer known lexemes over invented ones. The evaluation machinery is a three-step ablation that removes examples, then lexicon, then rules, scored with character-level chrF to quantify how much linguistic signal each layer actually contributes.
What would settle it
Construct a held-out set of historical Russenorsk sentences that appear in no published dictionary example and no prompt, run the full pipeline and the examples-removed ablation on that set, and have a Scandinavian contact-language expert score the reconstructions blind; if the full pipeline does not clearly beat the ablation, the claim that the dictionary and rules add real linguistic signal is refuted.
Extended reading notes
Core claim
The paper claims to present the first structured Russenorsk dictionary, with 691 entries grouped by synonyms and word origins, and to show that an LLM-driven agentic pipeline using this dictionary produces hypotheses about Russenorsk that largely align with a century of prior scholarship. The pipeline proposes source-language origins for borrowed terms, phonetic and morphological adaptation rules, and grammatical patterns; the paper's table shows which documented features each model recovers and which it misses, such as self-glossing and single-origin synonyms, and it notes one word-order contradiction where the models prefer SVO while the literature suggests SOV. The reconstruction agent, combining the dictionary, example sentences, and generated hypotheses, then renders contemporary texts into hypothetical Russenorsk; the paper reports incremental chrF gains from each resource layer but shows that removing the example sentences cuts ru-to-rn performance from 67.7 to 29.2, indicating that a substantial share of the apparent quality comes from copying prompt material.
Load-bearing premise
The load-bearing premise is that the 35-sentence evaluation benchmark is not effectively part of the example material already in the prompt; if the benchmark sentences and prompt examples overlap too much, the measured translation quality reflects copying rather than what the dictionary and rules actually contribute.
Editorial extensions
If this is right
- If the dictionary and agentic hypotheses are sound, the released Russenorsk vocabulary gives scholars a structured digital resource for studying a pidgin documented only in a small corpus.
- If the reconstruction pipeline works as claimed, the same method can be applied to other extinct or endangered contact languages with sparse records.
- The ablations imply that any future use of LLMs for low-resource language reconstruction must control for prompt leakage, since example removal alone drops ru-to-rn chrF from 67.7 to 29.2.
- Running two different LLMs in parallel recovers a broader set of documented Russenorsk features than either model alone, suggesting that model ensembles are a cheap way to widen hypothesis coverage.
- The word-order mismatch and the missed patterns like self-glossing give concrete targets for follow-up corpus-based verification.
Reading between the lines
- An inference the paper leaves implicit: the high ru-to-rn chrF with the full prompt likely measures the model's ability to rewrite prompt examples into Russenorsk rather than its command of the language; a cleaner benchmark built from source material withheld from both dictionary and examples would be needed to separate the two.
- A testable extension would be to run the same dictionary-and-agent pipeline on a better-documented pidgin with a known reference translation, which would let researchers calibrate how much of the reconstructed output is linguistic signal versus LLM fabrication.
- The models' failure to recover self-glossing and single-origin synonyms suggests that pure lexical statistics underrepresent pragmatic and sociolinguistic strategies; prompting with explicit information about communicative context might recover such patterns.
- If the reconstruction approach is applied to other dead languages, the not-strictly-falsifiable caveat could be partially addressed by asking multiple LLMs plus human experts to committee-edit the output and comparing against any later corpus finds.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper compiles a Russenorsk lexicon (691 entries) into a structured dictionary, uses LLM-based agents with this dictionary to generate hypotheses about Russenorsk grammar, phonetics, and lexicon, and compares the outputs to established findings in the literature (Table 1). It also proposes a 'reconstruction' translation pipeline that produces hypothetical Russenorsk renderings of modern texts, evaluated with chrF on a 35-sentence benchmark under ablations that remove examples, lexicon, and rules.
Significance. If the discovery pipeline is validated, this is a useful proof-of-concept for LLM-assisted historical linguistics on low-resource and dead contact languages. The paper's strengths are concrete: the authors release the dictionary, prompts, outputs, and benchmark; they explicitly acknowledge prompt leakage and run ablation studies; and they are honest about the speculative nature of the reconstruction. However, the central claim that 'nearly all hypotheses align' with the literature rests on an unspecified mapping of model outputs to feature labels, and the translation evaluation is too under-powered to support the incremental-gains interpretation. The resource itself may be valuable regardless of the discovery claims.
major comments (2)
- [Section 5, Table 1] The claim that 'nearly all hypotheses align' with observed features is not supported as stated because the mapping from raw model outputs to the checkmarks in Table 1 is never specified. There is no coding rubric, no definition of what counts as a match, no independent raters, and no inter-annotator agreement. The Fictitious Pidgin row is particularly problematic: the model is explicitly given no dictionary and no Russenorsk information (Section 3.2), yet it receives checkmarks for Word Order, Absence of Articles, Splitting up Russian Consonant Clusters, Synonyms of Dual Origin, and Founder Effect. At least the last two are Russenorsk-specific phenomena described in Section 4.3, and the consonant-cluster splitting is a Russenorsk-specific phonetic feature from Section 4.2; generic pidgin reasoning should not produce them. The only explanations are that the model is drawing on memorized Russenorsk content despite the instruction, or that the mapping is permissive enough to count any loosely related statement as a match. Either explanation undermines the ablations that were designed to isolate the dictionary's contribution. Please provide the coding protocol, report agreement statistics, and separately re-examine these specific cells before the discovery claim can be accepted.
- [Section 6.2.1, Table 2] The reconstruction evaluation is too under-powered to support the 'incremental, measurable gains' interpretation. With only 35 sentences and no confidence intervals or significance tests, the 11.4-point difference between the '– examples' condition (ru→rn chrF 29.2) and the lower-bound baseline (17.8) may well be sampling noise. Moreover, the benchmark sentences and the prompt examples are drawn from the same small historical corpus, so the ablation does not fully control for leakage; the paper acknowledges this but does not resolve it. The full-prompt score of 67.7 dropping to 29.2 when examples are removed demonstrates that the model largely copies prompt material, so the claim that the dictionary and rules add 'non-trivial signal' requires either a larger held-out benchmark, a human evaluation with experts in Scandinavian contact languages, or at least bootstrap confidence intervals on the chrF scores.
minor comments (5)
- [Section 2, heading and text] The heading 'Russenorsk V ocabulary' and the phrase 'new, extensive Russenorsk V ocabulary' contain a spacing typo; 'Vocabulary' should be one word.
- [Section 5, paragraph 2] The sentence 'by the agent that had to dictionary and no information on which specific pidgin we want to study' is ungrammatical; it should presumably read 'had no dictionary and no information'.
- [Table 2 caption] The caption contains the typo 'leakage of exemple sentences'; 'exemple' should be 'example'.
- [Section 6.2.1, 'Reference points'] The comparison of the reported chrF values with the OPUS-MT model-card scores is not apples-to-apples: the model cards report chrF2 (0.418 and 0.400), while Table 2 uses sacrebleu chrF defaults; chrF and chrF2 are different metrics. Please recompute with identical settings or remove the 'optimistic ceiling' comparison, since it currently understates the uncertainty.
- [Abstract] The phrase 'hypotheses previously proposed ones in the academic literature' is grammatically awkward; consider 'hypotheses previously proposed in the academic literature'.
Circularity Check
No circular derivation: disclosed leakage and an underspecified alignment rubric are validity concerns, not reductions to inputs.
full rationale
The claimed derivation chain—historical sources → dictionary → LLM hypotheses → judged alignment with literature, and dictionary/examples/rules → reconstructed translations → chrF check—does not contain a step in which an output is identical to an input by definition or in which a fitted parameter is relabeled as a prediction. The dictionary is compiled from Wiktionary and Broch and Jahr (1984) and manually cleaned; the hypotheses are then elicited from the LLM, not computed from the dictionary; the 'alignment' in Table 1 is an author judgment without a formal rubric. No equation in the paper sets these equal. The translation evaluation is the closest concern: Section 6.2.1 explicitly states the prompt artifacts are 'derived largely from the same historical sources as our 35-sentence benchmark,' and Table 2 shows ru→rn chrF dropping from 67.7 to 29.2 when examples are removed. This is disclosed prompt leakage and is discussed as such ('The dramatic gap ... indicates that the model often copies or lightly paraphrases example sentences'). Because the authors do not present the 67.7 figure as an independent predictive success—they call it a 'sanity check' and recommend human evaluation—the leakage is a methodological limitation, not a circular derivation. The Fictitious Pidgin condition’s unexpected checkmarks (e.g., Founder Effect, Splitting up Russian Consonant Clusters) do undermine the ablation’s isolating power, but that is a validity problem with the control, not an input-output identity. The self-citations in the introduction (Yamshchikov et al. 2022; Gorovaia et al. 2024; Schmidt et al. 2024; Sorokovikova et al. 2024) are contextual and are not load-bearing for the Russenorsk claims. The authors’ own Limitations section calls the output 'highly speculative' and the Ethics Statement acknowledges annotator bias, further reducing any impression that a hidden reduction is being passed off as discovery. Verdict: no significant circularity.
Assumptions & free parameters
free parameters (1)
- Semantic proximity threshold for dictionary clustering =
not reported
assumptions (4)
- domain assumption The surviving Russenorsk corpus (Broch & Jahr 1984 plus Wiktionary) is representative of the language as actually spoken.
- domain assumption The authors' qualitative mapping of LLM hypotheses to literature properties in Table 1 is reliable.
- domain assumption chrF on lowercased, diacritic-stripped text is a meaningful quality proxy for a heavily inflected pidgin.
- domain assumption The closed-source LLMs (Claude 3.5 Sonnet, OpenAI o1) respond consistently enough across runs for the findings to be interpretable.
Cite this review
Pith. "Pith review of Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study." pith.science (2026). https://pith.science/paper/MZRY6NE4
@misc{pith2026250611065,
author = {Pith},
title = {Pith review of: Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/MZRY6NE4}},
note = {Machine review of arXiv:2506.11065}
}
read the original abstract
Russenorsk, a pidgin language historically used in trade interactions between Russian and Norwegian speakers, represents a unique linguistic phenomenon. In this paper, we attempt to analyze its lexicon using modern large language models (LLMs), based on surviving literary sources. We construct a structured dictionary of the language, grouped by synonyms and word origins. Subsequently, we use this dictionary to formulate hypotheses about the core principles of word formation and grammatical structure in Russenorsk and show which hypotheses generated by large language models correspond to the hypotheses previously proposed ones in the academic literature. We also develop a "reconstruction" translation agent that generates hypothetical Russenorsk renderings of contemporary Russian and Norwegian texts.
Figures
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Vladimir Belikov. 2003. https://www.philol.msu.ru/ otipl/new/main/articles/belikov/vib-russenorsk.doc Some fragments of russenorsk grammar
work page 2003
-
[4]
Katar \' na Boldi z arov \'a . 2021. The contemporary norwegians' understanding of russenorsk
work page 2021
-
[5]
Ingvild Broch and Ernst Håkon Jahr. 1984. Russenorsk -- et pidginspråk i Norge [Russenorsk -- a Pidgin Language in Norway], 2nd edition. Novus, Oslo
work page 1984
-
[6]
Ingvlid Broch and Ernst Hakon Jahr. 1981. Russenorsk: Et pidginspråk i norge. Tromsø-Studier i Språkvitensk, 3. In Norwegian
work page 1981
-
[7]
Olaf Broch. 1927. Russenorsk. Archiv für Slavische Philologie, 41:81--130
work page 1927
-
[8]
Ioana Ciucă, Yuan-Sen Ting, Sandor Kruk, and Kartheik Iyer. 2023. https://arxiv.org/abs/2306.11648 Harnessing the power of adversarial prompting and large language models for robust hypothesis generation in astronomy . Preprint, arXiv:2306.11648
arXiv 2023
Show all 34 references
-
[9]
A. N. Davydov, V. N. Ponomarenko, and A. A. Kuratova. 1986. Russenorsk - arkticheskiy pidzhin evropy. In Russian
1986
-
[10]
José Andrés Alonso de la Fuente. 2020. On russenorsk -om in particular and on etymology and creolistics in general
2020
-
[11]
Qingxiu Dong, Li Dong, Ke Xu, Guangyan Zhou, Yaru Hao, Zhifang Sui, and Furu Wei. 2023. https://arxiv.org/abs/2309.05689 Large language model for science: A study on p vs. np . Preprint, arXiv:2309.05689
2023 arXiv
-
[12]
Noelia Ferruz, Steffen Schmidt, and Birte H \"o cker. 2022. Protgpt2 is a deep unsupervised language model for protein design. Nature communications, 13(1):4348
2022
-
[13]
Svetlana Gorovaia, Gleb Schmidt, and Ivan P Yamshchikov. 2024. Sui generis: Large language models for authorship attribution and verification in latin. In Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities, pages 398--412
2024
-
[14]
Bryan Hayden. 2019. Was russenorsk a continuum?
2019
-
[15]
opus-mt-ru-no model card
Helsinki\-NLP.a. opus-mt-ru-no model card . https://huggingface.co/Helsinki-NLP/opus-mt-ru-no. ChrF2 score = 0.418; accessed 19 May 2025
2025
-
[16]
opus-mt-no-ru model card
Helsinki\-NLP.b. opus-mt-no-ru model card . https://huggingface.co/Helsinki-NLP/opus-mt-no-ru. ChrF2 score = 0.400; accessed 19 May 2025
2025
-
[17]
- am beispiel von „russenorsk
Markus Hirnsperger. 2012. http://www.sub-arctic.ac.at/wp-content/uploads/2013/10/russ.pdf „pidgin - russisch" - am beispiel von „russenorsk" . Arbeitsgemeinschaft Arktis und Subarktis (A.A.S.)
2012
-
[18]
Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Oyvind Tafjord, and Peter Clark. 2024. https://arxiv.org/abs/2406.06769 DISCOVERYWORLD : A virtual environment for developing and evaluating automated scientific di...
2024 arXiv
-
[19]
Alamry, Vytautas Getautis, and Mohammad Khaja Nazeeruddin
Aistė Jegorovė, Jianxing Xia, Matas Steponaitis, Maryte Daskeviciene, Vygintas Jankauskas, Alytis Gruodis, Egidijus Kamarauskas, Tadas Malinauskas, Kasparas Rakstys, Khalid A. Alamry, Vytautas Getautis, and Mohammad Khaja Nazeeruddin. 2023. https://doi.org/10.1021/acs.chemmate...
2023 doi
-
[20]
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Z \'i dek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A. A. Kohl, Andrew J. Ballard, Andrew Cowie, Bernardino Romera-Paredes, S...
2021
-
[21]
a ge zur \
Frederik Kortlandt-Leiden. 2000. On russenorsk. Amsterdamer Beitr \"a ge zur \"a lteren Germanistik , 54:123
2000
-
[22]
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. https://openai.com/index/improving-mathematical-reasoning-with-process-supervision/ Improving mathematical reasoning with proces...
2023
-
[23]
Rodríguez Méndez, Thang Bui, Alyssa Goodman, Alberto Accomazzi, Jill Naiman, and Jesse Cranney
Tuan Dung Nguyen, Yuan-Sen Ting, Ioana Ciucă, Charlie O'Neill, Ze-Chang Sun, Maja Jabłońska, Sandor Kruk, Ernest Perkowski, Jack Miller, Jason Li, Josh Peek, Kartheik Iyer, Tomasz Różański, Pranav Khetarpal, Sharaf Zaman, David Brodrick, Sergio J. Rodríguez Méndez, Thang Bui, ...
2023 arXiv
-
[24]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei - Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the ACL, pages 311--318
2002
-
[25]
Perekhval'skaya
Elena V. Perekhval'skaya. 1987. Russenorsk kak primer pervonachal'nogo e'tapa formirovaniya pidzhina. In Russian
1987
-
[26]
Maja Popovi \' c . 2015. chrf: character n-gram f-score for automatic machine translation evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392--395
2015
-
[27]
Matt Post. 2018. A Call for Clarity in Reporting BLEU Scores . In Proceedings of the Third Conference on Machine Translation (WMT), pages 186--191
2018
-
[28]
Frederick Riemenschneider and Anette Frank. 2023. Exploring large language models for classical philology. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15181--15199
2023
-
[29]
Mustafa Safdari, Greg Serapio-Garc \' a, Cl \'e ment Crepy, Stephen Fitz, Peter Romero, Luning Sun, Marwa Abdulhai, Aleksandra Faust, and Maja Matari \'c . 2023. Personality traits in large language models. arXiv preprint arXiv:2307.00184
2023 arXiv
-
[30]
Yamshchikov
Gleb Schmidt, Veronica Vybornaya, and Ivan P. Yamshchikov. 2024. https://ceur-ws.org/Vol-3834/paper139.pdf Fine-tuning pre-trained language models for authorship attribution of the pseudo-dionysian ars rhetorica . In Proceedings of the Computational Humanities Research Confere...
2024
-
[31]
Aleksandra Sorokovikova, Sharwin Rezagholi, Natalia Fedorova, and Ivan Yamshchikov. 2024. Llms simulate big5 personality traits: Further evidence. In Proceedings of the 1st Workshop on Personalization of Generative AI Systems (PERSONALIZE 2024), pages 83--87
2024
-
[32]
Dieter Stern. 2020. Russian pidgin languages. Encyclopedia of Slavic languages and linguistics online
2020
-
[33]
Ivan Yamshchikov, Alexey Tikhonov, Yorgos Pantis, Charlotte Schubert, and J \"u rgen Jost. 2022. Bert in plutarch’s shadows. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 6071--6080
2022
-
[34]
Zonglin Yang, Xinya Du, Junxian Li, Jie Zheng, Soujanya Poria, and Erik Cambria. 2023. https://arxiv.org/abs/2309.02726 Large language models for automated open-domain scientific hypotheses discovery . Preprint, arXiv:2309.02726
2023 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.