Pith. sign in

REVIEW 2 major objections 5 minor 34 references

Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study

T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Russenorsk, a dead Russian-Norwegian trade pidgin, can be analyzed and reconstructed by an LLM pipeline built on a new structured dictionary.

desk verdict Useful artifacts and an honest ablation, but the central 'hypotheses align' claim lacks a defensible mapping procedure. read the letter →

arxiv 2506.11065 v1 pith:MZRY6NE4 submitted 2025-05-31 cs.CL

classification cs.CL
keywords Russenorskpidginlargelanguagemodelsagenticdiscoverylow-resourcelanguagestranslationreconstructionhistoricallinguisticscontact
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Drawing on surviving historical sources, the paper builds a structured dictionary of Russenorsk, a 19th-century Russian-Norwegian trade pidgin, and uses that dictionary to prompt large language models to generate hypotheses about the language's word origins, sound changes, and grammar. The central claim is that this agentic pipeline can serve as a Linguistic Discovery Assistant for dead contact languages: nearly all of the model-generated hypotheses match features documented in the academic literature, with a small set of exceptions. The paper also extends the pipeline into a reconstruction translator that produces hypothetical Russenorsk renderings of modern Russian and Norwegian texts, and it evaluates these renderings with character-level chrF scores under ablations that remove the dictionary, examples, and rules. The authors are careful to frame the reconstructed translations as principled speculation rather than verified historical language, and they make the dictionary, prompts, and benchmark publicly available.

What carries the argument

The central mechanism is the agentic discovery pipeline: a large language model is first prompted with the cleaned 691-entry Russenorsk lexicon to propose lexical-origin hypotheses; it is then given the same lexicon to infer phonetic and morphological transformations; and finally it receives dictionary entries and example sentences to infer grammatical rules. The reconstruction stage reuses all three outputs as context and instructs the model to prefer known lexemes over invented ones. The evaluation machinery is a three-step ablation that removes examples, then lexicon, then rules, scored with character-level chrF to quantify how much linguistic signal each layer actually contributes.

What would settle it

Construct a held-out set of historical Russenorsk sentences that appear in no published dictionary example and no prompt, run the full pipeline and the examples-removed ablation on that set, and have a Scandinavian contact-language expert score the reconstructions blind; if the full pipeline does not clearly beat the ablation, the claim that the dictionary and rules add real linguistic signal is refuted.

Watch

Extended reading notes

Core claim

The paper claims to present the first structured Russenorsk dictionary, with 691 entries grouped by synonyms and word origins, and to show that an LLM-driven agentic pipeline using this dictionary produces hypotheses about Russenorsk that largely align with a century of prior scholarship. The pipeline proposes source-language origins for borrowed terms, phonetic and morphological adaptation rules, and grammatical patterns; the paper's table shows which documented features each model recovers and which it misses, such as self-glossing and single-origin synonyms, and it notes one word-order contradiction where the models prefer SVO while the literature suggests SOV. The reconstruction agent, combining the dictionary, example sentences, and generated hypotheses, then renders contemporary texts into hypothetical Russenorsk; the paper reports incremental chrF gains from each resource layer but shows that removing the example sentences cuts ru-to-rn performance from 67.7 to 29.2, indicating that a substantial share of the apparent quality comes from copying prompt material.

Load-bearing premise

The load-bearing premise is that the 35-sentence evaluation benchmark is not effectively part of the example material already in the prompt; if the benchmark sentences and prompt examples overlap too much, the measured translation quality reflects copying rather than what the dictionary and rules actually contribute.

Editorial extensions

If this is right

  • If the dictionary and agentic hypotheses are sound, the released Russenorsk vocabulary gives scholars a structured digital resource for studying a pidgin documented only in a small corpus.
  • If the reconstruction pipeline works as claimed, the same method can be applied to other extinct or endangered contact languages with sparse records.
  • The ablations imply that any future use of LLMs for low-resource language reconstruction must control for prompt leakage, since example removal alone drops ru-to-rn chrF from 67.7 to 29.2.
  • Running two different LLMs in parallel recovers a broader set of documented Russenorsk features than either model alone, suggesting that model ensembles are a cheap way to widen hypothesis coverage.
  • The word-order mismatch and the missed patterns like self-glossing give concrete targets for follow-up corpus-based verification.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An inference the paper leaves implicit: the high ru-to-rn chrF with the full prompt likely measures the model's ability to rewrite prompt examples into Russenorsk rather than its command of the language; a cleaner benchmark built from source material withheld from both dictionary and examples would be needed to separate the two.
  • A testable extension would be to run the same dictionary-and-agent pipeline on a better-documented pidgin with a known reference translation, which would let researchers calibrate how much of the reconstructed output is linguistic signal versus LLM fabrication.
  • The models' failure to recover self-glossing and single-origin synonyms suggests that pure lexical statistics underrepresent pragmatic and sociolinguistic strategies; prompting with explicit information about communicative context might recover such patterns.
  • If the reconstruction approach is applied to other dead languages, the not-strictly-falsifiable caveat could be partially addressed by asking multiple LLMs plus human experts to committee-edit the output and comparing against any later corpus finds.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper compiles a Russenorsk lexicon (691 entries) into a structured dictionary, uses LLM-based agents with this dictionary to generate hypotheses about Russenorsk grammar, phonetics, and lexicon, and compares the outputs to established findings in the literature (Table 1). It also proposes a 'reconstruction' translation pipeline that produces hypothetical Russenorsk renderings of modern texts, evaluated with chrF on a 35-sentence benchmark under ablations that remove examples, lexicon, and rules.

Significance. If the discovery pipeline is validated, this is a useful proof-of-concept for LLM-assisted historical linguistics on low-resource and dead contact languages. The paper's strengths are concrete: the authors release the dictionary, prompts, outputs, and benchmark; they explicitly acknowledge prompt leakage and run ablation studies; and they are honest about the speculative nature of the reconstruction. However, the central claim that 'nearly all hypotheses align' with the literature rests on an unspecified mapping of model outputs to feature labels, and the translation evaluation is too under-powered to support the incremental-gains interpretation. The resource itself may be valuable regardless of the discovery claims.

major comments (2)
  1. [Section 5, Table 1] The claim that 'nearly all hypotheses align' with observed features is not supported as stated because the mapping from raw model outputs to the checkmarks in Table 1 is never specified. There is no coding rubric, no definition of what counts as a match, no independent raters, and no inter-annotator agreement. The Fictitious Pidgin row is particularly problematic: the model is explicitly given no dictionary and no Russenorsk information (Section 3.2), yet it receives checkmarks for Word Order, Absence of Articles, Splitting up Russian Consonant Clusters, Synonyms of Dual Origin, and Founder Effect. At least the last two are Russenorsk-specific phenomena described in Section 4.3, and the consonant-cluster splitting is a Russenorsk-specific phonetic feature from Section 4.2; generic pidgin reasoning should not produce them. The only explanations are that the model is drawing on memorized Russenorsk content despite the instruction, or that the mapping is permissive enough to count any loosely related statement as a match. Either explanation undermines the ablations that were designed to isolate the dictionary's contribution. Please provide the coding protocol, report agreement statistics, and separately re-examine these specific cells before the discovery claim can be accepted.
  2. [Section 6.2.1, Table 2] The reconstruction evaluation is too under-powered to support the 'incremental, measurable gains' interpretation. With only 35 sentences and no confidence intervals or significance tests, the 11.4-point difference between the '– examples' condition (ru→rn chrF 29.2) and the lower-bound baseline (17.8) may well be sampling noise. Moreover, the benchmark sentences and the prompt examples are drawn from the same small historical corpus, so the ablation does not fully control for leakage; the paper acknowledges this but does not resolve it. The full-prompt score of 67.7 dropping to 29.2 when examples are removed demonstrates that the model largely copies prompt material, so the claim that the dictionary and rules add 'non-trivial signal' requires either a larger held-out benchmark, a human evaluation with experts in Scandinavian contact languages, or at least bootstrap confidence intervals on the chrF scores.
minor comments (5)
  1. [Section 2, heading and text] The heading 'Russenorsk V ocabulary' and the phrase 'new, extensive Russenorsk V ocabulary' contain a spacing typo; 'Vocabulary' should be one word.
  2. [Section 5, paragraph 2] The sentence 'by the agent that had to dictionary and no information on which specific pidgin we want to study' is ungrammatical; it should presumably read 'had no dictionary and no information'.
  3. [Table 2 caption] The caption contains the typo 'leakage of exemple sentences'; 'exemple' should be 'example'.
  4. [Section 6.2.1, 'Reference points'] The comparison of the reported chrF values with the OPUS-MT model-card scores is not apples-to-apples: the model cards report chrF2 (0.418 and 0.400), while Table 2 uses sacrebleu chrF defaults; chrF and chrF2 are different metrics. Please recompute with identical settings or remove the 'optimistic ceiling' comparison, since it currently understates the uncertainty.
  5. [Abstract] The phrase 'hypotheses previously proposed ones in the academic literature' is grammatically awkward; consider 'hypotheses previously proposed in the academic literature'.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: disclosed leakage and an underspecified alignment rubric are validity concerns, not reductions to inputs.

full rationale

The claimed derivation chain—historical sources → dictionary → LLM hypotheses → judged alignment with literature, and dictionary/examples/rules → reconstructed translations → chrF check—does not contain a step in which an output is identical to an input by definition or in which a fitted parameter is relabeled as a prediction. The dictionary is compiled from Wiktionary and Broch and Jahr (1984) and manually cleaned; the hypotheses are then elicited from the LLM, not computed from the dictionary; the 'alignment' in Table 1 is an author judgment without a formal rubric. No equation in the paper sets these equal. The translation evaluation is the closest concern: Section 6.2.1 explicitly states the prompt artifacts are 'derived largely from the same historical sources as our 35-sentence benchmark,' and Table 2 shows ru→rn chrF dropping from 67.7 to 29.2 when examples are removed. This is disclosed prompt leakage and is discussed as such ('The dramatic gap ... indicates that the model often copies or lightly paraphrases example sentences'). Because the authors do not present the 67.7 figure as an independent predictive success—they call it a 'sanity check' and recommend human evaluation—the leakage is a methodological limitation, not a circular derivation. The Fictitious Pidgin condition’s unexpected checkmarks (e.g., Founder Effect, Splitting up Russian Consonant Clusters) do undermine the ablation’s isolating power, but that is a validity problem with the control, not an input-output identity. The self-citations in the introduction (Yamshchikov et al. 2022; Gorovaia et al. 2024; Schmidt et al. 2024; Sorokovikova et al. 2024) are contextual and are not load-bearing for the Russenorsk claims. The authors’ own Limitations section calls the output 'highly speculative' and the Ethics Statement acknowledges annotator bias, further reducing any impression that a hidden reduction is being passed off as discovery. Verdict: no significant circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claims rest on the representativeness of a very small historical corpus, the authors' subjective matching of model output to prior literature, and the validity of chrF as a quality metric for a heavily inflected pidgin. No new physical or linguistic entities are postulated.

free parameters (1)
  • Semantic proximity threshold for dictionary clustering = not reported
    The Sonnet-driven clustering groups entries by semantic proximity; the threshold is not quantified and required several days of manual correction, so the grouping depends on an undisclosed hand-tuned criterion.
assumptions (4)
  • domain assumption The surviving Russenorsk corpus (Broch & Jahr 1984 plus Wiktionary) is representative of the language as actually spoken.
    The dictionary, example sentences, and benchmark triplets all derive from this small corpus; if it is skewed (e.g., by Norwegian authorship), the hypotheses and translation scores inherit the bias. The paper itself notes this for greetings.
  • domain assumption The authors' qualitative mapping of LLM hypotheses to literature properties in Table 1 is reliable.
    No inter-annotator agreement or predefined matching protocol is given; the mapping is performed by the authors.
  • domain assumption chrF on lowercased, diacritic-stripped text is a meaningful quality proxy for a heavily inflected pidgin.
    The authors rely on chrF as the sole automatic metric and call the evaluation a 'sanity check'.
  • domain assumption The closed-source LLMs (Claude 3.5 Sonnet, OpenAI o1) respond consistently enough across runs for the findings to be interpretable.
    No temperature or determinism settings are reported; replication will vary.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study." pith.science (2026). https://pith.science/paper/MZRY6NE4

@misc{pith2026250611065,
  author       = {Pith},
  title        = {Pith review of: Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MZRY6NE4}},
  note         = {Machine review of arXiv:2506.11065}
}
read the original abstract

Russenorsk, a pidgin language historically used in trade interactions between Russian and Norwegian speakers, represents a unique linguistic phenomenon. In this paper, we attempt to analyze its lexicon using modern large language models (LLMs), based on surviving literary sources. We construct a structured dictionary of the language, grouped by synonyms and word origins. Subsequently, we use this dictionary to formulate hypotheses about the core principles of word formation and grammatical structure in Russenorsk and show which hypotheses generated by large language models correspond to the hypotheses previously proposed ones in the academic literature. We also develop a "reconstruction" translation agent that generates hypothetical Russenorsk renderings of contemporary Russian and Norwegian texts.

Figures

Figures reproduced from arXiv: 2506.11065 by the authors.

Figure 1
Figure 1. An example of a resulting vocabulary entry. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

34 extracted references · 24 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Vladimir Belikov. 2003. https://www.philol.msu.ru/ otipl/new/main/articles/belikov/vib-russenorsk.doc Some fragments of russenorsk grammar

  4. [4]

    Katar \' na Boldi z arov \'a . 2021. The contemporary norwegians' understanding of russenorsk

  5. [5]

    Ingvild Broch and Ernst Håkon Jahr. 1984. Russenorsk -- et pidginspråk i Norge [Russenorsk -- a Pidgin Language in Norway], 2nd edition. Novus, Oslo

  6. [6]

    Ingvlid Broch and Ernst Hakon Jahr. 1981. Russenorsk: Et pidginspråk i norge. Tromsø-Studier i Språkvitensk, 3. In Norwegian

  7. [7]

    Olaf Broch. 1927. Russenorsk. Archiv für Slavische Philologie, 41:81--130

  8. [8]

    Ioana Ciucă, Yuan-Sen Ting, Sandor Kruk, and Kartheik Iyer. 2023. https://arxiv.org/abs/2306.11648 Harnessing the power of adversarial prompting and large language models for robust hypothesis generation in astronomy . Preprint, arXiv:2306.11648

Show all 34 references
  1. [9]

    A. N. Davydov, V. N. Ponomarenko, and A. A. Kuratova. 1986. Russenorsk - arkticheskiy pidzhin evropy. In Russian

  2. [10]

    José Andrés Alonso de la Fuente. 2020. On russenorsk -om in particular and on etymology and creolistics in general

  3. [11]

    Qingxiu Dong, Li Dong, Ke Xu, Guangyan Zhou, Yaru Hao, Zhifang Sui, and Furu Wei. 2023. https://arxiv.org/abs/2309.05689 Large language model for science: A study on p vs. np . Preprint, arXiv:2309.05689

  4. [12]

    Noelia Ferruz, Steffen Schmidt, and Birte H \"o cker. 2022. Protgpt2 is a deep unsupervised language model for protein design. Nature communications, 13(1):4348

  5. [13]

    Svetlana Gorovaia, Gleb Schmidt, and Ivan P Yamshchikov. 2024. Sui generis: Large language models for authorship attribution and verification in latin. In Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities, pages 398--412

  6. [14]

    Bryan Hayden. 2019. Was russenorsk a continuum?

  7. [15]

    opus-mt-ru-no model card

    Helsinki\-NLP.a. opus-mt-ru-no model card . https://huggingface.co/Helsinki-NLP/opus-mt-ru-no. ChrF2 score = 0.418; accessed 19 May 2025

  8. [16]

    opus-mt-no-ru model card

    Helsinki\-NLP.b. opus-mt-no-ru model card . https://huggingface.co/Helsinki-NLP/opus-mt-no-ru. ChrF2 score = 0.400; accessed 19 May 2025

  9. [17]

    - am beispiel von „russenorsk

    Markus Hirnsperger. 2012. http://www.sub-arctic.ac.at/wp-content/uploads/2013/10/russ.pdf „pidgin - russisch" - am beispiel von „russenorsk" . Arbeitsgemeinschaft Arktis und Subarktis (A.A.S.)

  10. [18]

    Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Oyvind Tafjord, and Peter Clark. 2024. https://arxiv.org/abs/2406.06769 DISCOVERYWORLD : A virtual environment for developing and evaluating automated scientific di...

  11. [19]

    Alamry, Vytautas Getautis, and Mohammad Khaja Nazeeruddin

    Aistė Jegorovė, Jianxing Xia, Matas Steponaitis, Maryte Daskeviciene, Vygintas Jankauskas, Alytis Gruodis, Egidijus Kamarauskas, Tadas Malinauskas, Kasparas Rakstys, Khalid A. Alamry, Vytautas Getautis, and Mohammad Khaja Nazeeruddin. 2023. https://doi.org/10.1021/acs.chemmate...

  12. [20]

    John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Z \'i dek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A. A. Kohl, Andrew J. Ballard, Andrew Cowie, Bernardino Romera-Paredes, S...

  13. [21]

    a ge zur \

    Frederik Kortlandt-Leiden. 2000. On russenorsk. Amsterdamer Beitr \"a ge zur \"a lteren Germanistik , 54:123

  14. [22]

    Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. https://openai.com/index/improving-mathematical-reasoning-with-process-supervision/ Improving mathematical reasoning with proces...

  15. [23]

    Rodríguez Méndez, Thang Bui, Alyssa Goodman, Alberto Accomazzi, Jill Naiman, and Jesse Cranney

    Tuan Dung Nguyen, Yuan-Sen Ting, Ioana Ciucă, Charlie O'Neill, Ze-Chang Sun, Maja Jabłońska, Sandor Kruk, Ernest Perkowski, Jack Miller, Jason Li, Josh Peek, Kartheik Iyer, Tomasz Różański, Pranav Khetarpal, Sharaf Zaman, David Brodrick, Sergio J. Rodríguez Méndez, Thang Bui, ...

  16. [24]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei - Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the ACL, pages 311--318

  17. [25]

    Perekhval'skaya

    Elena V. Perekhval'skaya. 1987. Russenorsk kak primer pervonachal'nogo e'tapa formirovaniya pidzhina. In Russian

  18. [26]

    Maja Popovi \' c . 2015. chrf: character n-gram f-score for automatic machine translation evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, pages 392--395

  19. [27]

    Matt Post. 2018. A Call for Clarity in Reporting BLEU Scores . In Proceedings of the Third Conference on Machine Translation (WMT), pages 186--191

  20. [28]

    Frederick Riemenschneider and Anette Frank. 2023. Exploring large language models for classical philology. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15181--15199

  21. [29]

    Mustafa Safdari, Greg Serapio-Garc \' a, Cl \'e ment Crepy, Stephen Fitz, Peter Romero, Luning Sun, Marwa Abdulhai, Aleksandra Faust, and Maja Matari \'c . 2023. Personality traits in large language models. arXiv preprint arXiv:2307.00184

  22. [30]

    Yamshchikov

    Gleb Schmidt, Veronica Vybornaya, and Ivan P. Yamshchikov. 2024. https://ceur-ws.org/Vol-3834/paper139.pdf Fine-tuning pre-trained language models for authorship attribution of the pseudo-dionysian ars rhetorica . In Proceedings of the Computational Humanities Research Confere...

  23. [31]

    Aleksandra Sorokovikova, Sharwin Rezagholi, Natalia Fedorova, and Ivan Yamshchikov. 2024. Llms simulate big5 personality traits: Further evidence. In Proceedings of the 1st Workshop on Personalization of Generative AI Systems (PERSONALIZE 2024), pages 83--87

  24. [32]

    Dieter Stern. 2020. Russian pidgin languages. Encyclopedia of Slavic languages and linguistics online

  25. [33]

    Ivan Yamshchikov, Alexey Tikhonov, Yorgos Pantis, Charlotte Schubert, and J \"u rgen Jost. 2022. Bert in plutarch’s shadows. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 6071--6080

  26. [34]

    Zonglin Yang, Xinya Du, Junxian Li, Jie Zheng, Soujanya Poria, and Erik Cambria. 2023. https://arxiv.org/abs/2309.02726 Large language models for automated open-domain scientific hypotheses discovery . Preprint, arXiv:2309.02726

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.