Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Characterizing the Effects of Translation on Intertextuality using Multilingual Embedding Spaces

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Multilingual embedding ratios reveal that human Bible translations amplify intertextual links between testaments, while machine translations provide a neutral baseline.

desk verdict The intertextuality metric is clean and externally validated, but the paper's central human-vs-machine claim is contradicted by its own Table 4 and further undermined by a source-text confound admitted in the Limitations. read the letter →

arxiv 2501.10731 v1 pith:C6GGVSFW submitted 2025-01-18 cs.CL

classification cs.CL
keywords intertextualitymultilingualembeddingsmachinetranslationbiblicalcross-referencesrhetoricaldevicescosinesimilarityhumanvscorpus-levelmetric
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a simple geometric measure can characterise how translation changes intertextuality—the web of references between texts—and uses it to compare human and machine translations of the Bible. The measure is the ratio of average cosine similarity between verses known to be cross-linked to average similarity between random verse pairs, computed in a shared multilingual embedding space. Across five languages, human translations show consistently higher intertextuality ratios than machine translations of the same ancient manuscripts, which the authors read as evidence that human translators have a propensity to amplify literary characteristics such as continuity between testaments. The metric also returns a high ratio on a curated corpus of allusions in Latin literature, suggesting it detects genuine intertextual links. A sympathetic reader would take the paper's contribution to be a quantitative, corpus-level method for studying rhetorical effects in translation, plus evidence that human translation does not preserve but often heightens intertextual density.

What carries the argument

The mechanism is the cosine-similarity ratio between verse embeddings from a multilingual model. For a predetermined set of ground-truth cross-references, the method computes the mean cosine similarity of the linked verse pairs and divides it by the mean similarity of random pairs drawn from the same chapter, producing a ratio greater than one when intertextual verses are more similar than chance. Bootstrap resampling with 10,000 iterations yields 95% confidence intervals, allowing comparisons across translation conditions. The authors validate the metric on a benchmark of 945 expert-curated allusions in Latin literature, where it produces a ratio of 1.55, before applying it to biblical cross-references that are split into within-testament and across-testament sets.

What would settle it

Take a single source text—say, the Greek New Testament—and have both a human translator and a machine translation system render it into the same language. If the human translation's intertextuality ratio does not exceed the machine's for the same ground-truth cross-references, the amplification claim would be falsified. The paper's own limitation note makes this the natural control experiment.

Watch

Extended reading notes

Core claim

The paper's central claim is that multilingual embedding spaces can characterise intertextuality at the corpus level, and that human translations of the Hebrew and Christian testaments exhibit a higher degree of intertextuality than machine translations, which act as a neutral baseline. In the authors' words, 'human translations consistently show higher levels of intertextuality,' and this is quantified in the intertextuality ratios of Table 4 for English, Finnish, Turkish, Swedish, and Marathi. The authors additionally provide a qualitative example in which a human English translation nearly doubles the similarity between a verse in Hebrews and a verse in Isaiah by rendering both with the word 'sin,' while the machine translation restores distance. The paper frames this as support for existing scholarship proposing that human translators amplify certain literary characteristics of the original manuscripts.

Load-bearing premise

The paper's central comparison assumes that the human and machine translations are translations of the same source text; in fact, most human translations in the corpus were made from English versions, not from the ancient Hebrew and Greek manuscripts that were fed to the machine translator.

Editorial extensions

If this is right

  • If the amplification claim holds, readers of human-translated Bibles encounter a denser web of cross-testamental references than the ancient source texts themselves contain.
  • The ratio metric provides a quantitative tool for testing long-standing theories in translation studies about translator-driven literary amplification.
  • Machine translation output, by staying closer to the source's semantic surface, could serve as a controlled baseline for measuring stylistic intervention in human translation.
  • Intertextuality ratios are language-dependent: English shows the largest amplification and Marathi the smallest, so any account of translator behavior must explain cross-linguistic variation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • We think the paper's own limitation statement points to a decisive follow-up: translate the ancient Hebrew and Greek manuscripts directly into a language by both a professional human translator and a machine, then recompute the ratio; without that control, the human-versus-machine gap could reflect translation from an English intermediary rather than a human propensity.
  • The same ratio method could be used on other highly translated texts with known allusions, such as classical epic or legal corpora, to see whether amplification generalises beyond the Bible.
  • If amplification is a general human-translation effect, it may also affect other rhetorical devices (e.g., metaphor or repetition) measurable by embedding similarity, not just intertextuality.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper introduces a corpus-level intertextuality ratio computed as the average cosine similarity of ground-truth verse pairs divided by the average similarity of same-chapter random pairs in a multilingual embedding space. The authors validate the metric on a Latin benchmark (ratio 1.55, CI [1.53, 1.56]), then apply it to biblical cross-references across five languages, comparing human translations from the JHUBC with Aya23 machine translations of the ancient Hebrew and Greek manuscripts. They claim that human translations consistently show higher intertextuality than machine translations, which they interpret as evidence that human translators amplify literary characteristics. The paper also includes a qualitative case study of one overemphasized intertextual pair.

Significance. The metric validation on an external Latin benchmark and the release of code and data are genuine strengths; the paper gives a reproducible, parameter-light way to compare intertextuality at corpus level. If the human-machine comparison were valid, this would be a useful contribution to computational intertextuality and translation studies. However, the central comparative claim is undercut by two problems: the JHUBC human translations are mostly mediated by English rather than translated directly from the source manuscripts, and Table 4 contains a direct counterexample (Turkish). Thus the paper's main interpretive conclusion does not follow from the presented evidence.

major comments (3)
  1. [§5, Table 4] The sentence 'Human translations consistently show higher levels of intertextuality' is contradicted by the Turkish row of Table 4: the Aya23 NMT ratio exceeds the human ratio in the within-Jewish column (1.60 vs 1.50) and within-Christian column (1.71 vs 1.43) and is essentially tied across testaments (1.52 vs 1.51). In most other language rows the bootstrap 95% CIs overlap between human and NMT (e.g., Swedish within-Jewish 1.33 ± 0.15 vs 1.31 ± 0.22), so the data do not establish a consistent human advantage. This claim is the basis for the abstract's and Section 5's conclusion about human amplificatory propensity, so it is load-bearing.
  2. [Limitations; §3.2] The comparison of human and machine translation cannot isolate translator behavior because the two conditions are generated from different source texts. The machine translations are produced from the Hebrew Old Testament, Greek Old Testament, and Greek New Testament (Section 3.2), while the Limitations state that most JHUBC human translations were not translated directly from ancient manuscripts but instead work from English translations. Any observed human-vs-machine difference could therefore be a source-text effect rather than a translator effect. Since the paper's contribution 3 and the Section 5 interpretation rely on this contrast, this is a load-bearing confound.
  3. [§5, Table 5] The single hand-selected example of Hebrews 8:12 / Isaiah 43:25 is offered as evidence that human translation overemphasizes intertextuality while machine translation provides a neutral baseline, but one qualitative example cannot support a general claim about all human translations, especially when Table 4 already fails to show a consistent overall pattern. Moreover, 'neutral baseline' is not operationalized: the NMT ratios in Table 4 are substantially above 1 (e.g., Turkish NMT 1.60 and 1.71), so machine translations are not neutral by the paper's own metric.
minor comments (4)
  1. [§1] The text contains a typo: 'semiotic density of a any given text' should read 'semiotic density of any given text.'
  2. [§2] The intertextuality ratio depends on the vote threshold of 50 used to define ground-truth references; a short sensitivity analysis around this threshold would help establish that the main conclusions are not threshold artifacts.
  3. [§5, Table 4] The prose should distinguish point estimates from statistically distinguishable differences; several apparent human advantages have overlapping 95% CIs and therefore should not be described as consistent effects.
  4. [§5] The term 'neutral baseline' is used without a precise definition; please clarify whether it means a ratio near 1, a ratio close to the source manuscript, or something else, and report the corresponding values.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the intertextuality ratios are computed from external ground-truth references, external embeddings, and a fixed bootstrap procedure, and the comparative human-vs-machine claim does not reduce to the metric's definition.

full rationale

The paper's central measurement is a ratio of average cosine similarity between verses linked by an external cross-reference dataset to an average baseline similarity over randomized same-chapter pairs. This is a well-defined empirical statistic, not a fitted parameter or a renamed input. The benchmark validation on the Burns et al. (2021) Latin intertextuality corpus is independent of the Biblical data and is not used to tune the method. The comparative claim that human translations preserve or amplify intertextuality is an empirical reading of Table 4, and although the claim may be overstated given overlapping confidence intervals and the Turkish NMT rows, that is a correctness or interpretation concern, not circularity. The citation to McGovern et al. (2024) is self-citational but not load-bearing: the paper's own Table 4 and qualitative example carry the argument, and removing the citation does not change the derivation. The Limitations section explicitly cautions that human translations often derive from English rather than the ancient source manuscripts, but this affects the validity of the comparison, not whether the derivation reduces to its inputs. No equation equates the output to an input by construction, and no fitted value is relabeled as a prediction. The paper is therefore self-contained with respect to its stated metric and data, and no circular step is present.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

No new entities are introduced. The main free parameter is the vote threshold for ground-truth references. The key assumptions are about cross-lingual embedding comparability and the validity of the source texts for comparison.

free parameters (1)
  • vote_threshold = 50 votes
    Chosen as a threshold for counting a Bible cross-reference as valid; affects which verse pairs enter the intertextuality ratio.
assumptions (4)
  • domain assumption Multilingual embeddings place semantically related verses close across languages, so cosine similarity is comparable across translations.
    The entire method depends on cross-lingual semantic alignment.
  • domain assumption The Bible cross-reference dataset (Owens 2023) provides valid ground-truth intertextual links.
    Used as gold standard for which verse pairs are intertextual.
  • domain assumption Aya23 machine translations are faithful enough for the embedding metric to be meaningful.
    Translation quality varies (COMET 27-73), and low-quality outputs (Marathi) may distort similarities.
  • ad hoc to paper The JHUBC human translations are representative of human translation behavior despite being mediated by English.
    The paper relies on this despite acknowledging the mediation confound.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Characterizing the Effects of Translation on Intertextuality using Multilingual Embedding Spaces." pith.science (2026). https://pith.science/paper/C6GGVSFW

@misc{pith2026250110731,
  author       = {Pith},
  title        = {Pith review of: Characterizing the Effects of Translation on Intertextuality using Multilingual Embedding Spaces},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/C6GGVSFW}},
  note         = {Machine review of arXiv:2501.10731}
}
read the original abstract

Rhetorical devices are difficult to translate, but they are crucial to the translation of literary documents. We investigate the use of multilingual embedding spaces to characterize the preservation of intertextuality, one common rhetorical device, across human and machine translation. To do so, we use Biblical texts, which are both full of intertextual references and are highly translated works. We provide a metric to characterize intertextuality at the corpus level and provide a quantitative analysis of the preservation of this rhetorical device across extant human translations and machine-generated counterparts. We go on to provide qualitative analysis of cases wherein human translations over- or underemphasize the intertextuality present in the text, whereas machine translations provide a neutral baseline. This provides support for established scholarship proposing that human translators have a propensity to amplify certain literary characteristics of the original manuscripts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Modelling Intertextuality with N-gram Embeddings

    cs.CL 2025-09 conditional novelty 4.0 of 10

    Intertextuality is quantified as the thresholded average of pairwise cosine similarities between n-gram embeddings of two texts.

Reference graph

Works this paper leans on

26 extracted references · 20 canonical work pages · cited by 1 Pith paper

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Kelly Marchisio, Sebastian Ruder, Acyr Locatelli, Julia Kreutzer, Nick Frosst, Phil Blunsom, Marzieh Fadaee, Ahmet Üstün, and Sara Hooker. 2024. http://arxiv.org/abs/2405.15032 Aya 23: Open weight releases to further multilingual progress

  4. [4]

    Erich Auerbach. 1959. Scenes from the Drama of European Literature: six essays . Meridian Books

  5. [5]

    David Bamman and Gregory R. Crane. 2008. https://api.semanticscholar.org/CorpusID:11658482 The logic and discovery of textual allusion . In In Proceedings of the 2008 LREC Workshop on Language Technology for Cultural Heritage Data

  6. [6]

    Damien Broderick. 2017. Reading sf as a mega-text. Science fiction criticism: an anthology of essential writings, pages 139--48

  7. [7]

    Patrick J Burns, James A Brofos, Kyle Li, Pramit Chaudhuri, and Joseph P Dexter. 2021. Profiling of intertextuality in latin literature using word embeddings. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4900--4907

  8. [8]

    C Craig, K Goyal, G Crane, F Shamsian, and DA Smith. 2023. Testing the limits of neural sentence alignment models on classical greek and latin texts and translations. In Proceedings http://ceur-ws. org ISSN 1613 0073

Show all 26 references
  1. [9]

    Bruno Currie. 2019. https://doi.org/10.1080/00397679.2019.1648002 The iliad, the odyssey, and narratological intertextuality* . Symbolae Osloenses, 93(1):157--188

  2. [10]

    Dexter, Theodore Katz, Nilesh Tripuraneni, Tathagata Dasgupta, Ajay Kannan, James A

    Joseph P. Dexter, Theodore Katz, Nilesh Tripuraneni, Tathagata Dasgupta, Ajay Kannan, James A. Brofos, Jorge A. Bonilla Lopez, Lea A. Schroeder, Adriana Casarez, Maxim Rabinovich, Ayelet Haimson Lushkov, and Pramit Chaudhuri. 2017. https://doi.org/10.1073/pnas.1611910114 Quant...

  3. [11]

    Forstall and Walter J

    Christopher W. Forstall and Walter J. Scheirer. 2019. https://doi.org/10.1007/978-3-030-23415-7 Quantitative Intertextuality - Analyzing the Markers of Information Reuse . Springer

  4. [12]

    R.B. Hays. 1989. https://books.google.com/books?id=8faLhqRXH24C Echoes of Scripture in the Letters of Paul . Yale University Press. [ Kristeva(1986 [1969]) ] kristev Julia Kristeva. 1986 [1969]. Word, dialogue and novel, in: T Moi (ed), The Kristeva Reader. Columbia University...

  5. [13]

    John Lee. 2007. A Computational Model of Text Reuse in Ancient Literary Texts . In Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics , pages 472--479, Prague, Czech Republic. Association for Computational Linguistics

  6. [14]

    McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, and David Yarowsky

    Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, and David Yarowsky. 2020. The Johns Hopkins University Bible Corpus : 1600+ Tongues for Typological Exploration . In Proceedings of the Twelfth Language Resources ...

  7. [15]

    Hope McGovern, Hale Sirin, Tom Lippincott, and Andrew Caines. 2024. https://doi.org/10.18653/v1/2024.ml4al-1.26 Detecting narrative patterns in biblical H ebrew and G reek . In Proceedings of the 1st Workshop on Machine Learning for Ancient Languages (ML4AL 2024), pages 269--2...

  8. [16]

    Maria Moritz, Andreas Wiederhold, Barbara Pavlek, Yuri Bizzoni, and Marco B \"u chler. 2016. https://doi.org/10.18653/v1/D16-1190 Non- Literal Text Reuse in Historical Texts : An Approach to Identify Reuse Transformations and its Application to Bible Reuse . In Proceedings of ...

  9. [17]

    Steve Moyise. 2002. https://doi.org/10.4102/ve.v23i2.1211 Intertextuality and biblical studies: A review . Verbum et Ecclesia, 23

  10. [18]

    Conley Owens. 2023. https://www.openbible.info/labs/cross-references/ Bible cross references

  11. [19]

    Andrew Piper, Richard Jean So, and David Bamman. 2021. https://api.semanticscholar.org/CorpusID:237461067 Narrative theory for computational narrative understanding . In Conference on Empirical Methods in Natural Language Processing

  12. [20]

    Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.213 COMET : A neural framework for MT evaluation . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685--2702, ...

  13. [21]

    Yisi Sang, Xiangyang Mou, Jing Li, Jeffrey Stanton, and Mo Yu. 2022. https://api.semanticscholar.org/CorpusID:248496789 A survey of machine narrative reading comprehension assessments . In International Joint Conference on Artificial Intelligence

  14. [22]

    Hale Sirin. 2022. The Art of Scholarship: Auerbach, Tanpınar, and the Idea of Literary Knowledge. Ph.D. thesis, Johns Hopkins University

  15. [23]

    Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2020. http://arxiv.org/abs/2008.00401 Multilingual translation with extensible multilingual pretraining and finetuning

  16. [24]

    Torrey and J

    R.A. Torrey and J. Canne. 1982. The Treasury of Scripture Knowledge: Five-hundred Thousand Scripture References and Parallel Passages . Hendrickson Publishers Marketing, LLC

  17. [25]

    C van der Waal. 1980. http://www.jstor.org/stable/43047804 The continuity between the old and new testaments . Neotestamentica, 14:1--20

  18. [26]

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.