{"id":"f7db15aa-407d-4596-8671-e977d4643c67","arxiv_id":"2501.10731","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"Human-translated Bibles show higher intertextuality ratios than machine-translated ones in this embedding-based measure, but source-text mismatch limits the conclusion.","lead":"This paper gauges how well Biblical intertextual echoes survive translation by comparing cosine similarities of verse embeddings across human and machine translations. It reports that human translations often strengthen intertextuality, but the comparison is confounded because the human versions were not translated from the same source manuscripts as the machine ones.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline claim is contradicted by the paper's own Table 4: Turkish NMT ratios exceed or tie the human translation in every column, and most 95% CIs overlap; the source-text confound admitted in Limitations independently prevents a causal reading.","rationale":"The stress-test pass confirms the reader's REJECT but identifies a more direct problem than the one listed as weakest. Before worrying about whether human and machine translations come from different source texts, Table 4 already fails to show the claimed pattern. This is not a matter of external consensus; it is an internal inconsistency in the paper's key quantitative evidence. The metric itself has some support: the benchmark validation on Burns et al. (2021) yields a sensible ratio (1.55, CI [1.53, 1.56]), and the bootstrapped CIs are a reasonable attempt at uncertainty quantification. Credit is due for releasing code and data and for flagging the source-text limitation. However, the source-text limitation is not merely a caveat; it is a confound that breaks the causal interpretation of any remaining difference, and the Turkish row shows the asserted consistency is not even present. A matched-source replication would settle both issues, but on the evidence in the manuscript the central claim is unsupported. The verdict should remain REJECT; no change from the reader.","tokens_in":7740,"tokens_out":8557,"duration_ms":85918,"concrete_test":"Run a matched-source replication for all five languages: use Aya23 to translate the actual English source text underlying each JHUBC translation (e.g., the KJV/ASV base) into English, Finnish, Turkish, Swedish, and Marathi; recompute Table 4 on human and NMT outputs from identical sources, with the same 10,000-resample bootstrap procedure. Then test per language whether human ratios exceed NMT ratios. If the ordering does not persist across all five languages, or if Turkish/Swedish remain non-significant or reversed, the headline claim is unsupported; if it persists cleanly with separated CIs in every language, both the source-text confound and the internal consistency problem would be resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that human translations consistently show higher intertextuality than machine translations, supporting an amplificatory propensity in human translators—fails on the paper's own Table 4. Section 5 states 'Human translations consistently show higher levels of intertextuality,' but in Turkish the NMT ratios are higher than the human ratios in the within-Jewish column (1.60 vs 1.50), the within-Christian column (1.71 vs 1.43), and essentially tied across testaments (1.52 vs 1.51). Swedish and most other language rows have 95% bootstrap CIs that overlap between human and NMT; only Marathi within-testament shows a clearly separated difference. The paper's comparative claim is therefore not merely an overstatement of a trend; it is contradicted by the data as reported. A separate, independent flaw is flagged in the Limitations section: most JHUBC human translations were made from English versions, whereas the machine translations were generated from the Hebrew/Greek manuscripts. Even if the table had shown a consistent human advantage, that advantage could be a source-text effect rather than a translator effect. The single qualitative example in Table 5 is hand-selected and cannot substitute for the missing controlled comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces a corpus-level intertextuality ratio computed as the average cosine similarity of ground-truth verse pairs divided by the average similarity of same-chapter random pairs in a multilingual embedding space. The authors validate the metric on a Latin benchmark (ratio 1.55, CI [1.53, 1.56]), then apply it to biblical cross-references across five languages, comparing human translations from the JHUBC with Aya23 machine translations of the ancient Hebrew and Greek manuscripts. They claim that human translations consistently show higher intertextuality than machine translations, which they interpret as evidence that human translators amplify literary characteristics. The paper also includes a qualitative case study of one overemphasized intertextual pair.","tokens_in":7966,"tokens_out":4941,"duration_ms":46498,"significance":"The metric validation on an external Latin benchmark and the release of code and data are genuine strengths; the paper gives a reproducible, parameter-light way to compare intertextuality at corpus level. If the human-machine comparison were valid, this would be a useful contribution to computational intertextuality and translation studies. However, the central comparative claim is undercut by two problems: the JHUBC human translations are mostly mediated by English rather than translated directly from the source manuscripts, and Table 4 contains a direct counterexample (Turkish). Thus the paper's main interpretive conclusion does not follow from the presented evidence.","major_comments":[{"comment":"The sentence 'Human translations consistently show higher levels of intertextuality' is contradicted by the Turkish row of Table 4: the Aya23 NMT ratio exceeds the human ratio in the within-Jewish column (1.60 vs 1.50) and within-Christian column (1.71 vs 1.43) and is essentially tied across testaments (1.52 vs 1.51). In most other language rows the bootstrap 95% CIs overlap between human and NMT (e.g., Swedish within-Jewish 1.33 ± 0.15 vs 1.31 ± 0.22), so the data do not establish a consistent human advantage. This claim is the basis for the abstract's and Section 5's conclusion about human amplificatory propensity, so it is load-bearing.","section":"§5, Table 4"},{"comment":"The comparison of human and machine translation cannot isolate translator behavior because the two conditions are generated from different source texts. The machine translations are produced from the Hebrew Old Testament, Greek Old Testament, and Greek New Testament (Section 3.2), while the Limitations state that most JHUBC human translations were not translated directly from ancient manuscripts but instead work from English translations. Any observed human-vs-machine difference could therefore be a source-text effect rather than a translator effect. Since the paper's contribution 3 and the Section 5 interpretation rely on this contrast, this is a load-bearing confound.","section":"Limitations; §3.2"},{"comment":"The single hand-selected example of Hebrews 8:12 / Isaiah 43:25 is offered as evidence that human translation overemphasizes intertextuality while machine translation provides a neutral baseline, but one qualitative example cannot support a general claim about all human translations, especially when Table 4 already fails to show a consistent overall pattern. Moreover, 'neutral baseline' is not operationalized: the NMT ratios in Table 4 are substantially above 1 (e.g., Turkish NMT 1.60 and 1.71), so machine translations are not neutral by the paper's own metric.","section":"§5, Table 5"}],"minor_comments":[{"comment":"The text contains a typo: 'semiotic density of a any given text' should read 'semiotic density of any given text.'","section":"§1"},{"comment":"The intertextuality ratio depends on the vote threshold of 50 used to define ground-truth references; a short sensitivity analysis around this threshold would help establish that the main conclusions are not threshold artifacts.","section":"§2"},{"comment":"The prose should distinguish point estimates from statistically distinguishable differences; several apparent human advantages have overlapping 95% CIs and therefore should not be described as consistent effects.","section":"§5, Table 4"},{"comment":"The term 'neutral baseline' is used without a precise definition; please clarify whether it means a ratio near 1, a ratio close to the source manuscript, or something else, and report the corresponding values.","section":"§5"}],"recommendation":"reject","confidential_remarks":"The paper has a reusable metric and positive external validation, and the authors are transparent in the Limitations about the source-text mismatch. However, the central comparative claim fails on the paper's own Table 4, and the source-text confound cannot be repaired with the existing data. If the authors can re-run the comparison with human translations made directly from the Hebrew/Greek originals, or substantially restrict their claims, a revised metric-focused study could be resubmitted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick note on arXiv:2501.10731. The headline claim — that human translations amplify intertextuality relative to machine translations — doesn't survive contact with the paper's own Table 4. In Turkish the NMT ratios are higher than the human ones in both within-testament columns (1.60 vs 1.50 Jewish; 1.71 vs 1.43 Christian) and essentially tied across testaments (1.52 vs 1.51). Swedish human and NMT are within noise everywhere. Most 95% CIs overlap; only Marathi shows a clean separation. So the abstract's 'human translators consistently show higher levels of intertextuality' is not what the data show, and the 'neutral baseline' for machines is wishful.\n\nThe Limitations section does the paper credit: it admits that most JHUBC human translations were made from English rather than from the Hebrew/Greek originals, while the machine translations were generated from the ancient manuscripts. That confound alone prevents any causal reading about translator propensity, even if the numbers had lined up. The qualitative example in Table 5 is hand-picked.\n\nWhat's genuinely useful: the metric is refreshingly simple — a ratio of mean cosine similarities for known intertextual pairs over random same-chapter pairs — and it validates on Burns et al.'s Latin corpus (1.55, CI [1.53,1.56]). That's a real sanity check. The code and data are released, the five-language comparison is a nice resource, and the study opens a question that's worth asking. The paper is not a rehash of prior work; it's the first to apply embedding-based intertextuality scoring to translation, as far as I know.\n\nThe core issue is that the paper overreaches from an exploratory, confounded comparison to a strong theoretical claim. If the authors redid the human side using translations made directly from the originals (or at least compared like with like), the contribution would be solid. As it stands, the metric and the corpus are salvageable; the comparative claim is not.\n\nI'd send it to review — a serious referee could push the authors to fix the design or reframe the claims — but I wouldn't cite the human-vs-machine result. For a reading group, it's a useful cautionary tale about source-text confounds in cross-lingual embedding studies.","headline":"The intertextuality metric is clean and externally validated, but the paper's central human-vs-machine claim is contradicted by its own Table 4 and further undermined by a source-text confound admitted in the Limitations.","tokens_in":8465,"tokens_out":3784,"would_cite":false,"duration_ms":34351,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multilingual embedding ratios reveal that human Bible translations amplify intertextual links between testaments, while machine translations provide a neutral baseline.","keywords":["intertextuality","multilingual embeddings","machine translation","biblical cross-references","rhetorical devices","cosine similarity","human vs machine translation","corpus-level metric"],"falsifier":"Take a single source text—say, the Greek New Testament—and have both a human translator and a machine translation system render it into the same language. If the human translation's intertextuality ratio does not exceed the machine's for the same ground-truth cross-references, the amplification claim would be falsified. The paper's own limitation note makes this the natural control experiment.","tokens_in":7520,"feed_emoji":"🔗","tokens_out":8376,"duration_ms":76411,"temperature":0.7,"pith_summary":"This paper asks whether a simple geometric measure can characterise how translation changes intertextuality—the web of references between texts—and uses it to compare human and machine translations of the Bible. The measure is the ratio of average cosine similarity between verses known to be cross-linked to average similarity between random verse pairs, computed in a shared multilingual embedding space. Across five languages, human translations show consistently higher intertextuality ratios than machine translations of the same ancient manuscripts, which the authors read as evidence that human translators have a propensity to amplify literary characteristics such as continuity between testaments. The metric also returns a high ratio on a curated corpus of allusions in Latin literature, suggesting it detects genuine intertextual links. A sympathetic reader would take the paper's contribution to be a quantitative, corpus-level method for studying rhetorical effects in translation, plus evidence that human translation does not preserve but often heightens intertextual density.","feed_headline":"Human translators amplify cross-references; machines stay neutral","feed_subtitle":"A five-language cosine-similarity ratio shows human Bibles keep more intertextuality than machine output.","key_machinery":"The mechanism is the cosine-similarity ratio between verse embeddings from a multilingual model. For a predetermined set of ground-truth cross-references, the method computes the mean cosine similarity of the linked verse pairs and divides it by the mean similarity of random pairs drawn from the same chapter, producing a ratio greater than one when intertextual verses are more similar than chance. Bootstrap resampling with 10,000 iterations yields 95% confidence intervals, allowing comparisons across translation conditions. The authors validate the metric on a benchmark of 945 expert-curated allusions in Latin literature, where it produces a ratio of 1.55, before applying it to biblical cross-references that are split into within-testament and across-testament sets.","core_discovery":"The paper's central claim is that multilingual embedding spaces can characterise intertextuality at the corpus level, and that human translations of the Hebrew and Christian testaments exhibit a higher degree of intertextuality than machine translations, which act as a neutral baseline. In the authors' words, 'human translations consistently show higher levels of intertextuality,' and this is quantified in the intertextuality ratios of Table 4 for English, Finnish, Turkish, Swedish, and Marathi. The authors additionally provide a qualitative example in which a human English translation nearly doubles the similarity between a verse in Hebrews and a verse in Isaiah by rendering both with the word 'sin,' while the machine translation restores distance. The paper frames this as support for existing scholarship proposing that human translators amplify certain literary characteristics of the original manuscripts.","pith_inferences":["We think the paper's own limitation statement points to a decisive follow-up: translate the ancient Hebrew and Greek manuscripts directly into a language by both a professional human translator and a machine, then recompute the ratio; without that control, the human-versus-machine gap could reflect translation from an English intermediary rather than a human propensity.","The same ratio method could be used on other highly translated texts with known allusions, such as classical epic or legal corpora, to see whether amplification generalises beyond the Bible.","If amplification is a general human-translation effect, it may also affect other rhetorical devices (e.g., metaphor or repetition) measurable by embedding similarity, not just intertextuality."],"forward_implications":["If the amplification claim holds, readers of human-translated Bibles encounter a denser web of cross-testamental references than the ancient source texts themselves contain.","The ratio metric provides a quantitative tool for testing long-standing theories in translation studies about translator-driven literary amplification.","Machine translation output, by staying closer to the source's semantic surface, could serve as a controlled baseline for measuring stylistic intervention in human translation.","Intertextuality ratios are language-dependent: English shows the largest amplification and Marathi the smallest, so any account of translator behavior must explain cross-linguistic variation."],"supporting_citations":[{"why":"Supplies the ground-truth biblical cross-references that the ratio metric measures against.","marker":"(Owens, 2023)"},{"why":"Provides the five human Bible translations used as the human condition.","marker":"(McCarthy et al., 2020)"},{"why":"Provides the machine translation system that produces the neutral baseline.","marker":"(Aryabumi et al., 2024)"},{"why":"Supplies the benchmark corpus for intertextuality and demonstrates that neural embeddings capture intertextuality, which the present method extends.","marker":"Burns et al. (2021)"},{"why":"Prior hypothesis that human translators amplify literary characteristics, which this paper tests.","marker":"(McGovern et al., 2024)"},{"why":"Provides the COMET metric used to evaluate translation quality, supporting the adequacy of the machine translation baseline.","marker":"(Rei et al., 2020)"}],"fun_headline_variants":["Human translators amplify intertextuality; machines stay neutral","Multilingual embeddings show humans preserve literary echoes","Machine translations are neutral; humans intensify references","Biblical intertextuality: human translators boost, machines flatten"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's central comparison assumes that the human and machine translations are translations of the same source text; in fact, most human translations in the corpus were made from English versions, not from the ancient Hebrew and Greek manuscripts that were fed to the machine translator.","fun_headline_variants_meta":{"raw":{"variants":["Human translators amplify intertextuality; machines stay neutral","Multilingual embeddings show humans preserve literary echoes","Machine translations are neutral; humans intensify references","Biblical intertextuality: human translators boost, machines flatten"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000355,"raw_usage":{"total_tokens":1876,"prompt_tokens":838,"completion_tokens":1038,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":454,"completion_tokens_details":{"reasoning_tokens":978}},"tokens_in":454,"tokens_out":1038,"duration_ms":10073,"temperature":1.0,"reasoning_tokens":978,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:00:56.845087+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a single source text—say, the Greek New Testament—and have both a human translator and a machine translation system render it into the same language. If the human translation's intertextuality ratio does not exceed the machine's for the same ground-truth cross-references, the amplification claim would be falsified. The paper's own limitation note makes this the natural control experiment.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ground-truth biblical cross-references that the ratio metric measures against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the benchmark corpus for intertextuality and demonstrates that neural embeddings capture intertextuality, which the present method extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior hypothesis that human translators amplify literary characteristics, which this paper tests."}],"review_version":1}