{"id":"93469966-e7f4-4914-b431-9bf0911dae24","arxiv_id":"2504.19850","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Creative translation units increase readers' cognitive load, most strongly in human translation and least in machine translation, in a pilot with eight readers.","lead":"This pilot study measured Dutch readers' eye movements while they read an English short story in four versions: machine translation, post-edited, human-translated, and the original English. It finds that creative translation choices draw more reader attention, and this effect is strongest in human translation and weakest in machine translation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"With only two participants per modality, the HT>PE>MT interaction in Table 5 is not separable from participant identity; the ordering claim is the weakest link.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the modality-by-creativity ordering is estimated from two participants per condition, making it impossible to separate modality effects from participant idiosyncrasy. This is the correct point of maximum fragility in the argument. The paper is transparent about being a pilot and makes code and data available, so the concern is testable rather than fatal. The pooled UCP main effect (RQ1) is more defensible because it aggregates over participants, but the abstract's stronger claim about the ordering HT > PE > MT depends on interaction coefficients that cannot survive the n=2-per-cell design without additional evidence. A leave-one-participant-out or permutation check would settle whether the ordering is a stable signal or an artifact of two readers. Since the reader already conditioned the verdict on addressing the small-sample weakness, no verdict change is warranted; the condition should explicitly include the proposed stability check.","tokens_in":26961,"tokens_out":3847,"duration_ms":44051,"concrete_test":"Run a leave-one-participant-out stability analysis: refit the TFD GAM from Section 3.7 eight times, each time dropping one of the eight participants, and record the HT-vs-PE contrasts for CS and for Rep. If either contrast changes sign or loses significance in any fold, the interaction is driven by a single reader. Complement this with a permutation test that randomly reassigns the eight participants to the four modalities and recomputes the HT-vs-PE contrast; if the observed contrast is not in the tail of the permutation distribution, the ordering is not detectable beyond participant identity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two parts: UCPs increase cognitive load, and this effect is ordered HT > PE > MT. The first part is supported by a main effect pooling over many observations and is comparatively robust. The second, more novel part rests entirely on the Modality-by-Creativity interaction estimates in Section 4.2.2, Table 5. But the design in Section 3.2 assigns only two participants per modality (n=8 total), so any stable trait of those two readers is statistically confounded with the modality label. The GAM treats participants as random effects, yet with two clusters per condition the between-participant variance component is extremely unstable; the significant HT-vs-PE contrasts for CS (p=0.042) and Rep (p=0.001) could be produced by one atypical reader. The 'weakest for MT' leg is even thinner, since MT contributes only 26 CS segments (Table 1). The null error effect is also hard to interpret because Section 4.3 reports that MT readers skim-read sections with many errors, which could mask rather than refute an error effect. Thus the headline ordering is not supported by the data as analyzed, even though the pooled UCP effect and the qualitative RTA patterns are informative.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a pilot eye-tracking study of Dutch readers reading the same short story in four conditions: machine translation (MT), post-editing (PE), human translation (HT), and the original source text (ST). The authors annotate units of creative potential (UCPs) in the source and translations and classify them as creative shifts (CS) or reproductions, and they also annotate translation errors. They measure cognitive load via total fixation duration and other eye-tracking variables, complementing this with questionnaire data and retrospective think-aloud (RTA) interviews from eight participants (two per modality). The central claims are that UCPs increase cognitive load, that this increase is strongest for HT and weakest for MT, and that errors have no significant effect on cognitive load. The paper presents a transparent GAM analysis, a separate analysis of zero-fixation segments, non-parametric tests for other dependent variables, and a qualitative RTA analysis.","tokens_in":27194,"tokens_out":4957,"duration_ms":48317,"significance":"If the main effect of UCPs on cognitive load were established, it would be a meaningful contribution to translation reception research, supporting the idea that word-level creative choices in translation measurably shape reader attention. The paper is also valuable as a methodological pilot: it openly shares code and data, it distinguishes main effects from interactions, and it triangulates quantitative eye-tracking with qualitative RTA data, which is rare and instructive. The novel interaction claim (HT > PE > MT) is the most interesting element but also the least supported by the evidence as analyzed. The study's transparency and its explicit framing as a pilot make it a reasonable springboard for further research, though the strength of the abstract's conclusions currently exceeds what the data can bear.","major_comments":[{"comment":"With only two participants per modality (n=8 total), the significant Modality-by-Creativity interaction terms in Table 5 (HT-vs-PE contrasts: CS p=0.042, Rep p=0.001) are statistically confounded with participant identity. Because participant is treated as a random effect with only two clusters per condition, the between-participant variance component cannot be reliably estimated, and the reported ordering HT > PE > MT may be produced by a single atypical reader per condition. The abstract's claim that the UCP effect is 'highest for HT and lowest for MT' is therefore not supported by the data as analyzed. The authors should either present per-participant interaction estimates (e.g., an appendix showing the relevant slopes separately for each of the eight participants), or explicitly re-word the ordering claim as a hypothesis generated by the pilot rather than as a demonstrated result.","section":"Section 3.2 and Table 5"},{"comment":"The null effect of errors on TFD (Table 5, Errors Yes p=0.867) is presented as a finding in the abstract ('no effect of error was observed'), but the RTA data in Section 4.3 (Theme 4) indicate that MT participants engaged in skim-reading of error-dense passages, with participants 'giving up' and 'accepting that I wouldn't really get the thing' (Section 4.3; Section 5 RQ4). If reading strategy changes as a function of error density, the segment-level comparison of error vs. no-error units is not a clean test of the effect of errors on cognitive load: the expected increase in fixation duration may be cancelled by reduced effort in skimmed passages. The authors should either model reading strategy (e.g., separately analyzing first-pass reading and re-reading on error-dense vs. error-free passages) or explicitly qualify the null result as uninterpretable given the identified masking mechanism, rather than naming it as a finding in the abstract.","section":"Section 4.3, Theme 4 and Section 5, RQ4"},{"comment":"The GAM analysis excludes zero-fixation segments, yet Table 17 shows that CS segments have a higher zero rate (5%) than Reproductions (1.6%) and non-UCP units (2.6%). Excluding these skipped segments from the TFD model could bias the estimated effect of Creativity, because a segment that is skipped is recorded as zero TFD and is not merely missing at random. The separate chi-square analyses of zero rates were not significant, but the authors should still justify or model the zero-inflation (e.g., a zero-inflated model or a two-part analysis) or, at minimum, report whether the significance of the CS and Rep main effects (Table 5) survives a sensitivity analysis that treats skipped segments as zero-valued observations.","section":"Section 3.7 and Appendix C.3.1"},{"comment":"The authors state that the model was fitted 'after analysing the data' (Section 3.7), which indicates a post-hoc selection of the GAM specification and of the included predictors. While this is disclosed transparently, the multiple tests in Table 5 (and the additional tests in Appendix C.4) are not corrected for multiplicity, and several interaction p-values are in the 0.01-0.05 range that would not survive even mild correction. The manuscript should at least acknowledge this in the limitations and should avoid placing heavy interpretive weight on the least robust interaction terms.","section":"Section 3.7 and Table 5 (post-hoc model selection)"}],"minor_comments":[{"comment":"The text reports '28% (R2 = 0.241)' for FPT; 0.241 corresponds to 24.1%, not 28%, and the parenthetical should be corrected.","section":"Section 3.7"},{"comment":"The sentence 'Spearman's correlation shows again a low negative correlation (ρ = 0.23)' omits the negative sign; it should read ρ = -0.23 for consistency with the TFD and RP values.","section":"Section 4.2.3"},{"comment":"The three-way interaction HT:Rep:Error reaches significance (p=0.018) but is not discussed anywhere in the text; please comment on it or explicitly justify why it is not interpreted.","section":"Table 16"},{"comment":"There is a typo in 'Kurt V onnegut'; it should be 'Kurt Vonnegut'.","section":"Section 3.1"},{"comment":"In the Moorkens et al. (2018) entry, the co-author name 'Antonio Tora' appears to be missing the final 'l'; it should be 'Antonio Toral'.","section":"References"},{"comment":"The labels 'UCP*' and 'Not*' are not self-explanatory; the table caption should indicate that these rows refer to the source-text (ST) segments that correspond to UCPs in the original and non-UCP segments of the ST, respectively.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The paper's central contribution is the methodological combination of eye-tracking, word-level creativity annotation, and RTA interviews, and its full data/code transparency is a clear strength. The main obstacle is the mismatch between the abstract's causal-sounding claims and the n=2-per-modality design. The UCP main effect is likely robust, but the HT>PE>MT ordering and the null error finding are not supportable as stated. With a careful reframing of the results as hypothesis-generating pilot findings, additional per-participant diagnostics, and a more qualified abstract, the paper could be publishable as a methods-oriented pilot study. The editor may also want the authors to be explicit that the UCP/CS annotation is carried over from their prior work without re-assessed inter-annotator agreement in this study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: the headline result splits in two. The pooled finding that units of creative potential (UCPs) attract more cognitive load than surrounding words is reasonably supported and fits prior literary-reading work. The second half of the headline—that this effect is ordered HT > PE > MT—is not supported by this dataset as analyzed. With two readers per modality, any stable trait of those readers is confounded with the modality label, and the MT leg leans on only 26 CS segments. Treat the ordering as a hypothesis, not a result.\n\nWhat's new: word-level eye-tracking on the same texts from the authors' prior questionnaire study, with creativity annotations attached per segment, plus RTA interviews. That combination is genuinely new. The analysis is transparent: GAM with participant and UCP random effects, zero-fixation segments handled separately, frequency analysis included, code and data open. The RTA material is useful and gives a plausible reason for the null error effect (MT readers skim sections with many errors). The paper calls itself a pilot and does not hide the small n.\n\nSoft spots, in proportion. The main one is the interaction claim. Table 5 reports significant HT-vs-PE contrasts for CS and Rep, but at n=2 per modality those contrasts ride on participant identity. The authors acknowledge the small n but still write conclusions as if the ordering is empirical. A referee should push them to reframe RQ2 as exploratory and move the ordering into a preregistered follow-up. Second, the creativity annotation comes from the authors' own framework without inter-annotator agreement reported here; for word-level analysis that is a real gap, though minor for a pilot. Third, the null error effect is honestly discussed but under-interpreted: skim reading can mask an effect, so 'no effect' is not a clean null. The frequency analysis is a nice addition and shows UCP words are not simply rarer words.\n\nWho this is for: translation reception researchers and people working on literary MT evaluation. The methodology—pairing eye-tracking with segment-level creativity annotation—is worth reading even if the quantitative claims stay tentative. It deserves a serious referee: send it out, ask for reframing of the interaction as exploratory and a fuller participant-sensitivity check. My own verdict: conditional, with the pooled UCP effect and RTA analysis carrying the paper.","headline":"A transparent pilot whose pooled UCP effect is credible but whose HT>PE>MT ordering rests on two participants per condition.","tokens_in":27668,"tokens_out":2274,"would_cite":false,"duration_ms":22302,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that word-level creative solutions in translation increase readers' cognitive load, most in human translation, least in machine translation.","keywords":["eye-tracking","cognitive load","translation reception","literary translation","machine translation","post-editing","creativity","unit of creative potential"],"falsifier":"A replication with at least twenty readers per condition that finds no systematic increase in total fixation duration on creative shifts, or finds the increase to be equally large in machine translation as in human translation, would refute the central claim. The result would be strengthened if the replication also showed error-heavy segments producing longer dwell times, contradicting the authors' skim-reading explanation.","tokens_in":26783,"feed_emoji":"👁️","tokens_out":3724,"duration_ms":34408,"temperature":0.7,"pith_summary":"This pilot study claims that the creative solutions a translator chooses for difficult words and phrases measurably increase how long readers' eyes dwell on those words, and that this effect is strongest when the text is a human translation and weakest when it is raw machine translation. The authors use units of creative potential (UCPs) to locate words that force a translator to solve a problem creatively, and compare reader eye movements across four versions of the same short story: human translation, post-editing, machine translation, and the original source text. They find no significant effect of translation errors on cognitive load. If true, the result suggests that creative translation is not just an aesthetic preference but a measurable influence on reader attention, and that automated translation may flatten exactly the moments that engage readers.","feed_headline":"Translators' creative choices measurably steer readers' eyes","feed_subtitle":"Eye-tracking pilot finds creative units draw the most attention in human translations and least in raw machine translation.","key_machinery":"The central object is the unit of creative potential (UCP): a word or group of words in the source text that cannot be translated routinely and demands problem-solving creativity. Each UCP in the Dutch translations is annotated as a creative shift (CS, deviation from the original), a reproduction (no deviation), or an omission, and total fixation duration from an eye-tracker is the measure of cognitive load. The eye-tracking data are modelled with a generalized additive mixed-effects model, and retrospective think-aloud interviews with gaze replays provide the qualitative triangulation that links the measured attention to enjoyment and immersion.","core_discovery":"Readers show higher cognitive load on units of creative potential than on ordinary text, and the size of this effect depends on the translation modality: it is largest for human translation and smallest for machine translation, with post-editing in between. Errors, by contrast, show no significant effect on cognitive load. The retrospective think-aloud interviews lead the authors to hypothesise that the extra attention given to creative units reflects enjoyment and immersion rather than difficulty.","pith_inferences":["If the attention-on-creativity effect replicates with more readers, publishers using machine translation for literary genres could use eye-tracking on UCPs as a pre-screening tool to decide which texts need human translators.","The authors' skim-reading explanation for the null error effect is testable: compare fixation counts in the first versus second half of machine-translated texts, expecting a drop as errors accumulate.","The same design applied to a language pair with a different machine-translation quality baseline, or to a longer text, would show whether the human-translation advantage is about creativity per se or about overall textual quality.","Since creative shifts versus reproductions showed no cognitive-load difference, the load may be driven by UCP status (the source unit being problematic) rather than by the translator's degree of deviation; a word-frequency-matched control would clarify this."],"forward_implications":["If creative units reliably draw more fixation time, word-level creativity becomes a measurable dimension of translation reception.","Human translation's creative advantage shows up as measurable reader attention, not just preference ratings, with post-editing situated between human translation and machine translation.","Machine translation's flattening of creative solutions reduces the attention-grabbing potential of literary texts, even in places without overt errors.","Error-based metrics alone may miss what distinguishes translation modalities for readers; creativity annotations add signal that error counts do not capture.","The absence of an error effect, plausibly due to skim-reading of error-heavy machine translation, warns that whole-text error counts do not translate linearly into reading effort."],"supporting_citations":[{"why":"Supplies the source text, the four translation modalities, the UCP annotations, the questionnaire, and the earlier reception findings that this study extends.","marker":"Guerberof-Arenas and Toral (2024)"},{"why":"Defines units of creative potential and creative shifts versus reproductions, the annotation scheme used to classify creativity.","marker":"Guerberof-Arenas and Toral (2020)"},{"why":"Operationalises creative shifts as a measurable means of translation creativity, grounding the UCP concept.","marker":"Bayer-Hohenwarter (2011)"},{"why":"Provides the cognitive-control hypothesis linking eye movements to cognitive effort, the theoretical basis for using eye-tracking as a measure of cognitive load.","marker":"Rayner and Reingold (2015)"},{"why":"Shows that errors increase fixation duration and count in non-literary machine translation, the baseline that this study's null error effect runs against.","marker":"Kasperavičienė et al. (2020)"},{"why":"Earlier eye-tracking evidence that sections with translation errors draw longer fixations, another contrast for the paper's error findings.","marker":"Stymne et al. (2012)"},{"why":"Links increased cognitive load to immersion and engagement in reading, used to interpret the higher load on creative units as enjoyment-related.","marker":"Torres et al. (2021)"},{"why":"Finds that low-quality sentences and errors increase cognitive load in newspaper translations, a non-literary baseline the paper contrasts with its literary result.","marker":"Whyatt et al. (2024)"}],"fun_headline_variants":["Creative translation choices raise readers' cognitive load","Eye-tracking: creative units in translation draw more attention","Human translation's creative flair increases reader attention","Machine translation least taxing on readers' creative processing","Translation creativity: higher cognitive load for human versions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The statistical model must estimate the modality-by-creativity interaction from only two readers per condition; if either reader in a group is atypical, the claimed ordering of human, post-edited, and machine translation will not hold in a larger sample.","fun_headline_variants_meta":{"raw":{"variants":["Creative translation choices raise readers' cognitive load","Eye-tracking: creative units in translation draw more attention","Human translation's creative flair increases reader attention","Machine translation least taxing on readers' creative processing","Translation creativity: higher cognitive load for human versions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000221,"raw_usage":{"total_tokens":1390,"prompt_tokens":827,"completion_tokens":563,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":443,"completion_tokens_details":{"reasoning_tokens":492}},"tokens_in":443,"tokens_out":563,"duration_ms":5964,"temperature":1.0,"reasoning_tokens":492,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:41:35.642792+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication with at least twenty readers per condition that finds no systematic increase in total fixation duration on creative shifts, or finds the increase to be equally large in machine translation as in human translation, would refute the central claim. The result would be strengthened if the replication also showed error-heavy segments producing longer dwell times, contradicting the authors' skim-reading explanation.","supporting_citations":[{"cited_title":"These are good scores for reliability and shows the reliability of the scales","cited_arxiv_id":null,"evidence_quote":"Supplies the source text, the four translation modalities, the UCP annotations, the questionnaire, and the earlier reception findings that this study extends."}],"review_version":1}