Pith. sign in

REVIEW 4 major objections 6 minor 40 references

To MT or not to MT: An eye-tracking study on the reception by Dutch readers of different translation and creativity levels

T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper claims that word-level creative solutions in translation increase readers' cognitive load, most in human translation, least in machine translation.

desk verdict A transparent pilot whose pooled UCP effect is credible but whose HT>PE>MT ordering rests on two participants per condition. read the letter →

arxiv 2504.19850 v1 pith:YRNHXNVO submitted 2025-04-28 cs.CL

classification cs.CL
keywords eye-trackingcognitiveloadtranslationreceptionliterarymachinepost-editingcreativityunitofcreativepotential
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This pilot study claims that the creative solutions a translator chooses for difficult words and phrases measurably increase how long readers' eyes dwell on those words, and that this effect is strongest when the text is a human translation and weakest when it is raw machine translation. The authors use units of creative potential (UCPs) to locate words that force a translator to solve a problem creatively, and compare reader eye movements across four versions of the same short story: human translation, post-editing, machine translation, and the original source text. They find no significant effect of translation errors on cognitive load. If true, the result suggests that creative translation is not just an aesthetic preference but a measurable influence on reader attention, and that automated translation may flatten exactly the moments that engage readers.

What carries the argument

The central object is the unit of creative potential (UCP): a word or group of words in the source text that cannot be translated routinely and demands problem-solving creativity. Each UCP in the Dutch translations is annotated as a creative shift (CS, deviation from the original), a reproduction (no deviation), or an omission, and total fixation duration from an eye-tracker is the measure of cognitive load. The eye-tracking data are modelled with a generalized additive mixed-effects model, and retrospective think-aloud interviews with gaze replays provide the qualitative triangulation that links the measured attention to enjoyment and immersion.

What would settle it

A replication with at least twenty readers per condition that finds no systematic increase in total fixation duration on creative shifts, or finds the increase to be equally large in machine translation as in human translation, would refute the central claim. The result would be strengthened if the replication also showed error-heavy segments producing longer dwell times, contradicting the authors' skim-reading explanation.

Watch

Extended reading notes

Core claim

Readers show higher cognitive load on units of creative potential than on ordinary text, and the size of this effect depends on the translation modality: it is largest for human translation and smallest for machine translation, with post-editing in between. Errors, by contrast, show no significant effect on cognitive load. The retrospective think-aloud interviews lead the authors to hypothesise that the extra attention given to creative units reflects enjoyment and immersion rather than difficulty.

Load-bearing premise

The statistical model must estimate the modality-by-creativity interaction from only two readers per condition; if either reader in a group is atypical, the claimed ordering of human, post-edited, and machine translation will not hold in a larger sample.

Editorial extensions

If this is right

  • If creative units reliably draw more fixation time, word-level creativity becomes a measurable dimension of translation reception.
  • Human translation's creative advantage shows up as measurable reader attention, not just preference ratings, with post-editing situated between human translation and machine translation.
  • Machine translation's flattening of creative solutions reduces the attention-grabbing potential of literary texts, even in places without overt errors.
  • Error-based metrics alone may miss what distinguishes translation modalities for readers; creativity annotations add signal that error counts do not capture.
  • The absence of an error effect, plausibly due to skim-reading of error-heavy machine translation, warns that whole-text error counts do not translate linearly into reading effort.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the attention-on-creativity effect replicates with more readers, publishers using machine translation for literary genres could use eye-tracking on UCPs as a pre-screening tool to decide which texts need human translators.
  • The authors' skim-reading explanation for the null error effect is testable: compare fixation counts in the first versus second half of machine-translated texts, expecting a drop as errors accumulate.
  • The same design applied to a language pair with a different machine-translation quality baseline, or to a longer text, would show whether the human-translation advantage is about creativity per se or about overall textual quality.
  • Since creative shifts versus reproductions showed no cognitive-load difference, the load may be driven by UCP status (the source unit being problematic) rather than by the translator's degree of deviation; a word-frequency-matched control would clarify this.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper reports a pilot eye-tracking study of Dutch readers reading the same short story in four conditions: machine translation (MT), post-editing (PE), human translation (HT), and the original source text (ST). The authors annotate units of creative potential (UCPs) in the source and translations and classify them as creative shifts (CS) or reproductions, and they also annotate translation errors. They measure cognitive load via total fixation duration and other eye-tracking variables, complementing this with questionnaire data and retrospective think-aloud (RTA) interviews from eight participants (two per modality). The central claims are that UCPs increase cognitive load, that this increase is strongest for HT and weakest for MT, and that errors have no significant effect on cognitive load. The paper presents a transparent GAM analysis, a separate analysis of zero-fixation segments, non-parametric tests for other dependent variables, and a qualitative RTA analysis.

Significance. If the main effect of UCPs on cognitive load were established, it would be a meaningful contribution to translation reception research, supporting the idea that word-level creative choices in translation measurably shape reader attention. The paper is also valuable as a methodological pilot: it openly shares code and data, it distinguishes main effects from interactions, and it triangulates quantitative eye-tracking with qualitative RTA data, which is rare and instructive. The novel interaction claim (HT > PE > MT) is the most interesting element but also the least supported by the evidence as analyzed. The study's transparency and its explicit framing as a pilot make it a reasonable springboard for further research, though the strength of the abstract's conclusions currently exceeds what the data can bear.

major comments (4)
  1. [Section 3.2 and Table 5] With only two participants per modality (n=8 total), the significant Modality-by-Creativity interaction terms in Table 5 (HT-vs-PE contrasts: CS p=0.042, Rep p=0.001) are statistically confounded with participant identity. Because participant is treated as a random effect with only two clusters per condition, the between-participant variance component cannot be reliably estimated, and the reported ordering HT > PE > MT may be produced by a single atypical reader per condition. The abstract's claim that the UCP effect is 'highest for HT and lowest for MT' is therefore not supported by the data as analyzed. The authors should either present per-participant interaction estimates (e.g., an appendix showing the relevant slopes separately for each of the eight participants), or explicitly re-word the ordering claim as a hypothesis generated by the pilot rather than as a demonstrated result.
  2. [Section 4.3, Theme 4 and Section 5, RQ4] The null effect of errors on TFD (Table 5, Errors Yes p=0.867) is presented as a finding in the abstract ('no effect of error was observed'), but the RTA data in Section 4.3 (Theme 4) indicate that MT participants engaged in skim-reading of error-dense passages, with participants 'giving up' and 'accepting that I wouldn't really get the thing' (Section 4.3; Section 5 RQ4). If reading strategy changes as a function of error density, the segment-level comparison of error vs. no-error units is not a clean test of the effect of errors on cognitive load: the expected increase in fixation duration may be cancelled by reduced effort in skimmed passages. The authors should either model reading strategy (e.g., separately analyzing first-pass reading and re-reading on error-dense vs. error-free passages) or explicitly qualify the null result as uninterpretable given the identified masking mechanism, rather than naming it as a finding in the abstract.
  3. [Section 3.7 and Appendix C.3.1] The GAM analysis excludes zero-fixation segments, yet Table 17 shows that CS segments have a higher zero rate (5%) than Reproductions (1.6%) and non-UCP units (2.6%). Excluding these skipped segments from the TFD model could bias the estimated effect of Creativity, because a segment that is skipped is recorded as zero TFD and is not merely missing at random. The separate chi-square analyses of zero rates were not significant, but the authors should still justify or model the zero-inflation (e.g., a zero-inflated model or a two-part analysis) or, at minimum, report whether the significance of the CS and Rep main effects (Table 5) survives a sensitivity analysis that treats skipped segments as zero-valued observations.
  4. [Section 3.7 and Table 5 (post-hoc model selection)] The authors state that the model was fitted 'after analysing the data' (Section 3.7), which indicates a post-hoc selection of the GAM specification and of the included predictors. While this is disclosed transparently, the multiple tests in Table 5 (and the additional tests in Appendix C.4) are not corrected for multiplicity, and several interaction p-values are in the 0.01-0.05 range that would not survive even mild correction. The manuscript should at least acknowledge this in the limitations and should avoid placing heavy interpretive weight on the least robust interaction terms.
minor comments (6)
  1. [Section 3.7] The text reports '28% (R2 = 0.241)' for FPT; 0.241 corresponds to 24.1%, not 28%, and the parenthetical should be corrected.
  2. [Section 4.2.3] The sentence 'Spearman's correlation shows again a low negative correlation (ρ = 0.23)' omits the negative sign; it should read ρ = -0.23 for consistency with the TFD and RP values.
  3. [Table 16] The three-way interaction HT:Rep:Error reaches significance (p=0.018) but is not discussed anywhere in the text; please comment on it or explicitly justify why it is not interpreted.
  4. [Section 3.1] There is a typo in 'Kurt V onnegut'; it should be 'Kurt Vonnegut'.
  5. [References] In the Moorkens et al. (2018) entry, the co-author name 'Antonio Tora' appears to be missing the final 'l'; it should be 'Antonio Toral'.
  6. [Table 4] The labels 'UCP*' and 'Not*' are not self-explanatory; the table caption should indicate that these rows refer to the source-text (ST) segments that correspond to UCPs in the original and non-UCP segments of the ST, respectively.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the eye-tracking measurements are independent of the annotation framework, and the reported effects are ordinary statistical inferences rather than predictions derived from the model inputs.

full rationale

The paper is an empirical pilot study, not a derivation. Its central claims are that units of creative potential (UCPs) increase cognitive load, that this effect differs by translation modality, and that errors show no significant effect. These claims are based on eye-tracking data (TFD and related measures) collected in the experiment, then analyzed with a GAM. The UCP/CS/Reproduction annotations are taken from the authors' prior framework (Guerberof-Arenas and Toral 2020, 2024), but the target outcome — fixation durations — is measured independently with an eye-tracker. Nothing in the annotation definitions forces a particular TFD result; UCPs are defined as source-text units that pose translation problems, not as units that cause readers to fixate longer. The GAM is fitted to the collected data and reports coefficient estimates and p-values; this is inferential statistics, not a 'prediction' in the sense of an out-of-sample claim, and it is not circular to describe patterns in the same data one has modeled. The RTA-based suggestion that higher cognitive load in UCPs relates to enjoyment and immersion is explicitly introduced as a hypothesis ('leads us to hypothesize'), not as a derived result. The paper also transparently acknowledges the small sample size ('we only had two participants per modality, which does not allow for generalization'), which is a statistical limitation, not circularity. The self-citation to Guerberof-Arenas and Toral provides the materials, annotations, and questionnaire, but the novel contribution — eye-tracking and RTA data — is not contained in those inputs. No load-bearing step reduces to its own inputs by definition, and no fitted parameter is renamed as a prediction. Therefore the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The study rests on established eye-tracking assumptions, the authors' prior annotation framework for creativity, and the statistical premise that a mixed model can draw modality-level conclusions from two participants per condition. No new physical or conceptual entities are introduced.

assumptions (5)
  • domain assumption Total fixation duration (TFD) is a valid indicator of cognitive load during reading.
    Section 3.3 states TFD is the main variable for cognitive load, citing prior eye-tracking studies such as Skaramagkas et al. (2023) and Vanroy et al. (2022).
  • domain assumption The UCP, Creative Shift, and Reproduction annotations from Guerberof-Arenas and Toral (2020, 2022) reliably operationalize translation creativity across modalities.
    Section 3.1 borrows these annotations from the authors' prior study; this pilot does not report inter-annotator agreement.
  • domain assumption Two participants per condition are sufficient for the generalized additive mixed model to separate individual variability from modality effects.
    Sections 3.2 and 4.2.2 fit random effects for participants and UCPs despite having only two readers per modality.
  • domain assumption Excluding zero-fixation segments from the GAM and analyzing them separately does not bias the cognitive-load estimates.
    Section 3.7 describes this two-part analysis; the zero-fixation frequency table in Appendix C.3.1 shows no significant differences.
  • domain assumption Normalizing eye-tracking measures by words per segment makes segments of different lengths comparable.
    Section 3.6 states dataset II observations are normalized according to words per segment; this assumes a linear relationship between word count and reading time.

how reviews work

0 comments
Cite this review

Pith. "Pith review of To MT or not to MT: An eye-tracking study on the reception by Dutch readers of different translation and creativity levels." pith.science (2026). https://pith.science/paper/YRNHXNVO

@misc{pith2026250419850,
  author       = {Pith},
  title        = {Pith review of: To MT or not to MT: An eye-tracking study on the reception by Dutch readers of different translation and creativity levels},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YRNHXNVO}},
  note         = {Machine review of arXiv:2504.19850}
}
read the original abstract

This article presents the results of a pilot study involving the reception of a fictional short story translated from English into Dutch under four conditions: machine translation (MT), post-editing (PE), human translation (HT) and original source text (ST). The aim is to understand how creativity and errors in different translation modalities affect readers, specifically regarding cognitive load. Eight participants filled in a questionnaire, read a story using an eye-tracker, and conducted a retrospective think-aloud (RTA) interview. The results show that units of creative potential (UCP) increase cognitive load and that this effect is highest for HT and lowest for MT; no effect of error was observed. Triangulating the data with RTAs leads us to hypothesize that the higher cognitive load in UCPs is linked to increases in reader enjoyment and immersion. The effect of translation creativity on cognitive load in different translation modalities at word-level is novel and opens up new avenues for further research. All the code and data are available at https://github.com/INCREC/Pilot_to_MT_or_not_to_MT

Figures

Figures reproduced from arXiv: 2504.19850 by the authors.

Figure 1
Figure 1. An example of UCP, Reproduction, and CS from the experiment, including word-level glosses. 2.3 Reading in the Netherlands Dutch readership has some peculiarities worth men￾tioning regarding the cohabitation of English and Dutch languages. Recent market research shows that sales of foreign language books have increased 124% since 2020, accounting for 25% of all sales in 2024 (KVB Boekwerk, 2025), with the majority of… view at source ↗
Figure 3
Figure 3. Trial of PE version, divided in segments: en [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 2
Figure 2. the Trial Play Back Animation feature in [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Box plots for the eye-tracking data (dataset [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Scatter plot of word frequency and TFD (log [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Box plots of TFD (in ms.) for all independent [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: Scatter plot of word frequency and FPT (log [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: Scatter plot of word frequency and RP (log [PITH_FULL_IMAGE:figures/full_fig_p018_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

40 extracted references · 38 canonical work pages

  1. [1]

    At times, I struggled to understand what was happening in the story

  2. [2]

    My understanding of the character is unclear

  3. [3]

    I had a hard time recognizing the thread of the story

  4. [4]

    My mind wandered while reading the text

  5. [5]

    While reading, I found myself thinking about other things

  6. [6]

    I had a hard time keeping my mind on the text

  7. [7]

    While reading, my body was in the room, but my mind was inside the world created by the story

  8. [8]

    The text created a new world, and then that world suddenly disappeared when the story ended

Show all 40 references
  1. [9]

    At times when reading, the story world was closer to me than the real world

  2. [10]

    During the story, I felt sad when a main char- acter suffered in some way

  3. [11]

    The story affected me emotionally

  4. [12]

    I felt sorry for some of the characters

  5. [13]

    While reading the story I had a clear image of what the main character looked like

  6. [14]

    While reading the story I could envision the situations described

  7. [15]

    For enjoyment:

    I could imagine what the setting of the story looked like. For enjoyment:

  8. [19]

    Did you enjoy the text?

  9. [20]

    How likely is it that you would recommend the text to a friend?

  10. [21]

    Would you consider this text high literature? For translation reception:

  11. [22]

    The text was easy to understand

  12. [23]

    The text was well-written

  13. [24]

    I encountered words, sentences or paragraphs that were difficult to understand (including a box to write down which ones)

  14. [25]

    I encountered words, sentences or paragraphs that I found very beautiful (including a box to write down which ones)

  15. [26]

    I noticed I was reading a translation (including a box to indicate how people noticed)

  16. [27]

    What did you think of the translation?

  17. [28]

    Would you like to read a text by the same author and translator?

  18. [29]

    Would you like to read a text by the same author, but by a different translator?

  19. [30]

    These are good scores for reliability and shows the reliability of the scales

    Would you like to read a text by a different author, but the same translator? In Guerberof-Arenas and Toral (2024)’s study, the Cronbach’s alpha reliability coefficient (α) was 0.85 for narrative engagement, 0.87 for enjoyment and 0.79 for translation reception. These are good...

  20. [31]

    Confusion came from the narrative in HT, but from language use in MT

  21. [32]

    Engaging with and relating to narrative ele- ments occurred in HT, ST & PE

  22. [33]

    HT participants felt immersed in the story, the narrative, and the style

  23. [34]

    MT participants had difficulty understanding the text due to nonsensical words phrasing

  24. [35]

    PE participants were engaged in the narrative, but struggled with the style and characters at times

  25. [36]

    were just so weird

    Confusion came from the narrative in HT, but from language use in MT One of the very noticeable things is that across modalities all participants mentioned feeling con- fused multiple times throughout the narrative: feel- ings of confusion were mentioned 30 times in HT, 37 tim...

  26. [37]

    I was also very curious to see what would happen next

    Engaging with and relating to narrative elements occurred in HT, ST & PE HT, ST and PE participants mentioned feeling engaged in the narrative, including the events of the story and the moral issues at play, relating it to their own lives often. For these modalities, par- tici...

  27. [38]

    ironic and witty

    HT participants felt immersed in the story, the narrative, and the style Throughout the interviews, the HT participants made it clear that they liked the story in many of its facets and felt immersed in both the narrative and the style. P06_HT kept commenting about how im- mer...

  28. [39]

    [I] just didn’t really see what was happening here

    MT participants had difficulty understand- ing the text due to nonsensical words phrasing Translation errors in MT led to nonsensical phrasings, which caused the participants to strug- gle understanding the narrative and its events. Par- ticipants mentioned that they were not ...

  29. [40]

    strong imagery

    PE participants were engaged in the narra- tive, but struggled with the style and characters at times PE participants liked the story overall, thought it set-up the moral dilemma really well, and enjoyed themselves while reading the story. Both partici- pants related the situa...

  30. [325]

    Sheila Castilho and Natália Resende

    Routledge, New York City, US. Sheila Castilho and Natália Resende. 2022. MT-pese: Machine translation and post-editese. In Proceedings of the 23rd Annual Conference of the European As- sociation for Machine Translation, pages 305–306, Ghent, Belgium. European Association for M...

  31. [2020]

    In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 3790–3798, Marseille, France

    Literary machine translation under the mag- nifying glass: Assessing the quality of an NMT- translated detective novel on document level. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 3790–3798, Marseille, France. European Language Resources...

  32. [2022]

    good enough

    Comparing the effect of product-based metrics on the translation process. Frontiers in Psychology, 12. Callum Walker. 2021. Eye-tracking study of equivalent effect in translation : the reader experience of literary style. Palgrave Macmillan. Rebecca Webster, Margot Fonteyne, A...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.