REVIEW 4 major objections 6 minor 40 references
To MT or not to MT: An eye-tracking study on the reception by Dutch readers of different translation and creativity levels
T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper claims that word-level creative solutions in translation increase readers' cognitive load, most in human translation, least in machine translation.
desk verdict A transparent pilot whose pooled UCP effect is credible but whose HT>PE>MT ordering rests on two participants per condition. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the unit of creative potential (UCP): a word or group of words in the source text that cannot be translated routinely and demands problem-solving creativity. Each UCP in the Dutch translations is annotated as a creative shift (CS, deviation from the original), a reproduction (no deviation), or an omission, and total fixation duration from an eye-tracker is the measure of cognitive load. The eye-tracking data are modelled with a generalized additive mixed-effects model, and retrospective think-aloud interviews with gaze replays provide the qualitative triangulation that links the measured attention to enjoyment and immersion.
What would settle it
A replication with at least twenty readers per condition that finds no systematic increase in total fixation duration on creative shifts, or finds the increase to be equally large in machine translation as in human translation, would refute the central claim. The result would be strengthened if the replication also showed error-heavy segments producing longer dwell times, contradicting the authors' skim-reading explanation.
Extended reading notes
Core claim
Readers show higher cognitive load on units of creative potential than on ordinary text, and the size of this effect depends on the translation modality: it is largest for human translation and smallest for machine translation, with post-editing in between. Errors, by contrast, show no significant effect on cognitive load. The retrospective think-aloud interviews lead the authors to hypothesise that the extra attention given to creative units reflects enjoyment and immersion rather than difficulty.
Load-bearing premise
The statistical model must estimate the modality-by-creativity interaction from only two readers per condition; if either reader in a group is atypical, the claimed ordering of human, post-edited, and machine translation will not hold in a larger sample.
Editorial extensions
If this is right
- If creative units reliably draw more fixation time, word-level creativity becomes a measurable dimension of translation reception.
- Human translation's creative advantage shows up as measurable reader attention, not just preference ratings, with post-editing situated between human translation and machine translation.
- Machine translation's flattening of creative solutions reduces the attention-grabbing potential of literary texts, even in places without overt errors.
- Error-based metrics alone may miss what distinguishes translation modalities for readers; creativity annotations add signal that error counts do not capture.
- The absence of an error effect, plausibly due to skim-reading of error-heavy machine translation, warns that whole-text error counts do not translate linearly into reading effort.
Reading between the lines
- If the attention-on-creativity effect replicates with more readers, publishers using machine translation for literary genres could use eye-tracking on UCPs as a pre-screening tool to decide which texts need human translators.
- The authors' skim-reading explanation for the null error effect is testable: compare fixation counts in the first versus second half of machine-translated texts, expecting a drop as errors accumulate.
- The same design applied to a language pair with a different machine-translation quality baseline, or to a longer text, would show whether the human-translation advantage is about creativity per se or about overall textual quality.
- Since creative shifts versus reproductions showed no cognitive-load difference, the load may be driven by UCP status (the source unit being problematic) rather than by the translator's degree of deviation; a word-frequency-matched control would clarify this.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a pilot eye-tracking study of Dutch readers reading the same short story in four conditions: machine translation (MT), post-editing (PE), human translation (HT), and the original source text (ST). The authors annotate units of creative potential (UCPs) in the source and translations and classify them as creative shifts (CS) or reproductions, and they also annotate translation errors. They measure cognitive load via total fixation duration and other eye-tracking variables, complementing this with questionnaire data and retrospective think-aloud (RTA) interviews from eight participants (two per modality). The central claims are that UCPs increase cognitive load, that this increase is strongest for HT and weakest for MT, and that errors have no significant effect on cognitive load. The paper presents a transparent GAM analysis, a separate analysis of zero-fixation segments, non-parametric tests for other dependent variables, and a qualitative RTA analysis.
Significance. If the main effect of UCPs on cognitive load were established, it would be a meaningful contribution to translation reception research, supporting the idea that word-level creative choices in translation measurably shape reader attention. The paper is also valuable as a methodological pilot: it openly shares code and data, it distinguishes main effects from interactions, and it triangulates quantitative eye-tracking with qualitative RTA data, which is rare and instructive. The novel interaction claim (HT > PE > MT) is the most interesting element but also the least supported by the evidence as analyzed. The study's transparency and its explicit framing as a pilot make it a reasonable springboard for further research, though the strength of the abstract's conclusions currently exceeds what the data can bear.
major comments (4)
- [Section 3.2 and Table 5] With only two participants per modality (n=8 total), the significant Modality-by-Creativity interaction terms in Table 5 (HT-vs-PE contrasts: CS p=0.042, Rep p=0.001) are statistically confounded with participant identity. Because participant is treated as a random effect with only two clusters per condition, the between-participant variance component cannot be reliably estimated, and the reported ordering HT > PE > MT may be produced by a single atypical reader per condition. The abstract's claim that the UCP effect is 'highest for HT and lowest for MT' is therefore not supported by the data as analyzed. The authors should either present per-participant interaction estimates (e.g., an appendix showing the relevant slopes separately for each of the eight participants), or explicitly re-word the ordering claim as a hypothesis generated by the pilot rather than as a demonstrated result.
- [Section 4.3, Theme 4 and Section 5, RQ4] The null effect of errors on TFD (Table 5, Errors Yes p=0.867) is presented as a finding in the abstract ('no effect of error was observed'), but the RTA data in Section 4.3 (Theme 4) indicate that MT participants engaged in skim-reading of error-dense passages, with participants 'giving up' and 'accepting that I wouldn't really get the thing' (Section 4.3; Section 5 RQ4). If reading strategy changes as a function of error density, the segment-level comparison of error vs. no-error units is not a clean test of the effect of errors on cognitive load: the expected increase in fixation duration may be cancelled by reduced effort in skimmed passages. The authors should either model reading strategy (e.g., separately analyzing first-pass reading and re-reading on error-dense vs. error-free passages) or explicitly qualify the null result as uninterpretable given the identified masking mechanism, rather than naming it as a finding in the abstract.
- [Section 3.7 and Appendix C.3.1] The GAM analysis excludes zero-fixation segments, yet Table 17 shows that CS segments have a higher zero rate (5%) than Reproductions (1.6%) and non-UCP units (2.6%). Excluding these skipped segments from the TFD model could bias the estimated effect of Creativity, because a segment that is skipped is recorded as zero TFD and is not merely missing at random. The separate chi-square analyses of zero rates were not significant, but the authors should still justify or model the zero-inflation (e.g., a zero-inflated model or a two-part analysis) or, at minimum, report whether the significance of the CS and Rep main effects (Table 5) survives a sensitivity analysis that treats skipped segments as zero-valued observations.
- [Section 3.7 and Table 5 (post-hoc model selection)] The authors state that the model was fitted 'after analysing the data' (Section 3.7), which indicates a post-hoc selection of the GAM specification and of the included predictors. While this is disclosed transparently, the multiple tests in Table 5 (and the additional tests in Appendix C.4) are not corrected for multiplicity, and several interaction p-values are in the 0.01-0.05 range that would not survive even mild correction. The manuscript should at least acknowledge this in the limitations and should avoid placing heavy interpretive weight on the least robust interaction terms.
minor comments (6)
- [Section 3.7] The text reports '28% (R2 = 0.241)' for FPT; 0.241 corresponds to 24.1%, not 28%, and the parenthetical should be corrected.
- [Section 4.2.3] The sentence 'Spearman's correlation shows again a low negative correlation (ρ = 0.23)' omits the negative sign; it should read ρ = -0.23 for consistency with the TFD and RP values.
- [Table 16] The three-way interaction HT:Rep:Error reaches significance (p=0.018) but is not discussed anywhere in the text; please comment on it or explicitly justify why it is not interpreted.
- [Section 3.1] There is a typo in 'Kurt V onnegut'; it should be 'Kurt Vonnegut'.
- [References] In the Moorkens et al. (2018) entry, the co-author name 'Antonio Tora' appears to be missing the final 'l'; it should be 'Antonio Toral'.
- [Table 4] The labels 'UCP*' and 'Not*' are not self-explanatory; the table caption should indicate that these rows refer to the source-text (ST) segments that correspond to UCPs in the original and non-UCP segments of the ST, respectively.
Circularity Check
No circularity found: the eye-tracking measurements are independent of the annotation framework, and the reported effects are ordinary statistical inferences rather than predictions derived from the model inputs.
full rationale
The paper is an empirical pilot study, not a derivation. Its central claims are that units of creative potential (UCPs) increase cognitive load, that this effect differs by translation modality, and that errors show no significant effect. These claims are based on eye-tracking data (TFD and related measures) collected in the experiment, then analyzed with a GAM. The UCP/CS/Reproduction annotations are taken from the authors' prior framework (Guerberof-Arenas and Toral 2020, 2024), but the target outcome — fixation durations — is measured independently with an eye-tracker. Nothing in the annotation definitions forces a particular TFD result; UCPs are defined as source-text units that pose translation problems, not as units that cause readers to fixate longer. The GAM is fitted to the collected data and reports coefficient estimates and p-values; this is inferential statistics, not a 'prediction' in the sense of an out-of-sample claim, and it is not circular to describe patterns in the same data one has modeled. The RTA-based suggestion that higher cognitive load in UCPs relates to enjoyment and immersion is explicitly introduced as a hypothesis ('leads us to hypothesize'), not as a derived result. The paper also transparently acknowledges the small sample size ('we only had two participants per modality, which does not allow for generalization'), which is a statistical limitation, not circularity. The self-citation to Guerberof-Arenas and Toral provides the materials, annotations, and questionnaire, but the novel contribution — eye-tracking and RTA data — is not contained in those inputs. No load-bearing step reduces to its own inputs by definition, and no fitted parameter is renamed as a prediction. Therefore the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (5)
- domain assumption Total fixation duration (TFD) is a valid indicator of cognitive load during reading.
- domain assumption The UCP, Creative Shift, and Reproduction annotations from Guerberof-Arenas and Toral (2020, 2022) reliably operationalize translation creativity across modalities.
- domain assumption Two participants per condition are sufficient for the generalized additive mixed model to separate individual variability from modality effects.
- domain assumption Excluding zero-fixation segments from the GAM and analyzing them separately does not bias the cognitive-load estimates.
- domain assumption Normalizing eye-tracking measures by words per segment makes segments of different lengths comparable.
Cite this review
Pith. "Pith review of To MT or not to MT: An eye-tracking study on the reception by Dutch readers of different translation and creativity levels." pith.science (2026). https://pith.science/paper/YRNHXNVO
@misc{pith2026250419850,
author = {Pith},
title = {Pith review of: To MT or not to MT: An eye-tracking study on the reception by Dutch readers of different translation and creativity levels},
year = {2026},
howpublished = {\url{https://pith.science/paper/YRNHXNVO}},
note = {Machine review of arXiv:2504.19850}
}
read the original abstract
This article presents the results of a pilot study involving the reception of a fictional short story translated from English into Dutch under four conditions: machine translation (MT), post-editing (PE), human translation (HT) and original source text (ST). The aim is to understand how creativity and errors in different translation modalities affect readers, specifically regarding cognitive load. Eight participants filled in a questionnaire, read a story using an eye-tracker, and conducted a retrospective think-aloud (RTA) interview. The results show that units of creative potential (UCP) increase cognitive load and that this effect is highest for HT and lowest for MT; no effect of error was observed. Triangulating the data with RTAs leads us to hypothesize that the higher cognitive load in UCPs is linked to increases in reader enjoyment and immersion. The effect of translation creativity on cognitive load in different translation modalities at word-level is novel and opens up new avenues for further research. All the code and data are available at https://github.com/INCREC/Pilot_to_MT_or_not_to_MT
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
At times, I struggled to understand what was happening in the story
-
[2]
My understanding of the character is unclear
-
[3]
I had a hard time recognizing the thread of the story
-
[4]
My mind wandered while reading the text
-
[5]
While reading, I found myself thinking about other things
-
[6]
I had a hard time keeping my mind on the text
-
[7]
While reading, my body was in the room, but my mind was inside the world created by the story
-
[8]
The text created a new world, and then that world suddenly disappeared when the story ended
Show all 40 references
-
[9]
At times when reading, the story world was closer to me than the real world
-
[10]
During the story, I felt sad when a main char- acter suffered in some way
-
[11]
The story affected me emotionally
-
[12]
I felt sorry for some of the characters
-
[13]
While reading the story I had a clear image of what the main character looked like
-
[14]
While reading the story I could envision the situations described
-
[15]
For enjoyment:
I could imagine what the setting of the story looked like. For enjoyment:
-
[19]
Did you enjoy the text?
-
[20]
How likely is it that you would recommend the text to a friend?
-
[21]
Would you consider this text high literature? For translation reception:
-
[22]
The text was easy to understand
-
[23]
The text was well-written
-
[24]
I encountered words, sentences or paragraphs that were difficult to understand (including a box to write down which ones)
-
[25]
I encountered words, sentences or paragraphs that I found very beautiful (including a box to write down which ones)
-
[26]
I noticed I was reading a translation (including a box to indicate how people noticed)
-
[27]
What did you think of the translation?
-
[28]
Would you like to read a text by the same author and translator?
-
[29]
Would you like to read a text by the same author, but by a different translator?
-
[30]
These are good scores for reliability and shows the reliability of the scales
Would you like to read a text by a different author, but the same translator? In Guerberof-Arenas and Toral (2024)’s study, the Cronbach’s alpha reliability coefficient (α) was 0.85 for narrative engagement, 0.87 for enjoyment and 0.79 for translation reception. These are good...
2024
-
[31]
Confusion came from the narrative in HT, but from language use in MT
-
[32]
Engaging with and relating to narrative ele- ments occurred in HT, ST & PE
-
[33]
HT participants felt immersed in the story, the narrative, and the style
-
[34]
MT participants had difficulty understanding the text due to nonsensical words phrasing
-
[35]
PE participants were engaged in the narrative, but struggled with the style and characters at times
-
[36]
were just so weird
Confusion came from the narrative in HT, but from language use in MT One of the very noticeable things is that across modalities all participants mentioned feeling con- fused multiple times throughout the narrative: feel- ings of confusion were mentioned 30 times in HT, 37 tim...
-
[37]
I was also very curious to see what would happen next
Engaging with and relating to narrative elements occurred in HT, ST & PE HT, ST and PE participants mentioned feeling engaged in the narrative, including the events of the story and the moral issues at play, relating it to their own lives often. For these modalities, par- tici...
2004
-
[38]
ironic and witty
HT participants felt immersed in the story, the narrative, and the style Throughout the interviews, the HT participants made it clear that they liked the story in many of its facets and felt immersed in both the narrative and the style. P06_HT kept commenting about how im- mer...
-
[39]
[I] just didn’t really see what was happening here
MT participants had difficulty understand- ing the text due to nonsensical words phrasing Translation errors in MT led to nonsensical phrasings, which caused the participants to strug- gle understanding the narrative and its events. Par- ticipants mentioned that they were not ...
-
[40]
strong imagery
PE participants were engaged in the narra- tive, but struggled with the style and characters at times PE participants liked the story overall, thought it set-up the moral dilemma really well, and enjoyed themselves while reading the story. Both partici- pants related the situa...
2023
-
[325]
Sheila Castilho and Natália Resende
Routledge, New York City, US. Sheila Castilho and Natália Resende. 2022. MT-pese: Machine translation and post-editese. In Proceedings of the 23rd Annual Conference of the European As- sociation for Machine Translation, pages 305–306, Ghent, Belgium. European Association for M...
2022
-
[2020]
In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 3790–3798, Marseille, France
Literary machine translation under the mag- nifying glass: Assessing the quality of an NMT- translated detective novel on document level. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 3790–3798, Marseille, France. European Language Resources...
2023 arXiv
-
[2022]
good enough
Comparing the effect of product-based metrics on the translation process. Frontiers in Psychology, 12. Callum Walker. 2021. Eye-tracking study of equivalent effect in translation : the reader experience of literary style. Palgrave Macmillan. Rebecca Webster, Margot Fonteyne, A...
2024 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.