{"id":"68084ca8-5434-46cb-9adc-04505d7e7ea8","arxiv_id":"2504.18221","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A simple 'translate creatively' prompt at temperature 1.0 yields the most creative ChatGPT translations in Dutch, Spanish and Chinese, but all ChatGPT outputs remain less creative and more error-prone than human translations.","lead":"This study tests six ChatGPT configurations for translating a Kurt Vonnegut story into Dutch, Chinese, Catalan and Spanish, and finds that a simple prompt asking for creativity at temperature 1.0 produces the most creative outputs in three of the four languages. The result offers a practical, low-cost tip for using large language models in literary translation, while confirming that machine output still trails human translators.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claim that temperature 1.0 is optimal for the winning prompt is untested: the design never runs Prompt 3 at temperature 0.0, so the abstract's 'at the temperature of 1.0' fuses prompt and temperature effects without direct evidence.","rationale":"The reader correctly flags the single-annotator creativity index and in-sample selection as serious validity threats. My stress-test focuses on a narrower, more specific gap: the paper's central claim includes 'at the temperature of 1.0' as a property of the winning prompt, but the experimental design never varies temperature in the presence of that prompt. Phase 2 only establishes that, for Prompt 1, temperature 1.0 is better than 0.0 in ES/NL/ZH; Phase 3 then fixes the temperature and varies the prompt. This is an interaction effect that is asserted but not tested. A simple confirmatory run of Prompt 3 at temperature 0.0 would settle it. If the result shows no difference, the abstract should be corrected to delete the temperature claim or to phrase it as 'the temperature chosen in Phase 2.' If Prompt 3 at 0.0 is actually better, the finding changes. The single-draw stochasticity compounds this: at temperature 1.0, a single translation is one sample from a distribution, and selecting the best after the fact biases the comparison. The proposed test addresses both issues with repeated draws. I therefore recommend keeping the reader's CONDITIONAL verdict, since the paper is an honest exploratory case study with acknowledged limitations, but the conditions should include this confirmatory check.","tokens_in":15389,"tokens_out":6709,"duration_ms":62269,"concrete_test":"Run a small confirmatory experiment: for one language (e.g., Dutch), generate 5 independent ChatGPT outputs for each of four conditions—Prompt 1 and Prompt 3 at temperatures 0.0 and 1.0—and compute CI for each using the same annotation protocol, ideally with a second annotator blind to condition. Then compare the distributions (e.g., bootstrap or Mann-Whitney) of CI for Prompt 3 at 0.0 vs 1.0. If Prompt 3 at 0.0 is not significantly lower than at 1.0, the 'temperature of 1.0' component of the headline is not supported. This test also provides a replication check for the Prompt 3 advantage and a rough upper bound on annotation noise.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim asserts that Prompt 3 ('Translate the following text into [TG] creatively') at temperature 1.0 is the best configuration, but the experiment never tests that combination against Prompt 3 at another temperature. In Phase 2 (Section 3.3.2), temperature is varied only with Prompt 1; in Phase 3 (Section 3.3.3), prompts are varied at the single temperature already selected per language (1.0 for ES/NL/ZH, 0.0 for CA). Thus the 'at the temperature of 1.0' part of the claim is an untested interaction: Prompt 3 at 0.0 might be equally good or better, which would overturn the abstract's specific wording even if the creativity index were perfectly valid. The design also uses one stochastic ChatGPT output per configuration and selects the best configuration after inspecting results on the same 54 UCPs, so the ranking may reflect sampling luck and selection bias rather than a stable effect; the paper's own ANOVA found no significant main effect of modality on CSs (Section 5). The reader's concern about single-annotator scoring is legitimate, but the temperature confound is more directly load-bearing because it concerns what the experiment actually varied.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a case study evaluating ChatGPT (gpt-4o-2024-08-06) for literary translation from English into Dutch, Chinese, Catalan, and Spanish, using a short science-fiction story. Six ChatGPT configurations are compared, varying text granularity (paragraph vs. document), temperature (0.0 vs. 1.0), and zero-shot prompting strategy (minimal, genre/author-informed, direct creativity request), alongside DeepL (and for Catalan, Softcatalà and Google Translate) and human translations. Translations are manually annotated for creative shifts and errors, and a creativity index (CI) combines the two. The central claim is that the minimal prompt 'Translate the following text into [TG] creatively' at temperature 1.0 yields the most creative outputs and outperforms DeepL in Spanish, Dutch, and Chinese, although all ChatGPT outputs remain below human translation. The paper also reports automatic-metric correlations, an ANOVA on creative shifts and error points, and a sustainability statement.","tokens_in":15635,"tokens_out":5466,"duration_ms":57062,"significance":"If the central claim held, the paper would provide concrete, practical guidance for eliciting more creative literary translations from ChatGPT, with a cross-linguistic comparison (four target languages) and a clear negative result against human translation. The study is transparent: all code and data are released, the annotation procedure is described in detail, and the authors explicitly acknowledge the exploratory nature of the work. The finding that ChatGPT and other MT systems remain far below human translators in creativity, despite prompt engineering, is valuable and likely robust. However, the positive headline claim about the optimal prompt and temperature is not yet supported by the evidence as analyzed, because of the adaptive experimental design, the single-annotator evaluation without inter-annotator agreement, and the absence of inferential tests on the creativity index itself.","major_comments":[{"comment":"The abstract claims that the minimal creativity instruction 'at the temperature of 1.0' outperforms other configurations, but this exact combination was never directly tested. In Phase 2, temperature is varied only with Prompt 1; in Phase 3, prompts are varied at the single temperature already selected per language (1.0 for ES/NL/ZH, 0.0 for CA) using Prompt 1's results from Phase 2. Thus there is no comparison of Prompt 3 at temperature 0.0 versus 1.0, and the specific interaction between Prompt 3 and temperature is untested. This is load-bearing because the abstract's wording implies an optimal jointly tuned configuration, whereas the experiment only shows that, in this adaptive procedure, Prompt 3 happened to be the best among the prompts tested at the temperature inherited from Phase 2. The authors should either run Prompt 3 at both temperatures and report the comparison, or explicitly rephrase the claim as 'the best configuration among those tested was Prompt 3 at the temperature selected in Phase 2' without asserting an optimality for temperature 1.0.","section":"§3.3.2–§3.3.3 and Abstract"},{"comment":"The creativity index (CI) is computed from annotations by one single annotator per language, and no inter-annotator agreement is reported anywhere in the manuscript. This is a serious concern for the ranking of configurations, especially for languages where the CI differences between configurations are small. For example, in Table 4, ENZH Prompt 3 scores 1.03 versus Prompt 2 at -2.48, and in Table 3, ENZH T-1.0 scores -1.48 versus T-0.0 at -5.38. These gaps could easily be overturned by a different annotator's subjective judgments on a few units of creative potential or error severities. Since the CI is the sole criterion for the headline result, the absence of any reliability evidence makes the reported ranking fragile. The authors should provide at least a second annotation for a subset, report Cohen's kappa or a similar measure, and use annotation-based confidence intervals or a sensitivity analysis to show that the main conclusions are stable across plausible annotation noise.","section":"§3.4–§3.5"},{"comment":"The paper's own ANOVA undermines the central claim. The aligned-rank-transform ANOVA on the number of creative shifts (CSs) reports no significant main effect of Modality and no Modality×Language interaction, with only Language reaching significance. Creative shifts are the numerator of the CI and the operational definition of novelty, so this means the data do not demonstrate that prompting strategy affects the creativity component of the index. The significant Modality effect is found for Error Points, which is an acceptability component, not creativity per se. Moreover, the ANOVA treats individual sentences as independent observations even though each configuration is a single system output; the effective sample size is the number of configurations (7–8 per language), not the hundreds of sentences. The authors should either perform a permutation or bootstrap test on the CI differences across configurations, or frame the conclusion as 'differences in the combined index' without implying that the creative-shift component differs significantly across prompts.","section":"§5"},{"comment":"The adaptive experimental design selects the better granularity in Phase 1, then the better temperature in Phase 2, then the better prompt in Phase 3, all using the same 54 UCPs and 48 sentences for the final evaluation. This means the reported 'best configuration' is the result of maximizing the CI on the evaluation set itself, not on a held-out or confirmatory sample. When combined with the fact that each configuration is run once (a single stochastic output at temperature 1.0), the selection process can capitalize on sampling luck. The paper implicitly acknowledges variability in §6 ('a level of randomization in the output that is quite unpredictable'), but it does not address the statistical consequences for the ranking. The authors should either run multiple repetitions per configuration and report variance, or perform a small confirmatory study on a separate set of sentences, or explicitly label the result as an exploratory, within-sample optimum rather than a validated finding.","section":"§3.3 and §4"}],"minor_comments":[{"comment":"There are inconsistent renderings of the model name: 'Chat-GPT' in the abstract and 'ChatGPT' elsewhere; please standardize.","section":"Throughout"},{"comment":"The author name 'Vonnegut' is typeset as 'V onnegut' in multiple places, including the reference entry for the primary source text; correct these typos.","section":"§3.1, §3.2, References"},{"comment":"The text says that the 185 UCPs were annotated by 'two experienced translators and researchers' in the prior study, while the current study uses one annotator per language. Please make this distinction explicit to avoid confusion about the source of the UCP list versus the current annotation of the translations.","section":"§3.2"},{"comment":"The CI formula uses fixed severity weights (Minor=1, Major=5, Critical=15) but no sensitivity analysis is provided. Given the small CI gaps for ZH, a brief sensitivity check (e.g., reweighting or excluding Critical errors) would strengthen the robustness claims.","section":"§3.5"},{"comment":"The ANOVA results do not include effect sizes (e.g., partial eta-squared) or any measure of uncertainty for the pairwise comparisons beyond p-values; adding these would help the reader gauge the magnitude of the reported differences.","section":"§5"},{"comment":"The labels 'ENCA-S', 'ENCA-G', and '3d' in Table 10 are not fully defined in the captions; please explain the abbreviations and the meaning of the additional columns in the captions or in the main text.","section":"Table 5 and Table 10"},{"comment":"The sustainability statement contains a typo, 'GhatGPT' for 'ChatGPT', and the sentence 'To the best of our knowledge, the average CO2 emissions of GhatGPT models is not disclosed' has a number-agreement error; clean this up.","section":"Sustainability statement"}],"recommendation":"major_revision","confidential_remarks":"The paper relies on a creativity index developed by the authors in prior work, and the UCP annotations also come from that prior study. This is not in itself a reason for rejection, but it amplifies the need for independent reliability evidence, which is currently missing. The paper would be a good fit for a translation-studies or NLP venue if the authors either provide the missing validation or substantially weaken the headline claim. The temperature confound and the lack of statistical support for the CI ranking are the most load-bearing issues; they should be addressed head-on."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a useful empirical case study, with a real but narrower finding than the abstract claims. Confirming Du (2024)'s Chinese result, a minimal prompt \"Translate creatively\" produced the highest creativity score in Dutch and Spanish as well, and again in Chinese, while Catalan behaved differently—temperature 0.0 and no special prompt did best there. That multilingual pattern is new and worth having.\n\nThe paper does many things right. It ships code, data, and detailed per-language annotations; it compares against DeepL, Google, Softcatalà, and professional human translations; it runs automatic metrics and an ANOVA; and it is explicit about its exploratory nature. The conclusion that ChatGPT, even in its best configuration, falls well short of human translation is credible and consistent with prior work.\n\nThe soft spots are real and in part load-bearing. The headline claim couples prompt and temperature, but the design never tests Prompt 3 at temperature 0.0: temperature was selected in Phase 2 with Prompt 1 only, and Phase 3 then fixed temperature per language. So \"at the temperature of 1.0\" is an untested interaction. The stress-test note is correct on that point. Second, each configuration is a single stochastic ChatGPT run, evaluated by one annotator per language using a creativity index built by the authors themselves; there is no inter-annotator agreement and no uncertainty quantification. Third, the pipeline chooses granularity and then temperature by looking at results on the same 54 UCPs, which is in-sample optimization; the ANOVA actually finds no significant effect of modality on creative shifts, only on error points, so the ranking is fragile.\n\nNone of this makes the paper worthless. It is a transparent, reproducible case study with a plausible practical takeaway: a simple creativity instruction helps, and the gains appear mostly through fewer errors and more shifts, but the exact optimal temperature is not established. The limitations are acknowledged in the conclusions, though the abstract overstates the finding.\n\nI would send this to peer review rather than desk reject. A serious referee can ask for the missing Prompt 3×temperature cell, repeated generations, multi-annotator scoring, and a more cautious abstract. For readers working on literary MT or LLM prompting, this is a useful data point and a good example of how adaptive evaluation designs can overclaim.\n\nI would not hang my own conclusions on the temperature claim, but I'd cite it as evidence that minimal creative prompts help in multiple languages.","headline":"Useful multilingual replication of the 'creatively' prompt effect, but the abstract's temperature claim is untested and the evaluation rests on a single annotator.","tokens_in":16179,"tokens_out":3635,"would_cite":true,"duration_ms":36469,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Prompting ChatGPT with 'translate creatively' at temperature 1.0 yields the most creative literary outputs among tested settings and beats DeepL in three languages.","keywords":["literary translation","ChatGPT","machine translation","translation creativity","prompt engineering","temperature sampling","creativity index","large language models"],"falsifier":"Re-annotate the same ChatGPT and DeepL outputs with several independent annotators per language and check whether 'Translate the following text into [TG] creatively' at temperature 1.0 still yields the highest creativity index; as a stricter test, recompute the index on a larger, independently sampled set of creative-potential units from the same story and see whether the same configuration wins.","tokens_in":15188,"feed_emoji":"📖","tokens_out":8754,"duration_ms":75451,"temperature":0.7,"pith_summary":"This paper asks whether ChatGPT can be coaxed into producing more creative literary translations, and which of six configurations best does so. Using a short science fiction story translated from English into Dutch, Chinese, Catalan, and Spanish, it varies text granularity, temperature (0.0 vs 1.0), and prompting strategy, then scores every output with a creativity index that rewards novel departures from the source and penalizes errors. The headline finding is that the simplest instruction, 'Translate the following text into [TG] creatively,' at temperature 1.0, yields the most creative outputs in Spanish, Dutch, and Chinese, and it outperforms DeepL in those languages. The paper also reports that every ChatGPT configuration produces far fewer creative shifts and many more errors than professional human translations, so the practical takeaway is limited but real: an explicit one-word request for creativity measurably shifts the model's output, while richer genre and author context does not.","feed_headline":"Add one word to the prompt and ChatGPT beats DeepL","feed_subtitle":"Prompting 'translate creatively' at temperature 1.0 wins in Spanish, Dutch, and Chinese.","key_machinery":"The load-bearing instrument is the creativity index (CI), computed as $$\\text{CI} = \\left(\\frac{\\#\\text{CSs}}{\\#\\text{UCPs}} - \\frac{\\text{error points}}{\\#\\text{words in ST}}\\right) \\times 100.$$ Creative shifts (CSs) are annotated on 54 pre-selected units of creative potential from the source text, with each solution classified as abstraction, concretization, or modification, following Bayer-Hohenwarter's taxonomy. Error points come from a DQF-MQM-style severity scale (neutral, minor, major, critical), and the formula converts two qualitative judgments—novelty and acceptability—into a single number that ranks every configuration. The paper also uses automatic metrics (BLEU, chrF, TER, COMET, COMET-Kiwi), but the creativity index is what carries the central claim.","core_discovery":"On the paper's own terms, the discovery is that minimalism in prompting beats informativeness for eliciting creativity from ChatGPT. The configuration 'Translate the following text into [TG] creatively' at temperature 1.0 achieves the highest creativity index among all tested ChatGPT settings in English-to-Spanish, English-to-Dutch, and English-to-Chinese, and it also outscores DeepL in those three directions. In Catalan, the best setting is the even plainer 'Translate the following text into [TG]' at document level, because the creativity prompt and the genre prompt leave the story's invented nicknames untranslated, hurting the score. Temperature 1.0 generally increases both creative shifts and errors relative to 0.0, but the net index favors 1.0 in three languages. Across all languages and settings, the creativity index of the best ChatGPT output remains substantially below the professional human translations used as reference, and this gap is the paper's concluding caution about the model's creative ceiling.","pith_inferences":["If the one-word 'creatively' effect is robust, it may be a cheap, generalizable steering signal for other large language models and literary genres, without any fine-tuning.","Because the Catalan exception hinges on a handful of untranslated nicknames, the index is sensitive to a small number of units; a different weighting or a different set of units could plausibly reorder the configurations.","A testable extension is whether the same minimal-prompt advantage holds for non-literary but stylistically marked text (e.g., marketing copy, subtitles), where context prompts might matter more.","The large gap between machine and human creativity, even under the best prompt, offers a concrete benchmark for measuring future progress in generative translation systems."],"forward_implications":["A direct, short request for creativity is a more effective lever than supplying genre and author information when the goal is creative literary output in ChatGPT.","Raising temperature from 0.0 to 1.0 is net-positive for creativity in Spanish, Dutch, and Chinese, even though it adds errors, so users optimizing for creativity should not default to the lower setting.","The best ChatGPT output still trails professional human translations on the same index, so the model cannot currently replace a literary translator.","The optimal granularity is language-dependent: paragraph-level wins for Dutch and Chinese, document-level for Catalan and Spanish, meaning no single 'best context' setting generalizes.","Standard automatic metrics do not track the human creativity ranking, so creativity evaluation continues to require manual annotation."],"supporting_citations":[{"why":"Defines novelty and acceptability (skopos adequacy), the two conceptual pillars of the creativity criteria.","marker":"Bayer-Hohenwarter (2009)"},{"why":"Supplies the taxonomy of creative shifts (abstraction, modification, concretization) used to classify novel translation solutions.","marker":"Bayer-Hohenwarter (2011)"},{"why":"Introduces the creativity index formula that combines creative shifts with error points, the paper's core evaluation metric.","marker":"Guerberof-Arenas and Toral (2020)"},{"why":"Provides the pre-annotated units of creative potential in the source text and the professional human translations used as the reference baseline.","marker":"Guerberof-Arenas and Toral (2022)"},{"why":"Pilot study establishing the promising configuration (document level, temperature 1.0, explicit creativity prompt) that this paper replicates across more languages.","marker":"Du (2024)"},{"why":"Defines the MQM error severity levels used to weight errors in the creativity index.","marker":"Lommel et al. (2014)"}],"fun_headline_variants":["Minimal prompt 'translate creatively' beats DeepL in three languages","Creative prompt at temp 1.0 outscores DeepL in Spanish, Dutch, Chinese","One word 'creatively' makes ChatGPT beat DeepL","Minimal prompt beats DeepL for literary creativity","ChatGPT's creative edge: one word, three languages, beats DeepL"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking of configurations rests on a creativity index computed by one annotator per language on just 54 units of creative potential, so the results stand or fall on whether that annotation and unit selection faithfully measure translational creativity.","fun_headline_variants_meta":{"raw":{"variants":["Minimal prompt 'translate creatively' beats DeepL in three languages","Creative prompt at temp 1.0 outscores DeepL in Spanish, Dutch, Chinese","One word 'creatively' makes ChatGPT beat DeepL","Minimal prompt beats DeepL for literary creativity","ChatGPT's creative edge: one word, three languages, beats DeepL"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00127,"raw_usage":{"total_tokens":5146,"prompt_tokens":841,"completion_tokens":4305,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":457,"completion_tokens_details":{"reasoning_tokens":4213}},"tokens_in":457,"tokens_out":4305,"duration_ms":28558,"temperature":1.0,"reasoning_tokens":4213,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:21:34.354948+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate the same ChatGPT and DeepL outputs with several independent annotators per language and check whether 'Translate the following text into [TG] creatively' at temperature 1.0 still yields the highest creativity index; as a stricter test, recompute the index on a larger, independently sampled set of creative-potential units from the same story and see whether the same configuration wins.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines novelty and acceptability (skopos adequacy), the two conceptual pillars of the creativity criteria."},{"cited_title":"Creative Shifts","cited_arxiv_id":null,"evidence_quote":"Supplies the taxonomy of creative shifts (abstraction, modification, concretization) used to classify novel translation solutions."}],"review_version":1}