{"id":"aa608886-cc28-493e-a47f-9420c1e40bdc","arxiv_id":"1908.06140","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a small user study with three professional translators, CATaLog Online users selected machine translation output about 80% of the time, while rating the tool's color-coded TM highlighting positively.","lead":"The paper describes and evaluates CATaLog Online, a web-based translation tool, and reports that professional translators chose machine translation suggestions for about 80% of sentences in a small study. The tool's color-coded highlighting of matching fragments was rated positively, although post-editing took longer in CATaLog Online than in MateCat.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post-hoc removal of the near-duplicate APE option makes the 80% 'MT system' selection rate uninterpretable as an MT-specific preference.","rationale":"The reader's verdict of CONDITIONAL is appropriate and already captures the load-bearing concern: the post-hoc exclusion of APE directly undermines the interpretation of the central quantitative claim. The paper transparently discloses this design change in Section 4, so the issue is not hidden, but the disclosed change affects the strength of the conclusion. The 80% figure is accurate for the modified condition, yet the wording overstates it as an MT-specific preference. A replication with all four options would settle the question. A secondary issue is that the kappa interpretation in Section 4 appears internally reversed: the text says 'total number of edits (with a low κ)' and 'post-editing time (with a high κ),' while Table 3 shows the opposite (edits κ = 0.49/0.42/0.26; time κ = -0.16/-0.06/-0.13). This inconsistency affects the agreement analysis rather than the selection-rate claim and is not the primary basis for the verdict. The paper's detailed logging and explicit reporting of limitations are strengths, but the central claim needs the proposed test to be fully supported.","tokens_in":7339,"tokens_out":5398,"duration_ms":51866,"concrete_test":"Run a follow-up experiment with the same three professional translators (or a matched group) on the same 200-sentence set using all four options (TM, MT, APE, None), counterbalanced across sentences, and compute the MT-only selection rate. If MT-only remains ≈80% while APE is rarely chosen, the claim survives. If MT-only falls to ≈40% with APE chosen ≈40%, the reported 'MT system' rate is an artifact of removing APE.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 reports that the original three-option experiment (TM, MT, APE) with students was abandoned because 'the students' decision for the MT or APE system is based on chance, since the MT output and the output from the APE system are very similar to each other.' The professional study then offered only MT, TM, and None. The headline 'MT system achieves a selection rate of around 80%' therefore measures selection of the MT engine when no near-identical APE alternative is present. Since APE was removed precisely because it is nearly indistinguishable from MT, a substantial share of the 80% may represent 'machine-generated output' (MT or APE) rather than MT specifically. The paper provides no professional-translator data with all four options, so the MT-only rate under the original design is unknown. This matters because the abstract and conclusion credit 'the MT system,' not 'MT or APE.' The claim is internally consistent with Table 2 but is not robust to the design modification; the small sample (n=3) amplifies the uncertainty.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents CATaLog Online, a web-based CAT tool offering translation suggestions from translation memories (TM), machine translation (MT), and automatic post-editing (APE), together with color-coded intra-segment feedback and a Lucene-based retrieval backend. The authors report two user studies: a comparison of post-editing time between MateCat and CATaLog Online with 16 translation students, and a choice experiment in which three professional translators selected between MT, TM, and translating from scratch over 200 English-to-German news sentences. The central empirical claim is that the MT system achieves a selection rate of around 80% and that translators show a clear preference for MT output as the starting point for their translations. The paper also reports inter-rater agreement statistics, edit-type distributions, and qualitative user feedback on the tool's usability.","tokens_in":7516,"tokens_out":3543,"duration_ms":36476,"significance":"If the central claim were robust, the result would be a useful data point on translator behavior in CAT environments, namely that professional translators, when given MT, TM, and scratch options, predominantly choose MT. The paper's strength is its concrete system description: CATaLog Online is a real, web-based, open-source tool with logging capabilities and a color-coding scheme that is a plausible usability contribution. The authors also make an honest attempt to measure inter-rater agreement with Cohen's kappa, and they disclose negative feedback and limitations. However, the headline empirical claim rests on only three professional translators after the APE condition was removed post hoc, and no significance tests or confidence intervals are provided. The study therefore currently supports a much weaker statement: in a three-option design without APE, the three participating translators chose MT for roughly 80% of the sentences. The paper's significance as a system description is real, but its evaluative conclusions need substantial qualification.","major_comments":[{"comment":"The headline claim that the MT system achieves a selection rate of around 80% is based on data from only three professional translators (T1, T2, T3), and the paper reports no confidence intervals or significance tests. With n=3, the authors' own Table 3 shows that inter-rater agreement on suggestion selection is very low (pairwise Cohen's kappa between 0.05 and 0.20), which undercuts the statement in Sections 4 and 5 that 'translators have a clear preference in choosing the output of the MT system.' The claim should be reworded as a descriptive observation about these three translators, not a generalizable preference.","section":"Section 4, Table 2"},{"comment":"The APE condition was removed because, in the student experiment, 'the MT output and the output from the APE system are very similar to each other.' This means the reported 80% MT selection rate was measured in a design where no near-identical alternative was available. The data cannot distinguish a preference for MT specifically from a preference for machine-generated output, since a large share of the 80% might have gone to APE had it been offered. The abstract and conclusion credit 'the MT system' without acknowledging this design modification; the authors should either reframe the claim as 'MT output, in the absence of a near-duplicate APE option,' or provide professional-translator data from a four-option condition.","section":"Section 4, APE exclusion"},{"comment":"The first user study shows that CATaLog Online was slower than MateCat for all ten common sentences (for example, Stud1 took 1112 seconds in MateCat versus Stud9's 3079 seconds in CATaLog Online), and the paper does not report any statistical analysis of this difference. The paper's positive conclusions about user preference and the tool's usefulness do not directly address this performance gap. If the paper claims to 'improve CAT tools,' it should either temper the productivity-related claims or provide an analysis of the time difference, for instance by showing that the extra time is accompanied by better quality or a deliberate interface trade-off.","section":"Table 1, MateCat comparison"},{"comment":"The agreement analysis reports Cohen's kappa and Pearson's rho without confidence intervals, and with only three raters the kappa estimates in Table 3 are highly unstable. The sentence on Pearson's rho is also confusing: the text says the authors tested whether the total number of edits (with a low kappa) influences post-editing time (with a high kappa), but the kappa values in Table 3 do not cleanly separate these variables. This analysis should be either clarified with appropriate uncertainty measures or omitted from the main claims.","section":"Section 4, Table 3 and correlation"}],"minor_comments":[{"comment":"There is a typo: 'CATaLog Onlineof-fers' should be 'CATaLog Online offers.'","section":"Section 3, first paragraph"},{"comment":"The phrase 'enhanced with with a new interface layout' contains a duplicated 'with.'","section":"Section 1, CATaLog description"},{"comment":"The caption reads '200 sentenced' and should read '200 sentences.'","section":"Table 2 caption"},{"comment":"The user feedback section is anecdotal and does not describe a systematic procedure for collecting or coding the positive and negative impressions; adding a short method note would help readers interpret these results.","section":"Section 4, user feedback"},{"comment":"The third limitation bullet is acknowledged but not connected to the main 'clear preference' claim; the authors should state explicitly that the current experiment cannot assess whether choosing MT output improves final translation quality, coherence, or cohesion.","section":"Section 4.2, limitations"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a venue interested in translation technology and NLP evaluation. The main issue is that the central empirical claim is not robust to the small sample and the post-hoc removal of the APE condition, but this can be addressed by substantially qualifying the conclusions and making the limitations more prominent. The authors' prior work supplies the system components, but the evaluation in this manuscript would be much stronger if it included confidence intervals or a more careful framing of the 80% statistic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a modest systems contribution: it describes improvements to CATaLog Online, an existing web CAT tool, and reports two small user studies. The genuinely new result is the second study, where three professional translators chose the MT output about 80% of the time, with very low inter-translator agreement. That specific number isn't in the authors' prior work, and it is consistent with earlier findings that translators tend to prefer MT as a starting point.\n\nThe paper deserves credit for being transparent about its own design change. When the student pilot showed that MT and APE outputs were too similar to choose between reliably, the authors dropped APE and reran the study with professionals. That is honest reporting. The color-coding scheme for TM matches is a sensible idea, and the positive feedback from the three translators is a reasonable signal. The detailed XML logs are also useful for anyone doing translation process research.\n\nThe soft spots are the ones the reader and stress-test highlight. Removing APE because it is nearly indistinguishable from MT means the 80% figure measures MT only because the near-identical alternative is absent. A substantial share of those 80% may represent \"machine output\" generally rather than MT specifically. The paper says this clearly, but it does not solve the problem. With n=3 and no significance tests or confidence intervals, the word \"clear\" is doing more work than the data can support. Also, the kappa section contains a small but confusing inconsistency: the text says number of edits has low κ and editing time has high κ, while Table 3 shows the opposite (editing time near zero or negative, number of edits around 0.3–0.5). That should be fixed.\n\nProportionally, these are weaknesses, not fatal flaws. The central observation is internally consistent with Table 2, and the paper does not oversell beyond its evidence. It lists limitations and frames the findings as preliminary.\n\nWho is this for? Researchers and tool-builders working on CAT interfaces, TM/MT integration, and translation process effort. A reader in that space gets a small, honest data point plus the color-coding idea.\n\nRecommendation: send it to peer review. It's not a strong paper, but it is a legitimate empirical note with a real result and a clearly documented method. A serious referee should see it, and with revisions the kappa mix-up and the MT-vs-APE interpretation issue are fixable.","headline":"A small, clearly reported CAT tool study whose headline 80% MT selection rate rests on three translators and a post-hoc design change; still worth a look for the color-coding and log data.","tokens_in":8049,"tokens_out":2274,"would_cite":false,"duration_ms":22712,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Machine translation output becomes the starting point in about 80% of translator choices.","keywords":["computer-aided translation","translation memory","machine translation","post-editing","automatic post-editing","color coding","CAT tool evaluation","English-German translation"],"falsifier":"Run the same English-to-German task in a CAT tool with APE included, with MT, TM, and from-scratch options presented in randomized order, and with more than three professional translators; if the MT selection share falls well below 80% or splits evenly across options, the claimed clear MT preference is not supported.","tokens_in":7157,"feed_emoji":"🖥️","tokens_out":5917,"duration_ms":54840,"temperature":0.7,"pith_summary":"This paper claims that a web-based CAT tool can be improved by changing how translation suggestions are retrieved and presented, and that with these changes professional translators overwhelmingly take machine-translation output as their starting point. The authors report that in an English-to-German user study with three professional translators, the MT suggestion was selected for about 80% of segments, with translation-memory suggestions and translation from scratch making up the rest. They argue this shows a clear preference for MT as the basis of post-editing, even though the three translators did not agree with each other on which segments to choose MT. They also report that translators valued a color-coding scheme that highlights matching and mismatched parts of TM suggestions, helping them decide what to edit.","feed_headline":"MT output is the starting point for ~80% of translations","feed_subtitle":"A CAT-tool study finds professional translators choose machine translation over memory and scratch far more often.","key_machinery":"The central object is CATaLog Online, a web-based CAT tool whose suggestion pipeline combines three engines: a TM retriever that scores matches with a TER-based alignment and Needleman-Wunsch-style match/reward edit scoring, an integrated statistical MT system, and an automatic post-editing (APE) system built on an operation sequence model. Lucene/Nutch retrieval is used to keep TM search fast on large memories. The mechanism that carries the empirical claim is the suggestion-selection design: the tool presents MT, TM, APE, and from-scratch options side by side, logs which engine was chosen, and color-codes TM suggestions so matched text is green and mismatched text red. In the final experiment, APE was removed because its output was too similar to MT, leaving translators to choose between MT, TM, and scratch; the logged choices are the direct evidence for the roughly 80% MT selection rate.","core_discovery":"The paper's central discovery is the observed selection behavior in the CATaLog Online environment: when professional translators are offered MT output, a TM suggestion, or translation from scratch, they chose MT in roughly 160 of 200 sentences (about 80%), and this held over both the full set and the 100-sentence common set. The authors interpret this as evidence that MT should be treated as the primary post-editing starting point in CAT tools, with TM and scratch translation as fallbacks. They also found that translators did not reliably agree on which segments should use MT, and that editing time was translator-dependent and only weakly correlated with edit count. A supporting finding is that the interface's color coding of matched and mismatched words in TM suggestions was positively received, and was credited with making TM suggestion selection easier.","pith_inferences":["The reported ~80% selection rate was measured after APE was removed and with only three translators; if the experiment were repeated with APE included and suggestion order randomized, the MT share might be lower, since translators were forced to choose between MT and a very similar APE output in the earlier student study.","The dominance of MT may reflect the quality of the integrated MT system and the news-domain data; for domains where TM consistency matters more than fluency, translators might favor TM suggestions, so the preference should not be generalized without domain variation.","A testable extension would be an ablation of the color-coding feature: give translators the same MT and TM options with and without highlighting, and measure whether selection and post-editing speed change.","Because translators did not agree with each other on which segments should use MT, a per-segment MT-quality estimator could help by pre-selecting MT suggestions only where they are likely to be accepted, shifting from a global MT default to a segment-level recommendation."],"forward_implications":["If MT is the preferred starting point, CAT-tool interfaces should lead with MT output and treat TM as a secondary suggestion rather than the default.","Color-coding the matched and mismatched parts of TM suggestions gives translators a quick visual signal of which parts need editing, which can reduce the effort of deciding among suggestions.","Because editing time varies by translator and correlates weakly with the number of edits, post-editing time alone is not a reliable productivity metric in this setting.","APE output that closely resembles MT output adds little distinguishable value to the translator's choice set, supporting the decision to leave it out of the final evaluation."],"supporting_citations":[{"why":"Describes CATaLog Online as the web-based CAT tool that all experiments in this paper run in.","marker":"Pal et al. (2016a)"},{"why":"Defines the CATaLog TM retrieval and post-editing interface whose similarity metric this paper extends.","marker":"Nayek et al. (2015)"},{"why":"Supplies the integrated MT system whose output the translators selected in the central experiment.","marker":"Pal et al. (2015a)"},{"why":"Provides the APE system integrated into CATaLog Online and later excluded from the professional-translator experiment.","marker":"Pal et al. (2015b)"},{"why":"Describes MateCat, the CAT tool used as the baseline in the student post-editing comparison.","marker":"Federico et al. (2014)"},{"why":"Provides the WMT shared task context and data that motivate the APE component and evaluation.","marker":"Bojar et al. (2016)"},{"why":"Supplies Cohen's kappa, the agreement measure used for the three translators' selections, edit time, and edit counts.","marker":"Cohen (1960)"},{"why":"Supports the paper's caution that post-editing time is an imperfect measure of effort.","marker":"Koponen (2016)"}],"fun_headline_variants":["MT chosen as start for 80% of translations in CAT tool study","Translators pick MT over memory in 80% of segments","CAT tool study: MT becomes primary post-editing start","MT output wins as starting point in 80% of translations","Study: Translators favor MT start, not memory or scratch"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the three professional translators, tested after the APE option was removed because its output resembled MT too closely, give a representative picture of translator preference; if a small, modified study misrepresents that preference, the roughly 80% MT selection claim weakens.","fun_headline_variants_meta":{"raw":{"variants":["MT chosen as start for 80% of translations in CAT tool study","Translators pick MT over memory in 80% of segments","CAT tool study: MT becomes primary post-editing start","MT output wins as starting point in 80% of translations","Study: Translators favor MT start, not memory or scratch"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000142,"raw_usage":{"total_tokens":1105,"prompt_tokens":817,"completion_tokens":288,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":433,"completion_tokens_details":{"reasoning_tokens":201}},"tokens_in":433,"tokens_out":288,"duration_ms":3566,"temperature":1.0,"reasoning_tokens":201,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:54:04.300167+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same English-to-German task in a CAT tool with APE included, with MT, TM, and from-scratch options presented in randomized order, and with more than three professional translators; if the MT selection share falls well below 80% or splits evenly across options, the claimed clear MT preference is not supported.","supporting_citations":[{"cited_title":"K., Pal, S., Zampieri, M., Vela, M., and van Genabith, J","cited_arxiv_id":null,"evidence_quote":"Defines the CATaLog TM retrieval and post-editing interface whose similarity metric this paper extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes MateCat, the CAT tool used as the baseline in the student post-editing comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the WMT shared task context and data that motivate the APE component and evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Cohen's kappa, the agreement measure used for the three translators' selections, edit time, and edit counts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the paper's caution that post-editing time is an imperfect measure of effort."}],"review_version":1}