{"id":"e3e25e2f-f688-4da5-9ff3-30b023dfb08a","arxiv_id":"2412.09993","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new resource of 2,200 Persian idioms and two 200-sentence benchmarks show Claude-3.5-Sonnet leads idiom translation accuracy for Persian-English, with hybrid LLM-plus-NMT setups aiding weaker models in English-to-Persian.","lead":"This paper introduces Persian and English datasets of sentences containing idioms and compares how well different AI translation systems, from Google Translate to large language models, translate them. It finds that Claude-3.5-Sonnet is the strongest idiom translator in both directions, and that pairing weaker models with Google Translate helps for English-to-Persian.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The GPT-4o judge is validated only on GPT-3.5 and Google Translate outputs, yet it is used to rank Claude and all other models; if judge accuracy does not transfer to Claude's paraphrastic idiom translations, the headline ranking collapses.","rationale":"The reader's verdict is already CONDITIONAL, and the weakest assumption identified by the reader is exactly the gap I find most load-bearing: the GPT-4o judge is validated on only seven outputs from a narrow set of models and then used to rank all systems, including Claude, which is the centerpiece of the paper's strongest claim. My reading of Sections 4.3, 5.2, and 5.3 confirms this concern, and the paper's own admission that GPT-4o underlabels paraphrastic idiom rewrites makes the threat concrete for Claude's best-performing prompt settings. The concern is not that GPT-4o is intrinsically unreliable; it is that reliability established on GPT-3.5 and Google Translate outputs may not transfer to Claude's more fluent, idiomatic, and potentially more paraphrase-heavy outputs. A targeted manual evaluation of Claude's outputs against the judge would settle whether the ranking holds. I fully credit the paper's new Persian idiom resource, the carefully constructed 200-sentence benchmarks, and the honest reporting of inter-annotator agreement and correlation numbers; those are real contributions. But the headline ranking rests on an unvalidated transfer of judge reliability, so the conditional verdict is appropriate. I do not see a reason to move the verdict further: the concern is testable, and the paper's released data make the test feasible. I also note that the Limitations section acknowledges the small dataset size but not the judge-validation transfer gap; flagging that gap explicitly in a revision would strengthen the paper.","tokens_in":16327,"tokens_out":3721,"duration_ms":40570,"concrete_test":"Using the released benchmark and outputs (or regenerating them with the paper's prompts), manually score idiom translation with the same MQM-based binary protocol and at least two annotators on a targeted subset: Claude-3.5-Sonnet under its best prompt (CoT for En→Fa, MultiPrompt for Fa→En), the closest competitor in each direction (GPT-4o-mini for En→Fa, Command R+ or GPT-4o-mini for Fa→En), and one NMT baseline, for at least 100 of the 200 sentences per output. Compare human idiom-accuracy labels against GPT-4o labels separately for each model. If human-GPT4o agreement on Claude outputs is materially lower than the 0.76/0.71 reported in Table 5, or if Claude's rank falls behind another model under human labels, the judge-based ranking is not transportable and the paper's central claim must be re-evaluated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that Claude-3.5-Sonnet achieves the best idiom translation in both directions (Table 6, GPT-4o column). That claim depends on GPT-4o-as-judge labels being reliable for every system being ranked. It is currently calibrated only against manual labels on seven model outputs, all from GPT-3.5, Google Translate, or their combination, and only on the first 100 of 200 sentences (Sections 5.2 and 5.3, Table 4). The paper itself reports that GPT-4o tends to 'slightly underestimate model performance' and, for En→Fa, assigns a score of 1 only to translations closely resembling the gold standard (Section 4.3). Claude's highest En→Fa scores come from CoT and MultiPrompt settings, where the model is explicitly encouraged to replace idioms with natural Persian expressions; those outputs are the most likely to contain flexible paraphrases that a gold-reference-fixed judge would underlabel. No manual evidence is provided that GPT-4o labels are accurate on Claude, Qwen, Command R+, GPT-4o-mini, NLLB, or MADLAD outputs. If judge error is systematic, the En→Fa gap between Claude (94.0) and GPT-4o-mini (91.0) and the Fa→En ranking could be an artifact of the judge rather than of translation quality. Table 4's correlation of 0.88/0.79 is computed over n=7 aggregated points and does not establish per-system calibration. The limitations section notes dataset size but does not flag this transfer-of-validation gap, which is the more direct threat to the headline conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper contributes a new Persian idiom resource (PersianIdioms, 2,200 idioms, 700 with examples) and two parallel evaluation sets of 200 sentences each for English→Persian and Persian→English idiom translation, drawing English sentences from EPIE and MAGPIE and Persian sentences from PersianIdioms. It then evaluates five LLMs (GPT-3.5-turbo, GPT-4o-mini, Qwen-2.5-72B, Command R+, Claude-3.5-Sonnet), three NMT systems (NLLB-200, MADLAD-400, Google Translate), and hybrid LLM+NMT combinations under several prompting schemes (three single prompts, a chain-of-thought prompt, and a multi-step prompt). Translation quality is assessed with manual scores for idiom accuracy and fluency on a subset of outputs and with automatic metrics (COMET, BERTScore, BLEU, GPT-4o-as-judge). The central empirical claim is that Claude-3.5-Sonnet achieves the best GPT-4o idiom-accuracy scores in both translation directions, and that weak LLMs improve when combined with Google Translate in English→Persian, while Persian→English translation favors simpler prompts for weaker models and complex prompts for stronger ones.","tokens_in":16662,"tokens_out":8078,"duration_ms":82342,"significance":"The datasets and the PersianIdioms resource are potentially useful contributions to a low-resource language pair, and the paper is one of the few to compare prompting methods and LLM+NMT hybrids for idiom translation in Persian. The authors are transparent about data sources, manual annotation procedures, and inter-annotator agreement, and they release the data, which supports reproducibility. The comparison of automatic metrics against manual scores is a useful sanity check. However, the strength of the headline ranking is contingent on the validity of the GPT-4o judge across all evaluated systems; that validity is currently established only on a small, system-restricted sample. Consequently, the paper's contribution is significant conditional on the additional validation requested below.","major_comments":[{"comment":"The GPT-4o-as-judge is calibrated only against manual labels on seven outputs, all from GPT-3.5, Google Translate, or their combination, and only on the first 100 of 200 sentences. The paper then uses GPT-4o scores in Table 6 as the primary idiom-accuracy metric for every system and prompt, including Claude-3.5-Sonnet, Qwen-2.5, Command R+, GPT-4o-mini, NLLB, and MADLAD. Because §4.3 reports that GPT-4o tends to underlabel flexible paraphrases—precisely the kind of output Claude's CoT and MultiPrompt settings are designed to produce—the headline En→Fa gap between Claude (94.0) and GPT-4o-mini (91.0), and the Fa→En ranking, could be a judge artifact. Please add manual idiom-accuracy scores (or a substantial validation sample) for all models that are ranked, and report per-system judge agreement.","section":"§5.2, §5.3, §4.3, Table 4 and Table 6"},{"comment":"The Spearman correlations used to establish the reliability of GPT-4o (0.88 and 0.79) are computed over n=7 aggregated model outputs, with no confidence intervals, significance tests, or scatterplots. With n=7, the rank correlation is highly sensitive to a single output, and the statement \"GPT-4o performs comparably to humans\" is stronger than the evidence supports. Please report bootstrap or permutation intervals, or otherwise quantify the uncertainty around these correlations.","section":"§5.3, Table 4"},{"comment":"Table 6 gives point estimates without error bars, and Section 5.1 only states that GPT outputs were run \"multiple times\" without reporting the number of runs or the spread. Many of the comparisons in the GPT-4o column differ by only 1–3 points (for example, En→Fa Claude SinglePrompt2=93.0 vs. SinglePrompt3=93.5, or GPT-4o-mini SinglePrompt2=90.0 vs. MultiPrompt=91.0), so claims of superiority need variance estimates or significance tests. Please provide repeated-run statistics (mean ± std) or significance tests, or explicitly describe the results as exploratory.","section":"§5.4, §5.5, Table 6"},{"comment":"The limitations section acknowledges dataset size and language coverage but not the judge-transfer limitation described above. Because the central claim relies on GPT-4o scores for systems that were never manually evaluated, this missing limitation should be stated explicitly and discussed.","section":"§7 Limitations"}],"minor_comments":[{"comment":"MAGPIE is cited as (Xu et al., 2024), but that reference is a paper on alignment data synthesis, not the MAGPIE idiom corpus; the En→Fa dataset's provenance therefore needs a correct citation (Haagsma et al., 2020).","section":"§3.2 / References"},{"comment":"The choice of temperature 0.8 for translation and 0.1 for the judge is not justified; please add a sentence explaining the rationale.","section":"§5.1"},{"comment":"Row labels such as \"rSinglePrompt1\" and the lack of clear grouping make the table difficult to read; please reformat so that model, prompt, and metric columns are unambiguous.","section":"Table 6"},{"comment":"The source/back-translation uses \"Poor Mrs\" and then \"Poor mother\"; please clarify the source word and make the glosses consistent.","section":"Table 7"},{"comment":"The abstract states that BLEU and BERTScore are \"effective\" for evaluation, but Table 4 shows their correlation with idiom-translation scores is near zero (-0.03 and 0.18–0.25); please specify that they are useful for fluency only.","section":"Abstract / §5.3"}],"recommendation":"major_revision","confidential_remarks":"The paper has useful empirical scope but needs stronger statistical support and correct citations. The main fix is tractable: manually score a subset of outputs from the other models and show per-system judge agreement. If the authors do that and add uncertainty estimates, the paper could be suitable. I do not see circular reasoning or internal inconsistency; the issue is missing evidence, not a flawed derivation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The main thing this paper contributes is PersianIdioms: 2,200 Persian idioms with meanings, 700 with examples, plus two 200-sentence parallel idiom translation test sets with expert-validated gold translations. That is a real resource gap, and the paper is honest about its construction process. I'm not aware of another Persian idiom resource of this size, and the parallel sets are useful for future MT evaluation.\n\nWhat the paper does well beyond the resource: it runs a broad comparison of open and closed LLMs, NMT models, prompts, and LLM+NMT hybrids, and it reports manual evaluation with inter-annotator agreement on a subset. The manual scoring procedure is explained clearly, and the choice of Gwet's AC1 for skewed fluency scores is reasonable.\n\nThe soft spot is the one the stress-test flags, and I think it lands. GPT-4o is used as the primary idiom-accuracy judge for Table 6, but it is validated only on seven model outputs, all from GPT-3.5, Google Translate, or their combination, and on the first 100 sentences. The paper even says GPT-4o tends to underestimate paraphrastic translations, which is exactly what Claude's CoT and MultiPrompt outputs are most likely to be. The n=7 Spearman correlations are suggestive but not enough to establish per-system calibration. So the headline 'Claude is best in both directions' is plausible but not proven. The absence of error bars or significance tests on the Table 6 differences (e.g., 94.0 vs 91.0) compounds that. I'd like to see the judge validated on a sample of Claude and GPT-4o-mini outputs, or at least a sentence-level human evaluation on the systems that matter for the ranking.\n\nMinor notes: the dataset size is acknowledged in the limitations; the hybrid method is adapted from prior work, but that's a legitimate extension to a new language pair. The paper's tone is appropriately cautious except for the word 'outstanding' in the abstract and conclusion, which overstates what the evidence supports.\n\nWho this is for: anyone working on Persian MT, idiom translation, or LLM-as-judge reliability. I'd bring it to a reading group as an example of a resource paper with a somewhat fragile evaluation. It deserves peer review, but I'd want the judge-validation issue addressed before publication, or at least clearly acknowledged as a limitation. My verdict: send it to review, with a request for additional validation.","headline":"A valuable Persian idiom resource with a plausible but under-supported Claude ranking, because the GPT-4o judge is validated on only a fraction of the evaluated systems.","tokens_in":17176,"tokens_out":2651,"would_cite":true,"duration_ms":24691,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Claude-3.5-Sonnet achieves the highest idiom-translation accuracy in both English-to-Persian and Persian-to-English in this comparison of LLMs, NMT systems, and hybrids.","keywords":["idiom translation","Persian","English","large language models","machine translation","LLM-as-a-judge","prompt engineering","parallel datasets"],"falsifier":"Have multiple native Persian and English speakers annotate idiom-translation correctness for all 200 sentences across every model and prompt in Table 6, then rank systems by human majority vote and compare with the GPT-4o ranking; if the Spearman correlation drops below roughly 0.8 or the top system changes, the paper's central ranking is falsified.","tokens_in":16161,"feed_emoji":"🌐","tokens_out":8243,"duration_ms":71959,"temperature":0.7,"pith_summary":"This paper introduces the first parallel Persian–English datasets for idiom translation—200 sentences in each direction, with Persian idioms drawn from a new PersianIdioms resource of 2,200 idioms—and uses them to compare open and closed-source LLMs, NMT models, and hybrid LLM-then-NMT pipelines under several prompting strategies. Its central empirical claim is that Claude-3.5-Sonnet achieves the highest idiom-translation accuracy in both directions, reaching GPT-4o-judged scores of 94.0 for English-to-Persian and 75.0 for Persian-to-English. The paper also establishes that models translate English idioms far more accurately than Persian ones, that weaker LLMs improve in English-to-Persian when combined with Google Translate, and that the best prompting style depends on model strength and direction. If the results are correct, they provide a concrete, reproducible benchmark for idiom translation in this language pair and a validated evaluation recipe for future work.","feed_headline":"Claude-3.5-Sonnet tops Persian↔English idiom translation tests","feed_subtitle":"New datasets and a validated GPT-4o judge rank which models, prompts, and hybrids handle idioms best.","key_machinery":"The central objects are the two new parallel datasets (Fa→En and En→Fa, each 200 sentences with one idiom per sentence) and the evaluation protocol built around them: a binary idiom-translation metric defined within the MQM framework, fluency ratings on a 1–5 scale, and GPT-4o-as-judge with reference-guided grading, which the paper validates by Spearman correlation against human scores. The other load-bearing mechanism is the hybrid pipeline, in which an LLM first identifies idioms and replaces them with literal clauses and an NMT system then translates the resulting text; this pipeline is what produces the finding that weaker LLMs gain from combination with Google Translate in English→Persian.","core_discovery":"On two new parallel test sets of 200 sentences each, Claude-3.5-Sonnet obtains the highest GPT-4o idiom-accuracy scores in both translation directions among all tested systems: 94.0 for English→Persian with the chain-of-thought prompt and 75.0 for Persian→English with the multi-prompt setup, where GPT-4o scores are the paper's primary idiom-translation metric validated against human labels. The paper further claims that combining a weaker LLM with Google Translate significantly improves English→Persian idiom translation—raising Qwen-2.5-72B's GPT-4o score from 74.5 to 88.0 and GPT-3.5's from 72.0 to 79.0—while Persian→English outputs benefit from simple single prompts for GPT-3.5, GPT-4o-mini, and Qwen-2.5, and from complex CoT or MultiPrompt setups for Claude-3.5-Sonnet and Command R+. These findings are supported by manual annotation of 100 sentences from seven model outputs and by GPT-4o-as-judge scores on the full 200 sentences per direction.","pith_inferences":["If GPT-4o-as-judge reliability holds beyond the seven validation outputs, the same reference-guided binary-judge protocol could be exported to other low-resource language pairs, but the paper's own note that GPT-4o tends to underestimate paraphrased English→Persian translations suggests the judge may need reference-set expansion before ranking is trustworthy.","The hybrid gain pattern—an idiom-aware rewriter plus a fluent NMT backend—suggests a general complementarity principle that could be tested on other language pairs and other figurative-language phenomena such as metaphors and proverbs.","The 200-sentence sample and single-judge automation leave room for a stress test: re-annotating all outputs with multiple native-speaker judges could reveal whether the Claude-vs-GPT-4o-mini margins in Table 6 are stable or within annotation noise."],"forward_implications":["Claude-3.5-Sonnet becomes the default strong baseline for future Persian–English idiom-translation research, since it tops both directions on the paper's metrics.","For English→Persian, teams with limited compute can approximate strong-LLM performance by chaining an open-weight LLM such as Qwen-2.5-72B in front of Google Translate.","For Persian→English, prompt choice should be matched to model strength: single prompts for weaker models, chain-of-thought or multi-step prompts for larger ones.","The PersianIdioms resource and the two parallel datasets give the community a way to measure progress on a lower-resource figurative-language task that previously had no benchmark.","GPT-4o-as-judge, BLEU, and BERTScore can be used as a substitute for manual idiom-translation and fluency evaluation when ranking systems, provided the judge is re-validated on the target model outputs."],"supporting_citations":[{"why":"Supplies the EPIE English idiom sentences that the En→Fa dataset draws from.","marker":"Saxena and Paul, 2020"},{"why":"Supplies MAGPIE, the other source of English idiom sentences in the En→Fa dataset.","marker":"Xu et al., 2024"},{"why":"Provides the LLM-as-a-judge single-answer and reference-guided grading protocol used for GPT-4o idiom-accuracy scores.","marker":"Zheng et al., 2023"},{"why":"Establishes the non-literalness tendency of GPT models that the paper's LLM-versus-NMT comparison builds on.","marker":"Raunak et al., 2023"},{"why":"Supplies the hypothesis that LLM paraphrastic ability can help NMT translate figurative language, which the hybrid pipeline tests.","marker":"Hendy et al., 2023"},{"why":"Contributes one of the single prompts evaluated in the prompt-comparison experiments.","marker":"Yamada, 2024"},{"why":"Defines the MQM framework from which the paper's binary idiom-translation and fluency metrics are devised.","marker":"Lommel et al., 2014"}],"fun_headline_variants":["Claude-3.5-Sonnet wins Persian↔English idiom translation","Claude-3.5-Sonnet tops Persian-English idiom tests","Best idiom scores: Claude-3.5-Sonnet in both Persian↔English","Claude-3.5-Sonnet leads both ways on Persian idiom tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking of all models rests on GPT-4o-as-judge scores, but the judge's agreement with human annotators was tested on only seven outputs—all from GPT-3.5, Google Translate, or their combination—and on the first 100 sentences of those outputs; the paper assumes this reliability transfers to Claude, Qwen, Command R+, NLLB, and all prompt variants without direct validation.","fun_headline_variants_meta":{"raw":{"variants":["Claude-3.5-Sonnet wins Persian↔English idiom translation","Claude-3.5-Sonnet tops Persian-English idiom tests","Best idiom scores: Claude-3.5-Sonnet in both Persian↔English","Claude-3.5-Sonnet leads both ways on Persian idiom tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00053,"raw_usage":{"total_tokens":2580,"prompt_tokens":997,"completion_tokens":1583,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":1496}},"tokens_in":613,"tokens_out":1583,"duration_ms":13657,"temperature":1.0,"reasoning_tokens":1496,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:27:56.052238+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have multiple native Persian and English speakers annotate idiom-translation correctness for all 200 sentences across every model and prompt in Table 6, then rank systems by human majority vote and compare with the GPT-4o ranking; if the Spearman correlation drops below roughly 0.8 or the top system changes, the paper's central ranking is falsified.","supporting_citations":[],"review_version":1}