Pith. sign in

REVIEW 10 cited by

(Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.11804 v2 pith:C3V6RRYB submitted 2024-05-20 cs.CL

classification cs.CL
keywords translationhumanlanguagemulti-agentqualitytranslationscollaborationcultural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Literary translation remains one of the most challenging frontiers in machine translation due to the complexity of capturing figurative language, cultural nuances, and unique stylistic elements. In this work, we introduce TransAgents, a novel multi-agent framework that simulates the roles and collaborative practices of a human translation company, including a CEO, Senior Editor, Junior Editor, Translator, Localization Specialist, and Proofreader. The translation process is divided into two stages: a preparation stage where the team is assembled and comprehensive translation guidelines are drafted, and an execution stage that involves sequential translation, localization, proofreading, and a final quality check. Furthermore, we propose two innovative evaluation strategies: Monolingual Human Preference (MHP), which evaluates translations based solely on target language quality and cultural appropriateness, and Bilingual LLM Preference (BLP), which leverages large language models like GPT-4} for direct text comparison. Although TransAgents achieves lower d-BLEU scores, due to the limited diversity of references, its translations are significantly better than those of other baselines and are preferred by both human evaluators and LLMs over traditional human references and GPT-4} translations. Our findings highlight the potential of multi-agent collaboration in enhancing translation quality, particularly for longer texts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Steering Large Language Models for Machine Translation Personalization

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Contrastive steering of sparse autoencoder features personalizes literary machine translation to a target translator's style as well as twenty-shot prompting while keeping inference fast.

  2. New Evaluation Paradigm for Lexical Simplification

    cs.CL 2025-01 conditional novelty 6.0 of 10

    This paper proposes an all-in-one lexical simplification dataset with per-sentence complex word and substitute annotations, and a multi-LLM voting method that is claimed to outperform earlier baselines.

  3. Adaptive Few-shot Prompting for Machine Translation with Pre-trained Language Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Adaptive selection of translation examples through a hybrid LLM-embedding retriever and a self-supervised reranker improves few-shot machine translation across several LLMs.

  4. A 2-step Framework for Automated Literary Translation Evaluation: Its Promises and Pitfalls

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A rubric-plus-question-answering LLM framework for literary translation evaluation beats traditional MT metrics but still trails human agreement, especially on Korean honorifics.

  5. Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels

    cs.CL 2024-11 conditional novelty 5.0 of 10

    GPT-4 translates at the level of junior to mid-level professional translators but trails senior translators, with weaknesses in named entities and grammar.

  6. TACTIC: Translation Agents with Cognitive-Theoretic Interactive Collaboration

    cs.CL 2025-06 conditional novelty 4.0 of 10

    TACTIC, a cognitive-inspired six-agent workflow, improves LLM translation quality over direct prompting on FLORES-200 and WMT24, with the best DeepSeek-V3 setup reaching 96.19 XCOMET on English-to-X.

  7. Beyond the Sentence: A Survey on Context-Aware Machine Translation with Large Language Models

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A survey of context-aware machine translation with large language models, categorizing prompting, fine-tuning, and agent-based approaches.

  8. Agentic AI Systems Applied to tasks in Financial Services: Modeling and model risk management crews

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A CrewAI-based multi-agent system with human oversight built financial models and carried out model risk management checks on three public credit datasets, with results comparable to AutoML and Kaggle baselines.

  9. Findings of the WMT 2024 Shared Task on Discourse-Level Literary Translation

    cs.CL 2024-12 conditional novelty 4.0 of 10

    The WMT 2024 literary translation shared task finds that domain-enhanced systems lead in d-BLEU for Chinese-English, but human evaluators rank NLP2CT-UM and SJTU-LoveFiction at the top.

  10. TransBench: Benchmarking Machine Translation for Industrial-Scale Applications

    cs.CL 2025-05 reject novelty 2.0 of 10

    TransBench is a proposed e-commerce MT benchmark with a three-level evaluation framework and a fine-tuned quality-scoring model, but the paper contains no results and no released data or code.

Pith tools