Pith. sign in

REVIEW 7 cited by

Document-Level Machine Translation with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.02210 v2 pith:764ETW3L submitted 2023-04-05 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmstranslationdiscoursedocument-levelevaluationlanguagemodelingmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) such as ChatGPT can produce coherent, cohesive, relevant, and fluent answers for various natural language processing (NLP) tasks. Taking document-level machine translation (MT) as a testbed, this paper provides an in-depth evaluation of LLMs' ability on discourse modeling. The study focuses on three aspects: 1) Effects of Context-Aware Prompts, where we investigate the impact of different prompts on document-level translation quality and discourse phenomena; 2) Comparison of Translation Models, where we compare the translation performance of ChatGPT with commercial MT systems and advanced document-level MT methods; 3) Analysis of Discourse Modelling Abilities, where we further probe discourse knowledge encoded in LLMs and shed light on impacts of training techniques on discourse modeling. By evaluating on a number of benchmarks, we surprisingly find that LLMs have demonstrated superior performance and show potential to become a new paradigm for document-level translation: 1) leveraging their powerful long-text modeling capabilities, GPT-3.5 and GPT-4 outperform commercial MT systems in terms of human evaluation; 2) GPT-4 demonstrates a stronger ability for probing linguistic knowledge than GPT-3.5. This work highlights the challenges and opportunities of LLMs for MT, which we hope can inspire the future design and evaluation of LLMs.We release our data and annotations at https://github.com/longyuewangdcu/Document-MT-LLM.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning

    cs.CL 2025-09 conditional novelty 7.0 of 10

    AncientDoc is a new five-task benchmark for Chinese ancient documents, and it shows current vision-language models fail at page-level OCR but perform somewhat better on reasoning tasks.

  2. Leveraging Large Language Models for Accurate Sign Language Translation in Low-Resource Scenarios

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    A prompting method that links signs to short text descriptions lets large language models translate English and Italian into sign language glosses, beating prior models in low-data settings.

  3. LLMs are Introvert

    cs.AI 2025-07 conditional novelty 5.0 of 10

    A psychology-inspired prompting method (SIP-CoT with emotion-guided memory) makes LLM agents reproduce human-like attitudes and behaviors more closely in social simulations, but the evaluation lacks error bars, a name...

  4. Information Loss in LLMs' Multilingual Translation: The Role of Training Data, Language Proximity, and Language Family

    cs.CL 2025-06 reject novelty 5.0 of 10

    Round-trip translation quality in GPT-4 and Llama 2 is jointly shaped by training data volume and language distance from English, with orthographic, phylogenetic, syntactic, and geographic distances as the strongest p...

  5. Facts Do Care About Your Language: Assessing Answer Quality of Multilingual LLMs

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A small evaluation of Llama 3.1 shows factuality in school-level question answering degrades with decreasing language speaker count, though the statistical support is weakened by methodological issues.

  6. Demystifying ChatGPT: How It Masters Genre Recognition

    cs.CL 2025-07 reject novelty 3.0 of 10

    A benchmark reports that ChatGPT outperforms other LLMs and traditional classifiers on multi-label movie genre prediction, but the result is undermined by likely pretraining contamination and weak baselines.

  7. Robust pid sliding mode control for dc servo motor speed control

    eess.SY 2025-08 unverdicted novelty 2.0 of 10

    An abstract-only claim that SMC-PID outperforms PID for DC servo motor speed on the CE110 trainer; the submitted body text is an unrelated paper, so the result is unverifiable.

Pith tools