Pith. sign in

REVIEW 1 cited by

Don't Rank, Combine! Combining Machine Translation Hypotheses Using Quality Estimation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.06688 v2 pith:PZ2FKEJS submitted 2024-01-12 cs.CL cs.LG

classification cs.CLcs.LG
keywords qe-fusiontranslationcandidatesqualitycombiningconsistentlyestimationhuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural machine translation systems estimate probabilities of target sentences given source sentences, yet these estimates may not align with human preferences. This work introduces QE-fusion, a method that synthesizes translations using a quality estimation metric (QE), which correlates better with human judgments. QE-fusion leverages a pool of candidates sampled from a model, combining spans from different candidates using a QE metric such as CometKiwi. We compare QE-fusion against beam search and recent reranking techniques, such as Minimum Bayes Risk decoding or QE-reranking. Our method consistently improves translation quality in terms of COMET and BLEURT scores when applied to large language models (LLMs) used for translation (PolyLM, XGLM, Llama2, Mistral, ALMA, and Tower) and to multilingual translation models (NLLB), over five language pairs. Notably, QE-fusion exhibits larger improvements for LLMs due to their ability to generate diverse outputs. We demonstrate that our approach generates novel translations in over half of the cases and consistently outperforms other methods across varying numbers of candidates (5-200). Furthermore, we empirically establish that QE-fusion scales linearly with the number of candidates in the pool.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Quality-Aware Decoding: Unifying Quality Estimation and Decoding

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A uni-directional token-level QE model that scores partial translations is merged into beam search, improving NMT quality over N-best re-ranking on WMT23 English-German and Chinese-English.

Pith tools