Pith. sign in

REVIEW 1 cited by

Evaluating Machine Translation Performance on Chinese Idioms with a Blacklist Method

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.07646 v3 pith:374JLHP2 submitted 2017-11-21 cs.CL

classification cs.CL
keywords translationidiomsliteralblacklisterrorevaluationmethodchinese
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Idiom translation is a challenging problem in machine translation because the meaning of idioms is non-compositional, and a literal (word-by-word) translation is likely to be wrong. In this paper, we focus on evaluating the quality of idiom translation of MT systems. We introduce a new evaluation method based on an idiom-specific blacklist of literal translations, based on the insight that the occurrence of any blacklisted words in the translation output indicates a likely translation error. We introduce a dataset, CIBB (Chinese Idioms Blacklists Bank), and perform an evaluation of a state-of-the-art Chinese-English neural MT system. Our evaluation confirms that a sizable number of idioms in our test set are mistranslated (46.1%), that literal translation error is a common error type, and that our blacklist method is effective at identifying literal translation errors.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Chengyu-Bench is a 2,937-example human-verified benchmark showing LLMs are strong at idiom sentiment classification but weak at appropriateness and open cloze generation.

Pith tools