A framework that generates images from source sentences with a Stable Diffusion model fine-tuned by a scene-graph reward, then feeds them into a multimodal LLM, is claimed to improve machine translation, but key comparisons are confounded.
LLM-based Translation Inference with Iterative Bilingual Understanding
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The remarkable understanding and generation capabilities of large language models (LLMs) have greatly improved translation performance. However, incorrect understanding of the sentence to be translated can degrade translation quality. To address this issue, we proposed a novel Iterative Bilingual Understanding Translation (IBUT) method based on the cross-lingual capabilities of LLMs and the dual characteristics of translation tasks. The cross-lingual capability of LLMs enables the generation of contextual understanding for both the source and target languages separately. Furthermore, the dual characteristics allow IBUT to generate effective cross-lingual feedback, iteratively refining contextual understanding, thereby reducing errors and improving translation performance. Experimental results showed that the proposed IBUT outperforms several strong comparison methods, especially being generalized to multiple domains (e.g., news, commonsense, and cultural translation benchmarks).
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
A framework that generates images from source sentences with a Stable Diffusion model fine-tuned by a scene-graph reward, then feeds them into a multimodal LLM, is claimed to improve machine translation, but key comparisons are confounded.