Pith. sign in

REVIEW 3 cited by

A Survey on Multi-modal Machine Translation: Tasks, Methods and Challenges

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.12669 v2 pith:OLBPGK2O submitted 2024-05-21 cs.CL

classification cs.CL
keywords machinemulti-modaltranslationperformancesurveytypesvisualacademia
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, multi-modal machine translation has attracted significant interest in both academia and industry due to its superior performance. It takes both textual and visual modalities as inputs, leveraging visual context to tackle the ambiguities in source texts. In this paper, we begin by offering an exhaustive overview of 99 prior works, comprehensively summarizing representative studies from the perspectives of dominant models, datasets, and evaluation metrics. Afterwards, we analyze the impact of various factors on model performance and finally discuss the possible research directions for this task in the future. Over time, multi-modal machine translation has developed more types to meet diverse needs. Unlike previous surveys confined to the early stage of multi-modal machine translation, our survey thoroughly concludes these emerging types from different aspects, so as to provide researchers with a better understanding of its current state.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

  2. ConECT Dataset: Overcoming Data Scarcity in Context-Aware E-Commerce MT

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A new Czech-to-Polish e-commerce translation dataset is released, and the paper shows small improvements from visual and category context, with a negative result for image descriptions.

  3. ViDove: A Translation Agent System with Multimodal Context and Memory-Augmented Reasoning

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A multimodal, memory-augmented multi-agent system for video subtitling and translation, plus a new 17-hour benchmark, reports large BLEU/SubER gains on its own benchmark but not consistently on existing benchmarks.

Pith tools