Pith. sign in

REVIEW 3 cited by

Imagination improves Multimodal Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1705.04350 v2 pith:WDX5CI5F submitted 2017-05-11 cs.CL cs.CV

classification cs.CLcs.CV
keywords translationlearningdatasetexternalgroundedimageimproveslearned
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We decompose multimodal translation into two sub-tasks: learning to translate and learning visually grounded representations. In a multitask learning framework, translations are learned in an attention-based encoder-decoder, and grounded representations are learned through image representation prediction. Our approach improves translation performance compared to the state of the art on the Multi30K dataset. Furthermore, it is equally effective if we train the image prediction task on the external MS COCO dataset, and we find improvements if we train the translation model on the external News Commentary parallel text.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hindi Visual Genome: A Dataset for Multimodal English-to-Hindi Machine Translation

    cs.CL 2019-07 unverdicted novelty 7.0 of 10

    The paper releases the first multimodal English-Hindi machine translation dataset of 31,525 segments with images and a challenge test set of 1,400 segments selected via embedding similarity for image-resolvable ambiguities.

  2. TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries

    cs.CL 2025-05 conditional novelty 6.0 of 10

    This paper builds TopicVD, a topic-based documentary video-subtitle translation dataset, and shows with a cross-modal attention model that visual and contextual information improve BLEU scores.

  3. Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Migician is an instruction-tuned MLLM that performs free-form grounding across multiple images, with a new 630k dataset and a 10-task benchmark, but the evaluation is weakened by source overlap between training and be...

Pith tools