A 7B multimodal model trained with multi-task reinforcement learning beats much larger models on image-text translation benchmarks, though some out-of-distribution claims are contradicted by the paper's own tables.
Document image machine translation with dynamic multi-pre-trained models assembling
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MT$^{3}$: Scaling MLLM-based Text Image Machine Translation via Multi-Task Reinforcement Learning
A 7B multimodal model trained with multi-task reinforcement learning beats much larger models on image-text translation benchmarks, though some out-of-distribution claims are contradicted by the paper's own tables.