Machine-translated task data is the best average parallel-data source for cross-lingual transfer of vision-language encoders, but authentic caption-like data beats it in some languages, and multilingual training helps on average up to a point.
In: International Conference on Learning Representations (2020), https://openreview.net/forum?id=r1xCMyBtPS
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders
Machine-translated task data is the best average parallel-data source for cross-lingual transfer of vision-language encoders, but authentic caption-like data beats it in some languages, and multilingual training helps on average up to a point.