M2N2 evolves merging boundaries, uses resource competition for diversity and attraction-based pairing, achieving from-scratch evolution and specialized model fusion.
STAIR Captions: Constructing a Large-Scale Japanese Image Caption Dataset
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In recent years, automatic generation of image descriptions (captions), that is, image captioning, has attracted a great deal of attention. In this paper, we particularly consider generating Japanese captions for images. Since most available caption datasets have been constructed for English language, there are few datasets for Japanese. To tackle this problem, we construct a large-scale Japanese image caption dataset based on images from MS-COCO, which is called STAIR Captions. STAIR Captions consists of 820,310 Japanese captions for 164,062 images. In the experiment, we show that a neural network trained using STAIR Captions can generate more natural and better Japanese captions, compared to those generated using English-Japanese machine translation after generating English captions.
citation-role summary
citation-polarity summary
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
Competition and Attraction Improve Model Fusion
M2N2 evolves merging boundaries, uses resource competition for diversity and attraction-based pairing, achieving from-scratch evolution and specialized model fusion.