Pith. sign in

REVIEW 2 cited by

STAIR Captions: Constructing a Large-Scale Japanese Image Caption Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1705.00823 v1 pith:G5ZMBJV2 submitted 2017-05-02 cs.CL cs.CV

classification cs.CLcs.CV
keywords captionsjapaneseimagestaircaptionimagesdatasetdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, automatic generation of image descriptions (captions), that is, image captioning, has attracted a great deal of attention. In this paper, we particularly consider generating Japanese captions for images. Since most available caption datasets have been constructed for English language, there are few datasets for Japanese. To tackle this problem, we construct a large-scale Japanese image caption dataset based on images from MS-COCO, which is called STAIR Captions. STAIR Captions consists of 820,310 Japanese captions for 164,062 images. In the experiment, we show that a neural network trained using STAIR Captions can generate more natural and better Japanese captions, compared to those generated using English-Japanese machine translation after generating English captions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Competition and Attraction Improve Model Fusion

    cs.AI 2025-08 conditional novelty 6.0 of 10

    M2N2 evolves merging boundaries, uses resource competition for diversity and attraction-based pairing, achieving from-scratch evolution and specialized model fusion.

  2. Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A dynamic adapter generated from disentangled semantic and style features improves cross-lingual cross-modal retrieval over static adapters on image and video benchmarks.

Pith tools