Pith. sign in

REVIEW

iCap: Interactive Image Captioning with Predictive Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.11782 v3 pith:IW4GNR2M submitted 2020-01-31 cs.HC cs.CV

classification cs.HCcs.CV
keywords imagecaptioninginteractiveabd-capautomatedcompletionicapinput
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper we study a brand new topic of interactive image captioning with human in the loop. Different from automated image captioning where a given test image is the sole input in the inference stage, we have access to both the test image and a sequence of (incomplete) user-input sentences in the interactive scenario. We formulate the problem as Visually Conditioned Sentence Completion (VCSC). For VCSC, we propose asynchronous bidirectional decoding for image caption completion (ABD-Cap). With ABD-Cap as the core module, we build iCap, a web-based interactive image captioning system capable of predicting new text with respect to live input from a user. A number of experiments covering both automated evaluations and real user studies show the viability of our proposals.

Discussion (0). Sign in to comment.

Pith tools