Pith. sign in

REVIEW 1 cited by

Active Learning for Visual Question Answering: An Empirical Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.01732 v1 pith:THS6XVOD submitted 2017-11-06 cs.CV

classification cs.CV
keywords learningactivequestionsdeepgoal-drivenmodelansweringapproaches
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present an empirical study of active learning for Visual Question Answering, where a deep VQA model selects informative question-image pairs from a pool and queries an oracle for answers to maximally improve its performance under a limited query budget. Drawing analogies from human learning, we explore cramming (entropy), curiosity-driven (expected model change), and goal-driven (expected error reduction) active learning approaches, and propose a fast and effective goal-driven active learning scoring function to pick question-image pairs for deep VQA models under the Bayesian Neural Network framework. We find that deep VQA models need large amounts of training data before they can start asking informative questions. But once they do, all three approaches outperform the random selection baseline and achieve significant query savings. For the scenario where the model is allowed to ask generic questions about images but is evaluated only on specific questions (e.g., questions whose answer is either yes or no), our proposed goal-driven scoring function performs the best.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Maximally Separated Active Learning

    cs.LG 2024-11 conditional novelty 5.0 of 10

    MSAL scores uncertainty by cosine similarity to fixed equiangular class prototypes, and MSAL-D adds prototype-based diversity, reporting improved AUBC on MNIST, SVHN, and TinyImageNet.

Pith tools