Pith. sign in

REVIEW 1 cited by

Challenges and Prospects in Vision and Language Research

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.09317 v2 pith:67PFI5BK submitted 2019-04-19 cs.LG cs.CLcs.CVcs.NEstat.ML

classification cs.LGcs.CLcs.CVcs.NEstat.ML
keywords languagetasksunderstandingvisionachievingaffairsartificialbeen
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Language grounded image understanding tasks have often been proposed as a method for evaluating progress in artificial intelligence. Ideally, these tasks should test a plethora of capabilities that integrate computer vision, reasoning, and natural language understanding. However, rather than behaving as visual Turing tests, recent studies have demonstrated state-of-the-art systems are achieving good performance through flaws in datasets and evaluation procedures. We review the current state of affairs and outline a path forward.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Answering Questions about Data Visualizations using Efficient Bimodal Fusion

    cs.CV 2019-08 accept novelty 6.0 of 10

    PReFIL combines LSTM question embeddings with two levels of convolutional features via 1x1 convolutions and recurrent spatial aggregation, setting new state-of-the-art accuracy on FigureQA and DVQA.

Pith tools