REVIEW 1 cited by
Challenges and Prospects in Vision and Language Research
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Language grounded image understanding tasks have often been proposed as a method for evaluating progress in artificial intelligence. Ideally, these tasks should test a plethora of capabilities that integrate computer vision, reasoning, and natural language understanding. However, rather than behaving as visual Turing tests, recent studies have demonstrated state-of-the-art systems are achieving good performance through flaws in datasets and evaluation procedures. We review the current state of affairs and outline a path forward.
Forward citations
Cited by 1 Pith paper
-
Answering Questions about Data Visualizations using Efficient Bimodal Fusion
PReFIL combines LSTM question embeddings with two levels of convolutional features via 1x1 convolutions and recurrent spatial aggregation, setting new state-of-the-art accuracy on FigureQA and DVQA.
Discussion (0). Continue with ORCID to comment.