Pith. sign in

REVIEW 2 cited by

Multi-Image Visual Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.13706 v2 pith:TUAZKOL2 submitted 2021-12-27 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords imagequestionansweringvisualaccuracydatasetdifferentfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While a lot of work has been done on developing models to tackle the problem of Visual Question Answering, the ability of these models to relate the question to the image features still remain less explored. We present an empirical study of different feature extraction methods with different loss functions. We propose New dataset for the task of Visual Question Answering with multiple image inputs having only one ground truth, and benchmark our results on them. Our final model utilising Resnet + RCNN image features and Bert embeddings, inspired from stacked attention network gives 39% word accuracy and 99% image accuracy on CLEVER+TinyImagenet dataset.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GoVector: An I/O-Efficient Caching Strategy for High-Dimensional Vector Nearest Neighbor Search

    cs.DB 2025-08 reject novelty 5.0 of 10

    A GoVector abstract claims a hybrid static/dynamic cache plus disk reordering improves disk-based ANN search, but the manuscript body is an unrelated chart/table benchmark paper.

  2. LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Inserting a pixel-shuffle plus residual patch-merge layer inside the vision encoder compresses visual tokens more efficiently than post-encoder compression, at modest accuracy cost.

Pith tools