REVIEW 2 cited by
SAViR-T: Spatially Attentive Visual Reasoning with Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present a novel computational model, "SAViR-T", for the family of visual reasoning problems embodied in the Raven's Progressive Matrices (RPM). Our model considers explicit spatial semantics of visual elements within each image in the puzzle, encoded as spatio-visual tokens, and learns the intra-image as well as the inter-image token dependencies, highly relevant for the visual reasoning task. Token-wise relationship, modeled through a transformer-based SAViR-T architecture, extract group (row or column) driven representations by leveraging the group-rule coherence and use this as the inductive bias to extract the underlying rule representations in the top two row (or column) per token in the RPM. We use this relation representations to locate the correct choice image that completes the last row or column for the RPM. Extensive experiments across both synthetic RPM benchmarks, including RAVEN, I-RAVEN, RAVEN-FAIR, and PGM, and the natural image-based "V-PROM" demonstrate that SAViR-T sets a new state-of-the-art for visual reasoning, exceeding prior models' performance by a considerable margin.
Forward citations
Cited by 2 Pith papers
-
Johnny: Structuring Representation Space to Enhance Machine Abstract Reasoning Ability
Johnny tokenizes RPM images into a learned codebook, adds a self-referential 'sub-enumeration' loss to the reasoning module, and pairs it with a new Spin-Transformer layer; gains over strong baselines are modest, and ...
-
Learning Visual Abstract Reasoning through Dual-Stream Networks
A dual-stream network combining CNN and vision transformer features, followed by a rule extractor, achieves state-of-the-art accuracy on multiple Raven's Progressive Matrices benchmarks, including large gains on out-o...
Discussion (0). Continue with ORCID to comment.