REVIEW 5 cited by
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language. We extend the popular BERT architecture to a multi-modal two-stream model, pro-cessing both visual and textual inputs in separate streams that interact through co-attentional transformer layers. We pretrain our model through two proxy tasks on the large, automatically collected Conceptual Captions dataset and then transfer it to multiple established vision-and-language tasks -- visual question answering, visual commonsense reasoning, referring expressions, and caption-based image retrieval -- by making only minor additions to the base architecture. We observe significant improvements across tasks compared to existing task-specific models -- achieving state-of-the-art on all four tasks. Our work represents a shift away from learning groundings between vision and language only as part of task training and towards treating visual grounding as a pretrainable and transferable capability.
Forward citations
Cited by 5 Pith papers
-
Representations in vision and language converge in a shared, multidimensional space of perceived similarities
Similarity judgments of natural scene images and their sentence captions are aligned with each other, with visual brain responses, and with LLM-trained visual models.
-
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
LaVi encodes visual context into LayerNorm affine parameters, bypassing visual token concatenation, and reports LLaVA-comparable accuracy at a 94% FLOP reduction.
-
From Image Captioning to Visual Storytelling
Visual storytelling improves by treating it as image captioning followed by language-to-language story generation, with a new 'ideality' metric to gauge distance from an oracle.
-
On the Resilience of Underwater Semantic Wireless Communications
In a simulated underwater acoustic link, the SAGE semantic image system keeps semantic similarity around 50% up to 15-20% character error, indicating resilience to text corruption.
-
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
RDB improves remote sensing image-text retrieval mean recall by 1.15 to 2 percent over fully fine-tuned GeoRSCLIP using an asymmetric adapter and a dual-task consistency loss.
Discussion (0). Sign in to comment.