REVIEW 2 cited by
Neural Language Modeling with Visual Features
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multimodal language models attempt to incorporate non-linguistic features for the language modeling task. In this work, we extend a standard recurrent neural network (RNN) language model with features derived from videos. We train our models on data that is two orders-of-magnitude bigger than datasets used in prior work. We perform a thorough exploration of model architectures for combining visual and text features. Our experiments on two corpora (YouCookII and 20bn-something-something-v2) show that the best performing architecture consists of middle fusion of visual and text features, yielding over 25% relative improvement in perplexity. We report analysis that provides insights into why our multimodal language model improves upon a standard RNN language model.
Forward citations
Cited by 2 Pith papers
-
Vision-Assisted Foundation Model for Solving Multi-Task Vehicle Routing Problems
VaFM encodes constraint-specific VRP images via CNN into patch embeddings fused with graph nodes, using an auxiliary task to handle pixel imbalance, and reports better performance than prior methods on 16 VRP variants.
-
Cloud Removal With PolSAR-Optical Data Fusion Using A Two-Flow Residual Network
A two-flow residual network that fuses PolSAR polarization features with optical images outperforms six prior cloud-removal models on the authors' new airborne dataset, though data and code are not public.
Discussion (0). Continue with ORCID to comment.