Pith. sign in

REVIEW 2 cited by

Neural Language Modeling with Visual Features

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.02930 v1 pith:GJ4PHVIF submitted 2019-03-07 cs.CL

classification cs.CL
keywords languagefeaturesmodelvisualmodelingmodelsmultimodalneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal language models attempt to incorporate non-linguistic features for the language modeling task. In this work, we extend a standard recurrent neural network (RNN) language model with features derived from videos. We train our models on data that is two orders-of-magnitude bigger than datasets used in prior work. We perform a thorough exploration of model architectures for combining visual and text features. Our experiments on two corpora (YouCookII and 20bn-something-something-v2) show that the best performing architecture consists of middle fusion of visual and text features, yielding over 25% relative improvement in perplexity. We report analysis that provides insights into why our multimodal language model improves upon a standard RNN language model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vision-Assisted Foundation Model for Solving Multi-Task Vehicle Routing Problems

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    VaFM encodes constraint-specific VRP images via CNN into patch embeddings fused with graph nodes, using an auxiliary task to handle pixel imbalance, and reports better performance than prior methods on 16 VRP variants.

  2. Cloud Removal With PolSAR-Optical Data Fusion Using A Two-Flow Residual Network

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A two-flow residual network that fuses PolSAR polarization features with optical images outperforms six prior cloud-removal models on the authors' new airborne dataset, though data and code are not public.

Pith tools