Pith. sign in

REVIEW 2 cited by

Learning Video Context as Interleaved Multimodal Sequences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.21757 v2 pith:3GTEDSFJ submitted 2024-07-31 cs.CV cs.MM

classification cs.CVcs.MM
keywords videomultimodalvideosinterleavedmodelmovieseqchallengescontexts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Narrative videos, such as movies, pose significant challenges in video understanding due to their rich contexts (characters, dialogues, storylines) and diverse demands (identify who, relationship, and reason). In this paper, we introduce MovieSeq, a multimodal language model developed to address the wide range of challenges in understanding video contexts. Our core idea is to represent videos as interleaved multimodal sequences (including images, plots, videos, and subtitles), either by linking external knowledge databases or using offline models (such as whisper for subtitles). Through instruction-tuning, this approach empowers the language model to interact with videos using interleaved multimodal instructions. For example, instead of solely relying on video as input, we jointly provide character photos alongside their names and dialogues, allowing the model to associate these elements and generate more comprehensive responses. To demonstrate its effectiveness, we validate MovieSeq's performance on six datasets (LVU, MAD, Movienet, CMD, TVC, MovieQA) across five settings (video classification, audio description, video-text retrieval, video captioning, and video question-answering). The code will be public at https://github.com/showlab/MovieSeq.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation

    cs.CV 2024-11 conditional novelty 7.0 of 10

    MovieBench is the first public movie-level benchmark for long video generation, with hierarchical annotations, a character bank with audio, and new character-consistency metrics.

  2. DistinctAD: Distinctive Audio Description Generation in Contexts

    cs.CV 2024-11 conditional novelty 6.0 of 10

    DistinctAD improves automatic movie audio description quality and distinctiveness by adapting CLIP to movie data and using expectation-maximization attention plus a distinctive-word loss over consecutive clips.

Pith tools