Pith. sign in

REVIEW 8 cited by

MM-VID: Advancing Video Understanding with GPT-4V(ision)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.19773 v1 pith:FZSCVA4K submitted 2023-10-30 cs.CV

classification cs.CV
keywords videomm-vidgpt-4vunderstandingadvancedaudiocapabilitiescharacter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present MM-VID, an integrated system that harnesses the capabilities of GPT-4V, combined with specialized tools in vision, audio, and speech, to facilitate advanced video understanding. MM-VID is designed to address the challenges posed by long-form videos and intricate tasks such as reasoning within hour-long content and grasping storylines spanning multiple episodes. MM-VID uses a video-to-script generation with GPT-4V to transcribe multimodal elements into a long textual script. The generated script details character movements, actions, expressions, and dialogues, paving the way for large language models (LLMs) to achieve video understanding. This enables advanced capabilities, including audio description, character identification, and multimodal high-level comprehension. Experimental results demonstrate the effectiveness of MM-VID in handling distinct video genres with various video lengths. Additionally, we showcase its potential when applied to interactive environments, such as video games and graphic user interfaces.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    MESH, a three-layer video hallucination benchmark, shows LVMs ace basic objects and coarse traits but slip badly on fine character details and multi-subject actions in longer clips.

  2. Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Multimodal LLMs show systematic weaknesses in instance-level visual correspondence, and CoLVA, trained with a fine-grained vision expert and object-level contrastive learning, reaches 49.8% accuracy on the new MMVM be...

  3. DistinctAD: Distinctive Audio Description Generation in Contexts

    cs.CV 2024-11 conditional novelty 6.0 of 10

    DistinctAD improves automatic movie audio description quality and distinctiveness by adapting CLIP to movie data and using expectation-maximization attention plus a distinctive-word loss over consecutive clips.

  4. Generative Timelines for Instructed Visual Assembly

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A fine-tuned multimodal large language model that represents visual collections and timelines as token sequences can execute natural language timeline editing instructions more accurately than GPT-4o on synthetic benchmarks.

  5. Reason Before You Retrieve: Agentic Planning for Multi-modal RAG

    cs.AI 2026-06 reject novelty 5.0 of 10

    MM-R2 claims SOTA multimodal RAG accuracy on InfoSeek and Encyclopedic VQA via intent grounding plus a 10-topic KnowledgeMap, but its teacher trajectories leak the gold Wikipedia page and omit the image.

  6. VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos

    cs.IR 2025-02 conditional novelty 5.0 of 10

    VideoRAG combines graph-based text indexing with multimodal visual embeddings to answer questions across multi-hour video collections, supported by a new 134-hour benchmark.

  7. IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A question-adaptive visual compressor condenses each frame into 64 context tokens, improving long-term video QA accuracy while using fewer memory tokens than prior memory-augmented methods.

  8. How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A survey that categorizes pre-trained-model-based vision-language methods into four challenge-driven paradigms, with performance tables and a discussion of risks.

Pith tools