REVIEW 8 cited by
MM-VID: Advancing Video Understanding with GPT-4V(ision)
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present MM-VID, an integrated system that harnesses the capabilities of GPT-4V, combined with specialized tools in vision, audio, and speech, to facilitate advanced video understanding. MM-VID is designed to address the challenges posed by long-form videos and intricate tasks such as reasoning within hour-long content and grasping storylines spanning multiple episodes. MM-VID uses a video-to-script generation with GPT-4V to transcribe multimodal elements into a long textual script. The generated script details character movements, actions, expressions, and dialogues, paving the way for large language models (LLMs) to achieve video understanding. This enables advanced capabilities, including audio description, character identification, and multimodal high-level comprehension. Experimental results demonstrate the effectiveness of MM-VID in handling distinct video genres with various video lengths. Additionally, we showcase its potential when applied to interactive environments, such as video games and graphic user interfaces.
Forward citations
Cited by 8 Pith papers
-
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
MESH, a three-layer video hallucination benchmark, shows LVMs ace basic objects and coarse traits but slip badly on fine character details and multi-subject actions in longer clips.
-
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
Multimodal LLMs show systematic weaknesses in instance-level visual correspondence, and CoLVA, trained with a fine-grained vision expert and object-level contrastive learning, reaches 49.8% accuracy on the new MMVM be...
-
DistinctAD: Distinctive Audio Description Generation in Contexts
DistinctAD improves automatic movie audio description quality and distinctiveness by adapting CLIP to movie data and using expectation-maximization attention plus a distinctive-word loss over consecutive clips.
-
Generative Timelines for Instructed Visual Assembly
A fine-tuned multimodal large language model that represents visual collections and timelines as token sequences can execute natural language timeline editing instructions more accurately than GPT-4o on synthetic benchmarks.
-
Reason Before You Retrieve: Agentic Planning for Multi-modal RAG
MM-R2 claims SOTA multimodal RAG accuracy on InfoSeek and Encyclopedic VQA via intent grounding plus a 10-topic KnowledgeMap, but its teacher trajectories leak the gold Wikipedia page and omit the image.
-
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
VideoRAG combines graph-based text indexing with multimodal visual embeddings to answer questions across multi-hour video collections, supported by a new 134-hour benchmark.
-
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs
A question-adaptive visual compressor condenses each frame into 64 context tokens, improving long-term video QA accuracy while using fewer memory tokens than prior memory-augmented methods.
-
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
A survey that categorizes pre-trained-model-based vision-language methods into four challenge-driven paradigms, with performance tables and a discussion of risks.
Discussion (0). Continue with ORCID to comment.