REVIEW 34 cited by
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper introduces MiniGPT4-Video, a multimodal Large Language Model (LLM) designed specifically for video understanding. The model is capable of processing both temporal visual and textual data, making it adept at understanding the complexities of videos. Building upon the success of MiniGPT-v2, which excelled in translating visual features into the LLM space for single images and achieved impressive results on various image-text benchmarks, this paper extends the model's capabilities to process a sequence of frames, enabling it to comprehend videos. MiniGPT4-video does not only consider visual content but also incorporates textual conversations, allowing the model to effectively answer queries involving both visual and text components. The proposed model outperforms existing state-of-the-art methods, registering gains of 4.22%, 1.13%, 20.82%, and 13.1% on the MSVD, MSRVTT, TGIF, and TVQA benchmarks respectively. Our models and code have been made publicly available here https://vision-cair.github.io/MiniGPT4-video/
Forward citations
Cited by 34 Pith papers
-
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Consistent object IDs across video frames and a reconstructed bird's-eye view let large VLMs, and a fine-tuned 7B model, reach state-of-the-art results on ScanNet 3D question answering, captioning, and grounding.
-
SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs
SignLlama uses filtered pseudo-gloss pretraining and visual-prioritized distillation to make Llama-based sign language translation competitive on four benchmarks.
-
Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs
Coordinates global and local evidence views from a temporal hierarchy, with verification-guided routing, to improve long-video multiple-choice QA.
-
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models
GSTEP prunes visual tokens in VideoLLMs using a global spatio-temporal density score with farthest point sampling, retaining close to full benchmark accuracy at 75 to 90 percent token pruning.
-
ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding
Training-free latent-object memory anchors let frozen Video-LLMs retain object histories under a tight token budget and improve streaming and long-video QA.
-
Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding
Sign Language QA benchmarks are introduced from PHOENIX14T and CSL-Daily via template-generated questions, and a question-conditioned baseline outperforms video-language and cascaded baselines.
-
TimeThink: Reasoning with Time for Video LLMs
TimeThink adds step-wise temporal process rewards (max IoU of referenced intervals) to GRPO for Video-LLMs, improving grounding and reasoning over outcome-only RL baselines.
-
VUDG: A Dataset for Video Understanding Domain Generalization
VUDG is a domain-generalization benchmark for video understanding with 11 domains and 36,388 QA pairs, and it shows that current large video-language models lose accuracy across visual domains.
-
EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language
EmoSign is a 200-clip American Sign Language video dataset with native-signer sentiment and emotion labels plus baseline multimodal LLM results showing poor visual-only emotion recognition.
-
Domain Adaptation of VLM for Soccer Video Understanding
A three-stage curriculum (concept alignment, instruction tuning, downstream fine-tuning) adapted LLaVA-NeXT-Video to soccer, raising action classification accuracy from 11.8% to 63.5% and the VQA relative score from a...
-
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
A streaming video assistant drops temporally redundant tokens (over 80%) with minimal accuracy loss and uses the drop-ratio curve to trigger proactive responses at scene changes.
-
EventVAD: Training-Free Event-Aware Video Anomaly Detection
EventVAD improves training-free video anomaly detection by detecting event boundaries from CLIP and RAFT features and feeding coherent event segments to a 7B multimodal LLM, achieving 82.03 AUC on UCF-Crime and 64.04 ...
-
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
A graph-guided self-training method lets video-language models generate their own reasoning training data from raw videos, improving multi-step compositional reasoning accuracy.
-
VideoOrion: Tokenizing Object Dynamics in Videos
Encoding video as a small set of object tokens, produced by off-the-shelf detection, segmentation, and tracking models, improves video QA accuracy and enables video-based referring in a 7B video-LLM.
-
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.
-
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
LGSRR uses LLM-generated semantic descriptions and rankings to improve multimodal intent recognition, reporting SOTA results on MIntRec2.0 and IEMOCAP-DA with gains around 0.5-1.3%.
-
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
Replacing raw video and audio features with MLLM-generated natural-language captions improves hit rate and nDCG for two-tower and SASRec recommenders on MicroLens-100K.
-
Multi-modal brain encoding models for multi-modal stimuli
On movie-watching fMRI data, multi-modal vision-audio transformers predict brain activity better than unimodal video or speech models, with video dominating cross-modal alignment and video plus audio jointly contribut...
-
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
A chain-of-thought fine-tuned vision-language model can infer audio descriptions from silent videos, and using those descriptions as prompts improves video-to-audio generation.
-
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
Adding a trainable convolutional temporal encoder to LLaVA-style video models improves benchmark scores and allows heavy frame compression, but the controlled evidence for the causal claim is weak.
-
TinyLLaVA-Video: Towards Smaller LMMs for Video Understanding with Group Resampler
A small video model with a group resampler reduces video input to a few hundred tokens and beats several 7B models on common video benchmarks.
-
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
A single 3B-parameter end-to-end model with object-aware video perceiving and multi-granularity text fusion reports SOTA results across four instructed visual segmentation tasks.
-
LinVT: Empower Your Image-level Large Language Model to Understand Videos
A plug-and-play linear video tokenizer converts existing image-based LLMs into video-understanding LLMs by condensing frames into weighted-average tokens while preserving image capabilities.
-
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
DynFocus dynamically allocates a few tokens to selected frames and two tokens to the rest, reporting competitive video QA accuracy with lower token budgets.
-
MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models
Training a multimodal LLM on mostly text-only instructions with a small vision-language tail matches or beats vision-heavy instruction tuning on held-out text and vision tasks at about half the token cost.
-
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding
MCAM is a video captioning model combining 3DResNet and VidSwin features with a graph-inspired fusion module, reporting mixed gains on BDD-X and CoVLA but failing to implement the promised causal reasoning.
-
ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing
ViFusion combines dynamic tensor fusion with hierarchical AllReduce to speed up distributed video feature indexing, but the 8-22x throughput claim is an overstatement of bandwidth gains over a self-defined baseline.
-
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
LeanPO improves Video-LLM alignment by using a reference-free average-likelihood reward, self-generated winning/losing pairs, and dynamic label smoothing, yielding gains on six video benchmarks.
-
From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control
A new 124K-clip dataset with hierarchical text annotations, plus a pipeline that couples an LLM planner, a text-to-pose VAE, diffusion in-betweening, and physics control to generate long-horizon human behaviors.
-
Vision Generalist Model: A Survey
A structured review of vision generalist models, classifying them into encoding-based and sequence-to-sequence frameworks and summarizing datasets, benchmarks, techniques, and open problems.
-
PEFT A2Z: Parameter-Efficient Fine-Tuning Survey for Large Language and Vision Models
A survey that organizes PEFT methods into additive, selective, reparameterized, hybrid, and unified families, but with no new method or verified experiments.
-
Visual Large Language Models for Generalized and Specialized Applications
This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.
-
Do Language Models Understand Time?
A survey arguing that video-LLMs rely on pretrained encoders and short-biased datasets, leaving them weak at long-term temporal reasoning such as causality and event progression.
-
PolySmart @ TRECVid 2024 Video Captioning (VTT)
Fine-tuning LLaVA on prior TRECVid video-text pairs improves scores on the VTT24 captioning test set for BLEU, METEOR, CIDEr, and CIDEr-D, while SPICE and STS results are mixed.
Discussion (0). Continue with ORCID to comment.