REVIEW 11 cited by
LITA: Language Instructed Temporal-Localization Assistant
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
There has been tremendous progress in multimodal Large Language Models (LLMs). Recent works have extended these models to video input with promising instruction following capabilities. However, an important missing piece is temporal localization. These models cannot accurately answer the "When?" questions. We identify three key aspects that limit their temporal localization capabilities: (i) time representation, (ii) architecture, and (iii) data. We address these shortcomings by proposing Language Instructed Temporal-Localization Assistant (LITA) with the following features: (1) We introduce time tokens that encode timestamps relative to the video length to better represent time in videos. (2) We introduce SlowFast tokens in the architecture to capture temporal information at fine temporal resolution. (3) We emphasize temporal localization data for LITA. In addition to leveraging existing video datasets with timestamps, we propose a new task, Reasoning Temporal Localization (RTL), along with the dataset, ActivityNet-RTL, for learning and evaluating this task. Reasoning temporal localization requires both the reasoning and temporal localization of Video LLMs. LITA demonstrates strong performance on this challenging task, nearly doubling the temporal mean intersection-over-union (mIoU) of baselines. In addition, we show that our emphasis on temporal localization also substantially improves video-based text generation compared to existing Video LLMs, including a 36% relative improvement of Temporal Understanding. Code is available at: https://github.com/NVlabs/LITA
Forward citations
Cited by 11 Pith papers
-
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
TimeLens2 shows that a compact video MLLM can localize multiple evidence intervals in long videos by training on verified interval labels and a Wasserstein-based time-distance reward.
-
TimePLE: Rethinking Temporal Representation for Video Temporal Grounding
TimePLE predicts a whole video interval as a joint distribution over a position-duration square, rather than predicting start and end separately, and reports higher mIoU across four VTG benchmarks.
-
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
A new 28,497-sample dataset of 6DoF object manipulation trajectories is automatically extracted from egocentric video, and vision-language models are trained to generate these trajectories from action descriptions.
-
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
HandsOnVLM predicts future 2D hand trajectories from egocentric video and language instructions by treating trajectories as auto-regressive <HAND> tokens in a vision-language model.
-
Neptune: The Long Orbit to Benchmarking Long Video Understanding
Neptune is a 3,268-question, 2,405-video benchmark for long video understanding with a scalable LLM-based generation pipeline and an open-source answer-equivalence metric, GEM.
-
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding
Seq2Time improves video LLM temporal grounding by pretraining on image and clip sequences with a unified relative position token, yielding higher F1 and CIDEr on YouCook2 and higher recall on Charades-STA.
-
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries
ShotVL, a fine-tuned InternVL, improves zero-shot highlight-frame retrieval on the new BestShot benchmark, but its zero-shot claim is weakened by a small benchmark-derived training set.
-
TimeRefine: Temporal Grounding with Time Refining Video LLM
A training reformulation that turns timestamp prediction into iterative offset refinement, with an auxiliary L1 loss, improves Video LLM temporal grounding.
-
Foundation Models and Adaptive Feature Selection: A Synergistic Approach to Video Question Answering
A VideoQA pipeline using question-guided frame selection, MiniGPT-4 object grounding, and local-global graph fusion reports improved accuracy on five of six benchmarks, with one dataset contradicting the all-task claim.
-
Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation
MedRegion-CT integrates region-representative tokens, mask-driven segmentation tokens, and patient-specific attribute prompts into a multimodal LLM, reporting state-of-the-art scores on RadGenome-Chest CT report generation.
- Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
Discussion (0). Continue with ORCID to comment.