Pith. sign in

REVIEW 11 cited by

LITA: Language Instructed Temporal-Localization Assistant

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.19046 v1 pith:TY5U7XPG submitted 2024-03-27 cs.CV cs.AI

classification cs.CVcs.AI
keywords temporallocalizationlitavideolanguagellmsmodelsreasoning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There has been tremendous progress in multimodal Large Language Models (LLMs). Recent works have extended these models to video input with promising instruction following capabilities. However, an important missing piece is temporal localization. These models cannot accurately answer the "When?" questions. We identify three key aspects that limit their temporal localization capabilities: (i) time representation, (ii) architecture, and (iii) data. We address these shortcomings by proposing Language Instructed Temporal-Localization Assistant (LITA) with the following features: (1) We introduce time tokens that encode timestamps relative to the video length to better represent time in videos. (2) We introduce SlowFast tokens in the architecture to capture temporal information at fine temporal resolution. (3) We emphasize temporal localization data for LITA. In addition to leveraging existing video datasets with timestamps, we propose a new task, Reasoning Temporal Localization (RTL), along with the dataset, ActivityNet-RTL, for learning and evaluating this task. Reasoning temporal localization requires both the reasoning and temporal localization of Video LLMs. LITA demonstrates strong performance on this challenging task, nearly doubling the temporal mean intersection-over-union (mIoU) of baselines. In addition, we show that our emphasis on temporal localization also substantially improves video-based text generation compared to existing Video LLMs, including a 36% relative improvement of Temporal Understanding. Code is available at: https://github.com/NVlabs/LITA

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

    cs.CV 2026-07 conditional novelty 7.0 of 10

    TimeLens2 shows that a compact video MLLM can localize multiple evidence intervals in long videos by training on verified interval labels and a Wasserstein-based time-distance reward.

  2. TimePLE: Rethinking Temporal Representation for Video Temporal Grounding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    TimePLE predicts a whole video interval as a joint distribution over a position-duration square, rather than predicting start and end separately, and reports higher mIoU across four VTG benchmarks.

  3. Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new 28,497-sample dataset of 6DoF object manipulation trajectories is automatically extracted from egocentric video, and vision-language models are trained to generate these trajectories from action descriptions.

  4. HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction

    cs.CV 2024-12 conditional novelty 6.0 of 10

    HandsOnVLM predicts future 2D hand trajectories from egocentric video and language instructions by treating trajectories as auto-regressive <HAND> tokens in a vision-language model.

  5. Neptune: The Long Orbit to Benchmarking Long Video Understanding

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Neptune is a 3,268-question, 2,405-video benchmark for long video understanding with a scalable LLM-based generation pipeline and an open-source answer-equivalence metric, GEM.

  6. Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Seq2Time improves video LLM temporal grounding by pretraining on image and clip sequences with a unified relative position token, yielding higher F1 and CIDEr on YouCook2 and higher recall on Charades-STA.

  7. ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries

    cs.CV 2024-12 conditional novelty 5.0 of 10

    ShotVL, a fine-tuned InternVL, improves zero-shot highlight-frame retrieval on the new BestShot benchmark, but its zero-shot claim is weakened by a small benchmark-derived training set.

  8. TimeRefine: Temporal Grounding with Time Refining Video LLM

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A training reformulation that turns timestamp prediction into iterative offset refinement, with an auxiliary L1 loss, improves Video LLM temporal grounding.

  9. Foundation Models and Adaptive Feature Selection: A Synergistic Approach to Video Question Answering

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A VideoQA pipeline using question-guided frame selection, MiniGPT-4 object grounding, and local-global graph fusion reports improved accuracy on five of six benchmarks, with one dataset contradicting the all-task claim.

  10. Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation

    eess.IV 2025-06 conditional novelty 4.0 of 10

    MedRegion-CT integrates region-representative tokens, mask-driven segmentation tokens, and patient-specific attribute prompts into a multimodal LLM, reporting state-of-the-art scores on RadGenome-Chest CT report generation.

  11. Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs

    cs.CV 2025-01

Pith tools