Pith. sign in

REVIEW 2 cited by

Open-Vocabulary Temporal Action Localization using Multimodal Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15556 v1 pith:AJH2OIWB submitted 2024-06-21 cs.CV

classification cs.CV
keywords categoriesactiontrainingnovelopen-vocabularylocalizationmodeltemporal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open-Vocabulary Temporal Action Localization (OVTAL) enables a model to recognize any desired action category in videos without the need to explicitly curate training data for all categories. However, this flexibility poses significant challenges, as the model must recognize not only the action categories seen during training but also novel categories specified at inference. Unlike standard temporal action localization, where training and test categories are predetermined, OVTAL requires understanding contextual cues that reveal the semantics of novel categories. To address these challenges, we introduce OVFormer, a novel open-vocabulary framework extending ActionFormer with three key contributions. First, we employ task-specific prompts as input to a large language model to obtain rich class-specific descriptions for action categories. Second, we introduce a cross-attention mechanism to learn the alignment between class representations and frame-level video features, facilitating the multimodal guided features. Third, we propose a two-stage training strategy which includes training with a larger vocabulary dataset and finetuning to downstream data to generalize to novel categories. OVFormer extends existing TAL methods to open-vocabulary settings. Comprehensive evaluations on the THUMOS14 and ActivityNet-1.3 benchmarks demonstrate the effectiveness of our method. Code and pretrained models will be publicly released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A multi-grained network that recognizes seen actions with a supervised classifier and unseen actions via video-level coarse filtering plus proposal-level matching achieves state-of-the-art open-vocabulary temporal act...

  2. GASP: A Gradient-Aware Shortest Path Algorithm for Boundary-Confined Visualization of 2-Manifold Reeb Graphs

    cs.GR 2025-08 unverdicted novelty 5.0 of 10

    GASP is a new algorithm that draws Reeb graphs of 2-manifold scalar fields so they hug the shape's boundary, stay compact, and align with the function's gradient, beating TTK's barycenter layout in evaluation.

Pith tools