Text-only contrastive fine-tuning of an MLLM with hard negatives produces embeddings that handle temporal, negation, and multimodal nuances in video retrieval and achieves SOTA performance.
hub
Action- CLIP: A New Paradigm for Video Action Recognition
15 Pith papers cite this work. Polarity classification is still indexing.
hub tools
representative citing papers
M2R2 proposes a multimodal robotic representation for temporal action segmentation that combines proprioceptive and exteroceptive sensors with a novel training strategy enabling feature reuse across models, achieving new state-of-the-art results on three robotic datasets.
ICMPG combines LLM-based candidate generation with MPC-style physical simulation and semantic scoring to produce text-driven human motions that are both plausible and faithful.
Derives tractable optimal fair multi-class classifier and supplies in-processing and post-processing algorithms that converge to the accuracy-fairness Pareto frontier.
TAG-Head adds a lightweight transformer-plus-time-aligned-graph head to existing 3D backbones to reach new RGB-only state-of-the-art on FineGym and HAA500 while surpassing some multimodal baselines.
A three-stage pipeline uses few-shot VLM action parsing, sliding-window segmentation, and LLM sequence classification with peer context to measure student engagement from classroom videos.
A skeleton-based zero-shot VAD method distills LLM knowledge for action typicality during training and performs test-time context uniqueness analysis to derive scene-adaptive normality boundaries, claiming SOTA results on four datasets with over 100 unseen scenes.
TRUST is a scan-aware parameter-efficient image-to-video transfer learning framework for abdominal ultrasound trauma recognition featuring CFCA, MGMA, and VQSA modules that outperforms SOTA by 9.63% on in-house datasets.
TACO proposes Relative Structure Distillation and a lightweight specialization projection to mitigate inconsistency between fine-tuning and evaluation objectives in open-vocabulary video recognition, claiming state-of-the-art results on cross-dataset and base-to-novel benchmarks.
VidPrism introduces a heterogeneous temporal MoE with content-aware multi-rate sampling and bidirectional fusion for image-to-video transfer, claiming SOTA results on video benchmarks.
SimVA constructs a 4D similarity volume over video tokens and action classes then applies spatial, motion-aware, and Mamba-based temporal aggregation to achieve competitive zero-shot and few-shot performance on open-vocabulary action recognition benchmarks.
InternVideo combines masked video modeling and video-language contrastive learning into a single foundation model that reaches state-of-the-art results on 39 video datasets including 91.1% top-1 on Kinetics-400.
A rationale-informed pruning strategy for VLMs yields higher accuracy and more doubly-correct predictions than prior pruning methods on egocentric video benchmarks.
EV-CLIP introduces mask and context visual prompts to adapt CLIP for improved few-shot video action recognition under visual challenges such as low light and egocentric views, outperforming other efficient methods with backbone-scale-independent efficiency.
citing papers explorer
-
Adapting MLLMs for Nuanced Video Retrieval
Text-only contrastive fine-tuning of an MLLM with hard negatives produces embeddings that handle temporal, negation, and multimodal nuances in video retrieval and achieves SOTA performance.
-
M2R2: MultiModal Robotic Representation for Temporal Action Segmentation
M2R2 proposes a multimodal robotic representation for temporal action segmentation that combines proprioceptive and exteroceptive sensors with a novel training strategy enabling feature reuse across models, achieving new state-of-the-art results on three robotic datasets.
-
In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics
ICMPG combines LLM-based candidate generation with MPC-style physical simulation and semantic scoring to produce text-driven human motions that are both plausible and faithful.
-
Demystifying the Optimal Fair Classifier in Multi-Class Classification
Derives tractable optimal fair multi-class classifier and supplies in-processing and post-processing algorithms that converge to the accuracy-fairness Pareto frontier.
-
TAG-Head: Time-Aligned Graph Head for Plug-and-Play Fine-grained Action Recognition
TAG-Head adds a lightweight transformer-plus-time-aligned-graph head to existing 3D backbones to reach new RGB-only state-of-the-art on FineGym and HAA500 while surpassing some multimodal baselines.
-
Context Matters: Peer-Aware Student Behavioral Engagement Measurement via VLM Action Parsing and LLM Sequence Classification
A three-stage pipeline uses few-shot VLM action parsing, sliding-window segmentation, and LLM sequence classification with peer context to measure student engagement from classroom videos.
-
Action Hints: Semantic Typicality and Context Uniqueness for Generalizable Skeleton-based Video Anomaly Detection
A skeleton-based zero-shot VAD method distills LLM knowledge for action typicality during training and performs test-time context uniqueness analysis to derive scene-adaptive normality boundaries, claiming SOTA results on four datasets with over 100 unseen scenes.
-
TRUST: Efficient Abdominal Trauma Recognition via Image-to-Ultrasound-Video Transfer Learning
TRUST is a scan-aware parameter-efficient image-to-video transfer learning framework for abdominal ultrasound trauma recognition featuring CFCA, MGMA, and VQSA modules that outperforms SOTA by 9.63% on in-house datasets.
-
TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition
TACO proposes Relative Structure Distillation and a lightweight specialization projection to mitigate inconsistency between fine-tuning and evaluation objectives in open-vocabulary video recognition, claiming state-of-the-art results on cross-dataset and base-to-novel benchmarks.
-
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer
VidPrism introduces a heterogeneous temporal MoE with content-aware multi-rate sampling and bidirectional fusion for image-to-video transfer, claiming SOTA results on video benchmarks.
-
Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition
SimVA constructs a 4D similarity volume over video tokens and action classes then applies spatial, motion-aware, and Mamba-based temporal aggregation to achieve competitive zero-shot and few-shot performance on open-vocabulary action recognition benchmarks.
-
InternVideo: General Video Foundation Models via Generative and Discriminative Learning
InternVideo combines masked video modeling and video-language contrastive learning into a single foundation model that reaches state-of-the-art results on 39 video datasets including 91.1% top-1 on Kinetics-400.
-
Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding
A rationale-informed pruning strategy for VLMs yields higher accuracy and more doubly-correct predictions than prior pruning methods on egocentric video benchmarks.
-
EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges
EV-CLIP introduces mask and context visual prompts to adapt CLIP for improved few-shot video action recognition under visual challenges such as low light and egocentric views, outperforming other efficient methods with backbone-scale-independent efficiency.
- SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels