Pith. sign in

hub

Action- CLIP: A New Paradigm for Video Action Recognition

15 Pith papers cite this work. Polarity classification is still indexing.

15 Pith papers citing it

hub tools

representative citing papers

Adapting MLLMs for Nuanced Video Retrieval

cs.CV · 2025-12-15 · unverdicted · novelty 7.0

Text-only contrastive fine-tuning of an MLLM with hard negatives produces embeddings that handle temporal, negation, and multimodal nuances in video retrieval and achieves SOTA performance.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation

cs.RO · 2025-04-25 · unverdicted · novelty 7.0

M2R2 proposes a multimodal robotic representation for temporal action segmentation that combines proprioceptive and exteroceptive sensors with a novel training strategy enabling feature reuse across models, achieving new state-of-the-art results on three robotic datasets.

TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition

cs.CV · 2026-06-24 · unverdicted · novelty 5.0 · 2 refs

TACO proposes Relative Structure Distillation and a lightweight specialization projection to mitigate inconsistency between fine-tuning and evaluation objectives in open-vocabulary video recognition, claiming state-of-the-art results on cross-dataset and base-to-novel benchmarks.

citing papers explorer

Showing 15 of 15 citing papers.