SoccerLens benchmark shows state-of-the-art soccer VLMs achieve high classification accuracy yet fail to exceed 50% visual grounding performance and underutilize temporal information.
Clip surgery for better explainability with enhancement in open- vocabulary tasks
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Visual encoders leak identity information; a one-shot linear subspace removal method (ISP) reduces leakage to near-chance levels while retaining high non-biometric utility across datasets.
GCLIP improves training-free CLIP semantic segmentation by fusing intermediate attention maps and suppressing an abnormal FFN channel.
citing papers explorer
-
SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy
SoccerLens benchmark shows state-of-the-art soccer VLMs achieve high classification accuracy yet fail to exceed 50% visual grounding performance and underutilize temporal information.
-
From Measurement to Mitigation: Quantifying and Reducing Identity Leakage in Image Representation Encoders with Linear Subspace Removal
Visual encoders leak identity information; a one-shot linear subspace removal method (ISP) reduces leakage to near-chance levels while retaining high non-biometric utility across datasets.
-
Rethinking the Global Knowledge of CLIP in Training-Free Open-Vocabulary Semantic Segmentation
GCLIP improves training-free CLIP semantic segmentation by fusing intermediate attention maps and suppressing an abnormal FFN channel.
- Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP