TeD-Loc improves weakly supervised object localization by distilling CLIP text embeddings to patch embeddings through contrastive alignment plus a localization-guided classifier and QR orthogonalization of text embeddings.
An image is worth 16x16 words, what is a video worth?
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 2verdicts
UNVERDICTED 2representative citing papers
TemporalVLM adds timestamp-aware clip encoding and BiLSTM global aggregation to video LLMs, introduces the IndustryASM factory dataset, and reports outperformance on dense captioning, temporal grounding, highlight detection, and action segmentation.
citing papers explorer
-
TeD-Loc: Text Distillation for Weakly Supervised Object Localization
TeD-Loc improves weakly supervised object localization by distilling CLIP text embeddings to patch embeddings through contrastive alignment plus a localization-guided classifier and QR orthogonalization of text embeddings.
-
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
TemporalVLM adds timestamp-aware clip encoding and BiLSTM global aggregation to video LLMs, introduces the IndustryASM factory dataset, and reports outperformance on dense captioning, temporal grounding, highlight detection, and action segmentation.