REVIEW 5 cited by
Learning Video Representations using Contrastive Bidirectional Transformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper proposes a self-supervised learning approach for video features that results in significantly improved performance on downstream tasks (such as video classification, captioning and segmentation) compared to existing methods. Our method extends the BERT model for text sequences to the case of sequences of real-valued feature vectors, by replacing the softmax loss with noise contrastive estimation (NCE). We also show how to learn representations from sequences of visual features and sequences of words derived from ASR (automatic speech recognition), and show that such cross-modal training (when possible) helps even more.
Forward citations
Cited by 5 Pith papers
-
CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models
CLIP-CC-Bench is a 200-clip benchmark with expert paragraph references that ranks 17 video-language models via an ensemble of five embedding-based semantic judges.
-
VL-BERT: Pre-training of Generic Visual-Linguistic Representations
VL-BERT pre-trains a single-stream Transformer on image captions and text, and the resulting representation improves VCR, VQA, and RefCOCO+ benchmarks.
-
Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
TimesCLIP aligns image-based and text-based views of the same time series via contrastive learning to improve forecasting accuracy on several benchmarks, but the full multimodal model is not used on two of the six lon...
-
Reinforcement Learning meets Masked Video Modeling : Trajectory-Guided Adaptive Token Selection
A trajectory-attention token sampler trained jointly with a masked video autoencoder via PPO improves downstream action recognition accuracy over random- and activity-based masking baselines.
-
Kronecker Mask and Interpretive Prompts are Language-Action Video Learners
CLAVER adds a cross-frame temporal attention mask (Kronecker mask) and LLM-generated interpretive action prompts to CLIP, improving video action recognition.
Discussion (0). Continue with ORCID to comment.