REVIEW 5 cited by
Egocentric Video-Language Pretraining
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person video-text datasets, such as HowTo100M. In this work, we exploit the recently released Ego4D dataset to pioneer Egocentric VLP along three directions. (i) We create EgoClip, a 1st-person video-text pretraining dataset comprising 3.8M clip-text pairs well-chosen from Ego4D, covering a large variety of human daily activities. (ii) We propose a novel pretraining objective, dubbed EgoNCE, which adapts video-text contrastive learning to the egocentric domain by mining egocentric-aware positive and negative samples. (iii) We introduce EgoMCQ, a development benchmark that is close to EgoClip and hence can support effective validation and fast exploration of our design decisions in EgoClip and EgoNCE. Furthermore, we demonstrate strong performance on five egocentric downstream tasks across three datasets: video-text retrieval on EPIC-KITCHENS-100; action recognition on Charades-Ego; natural language query, moment query, and object state change classification on Ego4D challenge benchmarks. The dataset and code are available at https://github.com/showlab/EgoVLP.
Forward citations
Cited by 5 Pith papers
-
Improving Keystep Recognition in Ego-Video via Dexterous Focus
Hand-focused, stabilized cropping of ego-video improves Ego-Exo4D fine-grained keystep recognition accuracy from 39.18% to 45.81% (hands only) and 47.75% (hands plus ego).
-
DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos
DeCafNet reduces long-video temporal grounding cost by up to 47 percent while improving accuracy, using a lightweight sidekick encoder to select salient clips for a heavy expert encoder.
-
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
EAIL localizes a person in a 3D point cloud from head-mounted IMU signals by aligning short action segments with scene locations using vision-language training guidance.
-
Volume-Distance-Ratio Asymptote and Spacetime Inextendibility for FLRW Spacetimes
The paper derives conditions for past inextendibility of FLRW spacetimes with power-law scale factors using volume-distance-ratio asymptote criteria.
-
A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition
ProbRes uses a knowledge-guided stochastic search over activity labels to reduce VLM queries while matching or improving egocentric activity recognition accuracy.
Discussion (0). Sign in to comment.