BabyCL learns word-referent mappings from egocentric video in a single chronological pass via streaming visual learning, dual replay, and three contrastive losses, outperforming streaming baselines on the SAYCam 4AFC benchmark.
InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9777–9786
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
EgoEverything is a 5,000+ question benchmark that uses gaze as a proxy for human attention to generate long-context egocentric video QA pairs, where the best VLM scores 63.1% vs humans at 83.5%.
citing papers explorer
-
Continual Visual and Verbal Learning Through a Child's Egocentric Input
BabyCL learns word-referent mappings from egocentric video in a single chronological pass via streaming visual learning, dual replay, and three contrastive losses, outperforming streaming baselines on the SAYCam 4AFC benchmark.
-
EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment
EgoEverything is a 5,000+ question benchmark that uses gaze as a proxy for human attention to generate long-context egocentric video QA pairs, where the best VLM scores 63.1% vs humans at 83.5%.