EgoEverything is a 5,000+ question benchmark that uses gaze as a proxy for human attention to generate long-context egocentric video QA pairs, where the best VLM scores 63.1% vs humans at 83.5%.
In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment
EgoEverything is a 5,000+ question benchmark that uses gaze as a proxy for human attention to generate long-context egocentric video QA pairs, where the best VLM scores 63.1% vs humans at 83.5%.