Hand-object masked training and an HOI-dynamics-aware decoder yield more balanced cue-specific action recognition on a new inpainted DEHOI testbed and transfer to object-state and robot-manipulation tasks.
International Journal of Computer Vision130(1), 33–55 (2022)
4 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 4years
2026 4representative citing papers
OnPoint enables point-supervised online temporal action localization by distilling pseudo-segments, class-activation sequences, and anticipatory windows from an offline teacher to an online student.
GEST-Engine turns game engines into zero-cost dense ground-truth video generators; GTASA reveals frozen video encoders fail inter-entity spatial relation probes.
EgoSelf uses graph-based memory of user interactions to derive personalized profiles and predict future behaviors for egocentric assistants.
citing papers explorer
-
Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?
Hand-object masked training and an HOI-dynamics-aware decoder yield more balanced cue-specific action recognition on a new inpainted DEHOI testbed and transfer to object-state and robot-manipulation tasks.
-
OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization
OnPoint enables point-supervised online temporal action localization by distilling pseudo-segments, class-activation sequences, and anticipatory windows from an offline teacher to an online student.
-
GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models
GEST-Engine turns game engines into zero-cost dense ground-truth video generators; GTASA reveals frozen video encoders fail inter-entity spatial relation probes.
-
EgoSelf: From Memory to Personalized Egocentric Assistant
EgoSelf uses graph-based memory of user interactions to derive personalized profiles and predict future behaviors for egocentric assistants.