Hand trajectory encoding fused with video-text features via cross-attention improves Ego4D NLQ grounding performance, with largest gains on hand-object interaction and quantity/state queries.
ObjectNLQ@ Ego4D episodic mem- ory challenge 2024
3 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 3years
2026 3representative citing papers
A hybrid pipeline of OSGNet candidate generation followed by MLLM reranking secured first place in both the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge.
MARS converts long videos to captions and summaries, maintains modality-specific memories, and deploys an agent to select evidence or answer, placing second on the CASTLE Challenge leaderboard.
citing papers explorer
-
Hand Trajectory Fusion for Egocentric Natural Language Query Grounding
Hand trajectory encoding fused with video-text features via cross-attention improves Ego4D NLQ grounding performance, with largest gains on hand-object interaction and quantity/state queries.
-
OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026
A hybrid pipeline of OSGNet candidate generation followed by MLLM reranking secured first place in both the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge.
-
MARS: Technical Report for the CASTLE Challenge at EgoVis 2026
MARS converts long videos to captions and summaries, maintains modality-specific memories, and deploys an agent to select evidence or answer, placing second on the CASTLE Challenge leaderboard.