A multi-view model with head aggregation, uncertainty-based gaze selection, and epipolar scene attention outperforms single-view gaze target estimation and enables cross-view prediction.
Multimae: Multi-modal multi-task masked autoen- coders
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Multi-view Gaze Target Estimation
A multi-view model with head aggregation, uncertainty-based gaze selection, and epipolar scene attention outperforms single-view gaze target estimation and enables cross-view prediction.