Pith. sign in

REVIEW 2 cited by

GazeHTA: End-to-end Gaze Target Detection with Head-Target Association

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.10718 v3 pith:BU4LAQPX submitted 2024-04-16 cs.CV

classification cs.CV
keywords gazetargetdetectiongazehtaheadend-to-endhead-targetheads
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Precisely detecting which object a person is paying attention to is critical for human-robot interaction since it provides important cues for the next action from the human user. We propose an end-to-end approach for gaze target detection: predicting a head-target connection between individuals and the target image regions they are looking at. Most of the existing methods use independent components such as off-the-shelf head detectors or have problems in establishing associations between heads and gaze targets. In contrast, we investigate an end-to-end multi-person Gaze target detection framework with Heads and Targets Association (GazeHTA), which predicts multiple head-target instances based solely on input scene image. GazeHTA addresses challenges in gaze target detection by (1) leveraging a pre-trained diffusion model to extract scene features for rich semantic understanding, (2) re-injecting a head feature to enhance the head priors for improved head understanding, and (3) learning a connection map as the explicit visual associations between heads and gaze targets. Our extensive experimental results demonstrate that GazeHTA outperforms state-of-the-art gaze target detection methods and two adapted diffusion-based baselines on two standard datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GazeDETR: Gaze Detection using Disentangled Head and Gaze Representations

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    GazeDETR uses two disentangled decoders for head localization and gaze prediction, achieving state-of-the-art results on GazeFollow, VideoAttentionTarget, and ChildPlay.

  2. UniGaze: Towards Universal Gaze Estimation via Large-scale Pre-Training

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Curated MAE pre-training on normalized, pose-balanced face images improves gaze estimation generalization across datasets, outperforming semantic pre-training and prior domain-generalization methods.

Pith tools