REVIEW 5 cited by
State-of-the-Art in Human Scanpath Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The last years have seen a surge in models predicting the scanpaths of fixations made by humans when viewing images. However, the field is lacking a principled comparison of those models with respect to their predictive power. In the past, models have usually been evaluated based on comparing human scanpaths to scanpaths generated from the model. Here, instead we evaluate models based on how well they predict each fixation in a scanpath given the previous scanpath history. This makes model evaluation closely aligned with the biological processes thought to underly scanpath generation and allows to apply established saliency metrics like AUC and NSS in an intuitive and interpretable way. We evaluate many existing models of scanpath prediction on the datasets MIT1003, MIT300, CAT2000 train and CAT200 test, for the first time giving a detailed picture of the current state of the art of human scanpath prediction. We also show that the discussed method of model benchmarking allows for more detailed analyses leading to interesting insights about where and when models fail to predict human behaviour. The MIT/Tuebingen Saliency Benchmark will implement the evaluation of scanpath models as detailed here, allowing researchers to score their models on the established benchmark datasets MIT300 and CAT2000.
Forward citations
Cited by 5 Pith papers
-
Follow My Eyes: Backdoor Attacks on Goal-Directed Scanpath Prediction
Scene-conditioned spatial-misdirection and duration-inflation backdoors succeed at 2.5–10% poison ratios on multimodal scanpath predictors and resist five adapted defenses.
-
Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction
ScanDiff generates diverse, text-conditioned gaze scanpaths with a diffusion-ViT architecture and reports state-of-the-art results on three benchmarks.
-
Baseline behaviour in human vision
A memoryless Markov model based only on the previous saccade's length and angle outperforms a uniform random model across 36 eye-tracking datasets, which the authors interpret as evidence for a context-invariant motor prior.
-
Unified Attention Modeling for Efficient Free-Viewing and Visual Search via Shared Representations
Free-viewing attention features transfer to visual search with a 3.86% SemSS drop and large compute savings, supporting a shared representation between the two tasks.
-
Human Scanpath Prediction in Target-Present Visual Search with Semantic-Foveal Bayesian Attention
A semantic-foveal Bayesian model that uses YOLOv5 detections and no human scanpath training generates target-present search fixations that outperform IVSN and approach gaze-trained models on several COCO-Search18 metrics.
Discussion (0). Continue with ORCID to comment.