Pith. sign in

REVIEW 1 cited by

Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.16484 v2 pith:64IFJ7Y2 submitted 2023-11-26 cs.CV

classification cs.CV
keywords attentionvideomodelhumanmemorabilitygazegreaterhumans
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Understanding what makes a video memorable has important applications in advertising or education technology. Towards this goal, we investigate spatio-temporal attention mechanisms underlying video memorability. Different from previous works that fuse multiple features, we adopt a simple CNN+Transformer architecture that enables analysis of spatio-temporal attention while matching state-of-the-art (SoTA) performance on video memorability prediction. We compare model attention against human gaze fixations collected through a small-scale eye-tracking study where humans perform the video memory task. We uncover the following insights: (i) Quantitative saliency metrics show that our model, trained only to predict a memorability score, exhibits similar spatial attention patterns to human gaze, especially for more memorable videos. (ii) The model assigns greater importance to initial frames in a video, mimicking human attention patterns. (iii) Panoptic segmentation reveals that both (model and humans) assign a greater share of attention to things and less attention to stuff as compared to their occurrence probability.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. How to Take a Memorable Picture? Empowering Users with Actionable Feedback

    cs.CV 2026-02 conditional novelty 6.0 of 10

    MemCoach uses contrastive activation steering to make MLLMs suggest memorability-improving photo changes, outperforming zero-shot baselines on the new MemBench benchmark.

Pith tools