Pith. sign in

REVIEW 1 cited by

UniAR: A Unified model for predicting human Attention and Responses on visual content

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.10175 v3 pith:D2GDVPGV submitted 2023-12-15 cs.CV

classification cs.CV
keywords behaviorhumanattentioncontentuniarvisualmodelingacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Progress in human behavior modeling involves understanding both implicit, early-stage perceptual behavior, such as human attention, and explicit, later-stage behavior, such as subjective preferences or likes. Yet most prior research has focused on modeling implicit and explicit human behavior in isolation; and often limited to a specific type of visual content. We propose UniAR -- a unified model of human attention and preference behavior across diverse visual content. UniAR leverages a multimodal transformer to predict subjective feedback, such as satisfaction or aesthetic quality, along with the underlying human attention or interaction heatmaps and viewing order. We train UniAR on diverse public datasets spanning natural images, webpages, and graphic designs, and achieve SOTA performance on multiple benchmarks across various image domains and behavior modeling tasks. Potential applications include providing instant feedback on the effectiveness of UIs/visual content, and enabling designers and content-creation models to optimize their creation for human-centric improvements.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interpreting Radiologist's Intention from Eye Movements in Chest X-ray Diagnosis

    cs.CV 2025-07 reject novelty 6.0 of 10

    RadGazeIntent, a transformer model, predicts per-fixation diagnostic intention from radiologist gaze on chest X-rays, evaluated on three newly constructed intention-labeled datasets.

Pith tools