REVIEW 2 cited by
Human Attention in Fine-grained Classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The way humans attend to, process and classify a given image has the potential to vastly benefit the performance of deep learning models. Exploiting where humans are focusing can rectify models when they are deviating from essential features for correct decisions. To validate that human attention contains valuable information for decision-making processes such as fine-grained classification, we compare human attention and model explanations in discovering important features. Towards this goal, we collect human gaze data for the fine-grained classification dataset CUB and build a dataset named CUB-GHA (Gaze-based Human Attention). Furthermore, we propose the Gaze Augmentation Training (GAT) and Knowledge Fusion Network (KFN) to integrate human gaze knowledge into classification models. We implement our proposals in CUB-GHA and the recently released medical dataset CXR-Eye of chest X-ray images, which includes gaze data collected from a radiologist. Our result reveals that integrating human attention knowledge benefits classification effectively, e.g. improving the baseline by 4.38% on CXR. Hence, our work provides not only valuable insights into understanding human attention in fine-grained classification, but also contributes to future research in integrating human gaze with computer vision tasks. CUB-GHA and code are available at https://github.com/yaorong0921/CUB-GHA.
Forward citations
Cited by 2 Pith papers
-
CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
CT-ScanGaze, the first public eye-tracking dataset on CT volumes, contains 909 scans with radiologist gaze, reports, and findings, and CT-Searcher, a 3D scanpath model, beats adapted 2D baselines on it.
-
Multi-aspect Knowledge Distillation with Large Language Model
A multi-aspect knowledge distillation method that appends binary question-answer logits from an MLLM to a classifier's output improves fine-grained image classification accuracy by up to about 6 points.
Discussion (0). Continue with ORCID to comment.