Vision transformers, particularly LITv2, outperform ResNet baselines on a newly curated 3-class pornography classification dataset, but the evaluation is limited by dataset overlap and tuning issues.
Survey on the attention based RNN model and its applications in computer vision
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The recurrent neural networks (RNN) can be used to solve the sequence to sequence problem, where both the input and the output have sequential structures. Usually there are some implicit relations between the structures. However, it is hard for the common RNN model to fully explore the relations between the sequences. In this survey, we introduce some attention based RNN models which can focus on different parts of the input for each output item, in order to explore and take advantage of the implicit relations between the input and the output items. The different attention mechanisms are described in detail. We then introduce some applications in computer vision which apply the attention based RNN models. The superiority of the attention based RNN model is shown by the experimental results. At last some future research directions are given.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Sensitive Image Classification by Vision Transformers
Vision transformers, particularly LITv2, outperform ResNet baselines on a newly curated 3-class pornography classification dataset, but the evaluation is limited by dataset overlap and tuning issues.