On FBP5500, Diff-FBP reports a Pearson correlation of 0.9220 and MAE of 0.2110 using a frozen Diffusion Transformer feature extractor with generative pre-training.
Recent Advances in Vision Transformer: A Survey and Outlook of Recent Work
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Vision Transformers (ViTs) are becoming more popular and dominating technique for various vision tasks, compare to Convolutional Neural Networks (CNNs). As a demanding technique in computer vision, ViTs have been successfully solved various vision problems while focusing on long-range relationships. In this paper, we begin by introducing the fundamental concepts and background of the self-attention mechanism. Next, we provide a comprehensive overview of recent top-performing ViT methods describing in terms of strength and weakness, computational cost as well as training and testing dataset. We thoroughly compare the performance of various ViT algorithms and most representative CNN methods on popular benchmark datasets. Finally, we explore some limitations with insightful observations and provide further research direction. The project page along with the collections of papers are available at https://github.com/khawar512/ViT-Survey
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
REJECT 1roles
baseline 1polarities
baseline 1representative citing papers
citing papers explorer
-
Generative Pre-training for Subjective Tasks: A Diffusion Transformer-Based Framework for Facial Beauty Prediction
On FBP5500, Diff-FBP reports a Pearson correlation of 0.9220 and MAE of 0.2110 using a frozen Diffusion Transformer feature extractor with generative pre-training.