The paper claims a new efficient ViT architecture with 76.3%/79.6% ImageNet accuracy, but the provided full text is a hep-th paper and never describes or evaluates this architecture.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
A Lightweight Convolution and Vision Transformer integrated model with Multi-scale Self-attention Mechanism
The paper claims a new efficient ViT architecture with 76.3%/79.6% ImageNet accuracy, but the provided full text is a hep-th paper and never describes or evaluates this architecture.