Tri-FusionNet is a new ViT+RoBERTa+CLIP captioning model claiming state-of-the-art scores that do not survive internal consistency checks.
Contrastive language- image pre-training with knowledge graphs,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism
Tri-FusionNet is a new ViT+RoBERTa+CLIP captioning model claiming state-of-the-art scores that do not survive internal consistency checks.