RAR combines CLIP retrieval with MLLM ranking to improve few-shot and zero-shot fine-grained visual recognition on 5 benchmarks, 11 few-shot datasets, and 2 detection tasks.
arXiv preprint arXiv:2311.15732 , year=
4 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Releases M-EDESConv and M-TESC multilingual datasets and introduces MEGUMI model that fuses XLM-RoBERTa with emotion encoders for superior validation timing detection.
A formalized Minimal Cognitive Grid ranks computational models of analogy and metaphor by alignment with cognitive theories using Functional/Structural Ratio, Generality, and Performance Match dimensions.
citing papers explorer
-
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition
RAR combines CLIP retrieval with MLLM ranking to improve few-shot and zero-shot fine-grained visual recognition on 5 benchmarks, 11 few-shot datasets, and 2 detection tasks.
-
I Understand How You Feel: Enhancing Deeper Emotional Support Through Multilingual Emotional Validation in Dialogue System
Releases M-EDESConv and M-TESC multilingual datasets and introduces MEGUMI model that fuses XLM-RoBERTa with emotion encoders for superior validation timing detection.
-
Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid
A formalized Minimal Cognitive Grid ranks computational models of analogy and metaphor by alignment with cognitive theories using Functional/Structural Ratio, Generality, and Performance Match dimensions.
- AdaBoosting Text Prompts for Vision-Language Models