LoRA adapters fix collapsed visual CLS token attention in CLIP for superior cross-domain few-shot learning, and the new Semantic Probe framework revives prompt methods to reach state-of-the-art on four benchmarks.
Mitigate the gap: Investigating approaches for improving cross-modal alignment in clip
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3verdicts
UNVERDICTED 3roles
method 1polarities
use method 1representative citing papers
CLIP-RD adds VRD for cross-modality distillation consistency and XRD for bidirectional cross-modal symmetry to align student embedding geometry more closely with the teacher, yielding a 0.8 percentage point gain over prior distillation methods.
STAMBRIDGE uses STAM for robust EEG feature extraction and MFSB for cross-modal alignment, reporting 34.50% top-1 and 65.95% top-5 accuracy in 200-way zero-shot retrieval on THINGS-EEG plus diffusion-based image reconstructions.
citing papers explorer
-
Reviving In-domain Fine-tuning Methods for Source-Free Cross-domain Few-shot Learning
LoRA adapters fix collapsed visual CLS token attention in CLIP for superior cross-domain few-shot learning, and the new Semantic Probe framework revives prompt methods to reach state-of-the-art on four benchmarks.
-
CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
CLIP-RD adds VRD for cross-modality distillation consistency and XRD for bidirectional cross-modal symmetry to align student embedding geometry more closely with the teacher, yielding a 0.8 percentage point gain over prior distillation methods.
-
STAMBRIDGE: Spectral-Temporal Amplitude-aware Mid-Feature Bridge for EEG Visual Decoding
STAMBRIDGE uses STAM for robust EEG feature extraction and MFSB for cross-modal alignment, reporting 34.50% top-1 and 65.95% top-5 accuracy in 200-way zero-shot retrieval on THINGS-EEG plus diffusion-based image reconstructions.