A sequential-to-global SSL method based on DINO pretrains iterative foveal-inspired vision transformers to achieve competitive ImageNet-1K performance with constant compute regardless of input resolution.
In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers)
9 Pith papers cite this work. Polarity classification is still indexing.
years
2026 9verdicts
UNVERDICTED 9representative citing papers
LiteSemRAG delivers leading MRR@10 on three benchmarks using only lightweight semantic graph methods and zero LLM tokens.
Replacing selected attention heads in pretrained ViTs with depthwise convolutions, identified by simple strategies and recovered via fine-tuning, delivers 17-20% inference speedup on image tasks with minimal accuracy loss.
GeoFuse fuses aligned road maps with satellite imagery via token/channel interactions and dynamic gating plus contrastive learning, lifting Recall@1 by 3.46% on University-1652 and 23.18% on DenseUAV under weather degradation.
DBMF integrates scores from text-image and vision branches to improve out-of-distribution detection on endoscopic datasets by up to 24.84% over prior methods.
URMF uses learnable Gaussian posteriors to estimate modality-specific uncertainty and adjust fusion weights for improved multimodal sarcasm detection on MSD and MMSD2 benchmarks.
CA-GCL combines global contrastive learning with permutation-invariant text augmentation to deliver zero-shot 3D medical abnormality detection that is more robust to prompt changes than prior FVLP methods.
Four new Reddit-derived datasets for mental health detection tasks are presented with inter-annotator agreement above 0.8 and reported model F1 scores of 93-99%.
A cosine-similarity metric on SHAP feature attributions is proposed to quantify explanation stability for same-label inputs under perturbations in transformer-based sentiment classifiers.
citing papers explorer
-
Self-supervised pretraining for an iterative image size agnostic vision transformer
A sequential-to-global SSL method based on DINO pretrains iterative foveal-inspired vision transformers to achieve competitive ImageNet-1K performance with constant compute regardless of input resolution.
-
LiteSemRAG: Lightweight LLM-Free Semantic-Aware Graph Retrieval for Robust RAG
LiteSemRAG delivers leading MRR@10 on three benchmarks using only lightweight semantic graph methods and zero LLM tokens.
-
Accelerating Vision Foundation Models with Drop-in Depthwise Convolution
Replacing selected attention heads in pretrained ViTs with depthwise convolutions, identified by simple strategies and recovered via fine-tuning, delivers 17-20% inference speedup on image tasks with minimal accuracy loss.
-
Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse
GeoFuse fuses aligned road maps with satellite imagery via token/channel interactions and dynamic gating plus contrastive learning, lifting Recall@1 by 3.46% on University-1652 and 23.18% on DenseUAV under weather degradation.
-
DBMF: A Dual-Branch Multimodal Framework for Out-of-Distribution Detection
DBMF integrates scores from text-image and vision branches to improve out-of-distribution detection on endoscopic datasets by up to 24.84% over prior methods.
-
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
URMF uses learnable Gaussian posteriors to estimate modality-specific uncertainty and adjust fusion weights for improved multimodal sarcasm detection on MSD and MMSD2 benchmarks.
-
CA-GCL: Cross-Anatomy Global-Local Contrastive Learning for Robust 3D Medical Image Understanding
CA-GCL combines global contrastive learning with permutation-invariant text augmentation to deliver zero-shot 3D medical abnormality detection that is more robust to prompt changes than prior FVLP methods.
-
A Benchmark Suite of Reddit-Derived Datasets for Mental Health Detection
Four new Reddit-derived datasets for mental health detection tasks are presented with inter-annotator agreement above 0.8 and reported model F1 scores of 93-99%.
-
Empirical Characterization of Rationale Stability Under Controlled Perturbations for Explainable Pattern Recognition
A cosine-similarity metric on SHAP feature attributions is proposed to quantify explanation stability for same-label inputs under perturbations in transformer-based sentiment classifiers.