Lens adapts camera sensors in real time via the VisiT confidence-based quality indicator to improve vision model accuracy on domain-shifted images, shown on ImageNet-ES and a new diverse benchmark.
Self-ensembling for visual domain adaptation
4 Pith papers cite this work. Polarity classification is still indexing.
abstract
This paper explores the use of self-ensembling for visual domain adaptation problems. Our technique is derived from the mean teacher variant (Tarvainen et al., 2017) of temporal ensembling (Laine et al;, 2017), a technique that achieved state of the art results in the area of semi-supervised learning. We introduce a number of modifications to their approach for challenging domain adaptation scenarios and evaluate its effectiveness. Our approach achieves state of the art results in a variety of benchmarks, including our winning entry in the VISDA-2017 visual domain adaptation challenge. In small image benchmarks, our algorithm not only outperforms prior art, but can also achieve accuracy that is close to that of a classifier trained in a supervised fashion.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 4roles
method 1polarities
use method 1representative citing papers
Introduces a 9136-sample multi-view in-cabin dataset from a German city bus with RGB, depth, LiDAR, 3D annotations via pseudo-labeling, nuScenes conversion, and benchmarks on models like BEVFusion.
SVL uses vision-language alignment via scene-level shadow ratio regression and global-to-local coupling on a frozen DINOv3 encoder to disambiguate shadows from dark surfaces in dense prediction.
A new regularization approach for unsupervised domain adaptation that calibrates Renyi entropy of uncertainties estimated via variational Bayes.
citing papers explorer
-
Adaptive Camera Sensor for Vision Models
Lens adapts camera sensors in real time via the VisiT confidence-based quality indicator to improve vision model accuracy on domain-shifted images, shown on ImageNet-ES and a new diverse benchmark.
-
Multi-View In-Cabin Monitoring System for Public Transport Vehicles
Introduces a 9136-sample multi-view in-cabin dataset from a German city bus with RGB, depth, LiDAR, 3D annotations via pseudo-labeling, nuScenes conversion, and benchmarks on models like BEVFusion.
-
Revisiting Shadow Detection from a Vision-Language Perspective
SVL uses vision-language alignment via scene-level shadow ratio regression and global-to-local coupling on a frozen DINOv3 encoder to disambiguate shadows from dark surfaces in dense prediction.
-
Unsupervised Domain Adaptation via Calibrating Uncertainties
A new regularization approach for unsupervised domain adaptation that calibrates Renyi entropy of uncertainties estimated via variational Bayes.