A model-agreement metric combined with depth statistics finds real aerial scenes easier for vision transformers than synthetic ones, and quantifies the gap between the two domains.
Our methodology leverages both perceptual complexity, measured via multi-model consensus, and struc- tural complexity, captured through depth-based metrics
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Quantifying the synthetic and real domain gap in aerial scene understanding
A model-agreement metric combined with depth statistics finds real aerial scenes easier for vision transformers than synthetic ones, and quantifies the gap between the two domains.