Active Spatial Guidance replaces injected positional embeddings in ViTs with a training-only 2D coordinate regression loss on final-layer tokens, yielding better results than learned absolute or rotary embeddings on ImageNet-100, ADE20K, and Hypersim under matched training.
In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp
7 Pith papers cite this work, alongside 3,165 external citations. Polarity classification is still indexing.
fields
cs.CV 7years
2026 7verdicts
UNVERDICTED 7representative citing papers
TCG-AR is a real-time multi-view AR system for trading card games using only commodity RGB cameras and synthetic training data.
MaskAQ generates synthetic samples for ViT quantization by identifying and selectively aligning sparse informative regions in self-attention maps, combined with periodic sample refreshing.
SegRAG is a training-free retrieval-augmented framework that extracts class-specific point prompts from a filtered DINOv3 feature bank to boost SAM3 semantic segmentation performance on standard and agricultural benchmarks.
Open-source image-editing models show competitive zero-shot performance on monocular depth, surface normals, and semantic segmentation, sometimes matching tuned models.
GEAR-Seg decouples segmentation, semantic description, and LLM reasoning into an explicit chain for interpretable zero-shot reasoning segmentation while generating the GEAR-131K dataset.
PairWise is an open-source tool that finds visually aligned street-level image pairs by integrating feature matching with semantic segmentation mask alignment and outputs quantitative alignment metrics for filtering.
citing papers explorer
-
Active Spatial Guidance: Eliminating Injected Positional Mechanisms in Vision Transformers
Active Spatial Guidance replaces injected positional embeddings in ViTs with a training-only 2D coordinate regression loss on final-layer tokens, yielding better results than learned absolute or rotary embeddings on ImageNet-100, ADE20K, and Hypersim under matched training.
-
TCG-AR: Real-Time Multi-View Augmented Reality for Trading Card Game Streaming
TCG-AR is a real-time multi-view AR system for trading card games using only commodity RGB cameras and synthetic training data.
-
Selective Coupling of Decoupled Informative Regions: Masked Attention Alignment for Data-Free Quantization of Vision Transformers
MaskAQ generates synthetic samples for ViT quantization by identifying and selectively aligning sparse informative regions in self-attention maps, combined with periodic sample refreshing.
-
SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation
SegRAG is a training-free retrieval-augmented framework that extracts class-specific point prompts from a filtered DINOv3 feature bank to boost SAM3 semantic segmentation performance on standard and agricultural benchmarks.
-
Open-Source Image Editing Models Are Zero-Shot Vision Learners
Open-source image-editing models show competitive zero-shot performance on monocular depth, surface normals, and semantic segmentation, sometimes matching tuned models.
-
GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine
GEAR-Seg decouples segmentation, semantic description, and LLM reasoning into an explicit chain for interpretable zero-shot reasoning segmentation while generating the GEAR-131K dataset.
-
PairWise Image Finder: An Open-source Tool for Finding Visually Aligned Street-Level Image Pairs for Urban Perception Studies
PairWise is an open-source tool that finds visually aligned street-level image pairs by integrating feature matching with semantic segmentation mask alignment and outputs quantitative alignment metrics for filtering.