Urban-ImageNet is a 2-million-image multi-modal dataset with HUSIC 10-class taxonomy enabling benchmarks for urban scene classification, cross-modal retrieval, and instance segmentation.
hub
In: IEEE Conf
6 Pith papers cite this work, alongside 1,213 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 6roles
dataset 1polarities
background 1representative citing papers
Vesta is a unified embodied generalist model that outperforms specialist baselines by over 20% on average and improves real-world robotic task success by over 35%.
SegRAG is a training-free retrieval-augmented framework that extracts class-specific point prompts from a filtered DINOv3 feature bank to boost SAM3 semantic segmentation performance on standard and agricultural benchmarks.
GEAR-Seg decouples segmentation, semantic description, and LLM reasoning into an explicit chain for interpretable zero-shot reasoning segmentation while generating the GEAR-131K dataset.
Introduces the TSBOW dataset and benchmark for occluded vehicle detection in traffic surveillance under diverse and extreme weather conditions.
A literature review that categorizes deep learning approaches for visual hand gesture recognition, summarizes state-of-the-art methods across tasks, reviews datasets and metrics, and identifies challenges and future directions.
citing papers explorer
-
Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception
Urban-ImageNet is a 2-million-image multi-modal dataset with HUSIC 10-class taxonomy enabling benchmarks for urban scene classification, cross-modal retrieval, and instance segmentation.
-
Vesta: A Generalist Embodied Reasoning Model
Vesta is a unified embodied generalist model that outperforms specialist baselines by over 20% on average and improves real-world robotic task success by over 35%.
-
SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation
SegRAG is a training-free retrieval-augmented framework that extracts class-specific point prompts from a filtered DINOv3 feature bank to boost SAM3 semantic segmentation performance on standard and agricultural benchmarks.
-
GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine
GEAR-Seg decouples segmentation, semantic description, and LLM reasoning into an explicit chain for interpretable zero-shot reasoning segmentation while generating the GEAR-131K dataset.
-
TSBOW -- Traffic Surveillance Benchmark for Occluded Vehicles Under Various Weather Conditions
Introduces the TSBOW dataset and benchmark for occluded vehicle detection in traffic surveillance under diverse and extreme weather conditions.
-
Visual Hand Gesture Recognition with Deep Learning: A Comprehensive Review of Methods, Datasets, Challenges and Future Research Directions
A literature review that categorizes deep learning approaches for visual hand gesture recognition, summarizes state-of-the-art methods across tasks, reviews datasets and metrics, and identifies challenges and future directions.