KITScenes LongTail supplies multimodal driving data and multilingual expert reasoning traces to benchmark models on rare scenarios beyond basic safety metrics.
hub
Vision meets robotics: The KITTI dataset
14 Pith papers cite this work, alongside 9,724 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
polarities
background 3representative citing papers
INSANE releases multiple MAV datasets with cross-environment trajectories, rich multi-IMU and camera suites, high-rate vibration data, and sub-centimeter RTK GNSS ground truth for localization research.
FAT decomposes structured prediction into specialist hypothesis generation and foundation-model proxy reasoning, yielding consistent gains over baselines on detection, trajectory, and segmentation tasks.
HilDA pre-trains LiDAR backbones via multi-layer and global distillation from vision models plus temporal occupancy diffusion, yielding SOTA results on detection, flow, and occupancy tasks.
Modality Forcing lets a single DiT produce image and depth outputs in any order after training on sparse real-world depth, with larger image-pretrained models yielding better depth accuracy and a 57% AbsRel reduction versus prior joint generative baselines.
MR-LiDAR benchmark shows an 80-beam LiDAR with optimized distribution can match or exceed 128-beam uniform LiDAR for roadside vehicle and VRU detection.
GOLD-BEV learns dense BEV semantic maps including dynamic agents from ego-centric sensors by using synchronized aerial imagery for training supervision and pseudo-label generation.
InFlux++ introduces a synthetic training dataset and an extended real-world benchmark for per-frame camera intrinsics estimation, showing that finetuning on synthetic data improves focal length prediction.
Integrating DVS event data into InterFuser through token fusion yields a driving score of 77.2 and 100% route completion on CARLA benchmarks, indicating improved robustness in dynamic conditions.
SPARK applies keypoint detection with YOLO models to monocular images for low-latency 3D pose estimation of racing opponents, claiming better accuracy and speed than prior camera methods on real racing data.
Two radar odometry baselines improve trajectory estimates on challenging off-road routes in the GO dataset.
Introduces structured NuScenes-S dataset and 0.9B FastDrive VLM claiming 20% higher decision accuracy and over 10x inference speedup versus larger unstructured VLMs.
A survey that organizes methods for cross-domain object detection into a taxonomy, analyzes domain shift across detection stages, and outlines persistent challenges.
The paper presents a 5G terrestrial positioning system using multi-carrier carrier phase ranging, deep learning for NLOS identification, and IMU/camera sensor fusion via error-state EKF, achieving less than 5 meters error on simulated 5G signals over KITTI urban trajectories.
citing papers explorer
-
LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset
KITScenes LongTail supplies multimodal driving data and multilingual expert reasoning traces to benchmark models on rare scenarios beyond basic safety metrics.
-
INSANE: Cross-Domain UAV Data Sets with Increased Number of Sensors for developing Advanced and Novel Estimators
INSANE releases multiple MAV datasets with cross-environment trajectories, rich multi-IMU and camera suites, high-rate vibration data, and sub-centimeter RTK GNSS ground truth for localization research.
-
Rethinking Foundation Model Collaboration: Enhancing Specialized Models through Proxy Task Reasoning
FAT decomposes structured prediction into specialist hypothesis generation and foundation-model proxy reasoning, yielding consistent gains over baselines on detection, trajectory, and segmentation tasks.
-
HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training
HilDA pre-trains LiDAR backbones via multi-layer and global distillation from vision models plus temporal occupancy diffusion, yielding SOTA results on detection, flow, and occupancy tasks.
-
Modality Forcing for Scalable Spatial Generation
Modality Forcing lets a single DiT produce image and depth outputs in any order after training on sparse real-world depth, with larger image-pretrained models yielding better depth accuracy and a 57% AbsRel reduction versus prior joint generative baselines.
-
MR-LiDAR: A Multi-Resolution Roadside LiDAR Benchmark for Perception Diagnostics and Deployment Guidance
MR-LiDAR benchmark shows an 80-beam LiDAR with optimized distribution can match or exceed 128-beam uniform LiDAR for roadside vehicle and VRU detection.
-
GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes
GOLD-BEV learns dense BEV semantic maps including dynamic agents from ego-centric sensors by using synchronized aerial imagery for training supervision and pseudo-label generation.
-
InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics
InFlux++ introduces a synthetic training dataset and an extended real-world benchmark for per-frame camera intrinsics estimation, showing that finetuning on synthetic data improves focal length prediction.
-
InterFuserDVS: Event-Enhanced Sensor Fusion for Safe RL-Based Decision Making
Integrating DVS event data into InterFuser through token fusion yields a driving score of 77.2 and 100% route completion on CARLA benchmarks, indicating improved robustness in dynamic conditions.
-
SPARK: Low Latency Single-Camera 3D Pose Estimation for Autonomous Racing using Keypoints
SPARK applies keypoint detection with YOLO models to monocular images for low-latency 3D pose estimation of racing opponents, claiming better accuracy and speed than prior camera methods on real racing data.
-
Pushing Radar Odometry Beyond the Pavement: Current Capabilities and Challenges
Two radar odometry baselines improve trajectory estimates on challenging off-road routes in the GO dataset.
-
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
Introduces structured NuScenes-S dataset and 0.9B FastDrive VLM claiming 20% higher decision accuracy and over 10x inference speedup versus larger unstructured VLMs.
-
Generalization Under Scrutiny: Cross-Domain Detection Progresses, Pitfalls, and Persistent Challenges
A survey that organizes methods for cross-domain object detection into a taxonomy, analyzes domain shift across detection stages, and outlines persistent challenges.
-
A Robust 5G Terrestrial Positioning System with Sensor Fusion in GNSS-denied Scenarios
The paper presents a 5G terrestrial positioning system using multi-carrier carrier phase ranging, deep learning for NLOS identification, and IMU/camera sensor fusion via error-state EKF, achieving less than 5 meters error on simulated 5G signals over KITTI urban trajectories.