REVIEW 4 cited by
Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression Comprehension
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Embodied perception is essential for intelligent vehicles and robots in interactive environmental understanding. However, these advancements primarily focus on vision, with limited attention given to using 3D modeling sensors, restricting a comprehensive understanding of objects in response to prompts containing qualitative and quantitative queries. Recently, as a promising automotive sensor with affordable cost, 4D millimeter-wave radars provide denser point clouds than conventional radars and perceive both semantic and physical characteristics of objects, thereby enhancing the reliability of perception systems. To foster the development of natural language-driven context understanding in radar scenes for 3D visual grounding, we construct the first dataset, Talk2Radar, which bridges these two modalities for 3D Referring Expression Comprehension (REC). Talk2Radar contains 8,682 referring prompt samples with 20,558 referred objects. Moreover, we propose a novel model, T-RadarNet, for 3D REC on point clouds, achieving State-Of-The-Art (SOTA) performance on the Talk2Radar dataset compared to counterparts. Deformable-FPN and Gated Graph Fusion are meticulously designed for efficient point cloud feature modeling and cross-modal fusion between radar and text features, respectively. Comprehensive experiments provide deep insights into radar-based 3D REC. We release our project at https://github.com/GuanRunwei/Talk2Radar.
Forward citations
Cited by 4 Pith papers
-
MITO: A Millimeter-Wave Dataset and Simulator for Non-Line-of-Sight Perception
A new mmWave imaging dataset and simulator enable segmentation and classification of everyday objects hidden behind occluders.
-
MetaOcc: Spatio-Temporal Fusion of Surround-View 4D Radar and Camera for 3D Occupancy Prediction with Dual Training Strategies
A multi-modal 3D occupancy prediction framework that fuses 4D radar and cameras, with height-aware radar features and hierarchical spatio-temporal fusion, achieving state-of-the-art results on OmniHD-Scenes and Surrou...
-
DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-temporal Fusion
DVLO4D fuses sparse LiDAR queries with camera features, adds temporal memory and a sequence-level loss, and improves visual-LiDAR odometry accuracy to 0.73% translation error on KITTI 07-10.
-
RadarNeXt: Real-Time and Reliable 3D Object Detector Based On 4D mmWave Imaging Radar
A real-time 3D detector for 4D mmWave radar point clouds, built from a re-parameterizable MobileOne backbone and a deformable-convolution neck, achieves 50.48 mAP on VoD and 32.30 mAP on TJ4D.
Discussion (0). Continue with ORCID to comment.