REVIEW 2 cited by
V2X-ViT: Vehicle-to-Everything Cooperative Perception with Vision Transformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
V2X-ViT: Vehicle-to-Everything Cooperative Perception with Vision Transformer
read the original abstract
In this paper, we investigate the application of Vehicle-to-Everything (V2X) communication to improve the perception performance of autonomous vehicles. We present a robust cooperative perception framework with V2X communication using a novel vision Transformer. Specifically, we build a holistic attention model, namely V2X-ViT, to effectively fuse information across on-road agents (i.e., vehicles and infrastructure). V2X-ViT consists of alternating layers of heterogeneous multi-agent self-attention and multi-scale window self-attention, which captures inter-agent interaction and per-agent spatial relationships. These key modules are designed in a unified Transformer architecture to handle common V2X challenges, including asynchronous information sharing, pose errors, and heterogeneity of V2X components. To validate our approach, we create a large-scale V2X perception dataset using CARLA and OpenCDA. Extensive experimental results demonstrate that V2X-ViT sets new state-of-the-art performance for 3D object detection and achieves robust performance even under harsh, noisy environments. The code is available at https://github.com/DerrickXuNu/v2x-vit.
Forward citations
Cited by 2 Pith papers
-
Sarus: Privacy-Preserving Multi-Vendor Perception Fusion via Homomorphic Encryption
Sarus is an HE-based framework that fuses vendors' Gaussian-moment detection summaries in encrypted form, with linear-scaling server fusion and near-identical output to plaintext fusion.
-
SHLE: Devices Tracking and Depth Filtering for Stereo-based Height Limit Estimation
SHLE is a stereo pipeline with device tracking and temporal depth filtering that estimates height limits with under 10cm average error at 70m distance on the new Disparity Height dataset.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.