Pith. sign in

REVIEW 1 cited by

MV2DFusion: Leveraging Modality-Specific Object Semantics for Multi-Modal 3D Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.05945 v2 pith:HFTAQRZL submitted 2024-08-12 cs.CV

classification cs.CV
keywords detectionobjectmv2dfusionsemanticsframeworkfusiongeneratorimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rise of autonomous vehicles has significantly increased the demand for robust 3D object detection systems. While cameras and LiDAR sensors each offer unique advantages--cameras provide rich texture information and LiDAR offers precise 3D spatial data--relying on a single modality often leads to performance limitations. This paper introduces MV2DFusion, a multi-modal detection framework that integrates the strengths of both worlds through an advanced query-based fusion mechanism. By introducing an image query generator to align with image-specific attributes and a point cloud query generator, MV2DFusion effectively combines modality-specific object semantics without biasing toward one single modality. Then the sparse fusion process can be accomplished based on the valuable object semantics, ensuring efficient and accurate object detection across various scenarios. Our framework's flexibility allows it to integrate with any image and point cloud-based detectors, showcasing its adaptability and potential for future advancements. Extensive evaluations on the nuScenes and Argoverse2 datasets demonstrate that MV2DFusion achieves state-of-the-art performance, particularly excelling in long-range detection scenarios.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RoCA: Robust Cross-Domain End-to-End Autonomous Driving

    cs.CV 2025-06 conditional novelty 5.0 of 10

    RoCA, a Gaussian-process codebook over ego and agent tokens, improves cross-domain generalization and adaptation of end-to-end autonomous driving models without extra inference cost.

Pith tools