Pith. sign in

REVIEW 3 cited by

X-Drive: Cross-modality consistent multi-sensor data synthesis for driving scenarios

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.01123 v1 pith:JEYSYY6S submitted 2024-11-02 cs.CV

classification cs.CV
keywords x-drivecloudscross-modalitypointconditionsdatadrivingsynthesis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements have exploited diffusion models for the synthesis of either LiDAR point clouds or camera image data in driving scenarios. Despite their success in modeling single-modality data marginal distribution, there is an under-exploration in the mutual reliance between different modalities to describe complex driving scenes. To fill in this gap, we propose a novel framework, X-DRIVE, to model the joint distribution of point clouds and multi-view images via a dual-branch latent diffusion model architecture. Considering the distinct geometrical spaces of the two modalities, X-DRIVE conditions the synthesis of each modality on the corresponding local regions from the other modality, ensuring better alignment and realism. To further handle the spatial ambiguity during denoising, we design the cross-modality condition module based on epipolar lines to adaptively learn the cross-modality local correspondence. Besides, X-DRIVE allows for controllable generation through multi-level input conditions, including text, bounding box, image, and point clouds. Extensive results demonstrate the high-fidelity synthetic results of X-DRIVE for both point clouds and multi-view images, adhering to input conditions while ensuring reliable cross-modality consistency. Our code will be made publicly available at https://github.com/yichen928/X-Drive.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reference-Guided Diffusion Inpainting For Multimodal Counterfactual Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single reference image guides a diffusion model to insert coherent objects into camera-plus-lidar driving scenes and to insert mammographic anomalies into new scans.

  2. MultiEditor: Controllable Multimodal Object Editing for Driving Scenarios Using 3D Gaussian Splatting Priors

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A dual-branch diffusion framework jointly edits images and LiDAR point clouds in driving scenes using 3D Gaussian Splatting object priors, improving fidelity and boosting detection of rare vehicle classes.

  3. Solar Altitude Guided Scene Illumination

    cs.CV 2025-07 conditional novelty 6.0 of 10

    The paper conditions a latent diffusion camera-data generator on solar altitude, using a bin-and-residual encoding, to control daylight and image noise without manual labels.

Pith tools