Pith. sign in

REVIEW 2 cited by

GenMM: Geometrically and Temporally Consistent Multimodal Data Generation for Video and LiDAR

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10722 v1 pith:IZ7VTY3P submitted 2024-06-15 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords consistentlidargenmmobjectobjectssurfacevideobounding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal synthetic data generation is crucial in domains such as autonomous driving, robotics, augmented/virtual reality, and retail. We propose a novel approach, GenMM, for jointly editing RGB videos and LiDAR scans by inserting temporally and geometrically consistent 3D objects. Our method uses a reference image and 3D bounding boxes to seamlessly insert and blend new objects into target videos. We inpaint the 2D Regions of Interest (consistent with 3D boxes) using a diffusion-based video inpainting model. We then compute semantic boundaries of the object and estimate it's surface depth using state-of-the-art semantic segmentation and monocular depth estimation techniques. Subsequently, we employ a geometry-based optimization algorithm to recover the 3D shape of the object's surface, ensuring it fits precisely within the 3D bounding box. Finally, LiDAR rays intersecting with the new object surface are updated to reflect consistent depths with its geometry. Our experiments demonstrate the effectiveness of GenMM in inserting various 3D objects across video and LiDAR modalities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reference-Guided Diffusion Inpainting For Multimodal Counterfactual Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single reference image guides a diffusion model to insert coherent objects into camera-plus-lidar driving scenes and to insert mammographic anomalies into new scans.

  2. MultiEditor: Controllable Multimodal Object Editing for Driving Scenarios Using 3D Gaussian Splatting Priors

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A dual-branch diffusion framework jointly edits images and LiDAR point clouds in driving scenes using 3D Gaussian Splatting object priors, improving fidelity and boosting detection of rare vehicle classes.

Pith tools