Pith. sign in

REVIEW 2 cited by

Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.14494 v1 pith:ZUIVIIFG submitted 2024-12-19 cs.CV

classification cs.CV
keywords datarealaccountdrivingmodelsnovelsynthesissynthetic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The recent advent of large-scale 3D data, e.g. Objaverse, has led to impressive progress in training pose-conditioned diffusion models for novel view synthesis. However, due to the synthetic nature of such 3D data, their performance drops significantly when applied to real-world images. This paper consolidates a set of good practices to finetune large pretrained models for a real-world task -- harvesting vehicle assets for autonomous driving applications. To this end, we delve into the discrepancies between the synthetic data and real driving data, then develop several strategies to account for them properly. Specifically, we start with a virtual camera rotation of real images to ensure geometric alignment with synthetic data and consistency with the pose manifold defined by pretrained models. We also identify important design choices in object-centric data curation to account for varying object distances in real driving scenes -- learn across varying object scales with fixed camera focal length. Further, we perform occlusion-aware training in latent spaces to account for ubiquitous occlusions in real data, and handle large viewpoint changes by leveraging a symmetric prior. Our insights lead to effective finetuning that results in a $68.8\%$ reduction in FID for novel view synthesis over prior arts.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reference-Guided Diffusion Inpainting For Multimodal Counterfactual Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single reference image guides a diffusion model to insert coherent objects into camera-plus-lidar driving scenes and to insert mammographic anomalies into new scans.

  2. ArbiViewGen: Controllable Arbitrary Viewpoint Camera Data Generation for Autonomous Driving via Stable Diffusion Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    ArbiViewGen generates arbitrary-viewpoint driving camera images by stitching the six input views into pseudo-target views and training a Stable Diffusion model to reconstruct the original views, enabling self-supervis...

Pith tools