Pith. sign in

REVIEW 1 cited by

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.00707 v1 pith:LOOU4GYZ submitted 2025-07-01 cs.CV

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving

classification cs.CV
keywords generationbev-vaeimageautonomousconsistentdrivingmulti-viewscene
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Multi-view image generation in autonomous driving demands consistent 3D scene understanding across camera views. Most existing methods treat this problem as a 2D image set generation task, lacking explicit 3D modeling. However, we argue that a structured representation is crucial for scene generation, especially for autonomous driving applications. This paper proposes BEV-VAE for consistent and controllable view synthesis. BEV-VAE first trains a multi-view image variational autoencoder for a compact and unified BEV latent space and then generates the scene with a latent diffusion transformer. BEV-VAE supports arbitrary view generation given camera configurations, and optionally 3D layouts. Experiments on nuScenes and Argoverse 2 (AV2) show strong performance in both 3D consistent reconstruction and generation. The code is available at: https://github.com/Czm369/bev-vae.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Variational Inference for Bird's Eye View Segmentation in Autonomous Driving

    cs.CV 2026-07 conditional novelty 6.0

    TVB combines a conditional variational autoencoder, normalizing flows, and attention-based fusion to produce bird's-eye-view segmentation from multiple car cameras, reporting small but consistent IoU gains on nuScenes...