Pith. sign in

REVIEW 2 cited by

VideoPanda: Video Panoramic Diffusion with Multi-view Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.11389 v2 pith:6BGMQQI5 submitted 2025-04-15 cs.GR cs.AIcs.CV

classification cs.GRcs.AIcs.CV
keywords videovideopandamulti-viewpanoramicattentioncameracircconditions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

High resolution panoramic video content is paramount for immersive experiences in Virtual Reality, but is non-trivial to collect as it requires specialized equipment and intricate camera setups. In this work, we introduce VideoPanda, a novel approach for synthesizing 360$^\circ$ videos conditioned on text or single-view video data. VideoPanda leverages multi-view attention layers to augment a video diffusion model, enabling it to generate consistent multi-view videos that can be combined into immersive panoramic content. VideoPanda is trained jointly using two conditions: text-only and single-view video, and supports autoregressive generation of long-videos. To overcome the computational burden of multi-view video generation, we randomly subsample the duration and camera views used during training and show that the model is able to gracefully generalize to generating more frames during inference. Extensive evaluations on both real-world and synthetic video datasets demonstrate that VideoPanda generates more realistic and coherent 360$^\circ$ panoramas across all input conditions compared to existing methods. Visit the project website at https://research.nvidia.com/labs/toronto-ai/VideoPanda/ for results.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 360Anything: Geometry-Free Lifting of Images and Videos to 360{\deg}

    cs.CV 2026-01 conditional novelty 6.0 of 10

    360Anything lifts perspective images and videos to 360° panoramas with a diffusion transformer and sequence concatenation, requiring no camera metadata at test time.

  2. PanoWan: Lifting Diffusion Video Generation Models to 360{\deg} with Latitude/Longitude-aware Mechanisms

    cs.CV 2025-05 conditional novelty 6.0 of 10

    PanoWan adapts the Wan 2.1 text-to-video model to generate seamless 360-degree videos by remapping initial noise, rotating the latent grid during denoising, and padding the latent before VAE decoding, trained on a new...

Pith tools