REVIEW 3 cited by
PanoWan: Lifting Diffusion Video Generation Models to 360{deg} with Latitude/Longitude-aware Mechanisms
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
PanoWan: Lifting Diffusion Video Generation Models to 360{deg} with Latitude/Longitude-aware Mechanisms
read the original abstract
Panoramic video generation enables immersive 360{\deg} content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained generative priors from conventional text-to-video models for high-quality and diverse panoramic videos generation, due to limited dataset scale and the gap in spatial feature representations. In this paper, we introduce PanoWan to effectively lift pre-trained text-to-video models to the panoramic domain, equipped with minimal modules. PanoWan employs latitude-aware sampling to avoid latitudinal distortion, while its rotated semantic denoising and padded pixel-wise decoding ensure seamless transitions at longitude boundaries. To provide sufficient panoramic videos for learning these lifted representations, we contribute PanoVid, a high-quality panoramic video dataset with captions and diverse scenarios. Consequently, PanoWan achieves state-of-the-art performance in panoramic video generation and demonstrates robustness for zero-shot downstream tasks. Our project page is available at https://panowan.variantconst.com.
Forward citations
Cited by 3 Pith papers
-
ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space
A unified pipeline lifts any text/image/video input into a Spatial Generative Primitive, explores it with 3D-consistent panoramic video, and reconstructs photorealistic 3DGS worlds with stronger rich-input fidelity th...
-
EmoSpace: Immersive Affective Image Generation Guided by Fine-Grained Emotion Prototypes
EmoSpace generates emotion-controlled images and VR panoramas via a dynamic bank of 1,024 CLIP-space emotion prototypes, reporting higher fine-grained emotional alignment than baseline diffusion models.
-
ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space
A multimodal pipeline lifts text/image/video into a panorama–point-cloud primitive, generates a 3D-consistent panoramic video, and reconstructs a photorealistic 3DGS world, claiming open-source SOTA and better fidelit...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.