Video diffusion models can be adapted into permutation-invariant generators for sparse novel view synthesis by treating the problem as video completion and removing temporal order cues.
arXiv preprint arXiv:2601.16982 (2026)
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 4verdicts
UNVERDICTED 4roles
background 1polarities
background 1representative citing papers
RoboDream is a compositional world model that decouples robot trajectory execution from environment synthesis to generate scalable, diverse manipulation data, shown in abstract to improve downstream policies.
MVTrack4Gen uses multi-view point tracking as geometric and motion supervision for camera-conditioning-only novel-view video diffusion models to improve consistency.
A multi-view video diffusion model conditioned on relative camera poses via extended RoPE generates dense synchronized views from sparse inputs for 4D Gaussian splatting reconstruction, claiming SOTA results on human datasets and generalization to animals.
citing papers explorer
-
Novel View Synthesis as Video Completion
Video diffusion models can be adapted into permutation-invariant generators for sparse novel view synthesis by treating the problem as video completion and removing temporal order cues.
-
RoboDream: Compositional World Models for Scalable Robot Data Synthesis
RoboDream is a compositional world model that decouples robot trajectory execution from environment synthesis to generate scalable, diverse manipulation data, shown in abstract to improve downstream policies.
-
MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation
MVTrack4Gen uses multi-view point tracking as geometric and motion supervision for camera-conditioning-only novel-view video diffusion models to improve consistency.
-
Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction
A multi-view video diffusion model conditioned on relative camera poses via extended RoPE generates dense synchronized views from sparse inputs for 4D Gaussian splatting reconstruction, claiming SOTA results on human datasets and generalization to animals.