Pith. sign in

REVIEW 8 cited by

Controlling Space and Time with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07860 v2 pith:PLG27OOU submitted 2024-07-10 cs.CV

classification cs.CV
keywords posecameranoveltimedatadiffusiongenerationimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present 4DiM, a cascaded diffusion model for 4D novel view synthesis (NVS), supporting generation with arbitrary camera trajectories and timestamps, in natural scenes, conditioned on one or more images. With a novel architecture and sampling procedure, we enable training on a mixture of 3D (with camera pose), 4D (pose+time) and video (time but no pose) data, which greatly improves generalization to unseen images and camera pose trajectories over prior works that focus on limited domains (e.g., object centric). 4DiM is the first-ever NVS method with intuitive metric-scale camera pose control enabled by our novel calibration pipeline for structure-from-motion-posed data. Experiments demonstrate that 4DiM outperforms prior 3D NVS models both in terms of image fidelity and pose alignment, while also enabling the generation of scene dynamics. 4DiM provides a general framework for a variety of tasks including single-image-to-3D, two-image-to-video (interpolation and extrapolation), and pose-conditioned video-to-video translation, which we illustrate qualitatively on a variety of scenes. For an overview see https://4d-diffusion.github.io

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation

    cs.CV 2026-02 unverdicted novelty 7.0 of 10

    PerpetualWonder introduces a closed-loop generative simulator with a unified physical-visual representation for long-horizon action-conditioned 4D scene generation from one image.

  2. Feed-forward Motion In-betweening for Any 4D

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Proposes a feed-forward keyframe-conditioned in-betweening method for arbitrary 4D meshes using a topology-agnostic VAE and MMDiT-based rectified flow model.

  3. RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    RayDer is a unified transformer backbone for self-supervised static-scene novel view synthesis that absorbs dynamic content as a nuisance factor and shows power-law scaling with data and compute while matching supervi...

  4. GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Feature-space Gaussian Splat Feature Adapter (GS-Adapter) grounds camera-controlled video diffusion in 3D Gaussians, improving geometric consistency and controllability over SEVA and CameraCtrl without retraining geom...

  5. InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A training-free, near-zero-overhead inverse solver for novel-view video generation and inpainting that projects masks into continuous multi-channel latent masks and applies DDS with conjugate gradient in latent space.

  6. KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos

    cs.CV 2024-11 unverdicted novelty 5.0 of 10

    KFC-W is a self-supervised 3D-aware video model trained on videos and multiview internet photos that produces geometrically consistent interpolations between unposed input images without any 3D annotations.

  7. World Simulation with Video Foundation Models for Physical AI

    cs.CV 2025-10 unverdicted novelty 4.0 of 10

    Cosmos-Predict2.5 unifies text-to-world, image-to-world, and video-to-world generation in one model trained on 200M clips with RL post-training, delivering improved quality and control for physical AI.

  8. Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends

    cs.CV 2026-05 unverdicted novelty 2.0 of 10

    This survey reviews trends, challenges, benchmarks, and future directions in action-conditioned interactive world modeling for video and 3D generation.

Pith tools