Pith. sign in

REVIEW 8 cited by

Tora: Trajectory-oriented Diffusion Transformer for Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.21705 v4 pith:U6FFAMIG submitted 2024-07-31 cs.CV

classification cs.CV
keywords motiontoravideodiffusionaccuratelycontentgenerationintegrates
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in Diffusion Transformer (DiT) have demonstrated remarkable proficiency in producing high-quality video content. Nonetheless, the potential of transformer-based diffusion models for effectively generating videos with controllable motion remains an area of limited exploration. This paper introduces Tora, the first trajectory-oriented DiT framework that concurrently integrates textual, visual, and trajectory conditions, thereby enabling scalable video generation with effective motion guidance. Specifically, Tora consists of a Trajectory Extractor (TE), a Spatial-Temporal DiT, and a Motion-guidance Fuser (MGF). The TE encodes arbitrary trajectories into hierarchical spacetime motion patches with a 3D motion compression network. The MGF integrates the motion patches into the DiT blocks to generate consistent videos that accurately follow designated trajectories. Our design aligns seamlessly with DiT's scalability, allowing precise control of video content's dynamics with diverse durations, aspect ratios, and resolutions. Extensive experiments demonstrate that Tora excels in achieving high motion fidelity compared to the foundational DiT model, while also accurately simulating the complex movements of the physical world. Code is made available at https://github.com/alibaba/Tora .

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AnyI2V: Animating Any Conditional Image with Motion Control

    cs.CV 2025-07 conditional novelty 6.0 of 10

    AnyI2V animates arbitrary conditional images with user-defined trajectories by injecting debiased diffusion features and aligning attention queries across frames, without training.

  2. Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Vid-CamEdit re-synthesizes monocular videos along user-defined camera paths by conditioning a video diffusion model on 2D flows derived from estimated 3D geometry, without training on multi-view video data.

  3. Interactive Video Generation via Domain Adaptation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A training-free method combines mask normalization and temporal intrinsic denoising to improve trajectory control and perceptual quality in text-to-video diffusion.

  4. EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EPiC trains a 30M-parameter visibility-aware ControlNet on mask-based anchor videos from 5,000 in-the-wild videos and 500 steps, reaching SOTA camera accuracy on RealEstate10K and MiraData.

  5. EF-VI: Enhancing End-Frame Injection for Video Inbetweening

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EF-VI injects temporally expanded end-frame features into a transformer-based image-to-video diffusion model, improving video inbetweening quality over direct fine-tuning and bidirectional sampling baselines.

  6. Hybrid Neural-MPM for Interactive Fluid Simulations in Real-Time

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A hybrid neural-MPM solver with a chaos-triggered fallback and a diffusion-based sketch controller enables real-time interactive fluid simulation with user control.

  7. Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A retrieval-initialized symbolic regression method discovers equations of motion from video trajectories and uses them to guide image-to-video generation, improving physical alignment on classical mechanics scenes.

  8. Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Tora2 adds decoupled personalization embeddings, gated self-attention binding, and contrastive learning to Tora, enabling simultaneous appearance and trajectory customization for multiple entities in generated video.

Pith tools