Pith. sign in

Boximator: Generating rich and controllable motions for video synthesis

12 Pith papers cite this work. Polarity classification is still indexing.

12 Pith papers citing it
abstract

Generating rich and controllable motion is a pivotal challenge in video synthesis. We propose Boximator, a new approach for fine-grained motion control. Boximator introduces two constraint types: hard box and soft box. Users select objects in the conditional frame using hard boxes and then use either type of boxes to roughly or rigorously define the object's position, shape, or motion path in future frames. Boximator functions as a plug-in for existing video diffusion models. Its training process preserves the base model's knowledge by freezing the original weights and training only the control module. To address training challenges, we introduce a novel self-tracking technique that greatly simplifies the learning of box-object correlations. Empirically, Boximator achieves state-of-the-art video quality (FVD) scores, improving on two base models, and further enhanced after incorporating box constraints. Its robust motion controllability is validated by drastic increases in the bounding box alignment metric. Human evaluation also shows that users favor Boximator generation results over the base model.

citation-role summary

background 3

citation-polarity summary

fields

cs.CV 12

years

2026 12

roles

background 3

polarities

background 3

representative citing papers

MoRight: Motion Control Done Right

cs.CV · 2026-04-08 · unverdicted · novelty 7.0

MoRight disentangles object and camera motion via canonical-view specification and temporal cross-view attention, while decomposing motion into active user-driven and passive consequence components to learn and apply causality in video generation.

LooseControlVideo: Directorial Video Control using Spatial Blocking

cs.CV · 2026-06-17 · unverdicted · novelty 6.0

LooseControlVideo fine-tunes a video model on DNOCS-annotated data to enable layout and trajectory control via oriented 3D boxes, reporting 1.2-3x gains in trajectory accuracy over 2D baselines on nuScenes, HO-3D and BEHAVE.

Compositional Video Generation via Inference-Time Guidance

cs.CV · 2026-05-14 · unverdicted · novelty 6.0

CVG improves compositional faithfulness in frozen text-to-video diffusion models by steering early denoising steps with gradients from a classifier trained on the model's own cross-attention features.

PhyCo: Learning Controllable Physical Priors for Generative Motion

cs.CV · 2026-04-30 · unverdicted · novelty 6.0

PhyCo adds continuous physical control to video diffusion models via physics-supervised fine-tuning on a large simulation dataset and VLM-guided rewards, yielding measurable gains in physical realism on the Physics-IQ benchmark.

Evolution of Video Generative Foundations

cs.CV · 2026-04-07 · unverdicted · novelty 2.0

This survey traces video generation technology from GANs to diffusion models and then to autoregressive and multimodal approaches while analyzing principles, strengths, and future trends.

citing papers explorer

Showing 12 of 12 citing papers.