Pith. sign in

REVIEW 23 cited by

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.02492 v2 pith:LNS3Y4ZI submitted 2025-02-04 cs.CV

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

classification cs.CV
keywords motionvideomodelmodelsvideojamcoherencegenerationappearance
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models toward appearance fidelity at the expense of motion coherence. To address this, we introduce VideoJAM, a novel framework that instills an effective motion prior to video generators, by encouraging the model to learn a joint appearance-motion representation. VideoJAM is composed of two complementary units. During training, we extend the objective to predict both the generated pixels and their corresponding motion from a single learned representation. During inference, we introduce Inner-Guidance, a mechanism that steers the generation toward coherent motion by leveraging the model's own evolving motion prediction as a dynamic guidance signal. Notably, our framework can be applied to any video model with minimal adaptations, requiring no modifications to the training data or scaling of the model. VideoJAM achieves state-of-the-art performance in motion coherence, surpassing highly competitive proprietary models while also enhancing the perceived visual quality of the generations. These findings emphasize that appearance and motion can be complementary and, when effectively integrated, enhance both the visual quality and the coherence of video generation. Project website: https://hila-chefer.github.io/videojam-paper.github.io/

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking

    cs.CV 2026-05 unverdicted novelty 8.0

    TrackCraft3R is the first method to repurpose a video diffusion transformer as a feed-forward dense 3D tracker via dual-latent representations and temporal RoPE alignment, achieving SOTA performance with lower compute.

  2. HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation

    cs.CV 2026-04 unverdicted novelty 7.0

    HumANDiff improves motion consistency in human video generation by sampling diffusion noise on an articulated human body template and adding joint appearance-motion prediction plus a geometric consistency loss.

  3. Olaf-World: Orienting Latent Actions for Video World Modeling

    cs.CV 2026-02 conditional novelty 7.0

    Latent actions become transferable across visual contexts when aligned to temporal feature differences from a frozen video encoder (SeqΔ-REPA), improving zero-shot action transfer and data-efficient adaptation of vide...

  4. CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos

    cs.CV 2026-01 unverdicted novelty 7.0

    CoMoVi co-generates 3D human motions and 2D videos synchronously in a single diffusion denoising loop using 3D-to-2D projection and dual-branch diffusion with 3D-2D cross attentions.

  5. Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

    cs.CV 2026-07 conditional novelty 6.0

    A hierarchical global-keyframe then local-refinement diffusion pipeline produces stable 720p/30fps music-to-dance videos longer than one minute across five genres.

  6. Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

    cs.CV 2026-07 conditional novelty 6.0

    A hierarchical global-keyframe then local-refinement diffusion pipeline produces stable 720p/30fps music-to-dance videos longer than one minute across five genres.

  7. HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers

    cs.CV 2026-06 unverdicted novelty 6.0

    A framework for consistent long-horizon video relighting that propagates target latents across chunks and trains continuation via masked target-domain self-conditioning plus warm-start prompting.

  8. NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics

    cs.CV 2026-06 unverdicted novelty 6.0

    NEXUS introduces a graph-based neural energy-field model that derives forces from scalar energy and dissipation terms to achieve physically consistent contact-rich 3D dynamics.

  9. SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation

    cs.CV 2026-06 unverdicted novelty 6.0

    SpecLoR rectifies the amplitude spectrum of lookahead-estimated clean latents to natural-video priors during early ODE sampling steps, cutting physical artifacts with only four extra NFEs.

  10. LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    LaMo adds self-supervised latent motion priors via a motion drift loss during training and motion prior guidance during sampling to boost physical fidelity in video diffusion models like CogVideoX.

  11. PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation

    cs.CV 2026-05 conditional novelty 6.0

    PhyMotion scores generated human videos by grounding recovered 3D poses in a physics simulator across kinematic, contact, and dynamic axes, yielding stronger human correlation and larger RL post-training gains than pr...

  12. From Priors to Perception: Grounding Video-LLMs in Physical Reality

    cs.CV 2026-05 unverdicted novelty 6.0

    Video-LLMs fail physical reasoning due to semantic prior dominance rather than perception deficits; a new programmatic adversarial curriculum and visual-anchored reasoning chain enable substantial gains via standard L...

  13. CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    CoInteract adds a human-aware mixture-of-experts and spatially-structured co-generation to a diffusion transformer to synthesize videos with stable structures and physically plausible human-object contacts.

  14. REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion

    cs.CV 2025-12 conditional novelty 6.0

    Nonlinear multi-layer compression of frozen VFM patch semantics, jointly denoised with VAE latents, improves ImageNet 256x256 FID (12.9 vs 15.2 for REG at SiT-B/2, 400K) and accelerates convergence over REPA/ReDi/REG.

  15. Self-Forcing++: Towards Minute-Scale High-Quality Video Generation

    cs.CV 2025-10 conditional novelty 6.0

    Self-Forcing++ scales autoregressive video diffusion to over 4 minutes by using self-generated segments for guidance, reducing error accumulation and outperforming baselines in fidelity and consistency.

  16. Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

    cs.CV 2026-07 conditional novelty 5.0

    A hierarchical diffusion framework generates stable, music-synchronized 720p/30fps dance videos longer than a minute by planning sparse keyframes globally and filling them in locally.

  17. OptiWorld: Optimal Control for Video World Generation under Physical Constraints

    cs.CV 2026-05 unverdicted novelty 5.0

    OptiWorld inserts a classical optimal-control layer that extracts a world state, plans an optimal trajectory on a geometric manifold under physical constraints, and renders the video conditioned on that trajectory.

  18. Tempered Self-Similarity Alignment for Physically Plausible Video Generation

    cs.CV 2026-05 unverdicted novelty 5.0

    Tempered Self-similarity Alignment transfers relational structure from foundation-model STSS into video generators via probabilistic correspondence alignment, yielding reported gains in physical plausibility on VideoP...

  19. Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes

    cs.CV 2026-05 unverdicted novelty 5.0

    Introduces dual pose-image representation, cross-modal alignment, and iterative construction to improve prompt alignment and diversity in multi-person text-to-image generation.

  20. Reward-Forcing: Autoregressive Video Generation with Reward Feedback

    cs.CV 2026-01 unverdicted novelty 5.0

    Reward-Forcing guides autoregressive video generation with reward feedback to achieve performance comparable to teacher-dependent methods on benchmarks like VBench without relying on distillation.

  21. Motus: A Unified Latent Action World Model

    cs.CV 2025-12 unverdicted novelty 5.0

    Motus unifies understanding, video generation, and action in one latent world model via MoT experts and optical-flow latent actions, reporting gains over prior methods in simulation and real robots.

  22. MusicInfuser: Making Video Diffusion Listen and Dance

    cs.CV 2025-03 unverdicted novelty 5.0

    MusicInfuser uses a novel layer-wise adaptability criterion to adapt text-to-video diffusion models for generating music-synchronized dance videos with limited training on a single GPU.

  23. Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation

    cs.CV 2026-06 unverdicted novelty 4.0

    ST-DRC proposes latent in-context injection, TASS-RoPE, appearance-invariant augmentation, and three-stream guidance to improve identity preservation in text-to-video diffusion models built on LTX-2.3.