CoMoGen generates controllable interactive video from mask sequences and images by encoding masks into MMDiT via MaskAdapter and LoRA on motion layers, claiming SOTA motion fidelity.
In: International Con- ference on Learning Representations (2022),https://openreview.net/forum?id= nZeVKeeFYf9
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CV 2years
2026 2representative citing papers
Decoupling spatial basis refinement from low-rank channel mixing in convolutional layers yields better parameter-efficient fine-tuning for vision foundation models.
citing papers explorer
-
CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration
CoMoGen generates controllable interactive video from mask sequences and images by encoding masks into MMDiT via MaskAdapter and LoRA on motion layers, claiming SOTA motion fidelity.
-
LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models
Decoupling spatial basis refinement from low-rank channel mixing in convolutional layers yields better parameter-efficient fine-tuning for vision foundation models.