MagicVideo generates 256x256 text-conditioned video clips via latent diffusion with a custom 3D U-Net, achieving roughly 64 times lower compute than prior video diffusion models.
MagicMix: Semantic Mixing with Diffusion Models, Oct
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CV 2verdicts
UNVERDICTED 2representative citing papers
ELDiff integrates evidential learning into T2I diffusion via pixel evidence loss and token conflict loss to improve object-wise semantic consistency.
citing papers explorer
-
MagicVideo: Efficient Video Generation With Latent Diffusion Models
MagicVideo generates 256x256 text-conditioned video clips via latent diffusion with a custom 3D U-Net, achieving roughly 64 times lower compute than prior video diffusion models.
-
ELDiff: When Evidential Learning Meets Text-to-Image Diffusion
ELDiff integrates evidential learning into T2I diffusion via pixel evidence loss and token conflict loss to improve object-wise semantic consistency.