SwiftAudio performs caption-only distillation of a one-step TTA diffusion model by adapting VSD to audio with temporal smoothness regularization, achieving SOTA among one-step methods on AudioCaps and Clotho using ~45K captions.
A duality based approach for realtime tv-l 1 optical flow
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2representative citing papers
MEDN splits micro-expression video features into an Action-Unit-supervised motion branch and a sparse-attention emotion branch, then fuses them, reporting improved SAMM and CAS(ME)3 results—but the decoupling loss is mathematically flawed.
citing papers explorer
-
SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation
SwiftAudio performs caption-only distillation of a one-step TTA diffusion model by adapting VSD to audio with temporal smoothness regularization, achieving SOTA among one-step methods on AudioCaps and Clotho using ~45K captions.
-
MEDN: Motion-Emotion Feature Decoupling Network for Micro-Expression Recognition
MEDN splits micro-expression video features into an Action-Unit-supervised motion branch and a sparse-attention emotion branch, then fuses them, reporting improved SAMM and CAS(ME)3 results—but the decoupling loss is mathematically flawed.