A 1,200-prompt benchmark across six world-knowledge domains reports that ten state-of-the-art text-to-video models average below 0.70 on a 0 to 1 scale for producing videos consistent with real-world knowledge.
High-Order Matching for One-Step Shortcut Diffusion Models
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
One-step shortcut diffusion models [Frans, Hafner, Levine and Abbeel, ICLR 2025] have shown potential in vision generation, but their reliance on first-order trajectory supervision is fundamentally limited. The Shortcut model's simplistic velocity-only approach fails to capture intrinsic manifold geometry, leading to erratic trajectories, poor geometric alignment, and instability-especially in high-curvature regions. These shortcomings stem from its inability to model mid-horizon dependencies or complex distributional features, leaving it ill-equipped for robust generative modeling. In this work, we introduce HOMO (High-Order Matching for One-Step Shortcut Diffusion), a game-changing framework that leverages high-order supervision to revolutionize distribution transportation. By incorporating acceleration, jerk, and beyond, HOMO not only fixes the flaws of the Shortcut model but also achieves unprecedented smoothness, stability, and geometric precision. Theoretically, we prove that HOMO's high-order supervision ensures superior approximation accuracy, outperforming first-order methods. Empirically, HOMO dominates in complex settings, particularly in high-curvature regions where the Shortcut model struggles. Our experiments show that HOMO delivers smoother trajectories and better distributional alignment, setting a new standard for one-step generative models.
fields
cs.CV 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation
A 1,200-prompt benchmark across six world-knowledge domains reports that ten state-of-the-art text-to-video models average below 0.70 on a 0 to 1 scale for producing videos consistent with real-world knowledge.