REVIEW 4 major objections 5 minor 32 references
A small event gate on video diffusion can stop objects from moving before contact or drifting after placement.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 22:55 UTC pith:DBWNANPB
load-bearing objection A clean, practical DiT sampler fix for interaction failures; gains look real on their bench, but the causal-event claim still rests on a change-derived proxy and closed evaluation. the 4 major comments →
Event-Driven Video Generation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper claims that interaction failures in modern video DiTs are not inevitable scale problems but stem from unconstrained frame-first updates, and that a minimal event pathway—token-aligned activity prediction, event-grounded realization/consistency/ordering losses, and hysteresis-scheduled gated sampling—turns latent evolution into event-conditioned state transitions that substantially improve causal initiation, interaction realization, and stable postconditions without trading off appearance.
What carries the argument
Event-Driven Video Generation (EVD): a lightweight event head plus event-grounded losses and event-gated sampling (soft activation, hysteresis, early-step annealing) that modulates the DiT update field so latent state changes mainly where and when an interaction is active.
Load-bearing premise
That self-supervised pseudo-event targets from localized latent change, plus a learned activity head, are a faithful enough proxy for real causal interactions that gating them improves true contact and support structure rather than just suppressing non-local motion on an interaction-focused test set.
What would settle it
On held-out interaction prompts outside EVD-Bench, measure contact-before-motion rates, post-event drift, and support validity for DiT+EVD versus matched motion-masking and ungated baselines; if EVD no longer wins on those causal metrics while appearance stays flat, the event-grounding claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Event-Driven Video Generation (EVD), a DiT-compatible add-on that predicts token-aligned event activity, couples that activity to latent updates via realization/consistency/ordering losses, and applies hysteresis-scheduled event gating at sampling time. The central claim is that this modest event pathway reduces interaction hallucinations—state persistence, spatial accuracy, support relations, and contact stability—on an author-curated EVD-Bench of 150 interaction prompts, while preserving appearance relative to DiT-4B/30B baselines. Evidence includes matched-solver comparisons, component ablations (Table 3), motion-mask controls (Table 8), sensitivity sweeps (Table 5), human 2AFC preferences, and VBench Appearance/Dynamics scores, plus qualitative comparisons to closed external generators.
Significance. If the result holds beyond the custom interaction bench, EVD is a practical and transferable contribution: a small, solver-agnostic interface change that targets a widely observed failure mode of frame-first video diffusion without redesigning the backbone. Strengths include fully specified training/sampling algorithms (Algs. 1–2), explicit motion-masking controls, leakage audits for EVD-Bench, detailed human-eval protocol, and an honest limitations section. The work is significant as an engineering abstraction for interaction grounding, not as a full physics simulator; its value depends on whether the gains reflect causal event structure rather than localized motion suppression on prompts chosen for that signal.
major comments (4)
- [A.4, A.12; Eqs. (14)–(16); Table 8] Secs. A.4 and A.12 and Eqs. (14)–(16): pseudo-event targets are built from localized latent-change magnitude (with camera-mean subtraction), while Lreal and Lorder then penalize updates where activity is low. This makes part of the supervision definitionally close to “change only where change is.” Table 8’s motion-mask control is necessary and helpful, but still under-specifies what the learned head captures beyond change magnitude. Please report quantitative alignment of predicted ât with pseudo-targets vs. pure motion maps, and at least one analysis (e.g., event-map visualizations or failure cases where activity and motion diverge) showing that gating encodes interaction phase/causality rather than non-local motion suppression alone.
- [Sec. 5.1; A.10; Tables 1–2] Sec. 5.1 and A.10: the headline Dynamics gains (e.g., DiT-4B 78.9→94.8; human Dynamics ~96% favoring EVD) rest almost entirely on EVD-Bench, a 150-prompt author-built set filtered for single-event interaction structure and balanced to the paper’s four failure categories. Leakage audits reduce caption memorization risk, but do not address method–bench alignment: prompts were chosen for the same localized contact/placement/support phenomena that the event proxy is designed to detect. To support the general claim of reduced interaction hallucinations, add results on a public/general T2V suite (full VBench or equivalent) and a non-interaction control split showing that EVD does not merely trade motion diversity for score gains on interaction-centric prompts.
- [Tables 1–2; A.13] Tables 1–2 and A.13: human preference is reported as “% votes favoring EVD” against closed black-box APIs (Kling, Sora, Veo, etc.) under a normalization protocol that cannot match NFE, solver, or training compute. That protocol is documented carefully, but the tables currently read as head-to-head SOTA wins. Reframe these as black-box interaction-preference comparisons under standardized duration/resolution, and separate them from the matched DiT±EVD ablations that actually identify the method’s contribution. Without that separation, the strongest external claims overreach the experimental design.
- [Sec. 5; Fig. 2; Table 3] Sec. 5 and Fig. 2: the four failure categories (state persistence, spatial accuracy, support, contact) are the paper’s organizing taxonomy, yet quantitative results are only aggregate VBench Dynamics and overall human Dynamics preference. There is no per-category automatic or human breakdown on EVD-Bench. Given that ablations claim category-specific regressions (e.g., realization→contact; consistency→persistence), a per-category table is load-bearing for the claim that EVD systematically addresses those modes rather than improving a single global dynamics score.
minor comments (5)
- [Sec. 4.2; A.4] Notation drifts between main text and appendix: et vs ât, gbin vs gk, and Ce=1 main method vs multi-channel phase/type variants in A.4. Unify the primary instantiation used in all reported experiments.
- [Fig. 3; Sec. 4.2] Fig. 3 overview is dense; the soft-gate / hysteresis / schedule cascade would be clearer with a small numerical toy example of gate values across a contact onset.
- [Table 3] Table 3 “EVD wins %” against ablated variants is informative but asymmetric; also report absolute VBench/human scores for each ablation row for direct comparison to the baseline.
- [References; A.2] Several self-citations to concurrent arXiv notes (entropy-controlled flow matching, temporal pair consistency, corruption-aware training) appear in the backbone/notation sections without being necessary for EVD; trim or move to related work to avoid distraction.
- [Sec. 5.1; A.8] Clarify whether DiT-4B/30B are public checkpoints or internal reimplementations; reproducibility depends on this even if EVD modules are fully specified.
Circularity Check
Mild self-definitional coupling of “event activity” to latent change; empirical gains are not forced by construction and controls partially break the loop.
specific steps
-
self definitional
[Sec. 4.2–4.3 Eqs. (14),(16); App. A.3 (32)–(33); App. A.12 Eqs. (74)–(75)]
"Lreal = E[||(1−ãt)⊙Δt||²₂] … Lorder = E[||1[ãt<τon]⊙Δt||²₂ + ||1[ãt<τoff]⊙Δt||²₂]. … (i) No-event⇒no-update: et≈0⇒Δzt≈0. … We compute pseudo-event activity from token-level latent change … mτ,i = (1/C)||Tok(z^{τ+1}_1)_i − Tok(z^τ_1)_i||₁ … ˜mτ,i = max{0, mτ,i − mean_j mτ,j}."
Pseudo-event activity is defined from localized latent-change magnitude (with mean subtraction for camera motion). Realization and ordering losses then penalize updates wherever predicted activity is low. Under the base Flow-Matching target, the event head is therefore pushed to fire where ground-truth change is large, so “event-gated updates” partly reduce to “allow change where change-derived activity is high.” This is definitional coupling of the event variable to the quantity it is used to gate—not a free prediction of causal structure—though ablations and motion-mask controls show the full recipe is not identical to naive masking.
full rationale
EVD is an empirical methods paper, not a first-principles derivation that claims independent physical predictions. The only circularity-adjacent step is operational: event activity is tied to where latent state changes, and the realization/ordering losses then enforce no-event⇒no-update, so part of the training signal is close to “change where change is.” That is intentional method design, not a fitted constant renamed as a prediction, and it is not load-bearing via self-citation uniqueness theorems. The paper partially breaks pure tautology with (i) camera-suppressed localized pseudo-targets and motion-mask / inference-only controls that underperform full EVD (Tables 3 and 8), (ii) consistency, hysteresis, and early-step annealing that are not identical to the target definition, and (iii) external human/VBench comparisons. Self-citations of the author’s other arXiv notes are peripheral (notation/related work), not the central premise. Score 3 reflects one mild self-definitional step without a forced central result.
Axiom & Free-Parameter Ledger
free parameters (7)
- gate sharpness β
- hysteresis thresholds (τon, τoff)
- schedule cutoff t⋆ / t⋆_loss
- loss weights λreal, λcons, λorder
- event dropout pe
- time-weight decay κ
- CFG scale wcfg and NFE K
axioms (5)
- domain assumption Linear latent flow matching with velocity target vt = z1 − z0 is an adequate base generative objective for video DiTs.
- ad hoc to paper Meaningful interactions are localized in space-time and can be approximated by token-aligned activity from latent change (with camera suppression).
- ad hoc to paper Suppressing latent updates outside predicted events improves causal interaction fidelity rather than merely reducing motion diversity.
- domain assumption Early diffusion/flow steps determine coarse dynamics, so strong early event gating is preferable.
- domain assumption Human 2AFC on short clips and VBench Appearance/Dynamics are valid proxies for interaction realism.
invented entities (3)
-
Token-aligned event activity field ât / et
no independent evidence
-
Soft+hysteresis event gate with scheduled annealing
no independent evidence
-
EVD-Bench (150 interaction prompts, four failure categories)
no independent evidence
read the original abstract
Current text-to-video models can make individual frames look convincing while still getting simple interactions wrong: objects move before contact, an intended action is skipped, a placed object keeps drifting, or a support relation breaks. Our starting point is that standard frame-first denoising updates every latent region at every step, even when the prompt implies that only a local interaction should be active. We introduce Event-Driven Video Generation (EVD), a small DiT-compatible intervention that gives the sampler an explicit event signal. A lightweight head predicts token-level event activity; training losses tie that activity to latent state change; and event-gated sampling, with hysteresis and an early-step schedule, applies the update field mainly where an interaction is forming. On EVD-Bench, EVD improves human preference and VBench dynamics for state persistence, spatial accuracy, support relations, and contact stability, while keeping appearance quality comparable to the base model. The results suggest that a modest amount of event structure can correct several interaction failures that otherwise remain hidden behind good frame-level appearance.
Figures
Reference graph
Works this paper leans on
-
[1]
In: SIGGRAPH Asia 2024 Conference Papers
Bar-Tal, O., Chefer, H., Tov, O., Herrmann, C., Paiss, R., Zada, S., Ephrat, A., Hur, J., Liu, G., Raj, A., Li, Y., Rubinstein, M., Michaeli, T., Wang, O., Sun, D., Dekel, T., Mosseri, I.: Lumiere: A space-time diffusion model for video generation. In: SIGGRAPH Asia 2024 Conference Papers. SA ’24, Association for Computing Machinery, New York, NY, USA (20...
doi:10.1145/3680528 2024
-
[2]
ArXivabs/2311.15127(2023),https://api.semanticscholar.org/CorpusID: 26531255125
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D.: Stable video diffusion: Scaling latent video diffusion models to large datasets. ArXivabs/2311.15127(2023),https://api.semanticscholar.org/CorpusID: 26531255125
Pith/arXiv arXiv 2023
-
[3]
In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Brooks, T., Holynski, A., Efros, A.A.: InstructPix2Pix: Learning to Follow Im- age Editing Instructions . In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 18392–18402. IEEE Computer Society, Los Alamitos, CA, USA (Jun 2023).https://doi.org/10.1109/CVPR52729.2023. 01764,https : / / doi . ieeecomputersociety . org / 10 . 1...
-
[4]
In: Singh, A., Fazel, M., Hsu, D., Lacoste- Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J
Chefer, H., Singer, U., Zohar, A., Kirstain, Y., Polyak, A., Taigman, Y., Wolf, L., Sheynin, S.: VideoJAM: Joint appearance-motion representations for enhanced motion generation in video models. In: Singh, A., Fazel, M., Hsu, D., Lacoste- Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., Zhu, J. (eds.) Proceedings of the 42nd International Conference...
2025
-
[5]
In: NeurIPS 2021 Work- shop on Deep Generative Models and Downstream Applications (2021),https: //openreview.net/forum?id=qw8AKxfYbI25
Ho, J., Salimans, T.: Classifier-free diffusion guidance. In: NeurIPS 2021 Work- shop on Deep Generative Models and Downstream Applications (2021),https: //openreview.net/forum?id=qw8AKxfYbI25
2021
-
[6]
In: Chaudhuri, K., Salakhutdinov, R
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Ges- mundo, A., Attariyan, M., Gelly, S.: Parameter-efficient transfer learning for NLP. In: Chaudhuri, K., Salakhutdinov, R. (eds.) Proceedings of the 36th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 97, pp. 2790–2799. PMLR (09–15 ...
2019
-
[7]
In: International Con- ference on Learning Representations (2022),https://openreview.net/forum?id= nZeVKeeFYf932
Hu, E.J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: LoRA: Low-rank adaptation of large language models. In: International Con- ference on Learning Representations (2022),https://openreview.net/forum?id= nZeVKeeFYf932
2022
-
[8]
Huang, H., Ma, G., Duan, N., Chen, X., Wan, C., Ming, R., Wang, T., Wang, B., Lu, Z., Li, A., Zeng, X., Zhang, X., Yu, G., Yin, Y., Wu, Q., Sun, W., An, K., Han, X., Sun, D., Ji, W., Huang, B., Li, B., Wu, C., Huang, G., Xiong, H., He, J., Wu, J., Yuan, J., Wu, J., Liu, J., Guo, J., Tan, K., Chen, L., Chen, Q., Sun, R., Yuan, S., Yin, S., Liu, S., Chen, W...
Pith/arXiv arXiv 2025
-
[9]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., Wang, Y., Chen, X., Wang, L., Lin, D., Qiao, Y., Liu, Z.: Vbench: Comprehensive benchmark suite for video generative models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 21807–21818 (June 2024) 2, 25
2024
-
[10]
Huang, Z., Zhang, F., Xu, X., He, Y., Yu, J., Dong, Z., Ma, Q., Chanpaisit, N., Si, C., Jiang, Y., Wang, Y., Chen, X., Chen, Y.C., Wang, L., Lin, D., Qiao, Y., Liu, Z.: VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models . IEEE Transactions on Pattern Analysis & Machine Intelligence48(03), 3268–3285 (Mar 2026).https://doi.org...
-
[11]
In: The Thirteenth International Conference on Learning Representations (2025), https://openreview.net/forum?id=66NzcRQuOq2
Jin, Y., Sun, Z., Li, N., Xu, K., Xu, K., Jiang, H., Zhuang, N., Huang, Q., Song, Y., MU, Y., Lin, Z.: Pyramidal flow matching for efficient video generative modeling. In: The Thirteenth International Conference on Learning Representations (2025), https://openreview.net/forum?id=66NzcRQuOq2
2025
-
[12]
Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., Wu, K., Lin, Q., Yuan, J., Long, Y., Wang, A., Wang, A., Li, C., Huang, D., Yang, F., Tan, H., Wang, H., Song, J., Bai, J., Wu, J., Xue, J., Wang, J., Wang, K., Liu, M., Li, P., Li, S., Wang, W., Yu, W., Deng, X., Li, Y., Chen, Y., Cui, Y., Peng, Y., Yu, Z., H...
Pith/arXiv arXiv 2025
-
[13]
In: The Thirty-eighth Annual Conference on Neural Information Processing Systems (2024),https://openreview.net/forum?id=53daI9kbvf2
Li, J., Feng, W., Fu, T.J., Wang, X., Basu, S., Chen, W., Wang, W.Y.: T2v- turbo: Breaking the quality bottleneck of video consistency model with mixed reward feedback. In: The Thirty-eighth Annual Conference on Neural Information Processing Systems (2024),https://openreview.net/forum?id=53daI9kbvf2
2024
-
[14]
In: The Thirteenth International Conference on Learning Representations (2025),https://openreview.net/forum?id=BZwXMqu4zG2
Li, J., Long, Q., Zheng, J., Gao, X., Piramuthu, R., Chen, W., Wang, W.Y.: T2v- turbo-v2: Enhancing video model post-training through data, reward, and condi- tional guidance design. In: The Thirteenth International Conference on Learning Representations (2025),https://openreview.net/forum?id=BZwXMqu4zG2
2025
-
[15]
In: The Eleventh International Conference on Learning Representations (2023),https://openreview.net/forum?id=PqvMRDCJT9t4
Lipman, Y., Chen, R.T.Q., Ben-Hamu, H., Nickel, M., Le, M.: Flow matching for generative modeling. In: The Eleventh International Conference on Learning Representations (2023),https://openreview.net/forum?id=PqvMRDCJT9t4
2023
-
[16]
In: Avidan, S., Brostow, G., Cissé, M., Farinella, G.M., Hassner, T
Liu, N., Li, S., Du, Y., Torralba, A., Tenenbaum, J.B.: Compositional visual gen- eration with composable diffusion models. In: Avidan, S., Brostow, G., Cissé, M., Farinella, G.M., Hassner, T. (eds.) Computer Vision – ECCV 2022. pp. 423–439. Springer Nature Switzerland, Cham (2022) 25 18 C. Maduabuchi et al
2022
-
[17]
Ma, G., Huang, H., Yan, K., Chen, L., Duan, N., Yin, S., Wan, C., Ming, R., Song, X., Chen, X., Zhou, Y., Sun, D., Zhou, D., Zhou, J., Tan, K., An, K., Chen, M., Ji, W., Wu, Q., Sun, W., Han, X., Wei, Y., Ge, Z., Li, A., Wang, B., Huang, B., Wang, B., Li, B., Miao, C., Xu, C., Wu, C., Yu, C., Shi, D., Hu, D., Liu, E., Yu, G., Yang, G., Huang, G., Yan, G.,...
Pith/arXiv arXiv 2025
-
[18]
Maduabuchi, C.: Entropy-controlled flow matching (2026),https://arxiv.org/ abs/2602.222654
Pith/arXiv arXiv 2026
-
[19]
Maduabuchi, C., Chen, H., Han, Y., Wang, J.: Corruption-aware training of latent video diffusion models for robust text-to-video generation (2026),https://arxiv. org/abs/2505.2154525, 49
arXiv 2026
-
[20]
Maduabuchi, C., Wang, J.: Temporal pair consistency for variance-reduced flow matching (2026),https://arxiv.org/abs/2602.049084
arXiv 2026
-
[21]
In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV)
Peebles, W., Xie, S.: Scalable Diffusion Models with Transformers . In: 2023 IEEE/CVF International Conference on Computer Vision (ICCV). pp. 4172–4182. IEEE Computer Society, Los Alamitos, CA, USA (Oct 2023).https://doi.org/ 10.1109/ICCV51070.2023.00387,https://doi.ieeecomputersociety.org/10. 1109/ICCV51070.2023.0038732
-
[22]
arXiv preprint arXiv:2410.13720 (2024) 2, 3, 5, 24
Polyak, A., Zohar, A., Brown, A., Tjandra, A., Sinha, A., Lee, A., Vyas, A., Shi, B., Ma, C.Y., Chuang, C.Y., et al.: Movie gen: A cast of media foundation models. arXiv preprint arXiv:2410.13720 (2024) 2, 3, 5, 24
Pith/arXiv arXiv 2024
-
[23]
Qin, Y., Shi, Z., Yu, J., Wang, X., Zhou, E., Li, L., Yin, Z., Liu, X., Sheng, L., Shao, J., BAI, L., Ouyang, W., Zhang, R.: Worldsimbench: Towards video generation models as world simulators (2025),https://openreview.net/forum? id=ejGAytoWoe2
2025
-
[24]
In: 2022 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR)
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-Resolution Image Synthesis with Latent Diffusion Models . In: 2022 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR). pp. 10674–10685. IEEE Computer Society, Los Alamitos, CA, USA (Jun 2022).https://doi.org/ 10.1109/CVPR52688.2022.01042,https://doi.ieeecomputersociety...
-
[25]
Wan, T., Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.W., Chen, D., Yu, F., Zhao, H., Yang, J., Zeng, J., Wang, J., Zhang, J., Zhou, J., Wang, J., Chen, J., Zhu, K., Zhao, K., Yan, K., Huang, L., Feng, M., Zhang, N., Li, P., Wu, P., Chu, R., Feng, R., Zhang, S., Sun, S., Fang, T., Wang, T., Gui, T., Weng, T., Shen, T., Lin, W., Wang, W., Wang, W., Zhou, W.,...
Pith/arXiv arXiv 2025
-
[26]
Wu, B., Zou, C., Li, C., Huang, D., Yang, F., Tan, H., Peng, J., Wu, J., Xiong, J., Jiang, J., Linus, Patrol, Zhang, P., Chen, P., Zhao, P., Tian, Q., Liu, S., Kong, W., Wang, W., He, X., Li, X., Deng, X., Zhe, X., Li, Y., Long, Y., Peng, Y., Wu, Y., Liu, Y., Wang, Z., Dai, Z., Peng, B., Li, C., Gong, G., Xiao, G., Tian, J., Lin, J., Liu, J., Zhang, J., L...
Pith/arXiv arXiv 2025
-
[27]
In: The Thirteenth International Conference on Learning Represen- tations (2025),https://openreview.net/forum?id=LQzN6TRFg92, 3
Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., Yin, D., Yuxuan.Zhang, Wang, W., Cheng, Y., Xu, B., Gu, X., Dong, Y., Tang, J.: Cogvideox: Text-to-video diffusion models with an expert transformer. In: The Thirteenth International Conference on Learning Represen- tations (2025),https://openreview.net/fo...
2025
-
[28]
In: 2023 IEEE/CVF International Conference on Computer Vi- sion (ICCV)
Zhang, L., Rao, A., Agrawala, M.: Adding conditional control to text-to-image diffusion models. In: 2023 IEEE/CVF International Conference on Computer Vi- sion (ICCV). pp. 3813–3824 (2023).https://doi.org/10.1109/ICCV51070.2023. 0035532
-
[29]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
Zhang, L., Rao, A., Agrawala, M.: Adding conditional control to text-to-image diffusion models. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 3836–3847 (October 2023) 32, 33
2023
-
[30]
Zheng, D., Huang, Z., Liu, H., Zou, K., He, Y., Zhang, F., Gu, L., Zhang, Y., He, J., Zheng,W.S.,Qiao,Y.,Liu,Z.:Vbench-2.0:Advancingvideogenerationbenchmark suite for intrinsic faithfulness (2025),https://arxiv.org/abs/2503.217552
Pith/arXiv arXiv 2025
-
[31]
In: Thirty-seventh Conference on Neu- ral Information Processing Systems (2023),https://openreview.net/forum?id= 9fWKExmKa05, 40
Zheng, K., Lu, C., Chen, J., Zhu, J.: DPM-solver-v3: Improved diffusion ODE solver with empirical model statistics. In: Thirty-seventh Conference on Neu- ral Information Processing Systems (2023),https://openreview.net/forum?id= 9fWKExmKa05, 40
2023
-
[32]
Zheng, Z., Peng, X., Yang, T., Shen, C., Li, S., Liu, H., Zhou, Y., Li, T., You, Y.: Open-sora: Democratizing efficient video production for all (2024),https: //arxiv.org/abs/2412.204042, 3, 5 20 C. Maduabuchi et al. A Appendix Table of Contents Abbreviations and symbols.................................................21 Backbone and Notation................
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.