REVIEW 4 major objections 2 minor 96 references
Numerical Study of Oblique Detonation Initiation Assisted by Local Energy Deposition
T0 review · 4 major / 2 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper claims that pulsatile local energy deposition sustains oblique detonation on a finite wedge with less than 10% of the average power continuous deposition needs.
desk verdict The submission is unusable: the abstract is a fluid-dynamics study, but the full text is an unrelated AI image-generation paper, so there is nothing to assess. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the local energy deposition region placed near the finite wedge, which models the thermal effects of plasma-based initiation assistance. The mechanism being studied is how heat added ahead of or on the wedge couples with the oblique shock to form and sustain a detonation wave. For continuous deposition, the controlling parameter is the deposition power; for pulsatile deposition, the controlling parameters are single-pulse energy and pulse repetition frequency. The spatiotemporal evolution of the primary wave structures supplies the criterion for the minimum repetition frequency, and multi-pulse simulations are the verification step that closes the argument.
What would settle it
A multi-pulse simulation at the claimed minimum repetition frequency and average power, run with a different grid resolution or chemical mechanism, that either fails to sustain the on-wedge detonation or produces a longer initiation length would falsify the claim. The same test could be done experimentally by measuring the initiation length and detonation sustainability for pulsed deposition at an average power of 10% of the continuous threshold.
Extended reading notes
Core claim
The paper reports that without energy deposition, oblique detonation initiation fails on a finite wedge at low Mach numbers; with either continuous or pulsatile local energy deposition, a sustainable oblique detonation can be established on the wedge. As the continuous deposition power or the single-pulse energy increases, the initiations pass through a sequence of distinct modes, and evolution of the main wave structures under single-pulse deposition reveals a minimum pulse repetition frequency for maintaining the on-wedge detonation. Multi-pulse simulations confirm that this frequency keeps the detonation sustainable. The headline quantitative finding is that pulsatile deposition reaches the same initiation length as continuous deposition while consuming less than 10% of the average power, making pulsed energy the efficient route for initiation assistance on finite wedges under extreme flight conditions.
Load-bearing premise
The load-bearing premise is that the numerical setup—the finite wedge geometry, inflow conditions, and the way energy deposition is modelled—represents real oblique-detonation-engine flight well enough that the 10% power comparison carries over beyond the simulated cases.
Editorial extensions
If this is right
- If the result transfers to real engines, pulsed plasma deposition could replace continuous deposition for ODW initiation at low Mach numbers, cutting initiation-assistance power demand by more than an order of magnitude.
- The identified minimum pulse repetition frequency gives engine designers a concrete control parameter for keeping an on-wedge detonation sustainable.
- The sequential initiation modes with increasing deposition power or pulse energy provide a map for choosing operating points that avoid marginal or failed initiation.
- Because the same initiation length is preserved at a fraction of the average power, pulsatile assistance can be integrated without lengthening the wedge or the engine.
- Without any deposition, the finite wedge is predicted to fail initiation at low Mach, so energy assistance becomes a necessary rather than optional component in that flight regime.
Reading between the lines
- The 10% ratio is reported for the simulated conditions; an inference is that the gap between pulsed and continuous power may widen or shrink with deposition location, pulse shape, or wedge angle, since those parameters are not varied in the abstract's headline comparison.
- The minimum pulse repetition frequency criterion suggests a resonance-like coupling between pulse timing and detonation wave structure, which could be probed directly by measuring initiation length as a function of frequency.
- An experimental shock-tube test with pulsed laser or plasma deposition on a finite wedge could test whether the simulated 10% average-power advantage survives in a real reacting flow.
- For high-altitude low-Mach flight, the practical implication is that the engine's initiation system should be specified by time-averaged power plus repetition frequency, not just peak energy.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission, arXiv:2508.08943, presents an abstract in physics.flu-dyn claiming a numerical study of oblique detonation wave (ODW) initiation assisted by local energy deposition on a finite wedge, asserting that pulsatile energy deposition can sustain on-wedge detonation at less than 10% of the continuous-deposition power while maintaining the same initiation length. However, the full text of the manuscript is an unrelated paper titled "Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation," which is a computer vision manuscript about AI image generation. The body contains no equations, simulation setup, wedge geometry, inflow conditions, energy deposition model, grid resolution, chemistry or turbulence closure, convergence studies, or any results relevant to oblique detonation waves. The central claims of the abstract are therefore entirely unsupported by the submitted manuscript text.
Significance. If the claimed results were present and valid, the finding that pulsatile energy deposition consumes less than 10% of the continuous-deposition power for the same initiation length could be of practical interest for oblique detonation engine initiation at low Mach number and high altitude. The abstract also identifies a physically sensible mode sequence and a minimum repetition frequency estimation procedure using single-pulse simulations followed by multi-pulse verification. However, none of this content is present in the submitted manuscript. Because the scientific content described in the abstract is completely absent, the significance of the work cannot be assessed. The manuscript cannot be evaluated on its merits, and no strength such as reproducible code, machine-checked proofs, or falsifiable predictions can be credited.
major comments (4)
- [Full text (entire manuscript)] The full text of the submitted manuscript is the paper "Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation" (arXiv:2508.08949), which is a computer vision paper on diffusion transformers for image generation. This body text shares no content with the physics.flu-dyn abstract on oblique detonation wave initiation. There is no numerical solver description, no governing equations, no wedge geometry, no inflow conditions, no energy deposition model, no grid resolution, no chemistry or turbulence closure, no convergence studies, and no simulation results. The central claim of the abstract — that pulsatile energy deposition sustains on-wedge detonation at less than 10% of the continuous-deposition power with the same initiation length — is asserted without any supporting derivation, data, or methodology in the manuscript.
- [Abstract, final sentence] The claim that "sustainable on-wedge detonation can be achieved by pulsatile energy deposition with an average power consumption of less than 10% of that required for continuous energy deposition while maintaining a same initiation length" is a quantitative efficiency comparison that requires precise definitions of "average power consumption," "initiation length," and the parameter sets for both continuous and pulsatile deposition. None of these definitions appear in the manuscript, and no simulation data are presented that could substantiate the 10% figure. The claim is unverifiable from the submitted text.
- [Abstract, spatiotemporal evolution sentence] The abstract states that "Analysis of the spatiotemporal evolution of the primary wave structures under single-pulse energy deposition reveals the minimum pulse repetition frequency required for sustainable on-wedge detonation, which is subsequently verified through multi-pulse energy deposition simulations." This describes a two-step procedure that is central to the paper's methodology. The manuscript provides neither the single-pulse analysis nor the multi-pulse verification, nor any description of how the minimum repetition frequency is extracted from the spatiotemporal evolution. The absence of this material places the entire scientific argument outside the submitted text.
- [Full text (all sections)] The manuscript contains no limitations section, no error analysis, and no discussion of the validity of the simulation approach. Given that the body text is unrelated to the abstract, the work as submitted is internally inconsistent. No amount of revision to the present text could make the central claim assessable; the submission would require replacement by an entirely different manuscript containing the actual numerical study described in the abstract.
minor comments (2)
- [Abstract, first sentence] The abstract refers to "a fixed-angle wedge" and later to "on-wedge initiation," but the body text contains no wedge geometry or coordinate system, so these terms are undefined in the submitted document.
- [Abstract, methods sentence] The phrase "plasma-based initiation assistance techniques" suggests a specific physical model of energy deposition, but no plasma model, energy source term, or timescale is described anywhere in the manuscript.
Circularity Check
No circularity can be identified; the submitted full text does not contain the claimed study's methods or results.
full rationale
The abstract reports a numerical study of oblique detonation initiation with local energy deposition, but the provided full text is an unrelated paper on AI-based story image generation (Lay2Story). There is therefore no derivation chain, no equations, no fitted parameters, no self-citations, and no imported uniqueness theorem available to inspect. From the abstract alone, the described procedure—single-pulse analysis yields a minimum pulse repetition frequency, which is then verified in multi-pulse simulations—is an internal consistency check rather than a quantity defined to equal an input. No fitted constant is renamed as a prediction, and no result is forced by definition or by self-citation. The absence of the actual manuscript text and simulation data is a verifiability and integrity concern, not a circularity finding. Per the hard rules, circularity may only be claimed when the paper itself can be quoted to exhibit a specific reduction, so the honest outcome is no circularity, score 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The numerical governing equations and combustion chemistry models used in the simulations adequately capture oblique detonation initiation physics.
- domain assumption The energy deposition model represents the thermal effect of plasma-based initiation assistance.
Cite this review
Pith. "Pith review of Numerical Study of Oblique Detonation Initiation Assisted by Local Energy Deposition." pith.science (2026). https://pith.science/paper/GWXMXEKH
@misc{pith2026250808943,
author = {Pith},
title = {Pith review of: Numerical Study of Oblique Detonation Initiation Assisted by Local Energy Deposition},
year = {2026},
howpublished = {\url{https://pith.science/paper/GWXMXEKH}},
note = {Machine review of arXiv:2508.08943}
}
read the original abstract
Reliable initiation of oblique detonation waves (ODWs) is crucial for the stable operation of oblique detonation engines (ODEs), especially under flight conditions of low Mach numbers and/or high altitudes. In this case, conventional initiation approaches relying solely on a fixed-angle wedge may engender risks of initiation failure, which necessitates extra initiation assistance measures. In this study, ODW initiation over a finite wedge with local energy deposition is numerically investigated to assess the thermal effects of plasma-based initiation assistance techniques. Particular emphasis is put on the effects of forms and magnitudes of energy deposition on initiation modes and flow field structures of ODWs. The results demonstrate that on-wedge initiation of ODWs fails at a low Mach number without any energy depositions. In contrast, both continuous and pulsatile local energy depositions can effectively initiate ODWs, leading to sustainable detonation on the finite wedge. As continuous energy deposition power or pulsatile single pulse energy increases, several key detonation initiation modes emerge sequentially. Analysis of the spatiotemporal evolution of the primary wave structures under single-pulse energy deposition reveals the minimum pulse repetition frequency required for sustainable on-wedge detonation, which is subsequently verified through multi-pulse energy deposition simulations. Nevertheless, it is found that sustainable on-wedge detonation can be achieved by pulsatile energy deposition with an average power consumption of less than 10% of that required for continuous energy deposition while maintaining a same initiation length, suggesting that the pulsatile one is an efficient way of energy deposition for initiation assistance of ODWs on finite wedges under extreme flight conditions.
Reference graph
Works this paper leans on
-
[5]
Customttt: Motion and ap- pearance customized video generation via test-time training
Xiuli Bi, Jian Lu, Bo Liu, Xiaodong Cun, Yong Zhang, Weisheng Li, and Bin Xiao. Customttt: Motion and ap- pearance customized video generation via test-time training. In Proceedings of the AAAI Conference on Artificial Intelli- gence, pages 1871–1879, 2025. 2
2025
-
[6]
Relactrl: Relevance-guided efficient control for diffusion transformers
Ke Cao, Jing Wang, Ao Ma, Jiasong Feng, Zhanjie Zhang, Xuanhua He, Shanyuan Liu, Bo Cheng, Dawei Leng, Yuhui Yin, et al. Relactrl: Relevance-guided efficient control for diffusion transformers. arXiv preprint arXiv:2502.14377 ,
-
[7]
Gamegen-x: Interactive open-world game video generation
Haoxuan Che, Xuanhua He, Quande Liu, Cheng Jin, and Hao Chen. Gamegen-x: Interactive open-world game video generation. arXiv preprint arXiv:2411.00769, 2024. 3
arXiv 2024
-
[8]
Pixart- � : Fast training of diffusion transformer for photorealistic text-to-image synthesis
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al. Pixart- � : Fast training of diffusion transformer for photorealistic text-to-image synthesis. arXiv preprint arXiv:2310.00426, 2023. 2, 9, 11
-
[9]
Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, et al. Panda-70m: Captioning 70m videos with multiple cross-modality teachers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 13320–13331, 2024. 3
2024
-
[10]
Ctr-driven advertising image generation with multimodal large language models
Xingye Chen, Wei Feng, Zhenbang Du, Weizhen Wang, Yanyin Chen, Haohan Wang, Linkai Liu, Yaoyu Li, Jinyuan Zhao, Yu Li, et al. Ctr-driven advertising image generation with multimodal large language models. In Proceedings of the ACM on Web Conference 2025, pages 2262–2275, 2025. 11
2025
-
[11]
Scaling recti- fied flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling recti- fied flow transformers for high-resolution image synthesis. In Forty-first international conference on machine learning,
-
[12]
FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance
Jiasong Feng, Ao Ma, Jing Wang, Bo Cheng, Xiaodan Liang, Dawei Leng, and Yuhui Yin. Fancyvideo: Towards dynamic and consistent video generation via cross-frame textual guid- ance. arXiv preprint arXiv:2408.08189, 2024. 3
work page Pith review arXiv 2024
Show all 96 references
-
[13]
Dream- sim: Learning new dimensions of human visual similar- ity using synthetic data
Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dream- sim: Learning new dimensions of human visual similar- ity using synthetic data. arXiv preprint arXiv:2306.09344 ,
-
[14]
An image is worth one word: Personalizing text-to- image generation using textual inversion
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patash- nik, Amit H Bermano, Gal Chechik, and Daniel Cohen- Or. An image is worth one word: Personalizing text-to- image generation using textual inversion. arXiv preprint arXiv:2208.01618, 2022. 9
2022 arXiv
-
[15]
Check locate rectify: A training- free layout calibration system for text-to-image generation
Biao Gong, Siteng Huang, Yutong Feng, Shiwei Zhang, Yuyuan Li, and Yu Liu. Check locate rectify: A training- free layout calibration system for text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6624–6634, 2024. 9
2024
-
[16]
Variational au- toencoder: An unsupervised model for encoding and decod- ing fmri activity in visual cortex.NeuroImage, 198:125–136,
Kuan Han, Haiguang Wen, Junxing Shi, Kun-Han Lu, Yizhen Zhang, Di Fu, and Zhongming Liu. Variational au- toencoder: An unsupervised model for encoding and decod- ing fmri activity in visual cortex.NeuroImage, 198:125–136,
-
[17]
Anystory: Towards unified single and multiple subject personalization in text-to-image generation
Junjie He, Yuxiang Tuo, Binghui Chen, Chongyang Zhong, Yifeng Geng, and Liefeng Bo. Anystory: Towards unified single and multiple subject personalization in text-to-image generation. arXiv preprint arXiv:2501.09503, 2025. 2
2025 arXiv
-
[18]
Freeedit: Mask-free reference-based image editing with multi-modal instruction
Runze He, Kai Ma, Linjiang Huang, Shaofei Huang, Jialin Gao, Xiaoming Wei, Jiao Dai, Jizhong Han, and Si Liu. Freeedit: Mask-free reference-based image editing with multi-modal instruction. arXiv preprint arXiv:2409.18071 ,
-
[19]
Plangen: Towards unified layout planning and image generation in auto-regressive vision language models
Runze He, Bo Cheng, Yuhang Ma, Qingxiang Jia, Shanyuan Liu, Ao Ma, Xiaoyu Wu, Liebucha Wu, Dawei Leng, and Yuhui Yin. Plangen: Towards unified layout planning and image generation in auto-regressive vision language models. arXiv preprint arXiv:2503.10127, 2025. 3
2025 arXiv
-
[20]
Context- aware layout to image generation with enhanced object ap- pearance
Sen He, Wentong Liao, Michael Ying Yang, Yongxin Yang, Yi-Zhe Song, Bodo Rosenhahn, and Tao Xiang. Context- aware layout to image generation with enhanced object ap- pearance. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 15049– 1...
2021
-
[21]
Style aligned image generation via shared atten- tion
Amir Hertz, Andrey V oynov, Shlomi Fruchter, and Daniel Cohen-Or. Style aligned image generation via shared atten- tion. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 4775–4785,
-
[22]
Clipscore: A reference-free evaluation met- ric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation met- ric for image captioning. arXiv preprint arXiv:2104.08718,
-
[23]
Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. Advances in neural information processing systems , 30, 2017. 7
2017
-
[24]
Interactdiffusion: Interaction con- trol in text-to-image diffusion models
Jiun Tian Hoe, Xudong Jiang, Chee Seng Chan, Yap-Peng Tan, and Weipeng Hu. Interactdiffusion: Interaction con- trol in text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6180–6189, 2024. 9
2024
-
[25]
Learning disentangled iden- tifiers for action-customized text-to-image generation
Siteng Huang, Biao Gong, Yutong Feng, Xi Chen, Yuqian Fu, Yu Liu, and Donglin Wang. Learning disentangled iden- tifiers for action-customized text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 7797–7806, 2024. 9
2024
-
[26]
Reversion: Diffusion-based relation inversion from images
Ziqi Huang, Tianxing Wu, Yuming Jiang, Kelvin CK Chan, and Ziwei Liu. Reversion: Diffusion-based relation inversion from images. In SIGGRAPH Asia 2024 Conference Papers, pages 1–11, 2024. 9
2024
-
[27]
Gpt-4o system card
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perel- man, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Weli- hinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024. 4
2024 arXiv
-
[28]
Res-tuning: A flexible and efficient tuning paradigm via unbinding tuner from backbone
Zeyinzi Jiang, Chaojie Mao, Ziyuan Huang, Ao Ma, Yiliang Lv, Yujun Shen, Deli Zhao, and Jingren Zhou. Res-tuning: A flexible and efficient tuning paradigm via unbinding tuner from backbone. Advances in Neural Information Processing Systems, 36:42689–42716, 2023. 3
2023
-
[29]
Miradata: A large-scale video dataset with long durations and structured captions
Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xin- tao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. Miradata: A large-scale video dataset with long durations and structured captions. Advances in Neural Information Processing Systems, 37:48955–48970, 2025. 3
2025
-
[30]
Story generation with crowdsourced plot graphs
Boyang Li, Stephen Lee-Urban, George Johnston, and Mark Riedl. Story generation with crowdsourced plot graphs. In Proceedings of the AAAI Conference on Artificial Intelli- gence, pages 598–604, 2013. 2
2013
-
[31]
Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing
Dongxu Li, Junnan Li, and Steven Hoi. Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing. Advances in Neural Information Pro- cessing Systems, 36:30146–30166, 2023. 2, 6, 7, 8
2023
-
[32]
Planning and rendering: Towards prod- uct poster generation with diffusion models
Zhaochen Li, Fengheng Li, Wei Feng, Honghe Zhu, Yaoyu Li, Zheng Zhang, Jingjing Lv, Junjie Shen, Zhangang Lin, Jingping Shao, et al. Planning and rendering: Towards prod- uct poster generation with diffusion models. arXiv preprint arXiv:2312.08822, 2023. 11
2023 arXiv
-
[33]
Photomaker: Customizing re- alistic human photos via stacked id embedding
Zhen Li, Mingdeng Cao, Xintao Wang, Zhongang Qi, Ming- Ming Cheng, and Ying Shan. Photomaker: Customizing re- alistic human photos via stacked id embedding. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8640–8650, 2024. 9
2024
-
[34]
Ragar: Retrieval augment person- alized image generation guided by recommendation
Run Ling, Wenji Wang, Yuting Liu, Guibing Guo, Linying Jiang, and Xingwei Wang. Ragar: Retrieval augment person- alized image generation guided by recommendation. arXiv preprint arXiv:2505.01657, 2025. 2
2025 arXiv
-
[35]
Intelligent grimm-open-ended visual storytelling via latent diffusion models
Chang Liu, Haoning Wu, Yujie Zhong, Xiaoyun Zhang, Yan- feng Wang, and Weidi Xie. Intelligent grimm-open-ended visual storytelling via latent diffusion models. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6190–6200, 2024. 2, 3, 6, 7, 8
2024
-
[36]
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In European Conference on Computer Vision , pages 38–55. Springe...
2024
-
[37]
Bridge dif- fusion model: Bridge chinese text-to-image diffusion model with english communities
Shanyuan Liu, Bo Cheng, Yuhang Ma, Liebucha Wu, Ao Ma, Xiaoyu Wu, Dawei Leng, and Yuhui Yin. Bridge dif- fusion model: Bridge chinese text-to-image diffusion model with english communities. In Proceedings of the AAAI Con- ference on Artificial Intelligence, pages 5541–5549, 2025. 2
2025
-
[38]
One-prompt-one-story: Free-lunch consistent text-to-image generation using a single prompt
Tao Liu, Kai Wang, Senmao Li, Joost van de Weijer, Fa- had Shahbaz Khan, Shiqi Yang, Yaxing Wang, Jian Yang, and Ming-Ming Cheng. One-prompt-one-story: Free-lunch consistent text-to-image generation using a single prompt. arXiv preprint arXiv:2501.13554, 2025. 2, 6, 7, 8, 9
2025 arXiv
-
[39]
Recent advances in ood detection: Problems and approaches
Shuo Lu, Yingsheng Wang, Lijun Sheng, Aihua Zheng, Lingxiao He, and Jian Liang. Recent advances in ood detection: Problems and approaches. arXiv preprint arXiv:2409.11884, 2024. 2
2024 arXiv
-
[40]
Uni-layout: Integrating human feedback in unified layout generation and evaluation
Shuo Lu, Yanyin Chen, Wei Feng, Jiahao Fan, Fengheng Li, Zheng Zhang, Jingjing Lv, Junjie Shen, Ching Law, and Jian Liang. Uni-layout: Integrating human feedback in unified layout generation and evaluation. arXiv preprint arXiv:2508.02374, 2025. 2
2025 arXiv
-
[41]
Unified multi-modal latent diffusion for joint subject and text conditional image generation.arXiv preprint arXiv:2303.09319, 2023
Yiyang Ma, Huan Yang, Wenjing Wang, Jianlong Fu, and Jiaying Liu. Unified multi-modal latent diffusion for joint subject and text conditional image generation.arXiv preprint arXiv:2303.09319, 2023. 9
2023 arXiv
-
[42]
Hico: Hierarchical controllable diffu- sion model for layout-to-image generation.Advances in Neu- ral Information Processing Systems , 37:128886–128910,
Yuhang Ma, Shanyuan Liu, Ao Ma, Xiaoyu Wu, Dawei Leng, and Yuhui Yin. Hico: Hierarchical controllable diffu- sion model for layout-to-image generation.Advances in Neu- ral Information Processing Systems , 37:128886–128910,
-
[43]
Storynizor: Consistent story generation via inter-frame synchronized and shuffled id injection
Yuhang Ma, Wenting Xu, Chaoyi Zhao, Keqiang Sun, Qinfeng Jin, Zeng Zhao, Changjie Fan, and Zhipeng Hu. Storynizor: Consistent story generation via inter-frame synchronized and shuffled id injection. arXiv preprint arXiv:2409.19624, 2024. 2, 3
2024 arXiv
-
[44]
Story-adapter: A training-free iterative framework for long story visualization
Jiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang, Mude Hui, Bingjie Xu, and Yuyin Zhou. Story-adapter: A training-free iterative framework for long story visualization. arXiv preprint arXiv:2410.06244, 2024. 2
2024
-
[45]
Lego: Learning to disentangle and invert personalized con- cepts beyond object appearance in text-to-image diffusion models
Saman Motamed, Danda Pani Paudel, and Luc Van Gool. Lego: Learning to disentangle and invert personalized con- cepts beyond object appearance in text-to-image diffusion models. arXiv preprint arXiv:2311.13833, 2023. 9
2023 arXiv
-
[46]
Openvid-1m: A large-scale high-quality dataset for text-to- video generation
Kepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan, Zhen- heng Yang, Zhijie Chen, Xiang Li, Jian Yang, and Ying Tai. Openvid-1m: A large-scale high-quality dataset for text-to- video generation. arXiv preprint arXiv:2407.02371 , 2024. 3
2024 arXiv
-
[47]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF inter- national conference on computer vision , pages 4195–4205,
-
[48]
Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023. 11
2023 arXiv
-
[49]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...
2021
-
[50]
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020. 11
2020
-
[51]
Hierarchical text-conditional image gener- ation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image gener- ation with clip latents. arXiv preprint arXiv:2204.06125, 1 (2):3, 2022. 2
2022 arXiv
-
[52]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3, 9, 11
2022
-
[53]
U- net: Convolutional networks for biomedical image segmen- tation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, pa...
2015
-
[54]
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 2250...
2023
-
[55]
We also utilize the FID [23] metrics to as- sess the quality of the generated images
to remove the image background and replace it with random noise. We also utilize the FID [23] metrics to as- sess the quality of the generated images. Recall@1 mea- sures top-1 text-to-image matching accuracy, while human preference reflects averaged binary ratings from three ...
-
[56]
Carvekit: Automated high-quality back- ground removal framework
Nikita Selin. Carvekit: Automated high-quality back- ground removal framework. https://github.com/ OPHoperHPO/image- background- remove-tool,
-
[57]
Eventvad: Training-free event-aware video anomaly detection
Yihua Shao, Haojin He, Sijie Li, Siyu Chen, Xinwei Long, Fanhu Zeng, Yuxuan Fan, Muyang Zhang, Ziyang Yan, Ao Ma, et al. Eventvad: Training-free event-aware video anomaly detection. arXiv preprint arXiv:2504.13092, 2025. 2
2025 arXiv
-
[58]
Tr-dq: Time-rotation diffusion quantiza- tion
Yihua Shao, Deyang Lin, Fanhu Zeng, Minxi Yan, Muyang Zhang, Siyu Chen, Yuxuan Fan, Ziyang Yan, Haozhe Wang, Jingcai Guo, et al. Tr-dq: Time-rotation diffusion quantiza- tion. arXiv preprint arXiv:2503.06564, 2025. 11
2025 arXiv
-
[59]
In-context meta lora generation
Yihua Shao, Minxi Yan, Yang Liu, Siyu Chen, Wenjie Chen, Xinwei Long, Ziyang Yan, Lei Li, Chenyu Zhang, Nicu Sebe, et al. In-context meta lora generation. arXiv preprint arXiv:2501.17635, 2025. 11
2025 arXiv
-
[60]
Storybooth: Training-free multi-subject consistency for improved visual storytelling
Jaskirat Singh, Junshen K Chen, Jonas K Kohler, and Michael F Cohen. Storybooth: Training-free multi-subject consistency for improved visual storytelling. In The Thir- teenth International Conference on Learning Representa- tions. 2
-
[61]
Styledrop: Text-to-image synthesis of any style
Kihyuk Sohn, Lu Jiang, Jarred Barber, Kimin Lee, Nataniel Ruiz, Dilip Krishnan, Huiwen Chang, Yuanzhen Li, Irfan Essa, Michael Rubinstein, et al. Styledrop: Text-to-image synthesis of any style. Advances in Neural Information Pro- cessing Systems, 36:66860–66889, 2023. 9
2023
-
[62]
Instantx flux.1-dev ip-adapter page, 2024
InstantX Team. Instantx flux.1-dev ip-adapter page, 2024. 2, 6, 7, 8, 9
2024
-
[63]
Training-free consis- tent text-to-image generation
Yoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten, Lior Wolf, Gal Chechik, and Yuval Atzmon. Training-free consis- tent text-to-image generation. ACM Transactions on Graph- ics (TOG), 43(4):1–18, 2024. 2, 4, 6, 7, 8, 9
2024
-
[64]
Converting video formats with ffmpeg
Suramya Tomar. Converting video formats with ffmpeg. Linux journal, 2006(146):10, 2006. 3
2006
-
[65]
Face0: Instantaneously conditioning a text-to- image model on a face
Dani Valevski, Danny Lumen, Yossi Matias, and Yaniv Leviathan. Face0: Instantaneously conditioning a text-to- image model on a face. InSIGGRAPH Asia 2023 Conference Papers, pages 1–10, 2023. 9
2023
-
[66]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 11
2017
-
[67]
Is this loss informative? faster text-to-image customization by tracking objective dynamics
Anton V oronov, Mikhail Khoroshikh, Artem Babenko, and Max Ryabinin. Is this loss informative? faster text-to-image customization by tracking objective dynamics. Advances in Neural Information Processing Systems , 36:37491–37510,
-
[68]
p+: Extended textual conditioning in text-to- image generation
Andrey V oynov, Qinghao Chu, Daniel Cohen-Or, and Kfir Aberman. p+: Extended textual conditioning in text-to- image generation. arXiv preprint arXiv:2303.09522, 2023. 9
2023 arXiv
-
[69]
Qihoo-t2x: An efficiency-focused diffu- sion transformer via proxy tokens for text-to-any-task
Jing Wang, Ao Ma, Jiasong Feng, Dawei Leng, Yuhui Yin, and Xiaodan Liang. Qihoo-t2x: An efficiency-focused diffu- sion transformer via proxy tokens for text-to-any-task. arXiv e-prints, pages arXiv–2409, 2024. 3
2024
-
[70]
Wisa: World simulator assistant for physics-aware text-to-video generation
Jing Wang, Ao Ma, Ke Cao, Jun Zheng, Zhanjie Zhang, Jiasong Feng, Shanyuan Liu, Yuhang Ma, Bo Cheng, Dawei Leng, et al. Wisa: World simulator assistant for physics-aware text-to-video generation. arXiv preprint arXiv:2503.08153, 2025. 2
2025 arXiv
-
[71]
Videofactory: Swap at- tention in spatiotemporal diffusions for text-to-video gener- ation
Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, and Jiaying Liu. Videofactory: Swap at- tention in spatiotemporal diffusions for text-to-video gener- ation. 2023. 3
2023
-
[72]
Spnet: Learning stereo matching with slanted plane aggregation
Yun Wang, Longguang Wang, Hanyun Wang, and Yulan Guo. Spnet: Learning stereo matching with slanted plane aggregation. IEEE Robotics and Automation Letters , 7(3): 6258–6265, 2022. 11
2022
-
[73]
Internvid: A large-scale video-text dataset for multimodal understanding and generation
Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, et al. Internvid: A large-scale video-text dataset for multimodal understanding and generation. arXiv preprint arXiv:2307.06942, 2023. 3
2023 arXiv
-
[74]
Adstereo: Efficient stereo matching with adaptive downsampling and disparity alignment
Yun Wang, Kunhong Li, Longguang Wang, Junjie Hu, Dapeng Oliver Wu, and Yulan Guo. Adstereo: Efficient stereo matching with adaptive downsampling and disparity alignment. IEEE Transactions on Image Processing , 2025. 2
2025
-
[75]
Learning robust stereo matching in the wild with selective mixture-of-experts
Yun Wang, Longguang Wang, Chenghao Zhang, Yongjian Zhang, Zhanjie Zhang, Ao Ma, Chenyou Fan, Tin Lun Lam, and Junjie Hu. Learning robust stereo matching in the wild with selective mixture-of-experts. arXiv preprint arXiv:2507.04631, 2025. 11
2025 arXiv
-
[76]
Dualnet: Ro- bust self-supervised stereo matching with pseudo-label su- pervision
Yun Wang, Jiahao Zheng, Chenghao Zhang, Zhanjie Zhang, Kunhong Li, Yongjian Zhang, and Junjie Hu. Dualnet: Ro- bust self-supervised stereo matching with pseudo-label su- pervision. In Proceedings of the AAAI Conference on Artifi- cial Intelligence, pages 8178–8186, 2025. 2
2025
-
[77]
Styleadapter: A unified stylized image generation model
Zhouxia Wang, Xintao Wang, Liangbin Xie, Zhongang Qi, Ying Shan, Wenping Wang, and Ping Luo. Styleadapter: A unified stylized image generation model. arXiv preprint arXiv:2309.01770, 2023. 9
2023 arXiv
-
[78]
Dropoutgs: Dropping out gaus- sians for better sparse-view rendering
Yexing Xu, Longguang Wang, Minglin Chen, Sheng Ao, Li Li, and Yulan Guo. Dropoutgs: Dropping out gaus- sians for better sparse-view rendering. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 701–710, 2025. 11
2025
-
[79]
Freestyle layout-to-image synthesis
Han Xue, Zhiwu Huang, Qianru Sun, Li Song, and Wenjun Zhang. Freestyle layout-to-image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14256–14266, 2023. 9
2023
-
[80]
Facestudio: Put your face everywhere in seconds
Yuxuan Yan, Chi Zhang, Rui Wang, Yichao Zhou, Gege Zhang, Pei Cheng, Gang Yu, and Bin Fu. Facestudio: Put your face everywhere in seconds. arXiv preprint arXiv:2312.02663, 2023. 9
2023 arXiv
-
[81]
Seed-story: Multimodal long story generation with large language model
Shuai Yang, Yuying Ge, Yang Li, Yukang Chen, Yixiao Ge, Ying Shan, and Yingcong Chen. Seed-story: Multimodal long story generation with large language model. arXiv preprint arXiv:2407.08683, 2024. 2, 3, 7, 9
2024 arXiv
-
[82]
Reco: Region-controlled text-to-image genera- tion
Zhengyuan Yang, Jianfeng Wang, Zhe Gan, Linjie Li, Kevin Lin, Chenfei Wu, Nan Duan, Zicheng Liu, Ce Liu, Michael Zeng, et al. Reco: Region-controlled text-to-image genera- tion. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 14246–14255,
-
[83]
Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models. arXiv preprint arXiv:2308.06721,
-
[84]
Controlnet-xs: Rethinking the control of text-to- image diffusion models as feedback-control systems
Denis Zavadski, Johann-Friedrich Feiden, and Carsten Rother. Controlnet-xs: Rethinking the control of text-to- image diffusion models as feedback-control systems. In European Conference on Computer Vision, pages 343–362. Springer, 2024. 2
2024
-
[85]
Creatilayout: Siamese multimodal diffusion transformer for creative layout-to-image generation
Hui Zhang, Dexiang Hong, Tingwei Gao, Yitong Wang, Jie Shao, Xinglong Wu, Zuxuan Wu, and Yu-Gang Jiang. Creatilayout: Siamese multimodal diffusion transformer for creative layout-to-image generation. arXiv preprint arXiv:2412.03859, 2024. 9
2024 arXiv
-
[86]
A survey on personalized content synthesis with diffusion models
Xulu Zhang, Xiao-Yong Wei, Wengyu Zhang, Jinlin Wu, Zhaoxiang Zhang, Zhen Lei, and Qing Li. A survey on personalized content synthesis with diffusion models. arXiv preprint arXiv:2405.05538, 2024. 9
2024
-
[87]
Generative active learning for image synthesis personalization
Xulu Zhang, Wengyu Zhang, Xiaoyong Wei, Jinlin Wu, Zhaoxiang Zhang, Zhen Lei, and Qing Li. Generative active learning for image synthesis personalization. In Proceedings of the 32nd ACM International Conference on Multimedia , pages 10669–10677, 2024. 9
2024
-
[88]
Towards highly realistic artistic style transfer via stable diffusion with step-aware and layer-aware prompt
Zhanjie Zhang, Quanwei Zhang, Huaizhong Lin, Wei Xing, Juncheng Mo, Shuaicheng Huang, Jinheng Xie, Guangyuan Li, Junsheng Luan, Lei Zhao, et al. Towards highly realistic artistic style transfer via stable diffusion with step-aware and layer-aware prompt. In Proceedings of the ...
2024
-
[89]
Artbank: Artistic style transfer with pre-trained diffusion model and implicit style prompt bank
Zhanjie Zhang, Quanwei Zhang, Wei Xing, Guangyuan Li, Lei Zhao, Jiakai Sun, Zehua Lan, Junsheng Luan, Yiling Huang, and Huaizhong Lin. Artbank: Artistic style transfer with pre-trained diffusion model and implicit style prompt bank. In Proceedings of the AAAI Conference on Art...
2024
-
[90]
Lgast: Towards high- quality arbitrary style transfer with local–global style learn- ing
Zhanjie Zhang, Yuxiang Li, Ruichen Xia, Mengyuan Yang, Yun Wang, Lei Zhao, and Wei Xing. Lgast: Towards high- quality arbitrary style transfer with local–global style learn- ing. Neurocomputing, 623:129434, 2025. 2
2025
-
[91]
U- stydit: Ultra-high quality artistic style transfer using diffu- sion transformers
Zhanjie Zhang, Ao Ma, Ke Cao, Jing Wang, Shanyuan Liu, Yuhang Ma, Bo Cheng, Dawei Leng, and Yuhui Yin. U- stydit: Ultra-high quality artistic style transfer using diffu- sion transformers. arXiv preprint arXiv:2503.08157, 2025. 2
2025 arXiv
-
[92]
Spast: Arbitrary style trans- fer with style priors via pre-trained large-scale model.Neural Networks, page 107556, 2025
Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang, Yun Wang, and Lei Zhao. Spast: Arbitrary style trans- fer with style priors via pre-trained large-scale model.Neural Networks, page 107556, 2025. 2
2025
-
[93]
Vectorsketcher: Learning to create a vector-based free-hand sketch
Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang, Yun Wang, and Lei Zhao. Vectorsketcher: Learning to create a vector-based free-hand sketch. Engineering Appli- cations of Artificial Intelligence, 156:111005, 2025. 2
2025
-
[94]
Image generation from layout
Bo Zhao, Lili Meng, Weidong Yin, and Leonid Sigal. Image generation from layout. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 8584–8593, 2019. 9
2019
-
[95]
Uni-controlnet: All-in-one control to text-to-image diffusion models
Shihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao, Shaozhe Hao, Lu Yuan, and Kwan-Yee K Wong. Uni-controlnet: All-in-one control to text-to-image diffusion models. Advances in Neural Information Processing Sys- tems, 36:11127–11150, 2023. 2
2023
-
[96]
Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation
Guangcong Zheng, Xianpan Zhou, Xuewei Li, Zhongang Qi, Ying Shan, and Xi Li. Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 22490–22499, 2023. 9
2023
-
[97]
Enhanc- ing detail preservation for customized text-to-image gen- eration: A regularization-free approach
Yufan Zhou, Ruiyi Zhang, Tong Sun, and Jinhui Xu. Enhanc- ing detail preservation for customized text-to-image gen- eration: A regularization-free approach. arXiv preprint arXiv:2305.13579, 2023. 9
2023 arXiv
-
[98]
Storydiffusion: Consistent self- attention for long-range image and video generation
Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, and Qibin Hou. Storydiffusion: Consistent self- attention for long-range image and video generation. Ad- vances in Neural Information Processing Systems , 37: 110315–110340, 2025. 2, 6, 7, 8, 9
2025
-
[99]
Storymaker: Towards holistic consistent characters in text-to-image generation
Zhengguang Zhou, Jing Li, Huaxia Li, Nemo Chen, and Xu Tang. Storymaker: Towards holistic consistent characters in text-to-image generation. arXiv preprint arXiv:2409.12576,
-
[2023]
Accessed: March 6, 2025. 7
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.