REVIEW 3 major objections 2 minor 46 references
A Summer Meridional Subsurface Temperature Dipole Mode in the South China Sea
T0 review · 3 major / 2 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The submitted ocean-climate paper is missing: its full text is an unrelated video-generation manuscript, so the claimed South China Sea dipole mode has no supporting analysis in the submitted record.
desk verdict Abstract and body are two different papers; the dipole claim has no supporting evidence in the submitted text. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object claimed in the abstract is the summer subsurface temperature dipole mode in the South China Sea, defined as a pattern of warm temperature anomalies in the northern basin and cold anomalies in the southern basin during strong monsoon years, reversing in weak monsoon years. The proposed mechanism is vertical heat transport linked to opposite wind stress curl anomalies in the northern and southern basin, accompanied by a shallow meridional overturning circulation that redistributes heat between north and south. None of these elements—the dipole, the wind stress curl forcing, or the overturning circulation—appear anywhere in the submitted full text, which is a video-generation paper describing a Mixture of Cross-Attention architecture, a latent video perceptual loss, and a dataset called CelebIPVid.
What would settle it
Checking the submitted full text for any occurrence of 'South China Sea', 'reanalysis', 'heat budget', 'wind stress curl', or 'monsoon' would settle the matter; none of these appear in the supplied MoCA manuscript. Additionally, locating the claimed figures, data description, and heat-budget equations in the submitted body would be required to confirm the dipole claim; their absence falsifies the claim as presented.
Extended reading notes
Core claim
The paper as submitted claims to establish that the South China Sea exhibits a summer subsurface temperature dipole mode, characterized by warm anomalies in the north and cold anomalies in the south during strong monsoon years, with a reversed pattern in weak monsoon years, and that this mode controls interannual variability of upper-ocean heat content. It attributes the dipole to vertical heat transport associated with opposite wind stress curl anomalies in the northern and southern basin, alongside a shallow meridional overturning circulation that redistributes heat meridionally. The abstract also links the monsoon variability to El Niño–Southern Oscillation transitions. However, the supplied full-text manuscript is the MoCA text-to-video paper, which makes no mention of the South China Sea, ocean reanalysis, heat budget, or any related analysis, so the claimed discovery is not supported by any content in the submitted body.
Load-bearing premise
The load-bearing assumption is that the supplied full text is the manuscript corresponding to the abstract, namely an oceanographic study of the South China Sea; the full text actually describes a text-to-video generation model and contains no ocean analysis, so the central claim has no located evidence in the submitted record.
Editorial extensions
If this is right
- If the abstract's claim were supported, the summer subsurface temperature dipole mode would provide a new organizing description of interannual upper-ocean heat content variability in the South China Sea, with potential relevance to tropical cyclone intensity and regional climate prediction.
- A confirmed dipole would link South China Sea heat content variability to large-scale climate variability through El Niño–Southern Oscillation transitions via monsoon strength, offering a pathway for seasonal-to-interannual predictability.
- Attributing the dipole to wind stress curl-driven vertical heat transport and an accompanying shallow meridional overturning circulation would identify a specific dynamical mechanism that could be tested in ocean reanalyses and model experiments.
- The claimed mechanism, if verified, would connect basin-scale atmospheric forcing to a coherent ocean heat-redistribution pattern, refining understanding of how the South China Sea exchanges heat with the atmosphere and neighboring seas.
Reading between the lines
- A reader can test the abstract's claim directly with any high-resolution ocean reanalysis covering the South China Sea: compute summer subsurface temperature anomalies, separate north and south, and regress them against monsoon strength, wind stress curl, and vertical heat transport; the dipole and its forcing should appear if the claim is correct.
- The claimed link to El Niño–Southern Oscillation transitions suggests a concrete predictability test: build a regression of the dipole index on the phase and rate of change of ENSO indicators and check whether strong monsoon years cluster at specific ENSO transition phases.
- Because the supplied full text contains none of the oceanographic analysis, the abstract functions as a standalone claim; the responsible reading is that this submission, as given, does not yet constitute a verifiable scientific paper about the South China Sea.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The abstract announces a South China Sea (SCS) summer subsurface temperature dipole mode that controls interannual upper-ocean heat content, with warm north/cold south anomalies during strong monsoon years, linked to wind stress curl and vertical heat transport, and associated with a shallow meridional overturning circulation. The full text of the submitted manuscript, however, is arXiv:2508.03034v2, a paper on identity-preserving text-to-video generation (MoCA). The body contains no oceanographic data, no reanalysis product description, no heat budget equation, no dipole index definition, and no monsoon or ENSO analysis. The two parts of the submission are mutually unrelated.
Significance. If the abstract's claims were supported, the paper would potentially identify a previously undescribed regional ocean-climate mode that could be relevant for SCS heat content and tropical cyclone research. The actual body, however, provides no evidence, derivation, or analysis for any of these claims. The strengths visible in the body (the MoCA architecture, the CelebIPVid dataset, and the quantitative comparisons in Tables 1-4) pertain to video generation and are irrelevant to the abstract. Consequently, the scientific significance of the abstract's claims cannot be assessed from this manuscript, and the claims are not falsifiable from the material presented.
major comments (3)
- [Abstract vs. Full Text] The central claim of the abstract is wholly unsupported by the full text. The body is the MoCA text-to-video paper: Section 3 defines diffusion loss Eq. (1), the MoCA layer Eqs. (2)-(4), and training losses Eqs. (5)-(12); Section 4 reports experiments on the CelebIPVid dataset with Tables 1-4. Nowhere in the body do the words 'South China Sea,' 'reanalysis,' 'wind stress curl,' 'heat budget,' 'monsoon,' or 'ENSO' appear. The abstract's statements that 'we show...' and 'Heat budget analysis indicates...' therefore have no corresponding methods, data, or results in the manuscript.
- [Full Text] The manuscript lacks any definition or quantitative characterization of the claimed dipole mode. There is no specification of the reanalysis data product, the depth range or domain over which the dipole is defined, a formula for the dipole index, a heat budget equation, or any statistical significance test for the warm-north/cold-south pattern. Without these elements, the abstract's central assertions cannot be checked or reproduced. The omission is load-bearing because the dipole mode and its mechanism are the entire scientific content of the abstract.
- [References] The reference list is entirely that of the MoCA video-generation paper and contains no oceanographic citations. There is no reference to the ocean reanalysis product, to previous SCS temperature or heat-content variability studies, or to the ENSO-monsoon literature invoked in the abstract. The claimed results are therefore presented without any scientific context or comparison to prior work, further confirming that the abstract's oceanographic analysis is not present in this submission.
minor comments (2)
- [General] The manuscript's title, abstract, and body describe entirely different studies, so a reader cannot determine which part is intended as the submission; at minimum, the submission metadata and abstract/body correspondence need correction.
- [Abstract] The abstract uses 'high-resolution ocean reanalysis data' and 'Heat budget analysis' without naming the products or providing any equation or figure numbers; adding these identifiers would be essential even for a condensed abstract in a healthy submission.
Circularity Check
No circularity can be established: the supplied full text is an unrelated video-generation paper, so the abstract's dipole derivation has no equations or data to exhibit as circular.
full rationale
Circularity analysis requires quoting a specific reduction in which a claimed prediction is equivalent to its inputs by construction, or a load-bearing argument reduces to a self-citation. The supplied full text is the MoCA text-to-video paper (arXiv:2508.03034v2), which contains no South China Sea analysis, no ocean reanalysis description, no heat budget equation, and no dipole index definition. The abstract's mechanism claims, such as vertical heat transport linked to opposite wind stress curl anomalies, cannot be checked against any derivation chain because none is present. A missing derivation is a completeness failure, not circularity: there is no equation or fitted parameter to exhibit as equivalent to the conclusion, and no self-citation chain to trace. Therefore, under the hard rule that circularity may only be claimed when the specific reduction can be quoted, the correct finding is no significant circularity, with a score of 0 and an empty steps list.
Assumptions & free parameters
assumptions (1)
- ad hoc to paper The supplied full text corresponds to the oceanographic study described in the abstract.
Cite this review
Pith. "Pith review of A Summer Meridional Subsurface Temperature Dipole Mode in the South China Sea." pith.science (2026). https://pith.science/paper/JLIQM5KW
@misc{pith2026250803032,
author = {Pith},
title = {Pith review of: A Summer Meridional Subsurface Temperature Dipole Mode in the South China Sea},
year = {2026},
howpublished = {\url{https://pith.science/paper/JLIQM5KW}},
note = {Machine review of arXiv:2508.03032}
}
read the original abstract
The ocean heat content variability in the South China Sea (SCS) plays a pivotal role in regional climate and extreme weather events, such as tropical cyclones. Using high-resolution ocean reanalysis data, we show that the SCS exhibits a summer subsurface temperature dipole mode that controls the interannual variability of ocean heat content in the upper SCS. This dipole mode manifests as warm anomalies in the north and cold anomalies in the south during strong monsoon years, and a reversed pattern during weak monsoons years. The monsoon variability is linked to large-scale climate variability associated with El Ni\~no-Southern Oscillation transitions. Heat budget analysis indicates that this dipole pattern is primarily driven by vertical heat transport linked to opposite wind stress curl anomalies in the northern and southern basin. Accompanying the vertical heat transports is a shallow meridional overturning circulation that redistributes heat between the northern and southern SCS.
Reference graph
Works this paper leans on
-
[1]
Stable video diffusion: Scaling latent video diffusion models to large datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023. 1
arXiv 2023
-
[2]
Videocrafter1: Open diffusion models for high-quality video generation
Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, et al. Videocrafter1: Open diffusion models for high-quality video generation. arXiv preprint arXiv:2310.19512, 2023. 1
-
[3]
Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models
Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 7310– 7320, 2024. 1
work page 2024
-
[4]
Pixart- α: Fast training of diffusion transformer for photorealistic text-to-image synthesis
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al. Pixart- α: Fast training of diffusion transformer for photorealistic text-to-image synthesis. arXiv preprint arXiv:2310.00426, 2023. 1
-
[5]
Arcface: Additive angular margin loss for deep face recognition
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 4690–4699, 2019. 2, 5, 10, 11
work page 2019
-
[6]
Structure and content-guided video synthesis with diffusion models
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, pages 7346–7356, 2023. 1
work page 2023
-
[7]
Ingredients: Blending custom pho- tos with video diffusion transformers
Zhengcong Fei, Debang Li, Di Qiu, Changqian Yu, and Mingyuan Fan. Ingredients: Blending custom pho- tos with video diffusion transformers. arXiv preprint arXiv:2501.01790, 2025. 1, 2
arXiv 2025
-
[8]
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725, 2023. 1
arXiv 2023
Show all 46 references
-
[9]
Pulid: Pure and lightning id customization via con- trastive alignment
Zinan Guo, Yanze Wu, Zhuowei Chen, Lang Chen, and Qian He. Pulid: Pure and lightning id customization via con- trastive alignment. arXiv preprint arXiv:2404.16022, 2024. 2
2024 arXiv
-
[10]
Id-animator: Zero-shot identity-preserving human video generation.arXiv preprint arXiv:2404.15275, 2024
Xuanhua He, Quande Liu, Shengju Qian, Xin Wang, Tao Hu, Ke Cao, Keyu Yan, Man Zhou, and Jie Zhang. Id-animator: Zero-shot identity-preserving human video generation.arXiv preprint arXiv:2404.15275, 2024. 1, 2, 4, 5, 10
2024 arXiv
-
[11]
Latent video diffusion models for high-fidelity long video generation
Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for high-fidelity long video generation. arXiv preprint arXiv:2211.13221 ,
-
[12]
Clipscore: A reference-free evaluation met- ric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation met- ric for image captioning. arXiv preprint arXiv:2104.08718,
-
[13]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. NeurIPS, 33:6840–6851, 2020. 1
2020
-
[14]
Imagen video: High definition video generation with diffusion mod- els
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion mod- els. arXiv preprint arXiv:2210.02303, 2022. 1
-
[15]
Video dif- fusion models
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video dif- fusion models. In NeurIPS, 2022
2022
-
[16]
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868, 2022. 1
2022 arXiv
-
[17]
Vbench: Comprehensive bench- mark suite for video generative models
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. Vbench: Comprehensive bench- mark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and P...
2024
-
[18]
Hunyuanvideo: A systematic framework for large video generative models
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024. 1
2024 arXiv
-
[19]
Personalvideo: High id-fidelity video customization without dynamic and semantic degradation
Hengjia Li, Haonan Qiu, Shiwei Zhang, Xiang Wang, Yu- jie Wei, Zekun Li, Yingya Zhang, Boxi Wu, and Deng Cai. Personalvideo: High id-fidelity video customization without dynamic and semantic degradation. arXiv preprint arXiv:2411.17048, 2024. 1
2024 arXiv
-
[20]
Magicid: Hybrid preference optimization for id-consistent and dynamic-preserved video customization
Hengjia Li, Lifan Jiang, Xi Xiao, Tianyang Wang, Hongwei Yi, Boxi Wu, and Deng Cai. Magicid: Hybrid preference optimization for id-consistent and dynamic-preserved video customization. arXiv preprint arXiv:2503.12689, 2025. 1, 2
2025 arXiv
-
[21]
Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chi- nese understanding
Zhimin Li, Jianwei Zhang, Qin Lin, Jiangfeng Xiong, Yanxin Long, Xinchi Deng, Yingfang Zhang, Xingchao Liu, Minbin Huang, Zedong Xiao, et al. Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chi- nese understanding. arXiv preprint arXiv:2405.087...
2024 arXiv
-
[22]
Open-sora plan: Open-source large video generation model
Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Li- uhan Chen, et al. Open-sora plan: Open-source large video generation model. arXiv preprint arXiv:2412.00131, 2024
2024 arXiv
-
[23]
Sora: A review on background, technology, limitations, and opportunities of large vision models
Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jian- feng Gao, et al. Sora: A review on background, technology, limitations, and opportunities of large vision models. arXiv preprint arXiv:2402.17177, 2024
2024 arXiv
-
[24]
Tuning-free long video generation via global-local collaborative diffu- sion, 2025
Yongjia Ma, Junlin Chen, Donglin Di, Qi Xie, Lei Fan, Wei Chen, Xiaofei Gou, Na Zhao, and Xun Yang. Tuning-free long video generation via global-local collaborative diffu- sion, 2025. 1
2025
-
[25]
Adams bashforth moulton solver for inversion and editing in rectified flow, 2025
Yongjia Ma, Donglin Di, Xuan Liu, Xiaokai Chen, Lei Fan, Wei Chen, and Tonghua Su. Adams bashforth moulton solver for inversion and editing in rectified flow, 2025. 1
2025
-
[26]
Magic-me: Identity-specific video customized diffu- sion
Ze Ma, Daquan Zhou, Chun-Hsiao Yeh, Xue-She Wang, Xi- uyu Li, Huanrui Yang, Zhen Dong, Kurt Keutzer, and Jiashi Feng. Magic-me: Identity-specific video customized diffu- sion. arXiv preprint arXiv:2402.09368, 2024. 2
2024 arXiv
-
[27]
V oxceleb: a large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. V oxceleb: a large-scale speaker identification dataset. arXiv preprint arXiv:1706.08612, 2017. 2
2017 arXiv
-
[28]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 4195–4205,
-
[29]
State of the art on diffusion models for visual computing
Ryan Po, Wang Yifan, Vladislav Golyanik, Kfir Aberman, Jonathan T Barron, Amit Bermano, Eric Chan, Tali Dekel, Aleksander Holynski, Angjoo Kanazawa, et al. State of the art on diffusion models for visual computing. In Computer Graphics Forum, page e15063. Wiley Online Library, 2024. 1
2024
-
[30]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...
2021
-
[31]
High-resolution image syn- thesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. In CVPR, 2022. 1
2022
-
[32]
Mm-diffusion: Learning multi-modal diffusion mod- els for joint audio and video generation
Ludan Ruan, Yiyang Ma, Huan Yang, Huiguo He, Bei Liu, Jianlong Fu, Nicholas Jing Yuan, Qin Jin, and Baining Guo. Mm-diffusion: Learning multi-modal diffusion mod- els for joint audio and video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern...
2023
-
[33]
Mead: A large-scale audio-visual dataset for emotional talking-face generation
Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. Mead: A large-scale audio-visual dataset for emotional talking-face generation. In European conference on com- puter vision, pages 700–717. Springer, 2020. 2
2020
-
[34]
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024. 5
2024 arXiv
-
[35]
One-shot free-view neural talking-head synthesis for video conferenc- ing
Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. One-shot free-view neural talking-head synthesis for video conferenc- ing. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition , pages 10039–10049,
-
[36]
A survey on video dif- fusion models
Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, and Yu-Gang Jiang. A survey on video dif- fusion models. ACM Computing Surveys, 57(2):1–42, 2024. 1
2024
-
[37]
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiao- han Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 1, 2, 5, 10
2024 arXiv
-
[38]
Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models. arXiv preprint arXiv:2308.06721,
-
[39]
Identity- preserving text-to-video generation by frequency decompo- sition
Shenghai Yuan, Jinfa Huang, Xianyi He, Yunyuan Ge, Yu- jun Shi, Liuhan Chen, Jiebo Luo, and Li Yuan. Identity- preserving text-to-video generation by frequency decompo- sition. arXiv preprint arXiv:2411.17440 , 2024. 1, 2, 4, 5, 10
2024 arXiv
-
[40]
Magic mirror: Id-preserved video generation in video diffusion transformers
Yuechen Zhang, Yaoyang Liu, Bin Xia, Bohao Peng, Zexin Yan, Eric Lo, and Jiaya Jia. Magic mirror: Id-preserved video generation in video diffusion transformers. arXiv preprint arXiv:2501.03931, 2025. 1, 2
2025
-
[41]
Fantasyid: Face knowledge en- hanced id-preserving video generation
Yunpeng Zhang, Qiang Wang, Fan Jiang, Yaqi Fan, Mu Xu, and Yonggang Qi. Fantasyid: Face knowledge en- hanced id-preserving video generation. arXiv preprint arXiv:2502.13995, 2025. 2
2025 arXiv
-
[42]
Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan. Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3661–3670, 2021. 2
2021
-
[43]
Open-sora: Democratizing efficient video production for all
Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-sora: Democratizing efficient video production for all. arXiv preprint arXiv:2412.20404, 2024. 1
2024 arXiv
-
[44]
Concat-id: Towards univer- sal identity-preserving video synthesis
Yong Zhong, Zhuoyi Yang, Jiayan Teng, Xiaotao Gu, and Chongxuan Li. Concat-id: Towards univer- sal identity-preserving video synthesis. arXiv preprint arXiv:2503.14151, 2025. 1
2025 arXiv
-
[45]
Celebv- hq: A large-scale video facial attributes dataset
Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. Celebv- hq: A large-scale video facial attributes dataset. InEuropean conference on computer vision , pages 650–667. Springer,
-
[2022]
long, straight dark brown hair
2 Appendix of MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention (Supplementary Material) A. CelebIPVid Dataset Samples . . . . . . . . . . . . . . . . . . 10 B. Details of Evaluation Metrics . . . . . . . . . . . . . . . . . . 10 C. Additional E...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.