Pith. sign in

REVIEW 3 major objections 2 minor 46 references

A Summer Meridional Subsurface Temperature Dipole Mode in the South China Sea

T0 review · 3 major / 2 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The submitted ocean-climate paper is missing: its full text is an unrelated video-generation manuscript, so the claimed South China Sea dipole mode has no supporting analysis in the submitted record.

desk verdict Abstract and body are two different papers; the dipole claim has no supporting evidence in the submitted text. read the letter →

arxiv 2508.03032 v1 pith:JLIQM5KW submitted 2025-08-05 physics.ao-ph

classification physics.ao-ph
keywords SouthChinaSeasubsurfacetemperaturedipoleoceanheatcontentmonsoonvariabilitywindstresscurlreanalysismanuscriptmismatch
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This submission presents an abstract claiming that the South China Sea hosts a summer subsurface temperature dipole mode, with warm anomalies in the north and cold anomalies in the south during strong monsoon years, which controls interannual upper-ocean heat content variability. The abstract further attributes the dipole to vertical heat transport driven by opposite wind stress curl anomalies, accompanied by a shallow meridional overturning circulation. The full text supplied with the submission, however, is an entirely unrelated paper on identity-preserving text-to-video generation, containing no oceanographic data, figures, equations, or analysis of the South China Sea. As a result, the central claim of the abstract cannot be verified from the submitted manuscript, and the paper's own evidence base for the dipole mode is absent.

What carries the argument

The central object claimed in the abstract is the summer subsurface temperature dipole mode in the South China Sea, defined as a pattern of warm temperature anomalies in the northern basin and cold anomalies in the southern basin during strong monsoon years, reversing in weak monsoon years. The proposed mechanism is vertical heat transport linked to opposite wind stress curl anomalies in the northern and southern basin, accompanied by a shallow meridional overturning circulation that redistributes heat between north and south. None of these elements—the dipole, the wind stress curl forcing, or the overturning circulation—appear anywhere in the submitted full text, which is a video-generation paper describing a Mixture of Cross-Attention architecture, a latent video perceptual loss, and a dataset called CelebIPVid.

What would settle it

Checking the submitted full text for any occurrence of 'South China Sea', 'reanalysis', 'heat budget', 'wind stress curl', or 'monsoon' would settle the matter; none of these appear in the supplied MoCA manuscript. Additionally, locating the claimed figures, data description, and heat-budget equations in the submitted body would be required to confirm the dipole claim; their absence falsifies the claim as presented.

Watch

Extended reading notes

Core claim

The paper as submitted claims to establish that the South China Sea exhibits a summer subsurface temperature dipole mode, characterized by warm anomalies in the north and cold anomalies in the south during strong monsoon years, with a reversed pattern in weak monsoon years, and that this mode controls interannual variability of upper-ocean heat content. It attributes the dipole to vertical heat transport associated with opposite wind stress curl anomalies in the northern and southern basin, alongside a shallow meridional overturning circulation that redistributes heat meridionally. The abstract also links the monsoon variability to El Niño–Southern Oscillation transitions. However, the supplied full-text manuscript is the MoCA text-to-video paper, which makes no mention of the South China Sea, ocean reanalysis, heat budget, or any related analysis, so the claimed discovery is not supported by any content in the submitted body.

Load-bearing premise

The load-bearing assumption is that the supplied full text is the manuscript corresponding to the abstract, namely an oceanographic study of the South China Sea; the full text actually describes a text-to-video generation model and contains no ocean analysis, so the central claim has no located evidence in the submitted record.

Editorial extensions

If this is right

  • If the abstract's claim were supported, the summer subsurface temperature dipole mode would provide a new organizing description of interannual upper-ocean heat content variability in the South China Sea, with potential relevance to tropical cyclone intensity and regional climate prediction.
  • A confirmed dipole would link South China Sea heat content variability to large-scale climate variability through El Niño–Southern Oscillation transitions via monsoon strength, offering a pathway for seasonal-to-interannual predictability.
  • Attributing the dipole to wind stress curl-driven vertical heat transport and an accompanying shallow meridional overturning circulation would identify a specific dynamical mechanism that could be tested in ocean reanalyses and model experiments.
  • The claimed mechanism, if verified, would connect basin-scale atmospheric forcing to a coherent ocean heat-redistribution pattern, refining understanding of how the South China Sea exchanges heat with the atmosphere and neighboring seas.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A reader can test the abstract's claim directly with any high-resolution ocean reanalysis covering the South China Sea: compute summer subsurface temperature anomalies, separate north and south, and regress them against monsoon strength, wind stress curl, and vertical heat transport; the dipole and its forcing should appear if the claim is correct.
  • The claimed link to El Niño–Southern Oscillation transitions suggests a concrete predictability test: build a regression of the dipole index on the phase and rate of change of ENSO indicators and check whether strong monsoon years cluster at specific ENSO transition phases.
  • Because the supplied full text contains none of the oceanographic analysis, the abstract functions as a standalone claim; the responsible reading is that this submission, as given, does not yet constitute a verifiable scientific paper about the South China Sea.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The abstract announces a South China Sea (SCS) summer subsurface temperature dipole mode that controls interannual upper-ocean heat content, with warm north/cold south anomalies during strong monsoon years, linked to wind stress curl and vertical heat transport, and associated with a shallow meridional overturning circulation. The full text of the submitted manuscript, however, is arXiv:2508.03034v2, a paper on identity-preserving text-to-video generation (MoCA). The body contains no oceanographic data, no reanalysis product description, no heat budget equation, no dipole index definition, and no monsoon or ENSO analysis. The two parts of the submission are mutually unrelated.

Significance. If the abstract's claims were supported, the paper would potentially identify a previously undescribed regional ocean-climate mode that could be relevant for SCS heat content and tropical cyclone research. The actual body, however, provides no evidence, derivation, or analysis for any of these claims. The strengths visible in the body (the MoCA architecture, the CelebIPVid dataset, and the quantitative comparisons in Tables 1-4) pertain to video generation and are irrelevant to the abstract. Consequently, the scientific significance of the abstract's claims cannot be assessed from this manuscript, and the claims are not falsifiable from the material presented.

major comments (3)
  1. [Abstract vs. Full Text] The central claim of the abstract is wholly unsupported by the full text. The body is the MoCA text-to-video paper: Section 3 defines diffusion loss Eq. (1), the MoCA layer Eqs. (2)-(4), and training losses Eqs. (5)-(12); Section 4 reports experiments on the CelebIPVid dataset with Tables 1-4. Nowhere in the body do the words 'South China Sea,' 'reanalysis,' 'wind stress curl,' 'heat budget,' 'monsoon,' or 'ENSO' appear. The abstract's statements that 'we show...' and 'Heat budget analysis indicates...' therefore have no corresponding methods, data, or results in the manuscript.
  2. [Full Text] The manuscript lacks any definition or quantitative characterization of the claimed dipole mode. There is no specification of the reanalysis data product, the depth range or domain over which the dipole is defined, a formula for the dipole index, a heat budget equation, or any statistical significance test for the warm-north/cold-south pattern. Without these elements, the abstract's central assertions cannot be checked or reproduced. The omission is load-bearing because the dipole mode and its mechanism are the entire scientific content of the abstract.
  3. [References] The reference list is entirely that of the MoCA video-generation paper and contains no oceanographic citations. There is no reference to the ocean reanalysis product, to previous SCS temperature or heat-content variability studies, or to the ENSO-monsoon literature invoked in the abstract. The claimed results are therefore presented without any scientific context or comparison to prior work, further confirming that the abstract's oceanographic analysis is not present in this submission.
minor comments (2)
  1. [General] The manuscript's title, abstract, and body describe entirely different studies, so a reader cannot determine which part is intended as the submission; at minimum, the submission metadata and abstract/body correspondence need correction.
  2. [Abstract] The abstract uses 'high-resolution ocean reanalysis data' and 'Heat budget analysis' without naming the products or providing any equation or figure numbers; adding these identifiers would be essential even for a condensed abstract in a healthy submission.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity can be established: the supplied full text is an unrelated video-generation paper, so the abstract's dipole derivation has no equations or data to exhibit as circular.

full rationale

Circularity analysis requires quoting a specific reduction in which a claimed prediction is equivalent to its inputs by construction, or a load-bearing argument reduces to a self-citation. The supplied full text is the MoCA text-to-video paper (arXiv:2508.03034v2), which contains no South China Sea analysis, no ocean reanalysis description, no heat budget equation, and no dipole index definition. The abstract's mechanism claims, such as vertical heat transport linked to opposite wind stress curl anomalies, cannot be checked against any derivation chain because none is present. A missing derivation is a completeness failure, not circularity: there is no equation or fitted parameter to exhibit as equivalent to the conclusion, and no self-citation chain to trace. Therefore, under the hard rule that circularity may only be claimed when the specific reduction can be quoted, the correct finding is no significant circularity, with a score of 0 and an empty steps list.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The only identifiable axiom is the paper-abstract correspondence. Since the analysis is absent, no free parameters or invented entities can be enumerated.

assumptions (1)
  • ad hoc to paper The supplied full text corresponds to the oceanographic study described in the abstract.
    This assumption is required to treat any part of the manuscript as evidence for the dipole claim. The body text contradicts it, since the manuscript is a text-to-video generation paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Summer Meridional Subsurface Temperature Dipole Mode in the South China Sea." pith.science (2026). https://pith.science/paper/JLIQM5KW

@misc{pith2026250803032,
  author       = {Pith},
  title        = {Pith review of: A Summer Meridional Subsurface Temperature Dipole Mode in the South China Sea},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JLIQM5KW}},
  note         = {Machine review of arXiv:2508.03032}
}
read the original abstract

The ocean heat content variability in the South China Sea (SCS) plays a pivotal role in regional climate and extreme weather events, such as tropical cyclones. Using high-resolution ocean reanalysis data, we show that the SCS exhibits a summer subsurface temperature dipole mode that controls the interannual variability of ocean heat content in the upper SCS. This dipole mode manifests as warm anomalies in the north and cold anomalies in the south during strong monsoon years, and a reversed pattern during weak monsoons years. The monsoon variability is linked to large-scale climate variability associated with El Ni\~no-Southern Oscillation transitions. Heat budget analysis indicates that this dipole pattern is primarily driven by vertical heat transport linked to opposite wind stress curl anomalies in the northern and southern basin. Accompanying the vertical heat transports is a shallow meridional overturning circulation that redistributes heat between the northern and southern SCS.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

46 extracted references · 14 canonical work pages

  1. [1]

    Stable video diffusion: Scaling latent video diffusion models to large datasets

    Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023. 1

  2. [2]

    Videocrafter1: Open diffusion models for high-quality video generation

    Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, et al. Videocrafter1: Open diffusion models for high-quality video generation. arXiv preprint arXiv:2310.19512, 2023. 1

  3. [3]

    Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models

    Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 7310– 7320, 2024. 1

  4. [4]

    Pixart- α: Fast training of diffusion transformer for photorealistic text-to-image synthesis

    Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al. Pixart- α: Fast training of diffusion transformer for photorealistic text-to-image synthesis. arXiv preprint arXiv:2310.00426, 2023. 1

  5. [5]

    Arcface: Additive angular margin loss for deep face recognition

    Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 4690–4699, 2019. 2, 5, 10, 11

  6. [6]

    Structure and content-guided video synthesis with diffusion models

    Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, pages 7346–7356, 2023. 1

  7. [7]

    Ingredients: Blending custom pho- tos with video diffusion transformers

    Zhengcong Fei, Debang Li, Di Qiu, Changqian Yu, and Mingyuan Fan. Ingredients: Blending custom pho- tos with video diffusion transformers. arXiv preprint arXiv:2501.01790, 2025. 1, 2

  8. [8]

    Animatediff: Animate your personalized text-to-image diffusion models without specific tuning

    Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725, 2023. 1

Show all 46 references
  1. [9]

    Pulid: Pure and lightning id customization via con- trastive alignment

    Zinan Guo, Yanze Wu, Zhuowei Chen, Lang Chen, and Qian He. Pulid: Pure and lightning id customization via con- trastive alignment. arXiv preprint arXiv:2404.16022, 2024. 2

  2. [10]

    Id-animator: Zero-shot identity-preserving human video generation.arXiv preprint arXiv:2404.15275, 2024

    Xuanhua He, Quande Liu, Shengju Qian, Xin Wang, Tao Hu, Ke Cao, Keyu Yan, Man Zhou, and Jie Zhang. Id-animator: Zero-shot identity-preserving human video generation.arXiv preprint arXiv:2404.15275, 2024. 1, 2, 4, 5, 10

  3. [11]

    Latent video diffusion models for high-fidelity long video generation

    Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for high-fidelity long video generation. arXiv preprint arXiv:2211.13221 ,

  4. [12]

    Clipscore: A reference-free evaluation met- ric for image captioning

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation met- ric for image captioning. arXiv preprint arXiv:2104.08718,

  5. [13]

    Denoising dif- fusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. NeurIPS, 33:6840–6851, 2020. 1

  6. [14]

    Imagen video: High definition video generation with diffusion mod- els

    Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion mod- els. arXiv preprint arXiv:2210.02303, 2022. 1

  7. [15]

    Video dif- fusion models

    Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video dif- fusion models. In NeurIPS, 2022

  8. [16]

    Cogvideo: Large-scale pretraining for text-to-video generation via transformers

    Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. arXiv preprint arXiv:2205.15868, 2022. 1

  9. [17]

    Vbench: Comprehensive bench- mark suite for video generative models

    Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. Vbench: Comprehensive bench- mark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and P...

  10. [18]

    Hunyuanvideo: A systematic framework for large video generative models

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024. 1

  11. [19]

    Personalvideo: High id-fidelity video customization without dynamic and semantic degradation

    Hengjia Li, Haonan Qiu, Shiwei Zhang, Xiang Wang, Yu- jie Wei, Zekun Li, Yingya Zhang, Boxi Wu, and Deng Cai. Personalvideo: High id-fidelity video customization without dynamic and semantic degradation. arXiv preprint arXiv:2411.17048, 2024. 1

  12. [20]

    Magicid: Hybrid preference optimization for id-consistent and dynamic-preserved video customization

    Hengjia Li, Lifan Jiang, Xi Xiao, Tianyang Wang, Hongwei Yi, Boxi Wu, and Deng Cai. Magicid: Hybrid preference optimization for id-consistent and dynamic-preserved video customization. arXiv preprint arXiv:2503.12689, 2025. 1, 2

  13. [21]

    Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chi- nese understanding

    Zhimin Li, Jianwei Zhang, Qin Lin, Jiangfeng Xiong, Yanxin Long, Xinchi Deng, Yingfang Zhang, Xingchao Liu, Minbin Huang, Zedong Xiao, et al. Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chi- nese understanding. arXiv preprint arXiv:2405.087...

  14. [22]

    Open-sora plan: Open-source large video generation model

    Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Li- uhan Chen, et al. Open-sora plan: Open-source large video generation model. arXiv preprint arXiv:2412.00131, 2024

  15. [23]

    Sora: A review on background, technology, limitations, and opportunities of large vision models

    Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jian- feng Gao, et al. Sora: A review on background, technology, limitations, and opportunities of large vision models. arXiv preprint arXiv:2402.17177, 2024

  16. [24]

    Tuning-free long video generation via global-local collaborative diffu- sion, 2025

    Yongjia Ma, Junlin Chen, Donglin Di, Qi Xie, Lei Fan, Wei Chen, Xiaofei Gou, Na Zhao, and Xun Yang. Tuning-free long video generation via global-local collaborative diffu- sion, 2025. 1

  17. [25]

    Adams bashforth moulton solver for inversion and editing in rectified flow, 2025

    Yongjia Ma, Donglin Di, Xuan Liu, Xiaokai Chen, Lei Fan, Wei Chen, and Tonghua Su. Adams bashforth moulton solver for inversion and editing in rectified flow, 2025. 1

  18. [26]

    Magic-me: Identity-specific video customized diffu- sion

    Ze Ma, Daquan Zhou, Chun-Hsiao Yeh, Xue-She Wang, Xi- uyu Li, Huanrui Yang, Zhen Dong, Kurt Keutzer, and Jiashi Feng. Magic-me: Identity-specific video customized diffu- sion. arXiv preprint arXiv:2402.09368, 2024. 2

  19. [27]

    V oxceleb: a large-scale speaker identification dataset

    Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. V oxceleb: a large-scale speaker identification dataset. arXiv preprint arXiv:1706.08612, 2017. 2

  20. [28]

    Scalable diffusion models with transformers

    William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 4195–4205,

  21. [29]

    State of the art on diffusion models for visual computing

    Ryan Po, Wang Yifan, Vladislav Golyanik, Kfir Aberman, Jonathan T Barron, Amit Bermano, Eric Chan, Tali Dekel, Aleksander Holynski, Angjoo Kanazawa, et al. State of the art on diffusion models for visual computing. In Computer Graphics Forum, page e15063. Wiley Online Library, 2024. 1

  22. [30]

    Learning transferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...

  23. [31]

    High-resolution image syn- thesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. In CVPR, 2022. 1

  24. [32]

    Mm-diffusion: Learning multi-modal diffusion mod- els for joint audio and video generation

    Ludan Ruan, Yiyang Ma, Huan Yang, Huiguo He, Bei Liu, Jianlong Fu, Nicholas Jing Yuan, Qin Jin, and Baining Guo. Mm-diffusion: Learning multi-modal diffusion mod- els for joint audio and video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern...

  25. [33]

    Mead: A large-scale audio-visual dataset for emotional talking-face generation

    Kaisiyuan Wang, Qianyi Wu, Linsen Song, Zhuoqian Yang, Wayne Wu, Chen Qian, Ran He, Yu Qiao, and Chen Change Loy. Mead: A large-scale audio-visual dataset for emotional talking-face generation. In European conference on com- puter vision, pages 700–717. Springer, 2020. 2

  26. [34]

    Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024. 5

  27. [35]

    One-shot free-view neural talking-head synthesis for video conferenc- ing

    Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. One-shot free-view neural talking-head synthesis for video conferenc- ing. In Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition , pages 10039–10049,

  28. [36]

    A survey on video dif- fusion models

    Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, and Yu-Gang Jiang. A survey on video dif- fusion models. ACM Computing Surveys, 57(2):1–42, 2024. 1

  29. [37]

    Cogvideox: Text-to-video diffusion models with an expert transformer

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiao- han Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024. 1, 2, 5, 10

  30. [38]

    Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models

    Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models. arXiv preprint arXiv:2308.06721,

  31. [39]

    Identity- preserving text-to-video generation by frequency decompo- sition

    Shenghai Yuan, Jinfa Huang, Xianyi He, Yunyuan Ge, Yu- jun Shi, Liuhan Chen, Jiebo Luo, and Li Yuan. Identity- preserving text-to-video generation by frequency decompo- sition. arXiv preprint arXiv:2411.17440 , 2024. 1, 2, 4, 5, 10

  32. [40]

    Magic mirror: Id-preserved video generation in video diffusion transformers

    Yuechen Zhang, Yaoyang Liu, Bin Xia, Bohao Peng, Zexin Yan, Eric Lo, and Jiaya Jia. Magic mirror: Id-preserved video generation in video diffusion transformers. arXiv preprint arXiv:2501.03931, 2025. 1, 2

  33. [41]

    Fantasyid: Face knowledge en- hanced id-preserving video generation

    Yunpeng Zhang, Qiang Wang, Fan Jiang, Yaqi Fan, Mu Xu, and Yonggang Qi. Fantasyid: Face knowledge en- hanced id-preserving video generation. arXiv preprint arXiv:2502.13995, 2025. 2

  34. [42]

    Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

    Zhimeng Zhang, Lincheng Li, Yu Ding, and Changjie Fan. Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3661–3670, 2021. 2

  35. [43]

    Open-sora: Democratizing efficient video production for all

    Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-sora: Democratizing efficient video production for all. arXiv preprint arXiv:2412.20404, 2024. 1

  36. [44]

    Concat-id: Towards univer- sal identity-preserving video synthesis

    Yong Zhong, Zhuoyi Yang, Jiayan Teng, Xiaotao Gu, and Chongxuan Li. Concat-id: Towards univer- sal identity-preserving video synthesis. arXiv preprint arXiv:2503.14151, 2025. 1

  37. [45]

    Celebv- hq: A large-scale video facial attributes dataset

    Hao Zhu, Wayne Wu, Wentao Zhu, Liming Jiang, Siwei Tang, Li Zhang, Ziwei Liu, and Chen Change Loy. Celebv- hq: A large-scale video facial attributes dataset. InEuropean conference on computer vision , pages 650–667. Springer,

  38. [2022]

    long, straight dark brown hair

    2 Appendix of MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention (Supplementary Material) A. CelebIPVid Dataset Samples . . . . . . . . . . . . . . . . . . 10 B. Details of Evaluation Metrics . . . . . . . . . . . . . . . . . . 10 C. Additional E...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.