REVIEW 3 major objections 5 minor 5 cited by
Explicitly aligning a video model's reference features to a visual foundation model removes copy-paste artifacts and multi-subject confusion without slowing inference.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 18:00 UTC pith:MPHVRMY2
load-bearing objection Clean engineering paper: REPA-style alignment retargeted to multi-reference DiT tokens with a real push term, SOTA TotalScore, zero inference cost; metric-tradeoff caveats are real but not fatal. the 3 major comments →
RefAlign: Representation Alignment for Reference-to-Video Generation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper establishes that an explicit reference alignment loss, applied only to intermediate reference-image tokens inside the first several blocks of a diffusion transformer, is sufficient to give those tokens the identity consistency and inter-subject separability of a strong visual foundation model. The same positive-and-negative alignment removes the copy-paste and multi-subject failures that arise from unregularized VAE latents, while leaving inference identical to the base video model.
What carries the argument
Reference alignment (RA) loss: a cosine-similarity objective that pulls projected DiT reference tokens toward matching VFM tokens of the same subject and pushes them away from tokens of other subjects by a margin; averaged over the first K transformer blocks and added to the ordinary flow-matching objective.
Load-bearing premise
The method assumes that intermediate reference tokens in a chosen early depth of the transformer are the right place to force alignment to a particular frozen visual encoder, and that that encoder is identity-sensitive enough; if either choice is wrong the loss can hurt fidelity or collapse subject distinctions.
What would settle it
Train identical models with and without the negative (push-apart) term, or at the paper's preferred depth versus much shallower or deeper depths, then re-score FaceSim, NexusScore and multi-subject qualitative examples on the same OpenS2V-Eval set; if the claimed gains vanish or reverse, the alignment design is not load-bearing.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes RefAlign, a training-time representation alignment method for reference-to-video (R2V) generation. Building on a Wan2.1 DiT backbone, it introduces a reference alignment (RA) loss that projects intermediate reference-token features from the first K DiT blocks and aligns them to frozen visual foundation model (VFM) features: a positive cosine term pulls same-subject DiT and VFM tokens together, while a margin-based negative term pushes cross-subject pairs apart. The VFM and projector are discarded at inference, so there is no extra runtime cost. On OpenS2V-Eval the method reports state-of-the-art TotalScore for both 1.3B and 14B models (60.42% at 14B), with gains concentrated in FaceSim and NexusScore, supported by qualitative comparisons, a user study, and ablations on loss terms, alignment depth, and encoder choice.
Significance. If the reported gains hold under fuller evaluation, RefAlign is a practically useful contribution to controllable video generation: it targets well-known R2V failure modes (copy–paste leakage and multi-subject confusion) with a simple regularizer that adds no inference overhead and is compatible with a strong open backbone. The positive/negative RA design, the explicit contrast with REPA (reference-condition alignment vs. generation-target alignment), and the systematic encoder/depth ablations are concrete engineering insights that other R2V and video-editing pipelines can reuse. Code release is promised, which would further raise impact. The work is primarily empirical rather than theoretical, but the zero-overhead alignment idea is timely and transferable.
major comments (3)
- Table 1 and the abstract claim a better balance of text controllability and reference fidelity via SOTA TotalScore (60.42% at 14B). RefAlign-14B does lead FaceSim and NexusScore, but trails several baselines on GmeScore (text–video alignment; e.g., Phantom-14B 70.65% vs. 68.32%) and is not best on Aesthetics or NaturalScore (Kaleido 82.18%). The 1.3B row is similarly mixed. Because TotalScore is a composite whose weighting is not specified in the manuscript, the headline SOTA does not by itself establish a robust Pareto improvement. Please report how TotalScore is computed (or cite the exact OpenS2V formula), and add an explicit trade-off discussion—or a simple Pareto/radar view—so readers can judge whether identity gains come at a systematic cost to prompt following.
- §4.4 states that all ablations (Tables 2–3, Fig. 7) use 1800 training iterations, while the main models are trained for 3000 iterations (§4.1). Relative conclusions that are load-bearing for the method—necessity of L_neg (A vs. B), superiority of alignment over dual-encoder input (A vs. D), and the depth peak at K=9—may shift under the full schedule. Either re-run the critical ablations to 3000 iterations or provide evidence that rankings stabilize by 1800; otherwise the design choices that define RefAlign rest on a mismatched protocol.
- §3.3, Eq. (9): the negative term depends on a margin δ, and §4.1 reports λ=η=1.0 but never states the value of δ (nor a default when M>1). Without δ, L_neg is not reproducible. Please specify δ, how it was chosen, and whether results are sensitive to it. Relatedly, OpenS2V-Eval uses only 180 videos with no multi-seed error bars or significance tests; given that TotalScore gaps to strong open baselines are a few points, confidence intervals or repeated sampling would substantially strengthen the SOTA claim.
minor comments (5)
- §3.3 / Fig. 3: notation for projected features mixes ˆh^(l) and f; a short glossary of token shapes (M×N×D) would help readers implement the loss.
- Fig. 2(b) t-SNE is motivating but qualitative; stating the number of references/patches and whether features are taken before or after the MLP would make the entanglement claim more precise.
- Implementation: data-augmentation list for regular pairs is helpful; please also state reference resolution and whether VFM inputs are resized independently of the VAE path (Appendix A.3 hints at 480×832).
- Typos / polish: “V AE” spacing is inconsistent; “copy—paste” vs. “copy–paste”; “alleviating copy—paste” in §1; “we will make the model and code publicly available” appears mid-contribution list.
- User study (§4.5): report number of video pairs per comparison and whether raters saw the reference images, so preference rates can be interpreted.
Circularity Check
No circularity: empirical regularizer evaluated on an external benchmark; nothing reduces to its inputs by construction.
full rationale
RefAlign is a standard empirical methods paper. The load-bearing claim is that a training-only reference alignment loss (Eqs. 8–11: positive cosine pull of DiT reference tokens to same-subject VFM features plus a margin push against other subjects) improves identity consistency and multi-subject discriminability, yielding higher TotalScore on OpenS2V-Eval with zero inference cost. That claim is not a derivation: L_RA is an optimization regularizer whose form is stated independently of the evaluation metrics; TotalScore, FaceSim, NexusScore, etc. are computed by an external benchmark (OpenS2V-Eval) on generated videos, not algebraic rearrangements of the loss or of fitted constants. Hyperparameters (depth K=9, λ, η, DINOv3-L) are chosen via ablations and then frozen for the main comparison—normal ML practice, not “fitted input called prediction” of a related quantity. Inspiration from REPA is openly differentiated (reference-branch vs generation-target alignment; clean vs noisy features; added L_neg for multi-reference separability); REPA’s authors do not overlap with this paper, and no uniqueness theorem or self-cited ansatz is used to force the design. Self-citations (e.g., REG by co-author Ge Wu) appear only as related work and are not load-bearing. There is no self-definitional loop, no renaming of a known closed-form result as a first-principles prediction, and no step where Eq. X equals Eq. Y by construction. Score 0 is the correct honest finding.
Axiom & Free-Parameter Ledger
free parameters (5)
- RA loss weight η =
1.0
- negative-term weight λ =
1.0
- alignment depth K =
9
- margin δ in L_neg
- CFG scales μ1, μ2 =
5.0 / 7.5
axioms (4)
- domain assumption Rectified-flow training objective on Wan2.1 DiT is a valid base for fine-tuning R2V conditioning.
- domain assumption Frozen VFM patch features (especially DINOv3) supply identity-sensitive, appearance-robust semantic anchors suitable for aligning reference tokens.
- domain assumption OpenS2V-Eval TotalScore and its component metrics are adequate proxies for reference fidelity and text controllability.
- ad hoc to paper Cosine similarity (and hinge on 1−cos) is a sufficient metric for pull/push alignment of patch tokens.
invented entities (2)
-
Reference Alignment (RA) loss with positive and negative terms
no independent evidence
-
RefAlign training pipeline (VFM discarded at inference)
no independent evidence
read the original abstract
Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and reference images, enabling applications such as personalized advertising and virtual try-on. In practice, existing R2V methods typically introduce additional high-level semantic or cross-modal features alongside the VAE latent representation of the reference image and jointly feed them into the diffusion Transformer (DiT). These auxiliary representations provide semantic guidance and act as implicit alignment signals, which can partially alleviate pixel-level information leakage in the VAE latent space. However, they may still struggle to address copy--paste artifacts and multi-subject confusion caused by modality mismatch across heterogeneous encoder features. In this paper, we propose RefAlign, a representation alignment framework that explicitly aligns DiT reference-branch features to the semantic space of a visual foundation model (VFM). The core of RefAlign is a reference alignment loss that pulls the reference features and VFM features of the same subject closer to improve identity consistency, while pushing apart the corresponding features of different subjects to enhance semantic discriminability. This simple yet effective strategy is applied only during training, incurring no inference-time overhead, and achieves a better balance between text controllability and reference fidelity. Extensive experiments on the OpenS2V-Eval benchmark demonstrate that RefAlign outperforms current state-of-the-art methods in TotalScore, validating the effectiveness of explicit reference alignment for R2V tasks.
Forward citations
Cited by 5 Pith papers
-
Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment
Aura combines VLM meta-queries, T5-teacher alignment, subject-aware RoPE shifts, memory tokens, and a large AIGC-curated dataset to claim SOTA multi-element subject-to-video generation under OpenS2V-Eval Total score.
-
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping
Super-Linear Advantage Shaping (SLAS) introduces a non-linear geometric policy update for RL post-training of text-to-image models that reshapes the local policy space via advantage-dependent Fisher-Rao weighting to r...
-
SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models
SARA improves text alignment and motion quality in video diffusion models by routing token-relation distillation supervision to semantically salient pairs using a Stage-1 aligner trained with SAM masks and InfoNCE.
-
SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models
SARA introduces semantic saliency to guide relational alignment in video diffusion models, improving text following and motion quality over prior alignment methods.
-
Bernini: Latent Semantic Planning for Video Diffusion
Bernini is a framework that uses an MLLM planner to output semantic representations for a DiT renderer to generate or edit videos, reporting SOTA benchmark performance.
Reference graph
Works this paper leans on
-
[1]
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al. Video generation models as world simulators. OpenAI Blog, 1(8):1, 2024. 2
2024
-
[2]
Fan Bao, Chendong Xiang, Gang Yue, Guande He, Hongzhou Zhu, Kaiwen Zheng, Min Zhao, Shilong Liu, Yaole Wang, and Jun Zhu. Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models.arXiv preprint arXiv:2405.04233, 2024. 2, 7, 8
Pith/arXiv arXiv 2024
-
[3]
Kling-omni technical report.arXiv preprint arXiv:2512.16776,
Kling Team, Jialu Chen, Yuanzheng Ci, Xiangyu Du, Zipeng Feng, Kun Gai, Sainan Guo, Feng Han, Jingbin He, Kang He, et al. Kling-omni technical report.arXiv preprint arXiv:2512.16776,
-
[4]
Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025. 2, 3, 4, 5, 7, 8
Pith/arXiv arXiv 2025
-
[5]
Cogvideox: Text-to-video diffusion models with an expert transformer
Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. InICLR, 2025. 2
2025
-
[6]
Hunyuanvideo 1.5 technical report.arXiv preprint arXiv:2511.18870, 2025
Bing Wu, Chang Zou, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Jack Peng, Jianbing Wu, Jiangfeng Xiong, Jie Jiang, et al. Hunyuanvideo 1.5 technical report.arXiv preprint arXiv:2511.18870, 2025. 2
Pith/arXiv arXiv 2025
-
[7]
Multi- subject open-set personalization in video generation
Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Yuwei Fang, Kwot Sin Lee, Ivan Skorokhodov, Kfir Aberman, Jun-Yan Zhu, Ming-Hsuan Yang, and Sergey Tulyakov. Multi- subject open-set personalization in video generation. InCVPR, pages 6099–6110, 2025. 2, 3, 5
2025
-
[8]
Yuzhou Huang, Ziyang Yuan, Quande Liu, Qiulin Wang, Xintao Wang, Ruimao Zhang, Pengfei Wan, Di Zhang, and Kun Gai. Conceptmaster: Multi-concept video customization on diffusion transformer models without test-time tuning.arXiv preprint arXiv:2501.04698, 2025. 2, 3, 4
Pith/arXiv arXiv 2025
-
[9]
Phantom: Subject-consistent video generation via cross-modal alignment
Lijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen, Jiawei Liu, Gen Li, Siyu Zhou, Qian He, and Xinglong Wu. Phantom: Subject-consistent video generation via cross-modal alignment. InICCV, pages 14951–14961, October 2025. 2, 3, 5, 7, 8, 16 11
2025
-
[10]
Goku: Flow based video generative foundation models
Shoufa Chen, Chongjian Ge, Yuqi Zhang, Yida Zhang, Fengda Zhu, Hao Yang, Hongxiang Hao, Hui Wu, Zhichao Lai, Yifei Hu, et al. Goku: Flow based video generative foundation models. InCVPR, pages 23516–23527, 2025. 2
2025
-
[11]
Movie weaver: Tuning-free multi-concept video personalization with anchored prompts
Feng Liang, Haoyu Ma, Zecheng He, Tingbo Hou, Ji Hou, Kunpeng Li, Xiaoliang Dai, Felix Juefei-Xu, Samaneh Azadi, Animesh Sinha, et al. Movie weaver: Tuning-free multi-concept video personalization with anchored prompts. InCVPR, pages 13146–13156, 2025. 2
2025
-
[12]
Swifttry: Fast and consistent video virtual try-on with diffusion models
Hung Nguyen, Quang Qui-Vinh Nguyen, Khoi Nguyen, and Rang Nguyen. Swifttry: Fast and consistent video virtual try-on with diffusion models. InAAAI, volume 39, pages 6200–6208,
-
[13]
Pursuing temporal-consistent video virtual try-on via dynamic pose interaction
Dong Li, Wenqi Zhong, Wei Yu, Yingwei Pan, Dingwen Zhang, Ting Yao, Junwei Han, and Tao Mei. Pursuing temporal-consistent video virtual try-on via dynamic pose interaction. In CVPR, pages 22648–22657, 2025. 2
2025
-
[14]
Zhengcong Fei, Debang Li, Di Qiu, Jiahua Wang, Yikun Dou, Rui Wang, Jingtao Xu, Mingyuan Fan, Guibin Chen, Yang Li, et al. Skyreels-a2: Compose anything in video diffusion transform- ers.arXiv preprint arXiv:2504.02436, 2025. 2, 3, 7, 8
Pith/arXiv arXiv 2025
-
[15]
Yufan Deng, Xun Guo, Yizhi Wang, Jacob Zhiyuan Fang, Angtian Wang, Shenghai Yuan, Yiding Yang, Bo Liu, Haibin Huang, and Chongyang Ma. Cinema: Coherent multi-subject video generation via mllm-based guidance.arXiv preprint arXiv:2503.10391, 2025. 2, 3, 4
Pith/arXiv arXiv 2025
-
[16]
Bindweave: Subject-consistent video generation via cross- modal integration.ICLR, 2026
Zhaoyang Li, Dongjun Qian, Kai Su, Qishuai Diao, Xiangyang Xia, Chang Liu, Wenfei Yang, Tianzhu Zhang, and Zehuan Yuan. Bindweave: Subject-consistent video generation via cross- modal integration.ICLR, 2026. 2, 3, 4, 7, 8
2026
-
[17]
Id-crafter: Vlm-grounded online rl for compositional multi-subject video generation.CVPR, 2026
Panwang Pan, Jingjing Zhao, Yuchen Lin, Chenguo Lin, Chenxin Li, Hengyu Liu, Tingting Shen, and Yadong Mu. Id-crafter: Vlm-grounded online rl for compositional multi-subject video generation.CVPR, 2026. 2, 4
2026
-
[18]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv preprint arXiv:2409.12191, 2024. 2
Pith/arXiv arXiv 2024
-
[19]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report.arXiv preprint arXiv:2502.13923,
-
[20]
Image to video elements feature
kling. Image to video elements feature. https://klingai.com/image-to-video/ multi-id/new/, 2024. 3, 7, 8, 16
2024
-
[21]
Visualizing data using t-sne.JMLR, 9(86):2579– 2605, 2008
Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne.JMLR, 9(86):2579– 2605, 2008. 3
2008
-
[22]
Oriane Siméoni, Huy V V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, et al. Dinov3. arXiv preprint arXiv:2508.10104, 2025. 2, 6, 14
Pith/arXiv arXiv 2025
-
[23]
Opens2v-nexus: A detailed benchmark and million-scale dataset for subject-to-video generation
Shenghai Yuan, Xianyi He, Yufan Deng, Yang Ye, Jinfa Huang, Bin Lin, Jiebo Luo, and Li Yuan. Opens2v-nexus: A detailed benchmark and million-scale dataset for subject-to-video generation. NeurIPS, 2025. 3, 8, 14, 15
2025
-
[24]
Stand-in: A lightweight and plug-and-play identity control for video generation.CVPR, 2026
Bowen Xue, Zheng-Peng Duan, Qixin Yan, Wenjing Wang, Hao Liu, Chun-Le Guo, Chongyi Li, Chen Li, and Jing Lyu. Stand-in: A lightweight and plug-and-play identity control for video generation.CVPR, 2026. 3
2026
-
[25]
Concat-id: Towards universal identity-preserving video synthesis
Yong Zhong, Zhuoyi Yang, Jiayan Teng, Xiaotao Gu, and Chongxuan Li. Concat-id: Towards universal identity-preserving video synthesis. InICCVW, pages 1906–1915, 2025. 3
1906
-
[26]
Identity-preserving text-to-video generation by frequency decomposition
Shenghai Yuan, Jinfa Huang, Xianyi He, Yunyang Ge, Yujun Shi, Liuhan Chen, Jiebo Luo, and Li Yuan. Identity-preserving text-to-video generation by frequency decomposition. InCVPR, pages 12978–12988, 2025. 3 12
2025
-
[27]
Lynx: Towards high-fidelity personalized video generation
Shen Sang, Tiancheng Zhi, Tianpei Gu, Jing Liu, and Linjie Luo. Lynx: Towards high-fidelity personalized video generation. InCVPR, 2026. 3
2026
-
[28]
Zhenxing Zhang, Jiayan Teng, Zhuoyi Yang, Tiankun Cao, Cheng Wang, Xiaotao Gu, Jie Tang, Dan Guo, and Meng Wang. Kaleido: Open-sourced multi-subject reference video generation model.arXiv preprint arXiv:2510.18573, 2025. 3, 4, 7, 8
arXiv 2025
-
[29]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InICML, pages 8748–8763. PmLR, 2021. 3, 14
2021
-
[30]
Scaling zero-shot reference-to-video generation
Zijian Zhou, Shikun Liu, Haozhe Liu, Haonan Qiu, Zhaochong An, Weiming Ren, Zhiheng Liu, Xiaoke Huang, Kam Woh Ng, Tian Xie, et al. Scaling zero-shot reference-to-video generation. arXiv preprint arXiv:2512.06905, 2025. 4, 7, 8
arXiv 2025
-
[31]
Magref: Masked guidance for any-reference video generation.ICLR, 2026
Yufan Deng, Xun Guo, Yuanyang Yin, Jacob Zhiyuan Fang, Yiding Yang, Yizhi Wang, Shenghai Yuan, Angtian Wang, Bo Liu, Haibin Huang, et al. Magref: Masked guidance for any-reference video generation.ICLR, 2026. 4, 7, 8
2026
-
[32]
Polyvivid: Vivid multi-subject video generation with cross-modal interaction and enhancement
Teng Hu, Zhentao Yu, Zhengguang Zhou, Jiangning Zhang, Yuan Zhou, Qinglin Lu, and Ran Yi. Polyvivid: Vivid multi-subject video generation with cross-modal interaction and enhancement. NeurIPS, 2025. 4
2025
-
[33]
Teng Hu, Zhentao Yu, Zhengguang Zhou, Sen Liang, Yuan Zhou, Qin Lin, and Qinglin Lu. Hunyuancustom: A multimodal-driven architecture for customized video generation.arXiv preprint arXiv:2505.04512, 2025. 4
Pith/arXiv arXiv 2025
-
[34]
Junyi Chen, Tong He, Zhoujie Fu, Pengfei Wan, Kun Gai, and Weicai Ye. Vino: A unified visual generator with interleaved omnimodal context.arXiv preprint arXiv:2601.02358, 2026. 4, 7, 8, 16
arXiv 2026
-
[35]
Visual instruction tuning.NeurIPS, 36:34892–34916, 2023
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning.NeurIPS, 36:34892–34916, 2023. 4
2023
-
[36]
Vace: All-in-one video creation and editing
Zeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang, Yulin Pan, and Yu Liu. Vace: All-in-one video creation and editing. InICCV, pages 17191–17202, 2025. 4, 7, 8
2025
-
[37]
Representation alignment for generation: Training diffusion transformers is easier than you think
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representation alignment for generation: Training diffusion transformers is easier than you think. InICLR, 2025. 4, 5, 6
2025
-
[38]
Repa-e: Unlocking vae for end-to-end tuning of latent diffusion transformers
Xingjian Leng, Jaskirat Singh, Yunzhong Hou, Zhenchang Xing, Saining Xie, and Liang Zheng. Repa-e: Unlocking vae for end-to-end tuning of latent diffusion transformers. InICCV, pages 18262–18272, 2025. 4
2025
-
[39]
Ddt: Decoupled diffusion transformer
Shuai Wang, Zhi Tian, Weilin Huang, and Limin Wang. Ddt: Decoupled diffusion transformer. CVPR, 2026. 4
2026
-
[40]
Representation entanglement for generation: Training diffusion transformers is much easier than you think.NeurIPS, 2025
Ge Wu, Shen Zhang, Ruijing Shi, Shanghua Gao, Zhenyuan Chen, Lei Wang, Zhaowei Chen, Hongcheng Gao, Yao Tang, Jian Yang, et al. Representation entanglement for generation: Training diffusion transformers is much easier than you think.NeurIPS, 2025. 4
2025
-
[41]
Boosting generative image modeling via joint image-feature synthesis.NeurIPS,
Theodoros Kouzelis, Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, and Nikos Komodakis. Boosting generative image modeling via joint image-feature synthesis.NeurIPS,
-
[42]
Reconstruction vs
Jingfeng Yao, Bin Yang, and Xinggang Wang. Reconstruction vs. generation: Taming optimiza- tion dilemma in latent diffusion models. InCVPR, pages 15703–15712, 2025. 4
2025
-
[43]
Unleashing the potential of large language models for text-to-image generation through autore- gressive representation alignment.AAAI, 2026
Xing Xie, Jiawei Liu, Ziyue Lin, Huijie Fan, Zhi Han, Yandong Tang, and Liangqiong Qu. Unleashing the potential of large language models for text-to-image generation through autore- gressive representation alignment.AAAI, 2026. 4 13
2026
-
[44]
Videorepa: Learning physics for video generation through relational alignment with foundation models.NeurIPS, 2025
Xiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng, Xiangpeng Wan, Junchi Yan, and Yu Cheng. Videorepa: Learning physics for video generation through relational alignment with foundation models.NeurIPS, 2025. 4
2025
-
[45]
Scaling rectified flow transform- ers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transform- ers for high-resolution image synthesis. InICML, 2024. 4
2024
-
[46]
Dinov2: Learning robust visual features without supervision.TMLR, 2024
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision.TMLR, 2024. 4, 14
2024
-
[47]
Exploring the limits of transfer learning with a unified text-to-text transformer.JMLR, 21(140):1–67, 2020
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.JMLR, 21(140):1–67, 2020. 5
2020
-
[48]
Classifier-free diffusion guidance.arXiv preprint arXiv:2207.12598, 2022
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance.arXiv preprint arXiv:2207.12598, 2022. 5
Pith/arXiv arXiv 2022
-
[49]
Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alab- dulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, et al. Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features.arXiv preprint arXiv:2502.14786, 2025. 6, 14
Pith/arXiv arXiv 2025
-
[50]
Pikascenes.https://pika.art/ingredients/, 2024
Pika. Pikascenes.https://pika.art/ingredients/, 2024. 7, 8
2024
-
[51]
Phantom-data: Towards a general subject- consistent video generation dataset.ICLR, 2026
Zhuowei Chen, Bingchuan Li, Tianxiang Ma, Lijie Liu, Mingcong Liu, Yi Zhang, Gen Li, Xinghui Li, Siyu Zhou, Qian He, and Xinglong Wu. Phantom-data: Towards a general subject- consistent video generation dataset.ICLR, 2026. 8, 14
2026
-
[52]
Decoupled weight decay regularization.ICLR, 2019
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization.ICLR, 2019. 8
2019
-
[53]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. InCVPR, pages 16000–16009, 2022. 14
2022
-
[54]
Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024. 14 A Appendix A.1 Additional Training Details During training, we randomly drop the prompt, the reference, or both, each with a probability of 10%, for CFG. We train our model in two sta...
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.