Pith. sign in

REVIEW 3 major objections 5 minor 5 cited by

Explicitly aligning a video model's reference features to a visual foundation model removes copy-paste artifacts and multi-subject confusion without slowing inference.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 18:00 UTC pith:MPHVRMY2

load-bearing objection Clean engineering paper: REPA-style alignment retargeted to multi-reference DiT tokens with a real push term, SOTA TotalScore, zero inference cost; metric-tradeoff caveats are real but not fatal. the 3 major comments →

arxiv 2603.25743 v2 pith:MPHVRMY2 submitted 2026-03-26 cs.CV

RefAlign: Representation Alignment for Reference-to-Video Generation

classification cs.CV
keywords reference-to-video generationrepresentation alignmentdiffusion transformervisual foundation modelidentity consistencymulti-subject confusioncopy-paste artifactsOpenS2V-Eval
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Reference-to-video generation must follow a text prompt while keeping the identity and appearance of one or more reference subjects. Existing systems pack the reference into a VAE latent and often add extra semantic or multimodal features, hoping that joint injection will align everything inside the diffusion transformer. That hope fails when the features come from mismatched encoders: the model either pastes the reference almost pixel-for-pixel or mixes identities when several subjects appear. RefAlign shows that the fix is not more features at inference time but an explicit training-time alignment loss. The loss pulls the transformer's own reference-branch tokens toward the same-subject tokens of a frozen visual foundation model and pushes them away from other subjects. The foundation model and the projector are discarded after training, so generation cost is unchanged. On a public R2V benchmark the resulting model records the highest overall score reported so far, with clearer identity fidelity and better prompt following.

Core claim

The paper establishes that an explicit reference alignment loss, applied only to intermediate reference-image tokens inside the first several blocks of a diffusion transformer, is sufficient to give those tokens the identity consistency and inter-subject separability of a strong visual foundation model. The same positive-and-negative alignment removes the copy-paste and multi-subject failures that arise from unregularized VAE latents, while leaving inference identical to the base video model.

What carries the argument

Reference alignment (RA) loss: a cosine-similarity objective that pulls projected DiT reference tokens toward matching VFM tokens of the same subject and pushes them away from tokens of other subjects by a margin; averaged over the first K transformer blocks and added to the ordinary flow-matching objective.

Load-bearing premise

The method assumes that intermediate reference tokens in a chosen early depth of the transformer are the right place to force alignment to a particular frozen visual encoder, and that that encoder is identity-sensitive enough; if either choice is wrong the loss can hurt fidelity or collapse subject distinctions.

What would settle it

Train identical models with and without the negative (push-apart) term, or at the paper's preferred depth versus much shallower or deeper depths, then re-score FaceSim, NexusScore and multi-subject qualitative examples on the same OpenS2V-Eval set; if the claimed gains vanish or reverse, the alignment design is not load-bearing.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes RefAlign, a training-time representation alignment method for reference-to-video (R2V) generation. Building on a Wan2.1 DiT backbone, it introduces a reference alignment (RA) loss that projects intermediate reference-token features from the first K DiT blocks and aligns them to frozen visual foundation model (VFM) features: a positive cosine term pulls same-subject DiT and VFM tokens together, while a margin-based negative term pushes cross-subject pairs apart. The VFM and projector are discarded at inference, so there is no extra runtime cost. On OpenS2V-Eval the method reports state-of-the-art TotalScore for both 1.3B and 14B models (60.42% at 14B), with gains concentrated in FaceSim and NexusScore, supported by qualitative comparisons, a user study, and ablations on loss terms, alignment depth, and encoder choice.

Significance. If the reported gains hold under fuller evaluation, RefAlign is a practically useful contribution to controllable video generation: it targets well-known R2V failure modes (copy–paste leakage and multi-subject confusion) with a simple regularizer that adds no inference overhead and is compatible with a strong open backbone. The positive/negative RA design, the explicit contrast with REPA (reference-condition alignment vs. generation-target alignment), and the systematic encoder/depth ablations are concrete engineering insights that other R2V and video-editing pipelines can reuse. Code release is promised, which would further raise impact. The work is primarily empirical rather than theoretical, but the zero-overhead alignment idea is timely and transferable.

major comments (3)
  1. Table 1 and the abstract claim a better balance of text controllability and reference fidelity via SOTA TotalScore (60.42% at 14B). RefAlign-14B does lead FaceSim and NexusScore, but trails several baselines on GmeScore (text–video alignment; e.g., Phantom-14B 70.65% vs. 68.32%) and is not best on Aesthetics or NaturalScore (Kaleido 82.18%). The 1.3B row is similarly mixed. Because TotalScore is a composite whose weighting is not specified in the manuscript, the headline SOTA does not by itself establish a robust Pareto improvement. Please report how TotalScore is computed (or cite the exact OpenS2V formula), and add an explicit trade-off discussion—or a simple Pareto/radar view—so readers can judge whether identity gains come at a systematic cost to prompt following.
  2. §4.4 states that all ablations (Tables 2–3, Fig. 7) use 1800 training iterations, while the main models are trained for 3000 iterations (§4.1). Relative conclusions that are load-bearing for the method—necessity of L_neg (A vs. B), superiority of alignment over dual-encoder input (A vs. D), and the depth peak at K=9—may shift under the full schedule. Either re-run the critical ablations to 3000 iterations or provide evidence that rankings stabilize by 1800; otherwise the design choices that define RefAlign rest on a mismatched protocol.
  3. §3.3, Eq. (9): the negative term depends on a margin δ, and §4.1 reports λ=η=1.0 but never states the value of δ (nor a default when M>1). Without δ, L_neg is not reproducible. Please specify δ, how it was chosen, and whether results are sensitive to it. Relatedly, OpenS2V-Eval uses only 180 videos with no multi-seed error bars or significance tests; given that TotalScore gaps to strong open baselines are a few points, confidence intervals or repeated sampling would substantially strengthen the SOTA claim.
minor comments (5)
  1. §3.3 / Fig. 3: notation for projected features mixes ˆh^(l) and f; a short glossary of token shapes (M×N×D) would help readers implement the loss.
  2. Fig. 2(b) t-SNE is motivating but qualitative; stating the number of references/patches and whether features are taken before or after the MLP would make the entanglement claim more precise.
  3. Implementation: data-augmentation list for regular pairs is helpful; please also state reference resolution and whether VFM inputs are resized independently of the VAE path (Appendix A.3 hints at 480×832).
  4. Typos / polish: “V AE” spacing is inconsistent; “copy—paste” vs. “copy–paste”; “alleviating copy—paste” in §1; “we will make the model and code publicly available” appears mid-contribution list.
  5. User study (§4.5): report number of video pairs per comparison and whether raters saw the reference images, so preference rates can be interpreted.

Circularity Check

0 steps flagged

No circularity: empirical regularizer evaluated on an external benchmark; nothing reduces to its inputs by construction.

full rationale

RefAlign is a standard empirical methods paper. The load-bearing claim is that a training-only reference alignment loss (Eqs. 8–11: positive cosine pull of DiT reference tokens to same-subject VFM features plus a margin push against other subjects) improves identity consistency and multi-subject discriminability, yielding higher TotalScore on OpenS2V-Eval with zero inference cost. That claim is not a derivation: L_RA is an optimization regularizer whose form is stated independently of the evaluation metrics; TotalScore, FaceSim, NexusScore, etc. are computed by an external benchmark (OpenS2V-Eval) on generated videos, not algebraic rearrangements of the loss or of fitted constants. Hyperparameters (depth K=9, λ, η, DINOv3-L) are chosen via ablations and then frozen for the main comparison—normal ML practice, not “fitted input called prediction” of a related quantity. Inspiration from REPA is openly differentiated (reference-branch vs generation-target alignment; clean vs noisy features; added L_neg for multi-reference separability); REPA’s authors do not overlap with this paper, and no uniqueness theorem or self-cited ansatz is used to force the design. Self-citations (e.g., REG by co-author Ge Wu) appear only as related work and are not load-bearing. There is no self-definitional loop, no renaming of a known closed-form result as a first-principles prediction, and no step where Eq. X equals Eq. Y by construction. Score 0 is the correct honest finding.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 2 invented entities

The claim rests on standard diffusion/DiT training plus a small set of hand-chosen alignment hyperparameters and the modeling choice that VFM features are good teachers for reference identity. No new physical entities; free parameters are ordinary ML knobs whose values are reported and ablated.

free parameters (5)
  • RA loss weight η = 1.0
    Scalar balancing L_RF and L_RA; set to 1.0 by hand.
  • negative-term weight λ = 1.0
    Controls strength of inter-subject push; set to 1.0.
  • alignment depth K = 9
    Number of early DiT blocks whose reference tokens are aligned; chosen by ablation peak at 9.
  • margin δ in L_neg
    Hinge margin for pushing mismatched subject pairs; present in Eq. (9) but numerical value not explicitly tabulated.
  • CFG scales μ1, μ2 = 5.0 / 7.5
    Inference guidance weights for reference and text; set to 5.0 and 7.5.
axioms (4)
  • domain assumption Rectified-flow training objective on Wan2.1 DiT is a valid base for fine-tuning R2V conditioning.
    Assumed throughout §3.1–3.2; standard in recent video DiTs.
  • domain assumption Frozen VFM patch features (especially DINOv3) supply identity-sensitive, appearance-robust semantic anchors suitable for aligning reference tokens.
    Core premise of §3.3–3.4 and the t-SNE motivation in Fig. 2.
  • domain assumption OpenS2V-Eval TotalScore and its component metrics are adequate proxies for reference fidelity and text controllability.
    All quantitative claims rest on this external benchmark (§4).
  • ad hoc to paper Cosine similarity (and hinge on 1−cos) is a sufficient metric for pull/push alignment of patch tokens.
    Defines L_pos and L_neg in Eqs. (8)–(9); inherited from REPA-style practice but not derived.
invented entities (2)
  • Reference Alignment (RA) loss with positive and negative terms no independent evidence
    purpose: Explicitly regularize DiT reference-branch features toward same-subject VFM features and away from other subjects.
    Central technical contribution; defined in §3.3; independent evidence is the ablation and benchmark gains, not an external physical prediction.
  • RefAlign training pipeline (VFM discarded at inference) no independent evidence
    purpose: Apply alignment only at train time so reference controllability improves without inference overhead.
    Framework-level packaging of the RA loss on Wan2.1; evidence is empirical.

pith-pipeline@v1.1.0-grok45 · 21783 in / 2908 out tokens · 30778 ms · 2026-07-13T18:00:18.328395+00:00 · methodology

0 comments
read the original abstract

Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and reference images, enabling applications such as personalized advertising and virtual try-on. In practice, existing R2V methods typically introduce additional high-level semantic or cross-modal features alongside the VAE latent representation of the reference image and jointly feed them into the diffusion Transformer (DiT). These auxiliary representations provide semantic guidance and act as implicit alignment signals, which can partially alleviate pixel-level information leakage in the VAE latent space. However, they may still struggle to address copy--paste artifacts and multi-subject confusion caused by modality mismatch across heterogeneous encoder features. In this paper, we propose RefAlign, a representation alignment framework that explicitly aligns DiT reference-branch features to the semantic space of a visual foundation model (VFM). The core of RefAlign is a reference alignment loss that pulls the reference features and VFM features of the same subject closer to improve identity consistency, while pushing apart the corresponding features of different subjects to enhance semantic discriminability. This simple yet effective strategy is applied only during training, incurring no inference-time overhead, and achieves a better balance between text controllability and reference fidelity. Extensive experiments on the OpenS2V-Eval benchmark demonstrate that RefAlign outperforms current state-of-the-art methods in TotalScore, validating the effectiveness of explicit reference alignment for R2V tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment

    cs.CV 2026-07 conditional novelty 6.0

    Aura combines VLM meta-queries, T5-teacher alignment, subject-aware RoPE shifts, memory tokens, and a large AIGC-curated dataset to claim SOTA multi-element subject-to-video generation under OpenS2V-Eval Total score.

  2. Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping

    cs.CV 2026-05 unverdicted novelty 6.0

    Super-Linear Advantage Shaping (SLAS) introduces a non-linear geometric policy update for RL post-training of text-to-image models that reshapes the local policy space via advantage-dependent Fisher-Rao weighting to r...

  3. SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models

    cs.CV 2026-05 unverdicted novelty 6.0

    SARA improves text alignment and motion quality in video diffusion models by routing token-relation distillation supervision to semantically salient pairs using a Stage-1 aligner trained with SAM masks and InfoNCE.

  4. SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models

    cs.CV 2026-05 unverdicted novelty 6.0

    SARA introduces semantic saliency to guide relational alignment in video diffusion models, improving text following and motion quality over prior alignment methods.

  5. Bernini: Latent Semantic Planning for Video Diffusion

    cs.CV 2026-05 unverdicted novelty 5.0

    Bernini is a framework that uses an MLLM planner to output semantic representations for a DiT renderer to generate or edit videos, reporting SOTA benchmark performance.

Reference graph

Works this paper leans on

54 extracted references · 13 linked inside Pith · cited by 4 Pith papers

  1. [1]

    Video generation models as world simulators

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, et al. Video generation models as world simulators. OpenAI Blog, 1(8):1, 2024. 2

  2. [2]

    Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models.arXiv preprint arXiv:2405.04233, 2024

    Fan Bao, Chendong Xiang, Gang Yue, Guande He, Hongzhou Zhu, Kaiwen Zheng, Min Zhao, Shilong Liu, Yaole Wang, and Jun Zhu. Vidu: a highly consistent, dynamic and skilled text-to-video generator with diffusion models.arXiv preprint arXiv:2405.04233, 2024. 2, 7, 8

  3. [3]

    Kling-omni technical report.arXiv preprint arXiv:2512.16776,

    Kling Team, Jialu Chen, Yuanzheng Ci, Xiangyu Du, Zipeng Feng, Kun Gai, Sainan Guo, Feng Han, Jingbin He, Kang He, et al. Kling-omni technical report.arXiv preprint arXiv:2512.16776,

  4. [4]

    Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025. 2, 3, 4, 5, 7, 8

  5. [5]

    Cogvideox: Text-to-video diffusion models with an expert transformer

    Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. InICLR, 2025. 2

  6. [6]

    Hunyuanvideo 1.5 technical report.arXiv preprint arXiv:2511.18870, 2025

    Bing Wu, Chang Zou, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Jack Peng, Jianbing Wu, Jiangfeng Xiong, Jie Jiang, et al. Hunyuanvideo 1.5 technical report.arXiv preprint arXiv:2511.18870, 2025. 2

  7. [7]

    Multi- subject open-set personalization in video generation

    Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Yuwei Fang, Kwot Sin Lee, Ivan Skorokhodov, Kfir Aberman, Jun-Yan Zhu, Ming-Hsuan Yang, and Sergey Tulyakov. Multi- subject open-set personalization in video generation. InCVPR, pages 6099–6110, 2025. 2, 3, 5

  8. [8]

    Conceptmaster: Multi-concept video customization on diffusion transformer models without test-time tuning.arXiv preprint arXiv:2501.04698, 2025

    Yuzhou Huang, Ziyang Yuan, Quande Liu, Qiulin Wang, Xintao Wang, Ruimao Zhang, Pengfei Wan, Di Zhang, and Kun Gai. Conceptmaster: Multi-concept video customization on diffusion transformer models without test-time tuning.arXiv preprint arXiv:2501.04698, 2025. 2, 3, 4

  9. [9]

    Phantom: Subject-consistent video generation via cross-modal alignment

    Lijie Liu, Tianxiang Ma, Bingchuan Li, Zhuowei Chen, Jiawei Liu, Gen Li, Siyu Zhou, Qian He, and Xinglong Wu. Phantom: Subject-consistent video generation via cross-modal alignment. InICCV, pages 14951–14961, October 2025. 2, 3, 5, 7, 8, 16 11

  10. [10]

    Goku: Flow based video generative foundation models

    Shoufa Chen, Chongjian Ge, Yuqi Zhang, Yida Zhang, Fengda Zhu, Hao Yang, Hongxiang Hao, Hui Wu, Zhichao Lai, Yifei Hu, et al. Goku: Flow based video generative foundation models. InCVPR, pages 23516–23527, 2025. 2

  11. [11]

    Movie weaver: Tuning-free multi-concept video personalization with anchored prompts

    Feng Liang, Haoyu Ma, Zecheng He, Tingbo Hou, Ji Hou, Kunpeng Li, Xiaoliang Dai, Felix Juefei-Xu, Samaneh Azadi, Animesh Sinha, et al. Movie weaver: Tuning-free multi-concept video personalization with anchored prompts. InCVPR, pages 13146–13156, 2025. 2

  12. [12]

    Swifttry: Fast and consistent video virtual try-on with diffusion models

    Hung Nguyen, Quang Qui-Vinh Nguyen, Khoi Nguyen, and Rang Nguyen. Swifttry: Fast and consistent video virtual try-on with diffusion models. InAAAI, volume 39, pages 6200–6208,

  13. [13]

    Pursuing temporal-consistent video virtual try-on via dynamic pose interaction

    Dong Li, Wenqi Zhong, Wei Yu, Yingwei Pan, Dingwen Zhang, Ting Yao, Junwei Han, and Tao Mei. Pursuing temporal-consistent video virtual try-on via dynamic pose interaction. In CVPR, pages 22648–22657, 2025. 2

  14. [14]

    Skyreels-a2: Compose anything in video diffusion transform- ers.arXiv preprint arXiv:2504.02436, 2025

    Zhengcong Fei, Debang Li, Di Qiu, Jiahua Wang, Yikun Dou, Rui Wang, Jingtao Xu, Mingyuan Fan, Guibin Chen, Yang Li, et al. Skyreels-a2: Compose anything in video diffusion transform- ers.arXiv preprint arXiv:2504.02436, 2025. 2, 3, 7, 8

  15. [15]

    Cinema: Coherent multi-subject video generation via mllm-based guidance.arXiv preprint arXiv:2503.10391, 2025

    Yufan Deng, Xun Guo, Yizhi Wang, Jacob Zhiyuan Fang, Angtian Wang, Shenghai Yuan, Yiding Yang, Bo Liu, Haibin Huang, and Chongyang Ma. Cinema: Coherent multi-subject video generation via mllm-based guidance.arXiv preprint arXiv:2503.10391, 2025. 2, 3, 4

  16. [16]

    Bindweave: Subject-consistent video generation via cross- modal integration.ICLR, 2026

    Zhaoyang Li, Dongjun Qian, Kai Su, Qishuai Diao, Xiangyang Xia, Chang Liu, Wenfei Yang, Tianzhu Zhang, and Zehuan Yuan. Bindweave: Subject-consistent video generation via cross- modal integration.ICLR, 2026. 2, 3, 4, 7, 8

  17. [17]

    Id-crafter: Vlm-grounded online rl for compositional multi-subject video generation.CVPR, 2026

    Panwang Pan, Jingjing Zhao, Yuchen Lin, Chenguo Lin, Chenxin Li, Hengyu Liu, Tingting Shen, and Yadong Mu. Id-crafter: Vlm-grounded online rl for compositional multi-subject video generation.CVPR, 2026. 2, 4

  18. [18]

    Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv preprint arXiv:2409.12191, 2024

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv preprint arXiv:2409.12191, 2024. 2

  19. [19]

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al. Qwen2. 5-vl technical report.arXiv preprint arXiv:2502.13923,

  20. [20]

    Image to video elements feature

    kling. Image to video elements feature. https://klingai.com/image-to-video/ multi-id/new/, 2024. 3, 7, 8, 16

  21. [21]

    Visualizing data using t-sne.JMLR, 9(86):2579– 2605, 2008

    Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne.JMLR, 9(86):2579– 2605, 2008. 3

  22. [22]

    Oriane Siméoni, Huy V V o, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, et al. Dinov3. arXiv preprint arXiv:2508.10104, 2025. 2, 6, 14

  23. [23]

    Opens2v-nexus: A detailed benchmark and million-scale dataset for subject-to-video generation

    Shenghai Yuan, Xianyi He, Yufan Deng, Yang Ye, Jinfa Huang, Bin Lin, Jiebo Luo, and Li Yuan. Opens2v-nexus: A detailed benchmark and million-scale dataset for subject-to-video generation. NeurIPS, 2025. 3, 8, 14, 15

  24. [24]

    Stand-in: A lightweight and plug-and-play identity control for video generation.CVPR, 2026

    Bowen Xue, Zheng-Peng Duan, Qixin Yan, Wenjing Wang, Hao Liu, Chun-Le Guo, Chongyi Li, Chen Li, and Jing Lyu. Stand-in: A lightweight and plug-and-play identity control for video generation.CVPR, 2026. 3

  25. [25]

    Concat-id: Towards universal identity-preserving video synthesis

    Yong Zhong, Zhuoyi Yang, Jiayan Teng, Xiaotao Gu, and Chongxuan Li. Concat-id: Towards universal identity-preserving video synthesis. InICCVW, pages 1906–1915, 2025. 3

  26. [26]

    Identity-preserving text-to-video generation by frequency decomposition

    Shenghai Yuan, Jinfa Huang, Xianyi He, Yunyang Ge, Yujun Shi, Liuhan Chen, Jiebo Luo, and Li Yuan. Identity-preserving text-to-video generation by frequency decomposition. InCVPR, pages 12978–12988, 2025. 3 12

  27. [27]

    Lynx: Towards high-fidelity personalized video generation

    Shen Sang, Tiancheng Zhi, Tianpei Gu, Jing Liu, and Linjie Luo. Lynx: Towards high-fidelity personalized video generation. InCVPR, 2026. 3

  28. [28]

    Kaleido: Open-sourced multi-subject reference video generation model.arXiv preprint arXiv:2510.18573, 2025

    Zhenxing Zhang, Jiayan Teng, Zhuoyi Yang, Tiankun Cao, Cheng Wang, Xiaotao Gu, Jie Tang, Dan Guo, and Meng Wang. Kaleido: Open-sourced multi-subject reference video generation model.arXiv preprint arXiv:2510.18573, 2025. 3, 4, 7, 8

  29. [29]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InICML, pages 8748–8763. PmLR, 2021. 3, 14

  30. [30]

    Scaling zero-shot reference-to-video generation

    Zijian Zhou, Shikun Liu, Haozhe Liu, Haonan Qiu, Zhaochong An, Weiming Ren, Zhiheng Liu, Xiaoke Huang, Kam Woh Ng, Tian Xie, et al. Scaling zero-shot reference-to-video generation. arXiv preprint arXiv:2512.06905, 2025. 4, 7, 8

  31. [31]

    Magref: Masked guidance for any-reference video generation.ICLR, 2026

    Yufan Deng, Xun Guo, Yuanyang Yin, Jacob Zhiyuan Fang, Yiding Yang, Yizhi Wang, Shenghai Yuan, Angtian Wang, Bo Liu, Haibin Huang, et al. Magref: Masked guidance for any-reference video generation.ICLR, 2026. 4, 7, 8

  32. [32]

    Polyvivid: Vivid multi-subject video generation with cross-modal interaction and enhancement

    Teng Hu, Zhentao Yu, Zhengguang Zhou, Jiangning Zhang, Yuan Zhou, Qinglin Lu, and Ran Yi. Polyvivid: Vivid multi-subject video generation with cross-modal interaction and enhancement. NeurIPS, 2025. 4

  33. [33]

    Hunyuancustom: A multimodal-driven architecture for customized video generation.arXiv preprint arXiv:2505.04512, 2025

    Teng Hu, Zhentao Yu, Zhengguang Zhou, Sen Liang, Yuan Zhou, Qin Lin, and Qinglin Lu. Hunyuancustom: A multimodal-driven architecture for customized video generation.arXiv preprint arXiv:2505.04512, 2025. 4

  34. [34]

    Vino: A unified visual generator with interleaved omnimodal context.arXiv preprint arXiv:2601.02358, 2026

    Junyi Chen, Tong He, Zhoujie Fu, Pengfei Wan, Kun Gai, and Weicai Ye. Vino: A unified visual generator with interleaved omnimodal context.arXiv preprint arXiv:2601.02358, 2026. 4, 7, 8, 16

  35. [35]

    Visual instruction tuning.NeurIPS, 36:34892–34916, 2023

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning.NeurIPS, 36:34892–34916, 2023. 4

  36. [36]

    Vace: All-in-one video creation and editing

    Zeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang, Yulin Pan, and Yu Liu. Vace: All-in-one video creation and editing. InICCV, pages 17191–17202, 2025. 4, 7, 8

  37. [37]

    Representation alignment for generation: Training diffusion transformers is easier than you think

    Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representation alignment for generation: Training diffusion transformers is easier than you think. InICLR, 2025. 4, 5, 6

  38. [38]

    Repa-e: Unlocking vae for end-to-end tuning of latent diffusion transformers

    Xingjian Leng, Jaskirat Singh, Yunzhong Hou, Zhenchang Xing, Saining Xie, and Liang Zheng. Repa-e: Unlocking vae for end-to-end tuning of latent diffusion transformers. InICCV, pages 18262–18272, 2025. 4

  39. [39]

    Ddt: Decoupled diffusion transformer

    Shuai Wang, Zhi Tian, Weilin Huang, and Limin Wang. Ddt: Decoupled diffusion transformer. CVPR, 2026. 4

  40. [40]

    Representation entanglement for generation: Training diffusion transformers is much easier than you think.NeurIPS, 2025

    Ge Wu, Shen Zhang, Ruijing Shi, Shanghua Gao, Zhenyuan Chen, Lei Wang, Zhaowei Chen, Hongcheng Gao, Yao Tang, Jian Yang, et al. Representation entanglement for generation: Training diffusion transformers is much easier than you think.NeurIPS, 2025. 4

  41. [41]

    Boosting generative image modeling via joint image-feature synthesis.NeurIPS,

    Theodoros Kouzelis, Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, and Nikos Komodakis. Boosting generative image modeling via joint image-feature synthesis.NeurIPS,

  42. [42]

    Reconstruction vs

    Jingfeng Yao, Bin Yang, and Xinggang Wang. Reconstruction vs. generation: Taming optimiza- tion dilemma in latent diffusion models. InCVPR, pages 15703–15712, 2025. 4

  43. [43]

    Unleashing the potential of large language models for text-to-image generation through autore- gressive representation alignment.AAAI, 2026

    Xing Xie, Jiawei Liu, Ziyue Lin, Huijie Fan, Zhi Han, Yandong Tang, and Liangqiong Qu. Unleashing the potential of large language models for text-to-image generation through autore- gressive representation alignment.AAAI, 2026. 4 13

  44. [44]

    Videorepa: Learning physics for video generation through relational alignment with foundation models.NeurIPS, 2025

    Xiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng, Xiangpeng Wan, Junchi Yan, and Yu Cheng. Videorepa: Learning physics for video generation through relational alignment with foundation models.NeurIPS, 2025. 4

  45. [45]

    Scaling rectified flow transform- ers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transform- ers for high-resolution image synthesis. InICML, 2024. 4

  46. [46]

    Dinov2: Learning robust visual features without supervision.TMLR, 2024

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision.TMLR, 2024. 4, 14

  47. [47]

    Exploring the limits of transfer learning with a unified text-to-text transformer.JMLR, 21(140):1–67, 2020

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.JMLR, 21(140):1–67, 2020. 5

  48. [48]

    Classifier-free diffusion guidance.arXiv preprint arXiv:2207.12598, 2022

    Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance.arXiv preprint arXiv:2207.12598, 2022. 5

  49. [49]

    Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features.arXiv preprint arXiv:2502.14786, 2025

    Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alab- dulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, et al. Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features.arXiv preprint arXiv:2502.14786, 2025. 6, 14

  50. [50]

    Pikascenes.https://pika.art/ingredients/, 2024

    Pika. Pikascenes.https://pika.art/ingredients/, 2024. 7, 8

  51. [51]

    Phantom-data: Towards a general subject- consistent video generation dataset.ICLR, 2026

    Zhuowei Chen, Bingchuan Li, Tianxiang Ma, Lijie Liu, Mingcong Liu, Yi Zhang, Gen Li, Xinghui Li, Siyu Zhou, Qian He, and Xinglong Wu. Phantom-data: Towards a general subject- consistent video generation dataset.ICLR, 2026. 8, 14

  52. [52]

    Decoupled weight decay regularization.ICLR, 2019

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization.ICLR, 2019. 8

  53. [53]

    Masked autoencoders are scalable vision learners

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. InCVPR, pages 16000–16009, 2022. 14

  54. [54]

    Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

    Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024. 14 A Appendix A.1 Additional Training Details During training, we randomly drop the prompt, the reference, or both, each with a probability of 10%, for CFG. We train our model in two sta...