Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

A 4D try-on proxy lets video virtual try-on free the camera from the source trajectory.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 12:04 UTC pith:CJKG4PLM

load-bearing objection Solid systems paper that cleanly defines free-camera video try-on and ships a working 4D-proxy + DiT pipeline; the geometry chain is fragile under extreme views, but the evidence and honesty are good enough to engage. the 3 major comments →

arxiv 2606.26092 v2 pith:CJKG4PLM submitted 2026-06-24 cs.CV

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

classification cs.CV
keywords video virtual try-oncamera control4D reconstruction3D Gaussian splattingSMPL-XDiffusion TransformerCaM-VVT
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Existing video virtual try-on systems can only replay the camera path of the input monocular video, so a shopper cannot freely inspect the side or back of a garment. This paper defines Camera-controllable Video Virtual Try-on (CaM-VVT) and introduces TryOnCrafter as the first end-to-end Diffusion Transformer built for it. The key move is to stop editing pixels and instead build an explicit Renderable 4D Try-on Proxy: a 2D try-on result is distilled into a clothed 3D Gaussian avatar, animated by metric-aligned SMPL-X motion, and placed inside a reconstructed background point cloud. Rendering that proxy under any desired camera path then anchors a video DiT so the final frames stay geometrically consistent with both the new viewpoint and the original human motion. If the approach holds, digital fashion can offer interactive 360-degree inspection, bullet-time freezes, and subject relocation without cascading separate try-on and camera-control models.

Core claim

TryOnCrafter shows that camera-controllable video virtual try-on becomes tractable when a monocular source is first converted into an editable 4D proxy that decouples the clothed human from the background; anchoring a video Diffusion Transformer on renderings of that proxy produces photorealistic try-on under unconstrained novel trajectories while preserving garment identity and subject–background structure better than two-stage baselines.

What carries the argument

The Renderable 4D Try-on Proxy: a canonical clothed 3DGS avatar distilled from a single high-quality 2D try-on keyframe, LBS-deformed by a confidence-aligned SMPL-X sequence, and composited into a background point cloud so that any novel camera path can be rendered as a dense structural prior for the Proxy-Anchored Video DiT.

Load-bearing premise

The monocular geometry, body model, and single-keyframe 3D avatar stay accurate enough under radical new viewpoints that the rendered proxy can still enforce correct structure; when they break, the generator inherits the errors.

What would settle it

On the paper’s CaM-VVTBench orbit and zoom sequences, measure whether garment identity, limb integrity, and background coherence collapse once extreme parallax or SMPL-X hand/pose errors appear in the proxy (as already illustrated in the paper’s own failure case); systematic failure would refute the claim that the proxy is a sufficient geometric anchor.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper defines Camera-controllable Video Virtual Try-on (CaM-VVT) and proposes TryOnCrafter, a unified DiT framework that builds a Renderable 4D Try-on Proxy by distilling a single-keyframe 2D try-on into a clothed 3DGS avatar, animating it with metric-aligned SMPL-X (Eq. 1–2), and compositing it with a monocular MVS background point cloud. A Proxy-Anchored Video DiT then conditions on the rendered proxy, a Cross-view Reference Adapter, and multimodal garment/text cues to synthesize try-on videos under novel camera trajectories. The method reports SOTA VFID/LPIPS on ViViD and superior VBench scores on an author-built CaM-VVTBench versus two-stage VVT+camera-control cascades, and demonstrates applications such as relocalization, bullet time, and 360° orbits.

Significance. If the claims hold, the work opens a practically relevant interactive fashion setting beyond fixed-trajectory VVT and shows that an explicit human–background 4D proxy can stabilize large DiT generation under novel cameras better than cascaded baselines. Strengths include a clear task definition, a reusable proxy that supports multiple applications, external ViViD gains, component ablations that move metrics in the expected direction (Table 2, Fig. 7), and useful efficiency/reusability profiling in the supplement. The contribution is engineering-heavy but timely for digital fashion and camera-controllable video generation.

major comments (3)
  1. [Sec. 3.1, Eq. 1–2; Limitation §6; Fig. 7] The central claim of “strict structural synchronization” and “physically plausible deformations” under unconstrained trajectories rests on monocular MVS + SAM2 + SMPL-X + single-keyframe 3DGS (Sec. 3.1, Eq. 1–2). Limitation §6 and Fig. 7 already show misaligned hands and structural glitches under extreme parallax/SMPL-X error. The paper needs a quantitative stress evaluation (e.g., controlled SMPL-X noise, incomplete body, or large-orbit disocclusion) reporting how often and how severely novel-view garment identity and limb integrity fail, rather than only qualitative failure cases and proxy-noise robustness in Supp. S7.
  2. [Sec. 4.2; Table 3; Supp. S2] CaM-VVTBench is author-constructed (Sec. 4.2; Supp. S2/S5), and Stage-2 supervision uses synthetic multi-view pairs from the same dual-reprojection/proxy pipeline used at inference. Table 3 therefore risks overstating generalization: the DiT may learn to correct proxy artifacts that appear in both train and test. At minimum, report an external or held-out real multi-view/novel-trajectory subset, or a cross-domain protocol that does not reuse the training proxy synthesis machinery, and clarify how much of the Overall Score gain is proxy fidelity versus generative fill-in.
  3. [Sec. 4.4; Table 3; Fig. 5] Baselines for CaM-VVT are only Magic-Tryon cascaded with TrajectoryCrafter/ReCamMaster (and TryOnCrafter† + same controllers). Given that the paper’s own proxy is the main differentiator, a stronger control would be: (i) the same 4D proxy rendered without the Proxy-Anchored DiT refinement, and (ii) other recent geometry-aware V2V controllers conditioned on the authors’ proxy renders. Without these, it remains unclear how much of Table 3’s gain is the unified DiT versus simply better human geometry than fragmented point clouds.
minor comments (5)
  1. [Abstract] Abstract/Intro: “existing paradigms remains” → “remain”; several other minor grammar issues appear throughout.
  2. [Sec. 3.1, Eq. (2)] Eq. (2) uses R and R′_j without fully restating how world-space rotation is composed with LBS bone rotations; a short clarification would help reproducibility.
  3. [Table 1] Table 1: SSIM is not best while LPIPS/VFID are; a one-sentence discussion of this trade-off would avoid the appearance of selective emphasis.
  4. [Fig. 4–5] Fig. 4/5 captions and layout make it hard to match trajectories to columns; labeling orbit/zoom primitives on the figure would improve readability.
  5. [Sec. 6; Supp. S6] Inference cost (~1360s at 720P on one A100) is acknowledged in §6/Supp. S6; stating this more prominently in the main experiments would set expectations for interactive use.

Circularity Check

0 steps flagged

No definitional or fitted-input circularity: TryOnCrafter is a conditioned generative pipeline evaluated on external ViViD metrics and third-party baselines; author-built CaM-VVTBench and proxy-synthesized Stage-2 pairs are standard engineering practice, not Eq.X≡Eq.Y.

full rationale

Walked the load-bearing chain (monocular MVS + SAM2 split + confidence-aware SMPL-X alignment Eq.1 + single-keyframe 2D try-on → canonical 3DGS avatar → LBS deformation → hybrid render → Proxy-Anchored DiT). None of the six circularity patterns apply. The proxy is an explicit geometric conditioner built from external foundation models; the DiT is trained to map noisy rendered priors plus CRA/semantic cues to photorealistic video—it is not forced by construction to equal the proxy (ablations: w/o V_render collapses VFID). Stage-2 dual-reprojection synthetic pairs (Supp. S2) reuse the proxy machinery for multi-view supervision, which is ordinary novel-view training design, not a fitted parameter renamed as prediction. CaM-VVTBench is author-defined for a newly framed task, but results are also reported on the external ViViD benchmark against third-party methods (CatV2TON, Magic-Tryon, DreamVVT, TrajectoryCrafter, ReCamMaster). No uniqueness theorem, self-citation load-bearing premise, or ansatz smuggled via overlapping authors. Upstream fragility of MVS/SMPL-X under radical views (Limitation §6) is a correctness/assumption risk, not circularity. Score 0; steps empty.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 4 invented entities

The central claim rests on a long chain of external vision models and engineering choices rather than a closed-form derivation. Free parameters are training/inference hyperparameters and alignment weights. Domain axioms assume monocular reconstruction and parametric humans are good enough anchors. Invented entities are the task label, the 4D proxy construct, CRA, and the author benchmark—useful engineering objects, not independently measured physical entities.

free parameters (5)
  • Stage-1/Stage-2 learning rates and step counts = 1e-5 / 1e-6; 24000 steps
    1e-5 then 1e-6, 24k steps per stage (Table S1); chosen for training stability, not derived.
  • Inference CFG scale and denoising steps = CFG=6.0, steps=20
    CFG 6.0 and 20 steps set generation strength/quality tradeoff by hand.
  • Confidence weights w_i in point-to-surface alignment
    Multi-modal weights from MVS confidence and boundary proximity in Eq.1 prioritize torso over peripherals; functional form is design choice.
  • Keyframe viewing-score selection
    t* maximizes torso-forward alignment with camera axis (Eq.S1); heuristic for canonical avatar quality.
  • Resolution / clip-length training schedule = 256–1024; 81/49 frames
    256–1024 progressive resolutions and 81→49 frame clips (Table S1) are engineering knobs affecting reported fidelity.
axioms (5)
  • domain assumption Monocular MVS depth/cameras plus SAM2 human/background split yield a metric world-space scene usable for novel-view rendering.
    Invoked throughout §3.1 Scene Reconstruction; failures under parallax are only discussed in Limitations.
  • domain assumption SMPL-X sequences from GVHMR adequately capture non-rigid human dynamics for LBS deformation of a clothed 3DGS avatar.
    §3.1 Deformation and Rendering; hand/pose errors are a stated failure mode (Fig. 7).
  • domain assumption A single high-quality 2D try-on keyframe can be distilled into a view-consistent canonical 3DGS avatar that generalizes to unobserved angles.
    §3.1 Canonical 3DGS-based Avatar Generation; load-bearing for viewpoint-agnostic texture hallucination.
  • ad hoc to paper Pixel-aligned rendered proxy video is a sufficient structural prior for a pretrained I2V DiT to produce photorealistic, trajectory-faithful try-on under novel cameras.
    Core design of §3.2 Proxy-Anchored generation; supported by ablations but not independently proven outside this pipeline.
  • ad hoc to paper Synthetic multi-view pairs from dual-reprojection and randomized masking are valid supervision for real novel-trajectory try-on.
    Supplementary §S2 Stage-2 dataset construction; underpins CaM-VVT fine-tuning claims.
invented entities (4)
  • Camera-controllable Video Virtual Try-on (CaM-VVT) no independent evidence
    purpose: Task definition requiring free camera trajectories plus garment-consistent try-on and subject–background sync.
    Named as a 'pioneering research frontier' in Abstract/Intro; evaluation depends on author protocol.
  • Renderable 4D Try-on Proxy no independent evidence
    purpose: Decoupled clothed 3DGS avatar + SMPL-X motion + background point cloud used as geometric anchor and editable scene.
    Central construct of §3.1; no existence claim beyond this engineering representation.
  • Proxy-Anchored Video DiT with Cross-view Reference Adapter (CRA) no independent evidence
    purpose: Condition generative video synthesis on rendered priors and source features for identity/background fidelity.
    §3.2 architecture; CRA residual is a paper-specific module (Eq.3–4).
  • CaM-VVTBench no independent evidence
    purpose: Author training (~60K) and test (96) protocol with six camera primitives and VBench metrics for CaM-VVT.
    §4.2; primary evidence for camera-control claims is on this self-built set.

pith-pipeline@v1.1.0-grok45 · 23136 in / 4225 out tokens · 42712 ms · 2026-07-12T12:04:19.671879+00:00 · methodology

0 comments
read the original abstract

While Video Virtual Try-on (VVT) has achieved remarkable progress in synthesizing realistic garment overlays on dynamic subjects, existing paradigms remains fundamentally constrained by a passive dependency on source camera trajectories, failing to accommodate the requisite interactive freedom for omnidirectional viewpoint exploration. To address this limitation, we define a pioneering research frontier: Camera-controllable Video Virtual Try-on (CaM-VVT). Unlike conventional VVT, CaM-VVT not only necessitates viewpoint-agnostic texture hallucination but also strict structural synchronization between non-rigid human dynamics and background contexts under arbitrary, unconstrained camera movements. To tackle these challenges, we present TryOnCrafter, the first unified DiT-based framework specifically architected for the CaM-VVT task. Departing from implicit pixel-space manipulation, we introduce a Renderable 4D Try-on Proxy that explicitly decouples the human subject from the environment. This is achieved by distilling high-fidelity 2D try-on priors into a clothed 3DGS-based avatar, which is subsequently animated via SMPL-X sequences and metric-aligned into a reconstructed background point cloud. This proxy establishes a robust structural foundation with superior texture density and motion integrity. Our Proxy-Anchored Video DiT leverages this robust structural foundation as a primary geometric anchor, ensuring that the synthesized photorealistic videos are strictly constrained by prescribed trajectories and physically plausible deformations. Benefiting from the inherent editability of the 4D proxy, TryOnCrafter facilitates diverse downstream applications, including human relocalization, ``bullet time'' effects, and $360$-degree orbital viewing.

Figures

Figures reproduced from arXiv: 2606.26092 by Bo Zheng, Hao Sun, Hao Yan, Jinsong Lan, Juan Cao, Mengting Chen, Quanjian Song, Sheng Tang, Xiaoyong Zhu, Yu Li.

Figure 1
Figure 1. Figure 1: Examples synthesized by TryOnCrafter. We introduce a Renderable 4D Try-on Proxy (middle) as a geometric anchor to guide the Video Diffusion Transformer. This explicit 4D representation enables photorealistic try-on across unconstrained, novel camera trajectories (bottom) not present in the source video (top). Abstract. While Video Virtual Try-on (VVT) has achieved remark￾able progress in synthesizing reali… view at source ↗
Figure 2
Figure 2. Figure 2: Overview of TryOnCrafter. Top: 4D Try-on Proxy Construction. A metric￾aligned scene in world space is established via anchor-based alignment of dynamic point clouds and SMPL-X sequences. Integrating a canonical 3DGS-based avatar distilled from reference image, the resulting 4D Try-on Proxy supports interactive editing and rendering across arbitrary camera trajectories. Bottom: Proxy-Anchored Try-on Video G… view at source ↗
Figure 3
Figure 3. Figure 3: Details of 4D Try-on Proxy Construction and CRA Module. (a) A similarity transformation {Rt, tt, st} maps the SMPL-X sequence from camera space to the world space, utilizing the human point cloud as a geometric anchor. (b) The canonical 3DGS￾based avatar is dynamically warped via LBS driven by the aligned SMPL-X sequence, preserving structural integrity across the motion. (c) The CRA module facilitates int… view at source ↗
Figure 4
Figure 4. Figure 4: Qualitative comparison on the video try-on benchmark. The left two columns show results on ViViD-S [6], and the right column shows results on the in-the-wild test set. \mathbf {F}_{i}' = \mathrm {Attn}(\mathbf {F}_i, Q_i, K_i, V_i, O_i) + \mathbf {O}_l', (4) where Attn(·) denote the standard attention operation. By complementing the structural guidance of the 4D proxy with these fine-grained source feature… view at source ↗
Figure 5
Figure 5. Figure 5: Qualitative comparison results on CaM-VVTBench. Left trajectory: orbit right and zoom in, Right trajectory: orbit left. † denotes restricted to input trajectories. ating on the ViViD-S test set (180 samples). Following prior works [6,20,54], we adopt VFID with I3D [3] (VFIDI ) and ResNext [46] (VFIDR) under both paired and unpaired settings. For paired evaluation, we additionally report SSIM and LPIPS to m… view at source ↗
Figure 6
Figure 6. Figure 6: Versatile Applications of TryOnCrafter. Leveraging the decoupled 4D proxy, our model enables controllable synthesis: (a) Human Relocalization: Translating the motion sequence within S w while maintaining scene-geometry consistency. (b) 360- degree Orbital Viewing: Synthesizing full orbital trajectories with robust structural integrity in unobserved viewpoints. (c) Bullet Time: Rendering a frozen temporal m… view at source ↗
Figure 7
Figure 7. Figure 7: Left: Ablation studies of main component of TryOnCrafter. Right: Failure case caused by perspective-induced parallax and inaccuracies of SMPL-X estimation. observed angles. By synthesizing realistic, trajectory-aligned cloth deformations, our Proxy-Anchored Video DiT demonstrates superior structural robustness and high-fidelity appearance across unconstrained trajectories. Quantitative Comparison. As shown… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis

    cs.CV 2026-07 conditional novelty 6.0

    ExpertVerse is a new benchmark and training pipeline for knowledge-intensive image generation, and its KnowThinker model with BPPO reports state-of-the-art results on reasoning-editing tests.

Reference graph

Works this paper leans on

54 extracted references · 22 linked inside Pith · cited by 1 Pith paper

  1. [1]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision

    Bai,J., Xia, M., Fu, X.,Wang, X.,Mu, L., Cao, J.,Liu, Z., Hu, H.,Bai, X., Wan, P., et al.: Recammaster: Camera-controlled generative rendering from a single video. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 14834–14844 (2025)

  2. [2]

    In: Proceedings of the SIGGRAPH Asia 2025 Conference Papers

    Cao, C., Zhou, J., Li, S., Liang, J., Yu, C., Wang, F., Xue, X., Fu, Y.: Uni3c: Unifying precisely 3d-enhanced camera and human motion controls for video gen- eration. In: Proceedings of the SIGGRAPH Asia 2025 Conference Papers. pp. 1–12 (2025)

  3. [3]

    In: proceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 6299–6308 (2017)

  4. [4]

    In: European Conference on Computer Vision

    Choi, Y., Kwak, S., Lee, K., Choi, H., Shin, J.: Improving diffusion models for authentic virtual try-on in the wild. In: European Conference on Computer Vision. pp. 206–235. Springer (2024)

  5. [5]

    arXiv preprint arXiv:2501.11325 (2025)

    Chong, Z., Zhang, W., Zhang, S., Zheng, J., Dong, X., Li, H., Wu, Y., Jiang, D., Liang, X.: Catv2ton: Taming diffusion transformers for vision-based virtual try-on with temporal concatenation. arXiv preprint arXiv:2501.11325 (2025)

  6. [6]

    arXiv preprint arXiv:2405.11794 (2024)

    Fang, Z., Zhai, W., Su, A., Song, H., Zhu, K., Wang, M., Chen, Y., Liu, Z., Cao, Y., Zha, Z.J.: Vivid: Video virtual try-on using diffusion models. arXiv preprint arXiv:2405.11794 (2024)

  7. [7]

    arXiv preprint arXiv:2506.02528 (2025)

    Gong, Y., Song, Y., Li, Y., Li, C., Zhang, Y.: Relationadapter: Learning and trans- ferring visual relation with diffusion transformers. arXiv preprint arXiv:2506.02528 (2025)

  8. [8]

    Guo,H.,Zeng,B.,Song,Y.,Zhang,W.,Zhang,C.,Liu,J.:Any2anytryon:Leverag- ingadaptivepositionembeddingsforversatilevirtualclothingtasks.arXivpreprint arXiv:2501.15891 (2025) 16 Sun et al

  9. [9]

    He, H., Xu, Y., Guo, Y., Wetzstein, G., Dai, B., Li, H., Yang, C.: Cameractrl: En- ablingcameracontrolfortext-to-videogeneration.arXivpreprintarXiv:2404.02101 (2024)

  10. [10]

    In: European Con- ference on Computer Vision

    He, Z., Chen, P., Wang, G., Li, G., Torr, P.H., Lin, L.: Wildvidfit: Video virtual try-on in the wild via image-based controlled diffusion models. In: European Con- ference on Computer Vision. pp. 123–139. Springer (2024)

  11. [11]

    Hou,C.,Chen,Z.:Training-freecameracontrolforvideogeneration.arXivpreprint arXiv:2406.10126 (2024)

  12. [12]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision

    Huang, S., Song, Y., Zhang, Y., Guo, H., Wang, X., Liu, J.: Arteditor: Learning customized instructional image editor from few-shot examples. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 17651–17662 (2025)

  13. [13]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., et al.: Vbench: Comprehensive benchmark suite for video gener- ative models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 21807–21818 (2024)

  14. [14]

    arXiv preprint arXiv:2410.21276 (2024)

    Hurst, A., Lerer, A., Goucher, A.P., Perelman, A., Ramesh, A., Clark, A., Os- trow, A., Welihinda, A., Hayes, A., Radford, A., et al.: Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024)

  15. [15]

    In: Proceedings of the SIGGRAPH Asia 2025 Conference Papers

    Jiang, T., Ho, H.I., Kaufmann, M., Song, J.: Prioravatar: Efficient and robust avatar creation from monocular video using learned priors. In: Proceedings of the SIGGRAPH Asia 2025 Conference Papers. pp. 1–10 (2025)

  16. [16]

    In: SIG- GRAPH Asia 2024 Conference Papers

    Karras, J., Li, Y., Liu, N., Zhu, L., Yoo, I., Lugmayr, A., Lee, C., Kemelmacher- Shlizerman, I.: Fashion-vdm: Video diffusion model for virtual try-on. In: SIG- GRAPH Asia 2024 Conference Papers. pp. 1–11 (2024)

  17. [17]

    ACM Trans

    Kerbl, B., Kopanas, G., Leimkühler, T., Drettakis, G.: 3d gaussian splatting for real-time radiance field rendering. ACM Trans. Graph.42(4), 139–1 (2023)

  18. [18]

    In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

    Kim, J., Gu, G., Park, M., Park, S., Choo, J.: Stableviton: Learning semantic correspondence with latent diffusion model for virtual try-on. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 8176– 8185 (2024)

  19. [19]

    arXiv preprint arXiv:2412.09262 (2024)

    Li, C., Zhang, C., Xu, W., Lin, J., Xie, J., Feng, W., Peng, B., Chen, C., Xing, W.: Latentsync: Taming audio-conditioned latent diffusion models for lip sync with syncnet supervision. arXiv preprint arXiv:2412.09262 (2024)

  20. [20]

    arXiv preprint arXiv:2505.21325 (2025)

    Li, G., Zheng, S., Zhang, H., Chen, J., Luan, J., Ou, B., Zhao, L., Li, B., Jiang, P.T.: Magictryon: Harnessing diffusion transformer for garment-preserving video virtual try-on. arXiv preprint arXiv:2505.21325 (2025)

  21. [21]

    Advances in neural information processing systems37, 75125–75151 (2024)

    Li, X., Lai, Z., Xu, L., Qu, Y., Cao, L., Zhang, S., Dai, B., Ji, R.: Director3d: Real- world camera trajectory and 3d scene generation from text. Advances in neural information processing systems37, 75125–75151 (2024)

  22. [22]

    IEEE Signal Processing Letters (2024)

    Lin, J., Wu, Y., Wang, Z., Liu, X., Guo, Y.: Pair-id: A dual modal framework for identity preserving image generation. IEEE Signal Processing Letters (2024)

  23. [23]

    arXiv preprint arXiv:2411.14208 (2024)

    Liu, K., Shao, L., Lu, S.: Novel view extrapolation with video diffusion priors. arXiv preprint arXiv:2411.14208 (2024)

  24. [24]

    In: The Thirteenth International Conference on Learning Representations (2024)

    Meng, Y., Zhu, Z., Hui, L., Hou, J.: Nvs-solver: Video diffusion model as zero-shot novel view synthesizer. In: The Thirteenth International Conference on Learning Representations (2024)

  25. [25]

    In: Proceedings of the AAAI Conference on Artificial Intelligence

    Nguyen, H., Nguyen, Q.Q.V., Nguyen, K., Nguyen, R.: Swifttry: Fast and con- sistent video virtual try-on with diffusion models. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 39, pp. 6200–6208 (2025) TryOnCrafter 17

  26. [26]

    arXiv preprint arXiv:2503.10625 (2025)

    Qiu, L., Gu, X., Li, P., Zuo, Q., Shen, W., Zhang, J., Qiu, K., Yuan, W., Chen, G., Dong, Z., et al.: Lhm: Large animatable human reconstruction model from a single image in seconds. arXiv preprint arXiv:2503.10625 (2025)

  27. [27]

    In: Proceedings of the Computer Vision and Pattern Recognition Conference

    Qiu, L., Zhu, S., Zuo, Q., Gu, X., Dong, Y., Zhang, J., Xu, C., Li, Z., Yuan, W., Bo, L., et al.: Anigs: Animatable gaussian avatar from a single image with inconsistent gaussian reconstruction. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 21148–21158 (2025)

  28. [28]

    arXiv preprint arXiv:2408.00714 (2024)

    Ravi, N., Gabeur, V., Hu, Y.T., Hu, R., Ryali, C., Ma, T., Khedr, H., Rädle, R., Rolland, C., Gustafson, L., et al.: Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714 (2024)

  29. [29]

    In: SIGGRAPH Asia 2024 Conference Papers

    Shen,Z.,Pi,H.,Xia,Y.,Cen,Z.,Peng,S.,Hu,Z.,Bao,H.,Hu,R.,Zhou,X.:World- grounded human motion recovery via gravity-view coordinates. In: SIGGRAPH Asia 2024 Conference Papers. pp. 1–11 (2024)

  30. [30]

    IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)

    Song, Q., Lin, M., Zhan, W., Yan, S., Cao, L., Ji, R.: Univst: A unified frame- work for training-free localized video style transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)

  31. [31]

    arXiv preprint arXiv:2503.06508 (2025)

    Song, Q., Lin, Z., Zeng, Z., Zhang, Z., Cao, L., Ji, R.: Lightmotion: A light and tuning-free method for simulating camera motion in video generation. arXiv preprint arXiv:2503.06508 (2025)

  32. [32]

    arXiv preprint arXiv:2605.15824 (2026)

    Song, Q., Shen, Y., Chen, M., Sun, H., Lan, J., Zhu, X., Zheng, B., Cao, L.: Fashionchameleon: Towards real-time and interactive human-garment video cus- tomization. arXiv preprint arXiv:2605.15824 (2026)

  33. [33]

    arXiv preprint arXiv:2511.22098 (2025)

    Song, Q., Song, Y., Peng, K., Gao, Y., Shou, M.Z.: Worldwander: Bridging ego- centric and exocentric worlds in video generation. arXiv preprint arXiv:2511.22098 (2025)

  34. [34]

    arXiv preprint arXiv:2510.22994 (2025)

    Song, Q., Zhou, D., Lin, J., Shen, F., Wang, J., Hu, X., Chen, C., Heng, P.A.: Scenedecorator: Towards scene-oriented story generation with scene planning and scene consistency. arXiv preprint arXiv:2510.22994 (2025)

  35. [35]

    arXiv preprint arXiv:2502.01572 (2025)

    Song, Y., Liu, C., Shou, M.Z.: Makeanything: Harnessing diffusion transformers for multi-domain procedural sequence generation. arXiv preprint arXiv:2502.01572 (2025)

  36. [36]

    arXiv preprint arXiv:2503.20314 (2025)

    Wan, T., Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.W., Chen, D., Yu, F., Zhao, H., Yang, J., et al.: Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314 (2025)

  37. [37]

    In: Proceedings of the Computer Vision and Pattern Recognition Conference

    Wang, J., Chen, M., Karaev, N., Vedaldi, A., Rupprecht, C., Novotny, D.: Vggt: Visual geometry grounded transformer. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 5294–5306 (2025)

  38. [38]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Wang, S., Leroy, V., Cabon, Y., Chidlovskii, B., Revaud, J.: Dust3r: Geometric 3d vision made easy. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 20697–20709 (2024)

  39. [39]

    arXiv preprint arXiv:2507.13347 (2025)

    Wang, Y., Zhou, J., Zhu, H., Chang, W., Zhou, Y., Li, Z., Chen, J., Pang, J., Shen, C., He, T.:π 3: Permutation-equivariant visual geometry learning. arXiv preprint arXiv:2507.13347 (2025)

  40. [40]

    In: Proceedings of the 32nd ACM Interna- tional Conference on Multimedia

    Wang, Y., Dai, W., Chan, L., Zhou, H., Zhang, A., Liu, S.: Gpd-vvto: Preserving garment details in video virtual try-on. In: Proceedings of the 32nd ACM Interna- tional Conference on Multimedia. pp. 7133–7142 (2024)

  41. [41]

    In: ACM SIGGRAPH 2024 Conference Papers

    Wang, Z., Yuan, Z., Wang, X., Li, Y., Chen, T., Xia, M., Luo, P., Shan, Y.: Mo- tionctrl: A unified and flexible motion controller for video generation. In: ACM SIGGRAPH 2024 Conference Papers. pp. 1–11 (2024)

  42. [42]

    arXiv preprint arXiv:2508.02324 (2025) 18 Sun et al

    Wu, C., Li, J., Zhou, J., Lin, J., Gao, K., Yan, K., Yin, S.m., Bai, S., Xu, X., Chen, Y., et al.: Qwen-image technical report. arXiv preprint arXiv:2508.02324 (2025) 18 Sun et al

  43. [43]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Wu, J.Z., Zhang, Y., Turki, H., Ren, X., Gao, J., Shou, M.Z., Fidler, S., Gojcic, Z., Ling, H.: Difix3d+: Improving 3d reconstructions with single-step diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 26024–26035 (2025)

  44. [44]

    In: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Wu, R., Gao, R., Poole, B., Trevithick, A., Zheng, C., Barron, J.T., Holynski, A.: Cat4d: Create anything in 4d with multi-view video diffusion models. In: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 26057–26068 (2025)

  45. [45]

    arXiv preprint arXiv:2411.19324 (2024)

    Xiao, Z., Ouyang, W., Zhou, Y., Yang, S., Yang, L., Si, J., Pan, X.: Trajectory attention for fine-grained video motion control. arXiv preprint arXiv:2411.19324 (2024)

  46. [46]

    In: Proceedings of the IEEE conference on computer vision and pattern recognition

    Xie,S.,Girshick,R.,Dollár,P.,Tu,Z.,He,K.:Aggregatedresidualtransformations for deep neural networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 1492–1500 (2017)

  47. [47]

    In: Proceedings of the AAAI Conference on Artificial Intelligence

    Xu, Y., Gu, T., Chen, W., Chen, A.: Ootdiffusion: Outfitting fusion based latent diffusion for controllable virtual try-on. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 39, pp. 8996–9004 (2025)

  48. [48]

    In: Proceedings of the 32nd ACM International Conference on Multimedia

    Xu, Z., Chen, M., Wang, Z., Xing, L., Zhai, Z., Sang, N., Lan, J., Xiao, S., Gao, C.: Tunnel try-on: Excavating spatial-temporal tunnels for high-quality virtual try-on in videos. In: Proceedings of the 32nd ACM International Conference on Multimedia. pp. 3199–3208 (2024)

  49. [49]

    In: Proceedings of the IEEE/CVF international conference on computer vision

    Yu, M., Hu, W., Xing, J., Shan, Y.: Trajectorycrafter: Redirecting camera trajec- tory for monocular videos via diffusion models. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 100–111 (2025)

  50. [50]

    arXiv preprint arXiv:2409.02048 (2024)

    Yu, W., Xing, J., Yuan, L., Hu, W., Li, X., Huang, Z., Gao, X., Wong, T.T., Shan, Y., Tian, Y.: Viewcrafter: Taming video diffusion models for high-fidelity novel view synthesis. arXiv preprint arXiv:2409.02048 (2024)

  51. [51]

    arXiv preprint arXiv:2410.03825 (2024)

    Zhang, J., Herrmann, C., Hur, J., Jampani, V., Darrell, T., Cole, F., Sun, D., Yang, M.H.: Monst3r: A simple approach for estimating geometry in the presence of motion. arXiv preprint arXiv:2410.03825 (2024)

  52. [52]

    arXiv preprint arXiv:2503.07027 (2025)

    Zhang, Y., Yuan, Y., Song, Y., Wang, H., Liu, J.: Easycontrol: Adding efficient and flexible control for diffusion transformer. arXiv preprint arXiv:2503.07027 (2025)

  53. [53]

    In: Proceedings of the AAAI Conference on Artificial Intelligence

    Zhang, Y., Zhang, Q., Song, Y., Zhang, J., Tang, H., Liu, J.: Stable-hair: Real- world hair transfer via diffusion model. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 39, pp. 10348–10356 (2025)

  54. [54]

    Zuo, T., Huang, Z., Ning, S., Lin, E., Liang, C., Zheng, Z., Jiang, J., Zhang, Y., Gao, M., Dong, X.: Dreamvvt: Mastering realistic video virtual try-on in the wild via a stage-wise diffusion transformer framework. arXiv preprint arXiv:2508.02807 (2025) TryOnCrafter 19 Supplementary Materials S1 Overview In the supplementary materials, we provide addition...