Pith. sign in

REVIEW 19 cited by

ViViD: Video Virtual Try-on using Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.11794 v2 pith:HYUVHGAL submitted 2024-05-20 cs.CV

classification cs.CV
keywords videotry-onvirtualclothingdiffusionmodelvividdataset
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Video virtual try-on aims to transfer a clothing item onto the video of a target person. Directly applying the technique of image-based try-on to the video domain in a frame-wise manner will cause temporal-inconsistent outcomes while previous video-based try-on solutions can only generate low visual quality and blurring results. In this work, we present ViViD, a novel framework employing powerful diffusion models to tackle the task of video virtual try-on. Specifically, we design the Garment Encoder to extract fine-grained clothing semantic features, guiding the model to capture garment details and inject them into the target video through the proposed attention feature fusion mechanism. To ensure spatial-temporal consistency, we introduce a lightweight Pose Encoder to encode pose signals, enabling the model to learn the interactions between clothing and human posture and insert hierarchical Temporal Modules into the text-to-image stable diffusion model for more coherent and lifelike video synthesis. Furthermore, we collect a new dataset, which is the largest, with the most diverse types of garments and the highest resolution for the task of video virtual try-on to date. Extensive experiments demonstrate that our approach is able to yield satisfactory video try-on results. The dataset, codes, and weights will be publicly available. Project page: https://becauseimbatman0.github.io/ViViD.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    TryOnCrafter is the first DiT-based framework for camera-controllable video virtual try-on via a renderable 4D try-on proxy distilled from 2D priors into 3DGS avatar animated with SMPL-X.

  2. UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on

    cs.CV 2026-08 conditional novelty 6.0 of 10

    UniVVT reports state-of-the-art video and image virtual try-on by conditioning a diffusion video generator on task tokens from a multimodal language model, with no masks, poses, or warping at inference.

  3. Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Vid-CamEdit re-synthesizes monocular videos along user-defined camera paths by conditioning a video diffusion model on 2D flows derived from estimated 3D geometry, without training on multi-view video data.

  4. Low-Barrier Dataset Collection with Real Human Body for Interactive Per-Garment Virtual Try-On

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A per-garment virtual try-on pipeline that trains a GAN from a two-minute real-human video capture and uses a hybrid pose-plus-DensePose input to synthesize the garment with accurate alignment.

  5. Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DPIDM, a diffusion model with pose-aware spatial and temporal attention plus a temporal attention loss, reports state-of-the-art video virtual try-on and cuts VFID on VVT from 1.280 to 0.506.

  6. 3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion Models

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A diffusion video try-on framework that uses animated textured 3D meshes as frame-level guidance, plus a new high-resolution benchmark, achieves stronger temporal consistency and garment fidelity than two released baselines.

  7. VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A zero-shot diffusion framework that inserts a reference object into a video with high-fidelity appearance preservation and precise key-point trajectory motion control.

  8. SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SwiftTry makes diffusion-based video virtual try-on faster and more consistent by shifting non-overlapping video chunks during sampling and caching features across denoising steps.

  9. Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    UES adds a self-supervised video condition to text-to-video diffusion models, enabling them to edit videos from delta prompts without paired supervision.

  10. DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A two-stage diffusion pipeline that enriches 2D pose guidance with generated depth and normal maps to achieve state-of-the-art human image animation.

  11. Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image -- Technical Preview

    cs.CV 2025-09 conditional novelty 5.0 of 10

    VFR generates minute-long virtual try-on videos by auto-regressively chaining short diffusion-generated segments that are kept consistent with a 360-degree anchor video of the user.

  12. Real-Time Per-Garment Virtual Try-On with Temporal Consistency for Loose-Fitting Garments

    cs.GR 2025-06 conditional novelty 5.0 of 10

    A per-garment virtual try-on method for loose-fitting garments uses a garment-invariant pose representation and a recurrent ConvLSTM synthesis network to achieve temporally smoother try-on video at about 10 fps.

  13. ChronoTailor: Harnessing Attention Guidance for Fine-Grained Video Virtual Try-On

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ChronoTailor combines region-aware attention guidance, temporal feature fusion, and multi-scale garment-pose alignment to produce state-of-the-art video virtual try-on results, and contributes the StyleDress dataset.

  14. OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    OmniV2V is one diffusion-transformer model that performs eight video generation and editing tasks by combining mask, pose, image, and text-instruction conditions.

  15. RealVVT: Towards Photorealistic Video Virtual Try-on via Spatio-Temporal Consistency

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A diffusion-based framework improves spatial and temporal consistency of clothes in virtual try-on videos, reporting the best FID/KID and several video metrics on four public datasets.

  16. 1-2-1: Renaissance of Single-Network Paradigm for Virtual Try-On

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A single-network virtual try-on model with modality-specific normalization and shared attention matches or beats dual-network reference-based models on image and video try-on benchmarks.

  17. PEMF-VTO: Point-Enhanced Video Virtual Try-on via Mask-free Paradigm

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A mask-free video virtual try-on model that uses sparse point correspondences between garment and frames, plus frame-to-frame tracking, to improve garment transfer and temporal coherence.

  18. Human Motion Video Generation: A Survey

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A comprehensive survey with a five-phase pipeline model for human motion video generation, covering over 200 papers and adding a new benchmark comparison of nine pose-guided methods.

  19. TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    TalkFashion, a text-driven virtual try-on assistant, reports better semantic consistency and visual quality than four baselines on VITON-HD by combining an LLM router, catalog matching, and automatic mask generation.

Pith tools