Pith. sign in

REVIEW 29 cited by

HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.04512 v2 pith:SCPKYW2D submitted 2025-05-07 cs.CV

HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

classification cs.CV
keywords generationvideocustomizedhunyuancustommoduleconsistencymulti-modalacross
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities. In this paper, we propose HunyuanCustom, a multi-modal customized video generation framework that emphasizes subject consistency while supporting image, audio, video, and text conditions. Built upon HunyuanVideo, our model first addresses the image-text conditioned generation task by introducing a text-image fusion module based on LLaVA for enhanced multi-modal understanding, along with an image ID enhancement module that leverages temporal concatenation to reinforce identity features across frames. To enable audio- and video-conditioned generation, we further propose modality-specific condition injection mechanisms: an AudioNet module that achieves hierarchical alignment via spatial cross-attention, and a video-driven injection module that integrates latent-compressed conditional video through a patchify-based feature-alignment network. Extensive experiments on single- and multi-subject scenarios demonstrate that HunyuanCustom significantly outperforms state-of-the-art open- and closed-source methods in terms of ID consistency, realism, and text-video alignment. Moreover, we validate its robustness across downstream tasks, including audio and video-driven customized video generation. Our results highlight the effectiveness of multi-modal conditioning and identity-preserving strategies in advancing controllable video generation. All the code and models are available at https://hunyuancustom.github.io.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 29 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation

    cs.CV 2026-05 unverdicted novelty 7.0

    UniCustom fuses ViT and VAE features before VLM encoding and uses two-stage training plus slot-wise regularization to improve subject consistency in multi-reference diffusion-based image generation.

  2. Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding

    cs.CV 2026-04 unverdicted novelty 7.0

    MTSS replaces monolithic video captions with factorized streams and relational grounding, yielding reported gains in understanding benchmarks and generation consistency.

  3. Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation

    cs.CV 2026-04 unverdicted novelty 7.0

    Prompt Relay is an inference-time plug-and-play method that penalizes cross-attention to enforce temporal prompt alignment and reduce semantic entanglement in multi-event video generation.

  4. UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

    cs.CV 2026-08 conditional novelty 6.0

    A visual proxy that renders human motion under the driving camera and overlays camera trajectory markers lets a video diffusion model control both body motion and camera movement from a single visual conditioning space.

  5. HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

    cs.CV 2026-07 conditional novelty 6.0

    HOMIE unifies inter- and intra-subject video personalization by injecting MLLM-derived relational features into DiT self-attention (GMG) and tagging tokens with modality/reference embeddings (MRE), reporting SOTA on a...

  6. Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment

    cs.CV 2026-07 conditional novelty 6.0

    Aura combines VLM meta-queries, T5-teacher alignment, subject-aware RoPE shifts, memory tokens, and a large AIGC-curated dataset to claim SOTA multi-element subject-to-video generation under OpenS2V-Eval Total score.

  7. Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models

    cs.CV 2026-07 unverdicted novelty 6.0

    Ink3D decouples geometry from texture by generating dense orbit videos with a conditional video model and baking them via a neural optimizer to produce complex 3D textures.

  8. ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

    cs.CV 2026-06 unverdicted novelty 6.0

    ARGUS converts MLLM-selected identity evidence into a synchronized 3x3 mosaic injected as negative-time memory in a diffusion model, plus supporting training techniques, to achieve SOTA subject preservation on human v...

  9. MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data

    cs.CV 2026-06 unverdicted novelty 6.0

    MetaWorld scales multi-agent video world models from single-view videos using monocular decomposition into ego-motion and trajectories, subject-aware generation, and cross-attention alignment for consistency.

  10. UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    A unified visual conditioning approach fuses semantic and appearance features before VLM processing, with two-stage training and slot-wise regularization, to improve consistency in multi-reference image generation.

  11. FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    FaithfulFaces introduces a pose-faithful identity aligner with a shared dictionary and invariance constraint to maintain facial identity in text-to-video generation under large pose changes and occlusions.

  12. MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    MMControl adds multi-modal controls for identity, timbre, pose, and layout to unified audio-video diffusion models via dual-stream injection and adjustable guidance scaling.

  13. MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation

    cs.CV 2026-04 conditional novelty 6.0

    Translation function vectors extracted from a single English→X direction transfer across unseen target languages in three multilingual LLMs, extending language-agnosticity findings to task-level representations.

  14. TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    TS-Attn dynamically separates and rearranges attention in existing text-to-video models to improve temporal consistency and prompt adherence for videos with multiple sequential actions.

  15. OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    OmniShow unifies text, image, audio, and pose conditions into an end-to-end model for high-quality human-object interaction video generation and introduces the HOIVG-Bench benchmark, claiming state-of-the-art results.

  16. Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation

    cs.CV 2026-04 conditional novelty 6.0

    SideInfo-RoPE encodes reference-identity agreement as an extra rotary axis, disambiguating similar characters in multi-reference multi-shot video generation while keeping full semantic attention.

  17. RefAlign: Representation Alignment for Reference-to-Video Generation

    cs.CV 2026-03 conditional novelty 6.0

    Explicit training-time alignment of DiT reference features to a VFM (with pull/push loss) raises OpenS2V-Eval TotalScore over prior R2V methods with no inference cost.

  18. MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model

    cs.CV 2026-03 conditional novelty 6.0

    Using a 3D foundation model to produce viewpoint-aware anchors plus multi-view reference textures enables realistic human-object-interaction reenactment with large out-of-plane rotations.

  19. OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

    cs.SD 2026-02 conditional novelty 6.0

    A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.

  20. InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion

    cs.CV 2025-12 conditional novelty 6.0

    InsertAnywhere inserts a reference object into arbitrary videos by reconstructing 4D geometry to propagate a user-given placement across frames and fine-tuning video diffusion on ROSE++, a removal-to-insertion dataset...

  21. CustomX: Unified Character, Action, and Scene Customization in Video World Models

    cs.CV 2025-12 conditional novelty 6.0

    AniX generates controllable videos of a user-supplied character performing typed actions inside a user-supplied 3D scene by fine-tuning a pre-trained video generator on small locomotion datasets.

  22. Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation

    cs.CV 2026-07 conditional novelty 5.0

    A keyframe-anchored, training-free pipeline—terminal-state prompts, chained keyframe generation, and identity-aware sampling—ranks third on the IPVG26 Track 2 leaderboard.

  23. DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    DomainShuttle introduces domain-aware modeling and token separation techniques to achieve high subject fidelity with generative flexibility in open-domain subject-driven text-to-video tasks.

  24. HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    HarmoView proposes Multi-level Feature Injection, learnable proxy tokens, Jump-RoPE, and Progressive View Curriculum plus a new multi-view dataset to achieve state-of-the-art identity-consistent video generation from ...

  25. CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    Introduces CineDance-1M dataset for multi-shot long-form text-to-audio-video generation along with CineBench and a model adaptation.

  26. Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation

    cs.CV 2026-05 unverdicted novelty 5.0

    Omni-Customizer proposes an end-to-end framework using Omni-Context Fusion, Masked TTS Cross-Attention, Semantic-Anchored Multimodal RoPE, and specialized training curricula to achieve precise multimodal identity bind...

  27. Controllable Video Object Insertion via Multiview Priors

    cs.CV 2026-04 unverdicted novelty 5.0

    A multi-view prior-based framework for video object insertion that uses dual-path conditioning and an integration-aware consistency module to improve appearance stability and occlusion handling.

  28. Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation

    cs.CV 2026-06 unverdicted novelty 4.0

    ST-DRC proposes latent in-context injection, TASS-RoPE, appearance-invariant augmentation, and three-stream guidance to improve identity preservation in text-to-video diffusion models built on LTX-2.3.

  29. Evolution of Video Generative Foundations

    cs.CV 2026-04 unverdicted novelty 2.0

    This survey traces video generation technology from GANs to diffusion models and then to autoregressive and multimodal approaches while analyzing principles, strengths, and future trends.