Pith. sign in

REVIEW 6 cited by

CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.12035 v1 pith:M7YNJDSN submitted 2024-03-18 cs.CV

classification cs.CV
keywords modelbetterconsistencyvideocompatibilitycontrollabilityinpaintingtext-guided
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in video generation have been remarkable, yet many existing methods struggle with issues of consistency and poor text-video alignment. Moreover, the field lacks effective techniques for text-guided video inpainting, a stark contrast to the well-explored domain of text-guided image inpainting. To this end, this paper proposes a novel text-guided video inpainting model that achieves better consistency, controllability and compatibility. Specifically, we introduce a simple but efficient motion capture module to preserve motion consistency, and design an instance-aware region selection instead of a random region selection to obtain better textual controllability, and utilize a novel strategy to inject some personalized models into our CoCoCo model and thus obtain better model compatibility. Extensive experiments show that our model can generate high-quality video clips. Meanwhile, our model shows better motion consistency, textual controllability and model compatibility. More details are shown in [cococozibojia.github.io](cococozibojia.github.io).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OutDreamer: Video Outpainting with a Diffusion Transformer

    cs.CV 2025-06 conditional novelty 6.0 of 10

    OutDreamer couples a diffusion transformer with mask-driven self-attention and a latent alignment loss to outpaint videos in a zero-shot manner, exceeding prior zero-shot baselines on standard benchmarks.

  2. Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Vid-CamEdit re-synthesizes monocular videos along user-defined camera paths by conditioning a video diffusion model on 2D flows derived from estimated 3D geometry, without training on multi-view video data.

  3. MiniMax-Remover: Taming Bad Noise Helps Video Object Removal

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A two-stage video object remover that removes text conditioning and uses minimax adversarial noise to achieve high-quality removal in 6 sampling steps without classifier-free guidance.

  4. DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DreamDance animates a single character artwork by reconstructing its background as a 3D Gaussian scene and then inpainting the animated character into the rendered video.

  5. Se\~norita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A 2M-pair instruction-based video editing dataset built from real videos and specialist models, demonstrated to train editors that beat prior methods.

  6. Follow-Your-Creation: Empowering 4D Creation through Video Inpainting

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Follow-Your-Creation fine-tunes the Wan2.1 video inpainting model on composite point-cloud and editing masks so a single monocular video can be converted into editable 4D video with new camera motion.

Pith tools