Pith. sign in

REVIEW 10 cited by

CVPR 2023 Text Guided Video Editing Competition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.16003 v1 pith:I2WAQFT2 submitted 2023-10-24 cs.CV

classification cs.CV
keywords videocompetitiondatasetcvpreditingevaluatemodelstasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans watch more than a billion hours of video per day. Most of this video was edited manually, which is a tedious process. However, AI-enabled video-generation and video-editing is on the rise. Building on text-to-image models like Stable Diffusion and Imagen, generative AI has improved dramatically on video tasks. But it's hard to evaluate progress in these video tasks because there is no standard benchmark. So, we propose a new dataset for text-guided video editing (TGVE), and we run a competition at CVPR to evaluate models on our TGVE dataset. In this paper we present a retrospective on the competition and describe the winning method. The competition dataset is available at https://sites.google.com/view/loveucvpr23/track4.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Audio-Guided Visual Editing with Complex Multi-Modal Prompts

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A training-free framework that maps audio embeddings into Stable Diffusion's text space and fuses multiple audio/text prompts via per-patch residual noise selection, outperforming text-only editors on new audio-visual...

  2. STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    STR-Match edits videos without retraining by matching source and target 'spatiotemporal relevance scores' derived from attention maps during latent optimization.

  3. Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A temporal attention purification and skip-connection rerouting method that separates motion learning from appearance learning in text-to-video diffusion model customization.

  4. Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A training-free method that matches sparse-point inter-frame feature correlations to transfer reference motion to generated videos with improved temporal consistency.

  5. PRIMEdit: Probability Redistribution for Instance-aware Multi-object Video Editing with Benchmark Dataset

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PRIMEdit edits multiple video objects independently using per-object masks and captions, and contributes a benchmark dataset and a leakage metric.

  6. DIVE: Taming DINO for Subject-Driven Video Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DIVE uses DINOv2 feature maps as automatic video correspondences to carry source motion, while LoRA adapters carry the target identity.

  7. ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport

    cs.CV 2026-08 conditional novelty 5.0 of 10

    Motion-aligned causal averaging of one-step Chord edit fields, plus shared noise, cuts temporal flicker and warping error by about half to three quarters at 2 NFE/frame with no training or inversion.

  8. Consistent and Editable: A Balanced Framework for Text-Guided Video Editing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    EquiEdit balances temporal consistency and editability in diffusion-based text-guided video editing via a temporal Mamba module and spectral noise injection on initial latents.

  9. FlowDrag: 3D-aware Drag-based Image Editing with Mesh-guided Deformation Vector Flow Fields

    cs.GR 2025-07 conditional novelty 5.0 of 10

    FlowDrag combines 3D mesh deformation with diffusion-based drag editing, using the resulting 2D vector flow to steer the denoising process, and adds a ground-truth benchmark built from video frames.

  10. Re-Attentional Controllable Video Diffusion Editing

    cs.CV 2024-12 conditional novelty 5.0 of 10

    ReAtCo improves text-guided video editing by using attention-map gradients to place edited objects in user-specified regions and by re-injecting the original background during diffusion sampling.

Pith tools