Pith. sign in

REVIEW 2 cited by

Motion by Queries: Identity-Motion Trade-offs in Text-to-Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.07750 v3 pith:RGGC6JDB submitted 2024-12-10 cs.CV

classification cs.CV
keywords motionidentityinjectionvideogenerationmodelsqueryrepresentations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-video diffusion models have shown remarkable progress in generating coherent video clips from textual descriptions. However, the interplay between motion, structure, and identity representations in these models remains under-explored. Here, we investigate how self-attention query (Q) features simultaneously govern motion, structure, and identity and examine the challenges arising when these representations interact. Our analysis reveals that Q affects not only layout, but that during denoising Q also has a strong effect on subject identity, making it hard to transfer motion without the side-effect of transferring identity. Understanding this dual role enabled us to control query feature injection (Q injection) and demonstrate two applications: (1) a zero-shot motion transfer method - implemented with VideoCrafter2 and WAN 2.1 - that is 10 times more efficient than existing approaches, and (2) a training-free technique for consistent multi-shot video generation, where characters maintain identity across multiple video shots while Q injection enhances motion fidelity.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    EmoWorld adds three training-free steering operators to a frozen video diffusion transformer that separately control atmosphere, affect-bearing cues, and temporal emotion transitions in generated videos.

  2. Controlling Motion Transfer in Diffusion Transformers via Attention Heads

    cs.CV 2026-07 accept novelty 6.0 of 10

    Video DiTs encode motion and structure in separate attention-head subsets; selecting and guiding those heads yields training-free motion transfer with higher fidelity and structural alignment than existing methods.

Pith tools