Pith. sign in

REVIEW 2 cited by

Compositional Video Generation as Flow Equalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.06182 v1 pith:VPTBQQXR submitted 2024-06-10 cs.CV

classification cs.CV
keywords videovicocompositionalmodelsattentionconceptsflowinfluence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale Text-to-Video (T2V) diffusion models have recently demonstrated unprecedented capability to transform natural language descriptions into stunning and photorealistic videos. Despite the promising results, a significant challenge remains: these models struggle to fully grasp complex compositional interactions between multiple concepts and actions. This issue arises when some words dominantly influence the final video, overshadowing other concepts.To tackle this problem, we introduce \textbf{Vico}, a generic framework for compositional video generation that explicitly ensures all concepts are represented properly. At its core, Vico analyzes how input tokens influence the generated video, and adjusts the model to prevent any single concept from dominating. Specifically, Vico extracts attention weights from all layers to build a spatial-temporal attention graph, and then estimates the influence as the \emph{max-flow} from the source text token to the video target token. Although the direct computation of attention flow in diffusion models is typically infeasible, we devise an efficient approximation based on subgraph flows and employ a fast and vectorized implementation, which in turn makes the flow computation manageable and differentiable. By updating the noisy latent to balance these flows, Vico captures complex interactions and consequently produces videos that closely adhere to textual descriptions. We apply our method to multiple diffusion-based video models for compositional T2V and video editing. Empirical results demonstrate that our framework significantly enhances the compositional richness and accuracy of the generated videos. Visit our website at~\href{https://adamdad.github.io/vico/}{\url{https://adamdad.github.io/vico/}}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations

    cs.CV 2025-01 conditional novelty 5.0 of 10

    BlobGEN-Vid uses blob video representations, masked 3D attention, and context interpolation to achieve state-of-the-art layout controllability in text-to-video generation.

  2. Mapping the Mind of an Instruction-based Image Editing using SMILE

    cs.AI 2024-12 reject novelty 4.0 of 10

    SMILE applies LIME-style prompt perturbation with image-embedding distances to create word-level heatmaps for instruction-based image editing models.

Pith tools