Pith. sign in

REVIEW 3 cited by

Collaborative Score Distillation for Consistent Visual Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.04787 v1 pith:7X57MKO5 submitted 2023-07-04 cs.CV cs.LG

classification cs.CVcs.LG
keywords imagesvisualacrossmultiplepriorsscorecollaborativeconsistency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative priors of large-scale text-to-image diffusion models enable a wide range of new generation and editing applications on diverse visual modalities. However, when adapting these priors to complex visual modalities, often represented as multiple images (e.g., video), achieving consistency across a set of images is challenging. In this paper, we address this challenge with a novel method, Collaborative Score Distillation (CSD). CSD is based on the Stein Variational Gradient Descent (SVGD). Specifically, we propose to consider multiple samples as "particles" in the SVGD update and combine their score functions to distill generative priors over a set of images synchronously. Thus, CSD facilitates seamless integration of information across 2D images, leading to a consistent visual synthesis across multiple samples. We show the effectiveness of CSD in a variety of tasks, encompassing the visual editing of panorama images, videos, and 3D scenes. Our results underline the competency of CSD as a versatile method for enhancing inter-sample consistency, thereby broadening the applicability of text-to-image diffusion models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Stroke of Surprise: Progressive Semantic Illusions in Vector Sketching

    cs.CV 2026-02 unverdicted novelty 7.0 of 10

    Stroke of Surprise is a framework that generates vector sketches undergoing semantic transformation from one concept to another by adding strokes, using dual-branch SDS and overlay loss for optimization.

  2. JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    A training-free two-stage pipeline uses cross-space dual-branch denoising with CLIP-guided voxel alignment and SDF blending for geometry, followed by view-conditioned 2D diffusion texture projection, to produce dual-s...

  3. VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    VISTA introduces a new synthetic triplet dataset and diffusion-transformer framework with style adapter that jointly models style, content, and motion to achieve state-of-the-art video style transfer.

Pith tools