Pith. sign in

VideoSketcher: Sequential Sketch Generation Using Video Model Priors

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Sketching is inherently sequential: strokes are drawn progressively to explore and refine ideas. Yet most generative approaches treat sketches as static images, ignoring the temporal process underlying creative exploration. Modeling this sequential structure remains challenging: prior methods either rely on large-scale human-drawn datasets with limited diversity, or use large language models (LLMs) to produce drawing instructions, often at the cost of visual fidelity. We present VideoSketcher, a method for generating high-quality sketching processes by adapting pretrained text-to-video diffusion models to the sparse, continuous nature of sketch formation. Our key insight is that LLMs and video diffusion models offer complementary strengths: LLMs act as semantic planners that decompose concepts into step-by-step instructions, while video diffusion models serve as powerful "renderers" that translate them into temporally coherent sketch sequences. We introduce a two-stage fine-tuning strategy that decouples temporal structure from visual appearance: stroke ordering is learned from synthetic shape compositions, while style is distilled from as few as seven hand-drawn examples. Despite minimal supervision, our method can generate diverse, high-quality sequential sketches that faithfully follow specified drawing orders. Our framework naturally extends to brush style control and autoregressive generation, supporting artistic applications.

fields

cs.CV 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Draw This First

cs.CV · 2026-08-12 · conditional · novelty 7.0

Draw order is encoded as color in an image, generated by a pretrained diffusion transformer, then decoded into ordered vector strokes whose order follows language instructions.

citing papers explorer

Showing 1 of 1 citing paper.

  • Draw This First cs.CV · 2026-08-12 · conditional · none · ref 11 · internal anchor

    Draw order is encoded as color in an image, generated by a pretrained diffusion transformer, then decoded into ordered vector strokes whose order follows language instructions.