Pith. sign in

REVIEW 6 cited by

StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.17093 v1 pith:SGBSTZTM submitted 2024-01-30 cs.CV cs.CL

classification cs.CVcs.CL
keywords visualstrokenuwavectormethodsrepresentationtokensgraphicgraphics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To leverage LLMs for visual synthesis, traditional methods convert raster image information into discrete grid tokens through specialized visual modules, while disrupting the model's ability to capture the true semantic representation of visual scenes. This paper posits that an alternative representation of images, vector graphics, can effectively surmount this limitation by enabling a more natural and semantically coherent segmentation of the image information. Thus, we introduce StrokeNUWA, a pioneering work exploring a better visual representation ''stroke tokens'' on vector graphics, which is inherently visual semantics rich, naturally compatible with LLMs, and highly compressed. Equipped with stroke tokens, StrokeNUWA can significantly surpass traditional LLM-based and optimization-based methods across various metrics in the vector graphic generation task. Besides, StrokeNUWA achieves up to a 94x speedup in inference over the speed of prior methods with an exceptional SVG code compression ratio of 6.9%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SVGenius benchmarks 22 LLMs on 2,377 SVG tasks across understanding, editing, and generation, finding universal degradation with complexity.

  2. AutoSketch: VLM-assisted Style-Aware Vector Sketch Completion

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A two-stage method that uses VLM-generated style descriptions and style-adjustment code to complete partial vector sketches in a style-consistent, prompt-aligned way.

  3. LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A diffusion transformer trained on SVG construction sequences generates and vectorizes layered SVG graphics, breaking creation into editable steps.

  4. Empowering LLMs to Understand and Generate Complex Vector Graphics

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LLM4SVG adds learnable SVG tokens and SFT data so LLMs can generate and describe scalable vector graphics much better than general-purpose LLMs.

  5. VQ-SGen: A Vector Quantized Stroke Representation for Creative Sketch Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    VQ-SGen encodes each stroke as discrete shape and location codes and generates them autoregressively, producing more coherent and diverse creative sketches than prior methods.

  6. SVGDreamer++: Advancing Editability and Diversity in Text-Guided SVG Generation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    SVGDreamer++ uses SAM-based hierarchical masks and adaptive path control to generate text-guided SVGs that are more editable and visually detailed.

Pith tools