REVIEW 6 cited by
StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
To leverage LLMs for visual synthesis, traditional methods convert raster image information into discrete grid tokens through specialized visual modules, while disrupting the model's ability to capture the true semantic representation of visual scenes. This paper posits that an alternative representation of images, vector graphics, can effectively surmount this limitation by enabling a more natural and semantically coherent segmentation of the image information. Thus, we introduce StrokeNUWA, a pioneering work exploring a better visual representation ''stroke tokens'' on vector graphics, which is inherently visual semantics rich, naturally compatible with LLMs, and highly compressed. Equipped with stroke tokens, StrokeNUWA can significantly surpass traditional LLM-based and optimization-based methods across various metrics in the vector graphic generation task. Besides, StrokeNUWA achieves up to a 94x speedup in inference over the speed of prior methods with an exceptional SVG code compression ratio of 6.9%.
Forward citations
Cited by 6 Pith papers
-
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
SVGenius benchmarks 22 LLMs on 2,377 SVG tasks across understanding, editing, and generation, finding universal degradation with complexity.
-
AutoSketch: VLM-assisted Style-Aware Vector Sketch Completion
A two-stage method that uses VLM-generated style descriptions and style-adjustment code to complete partial vector sketches in a style-consistent, prompt-aligned way.
-
LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer
A diffusion transformer trained on SVG construction sequences generates and vectorizes layered SVG graphics, breaking creation into editable steps.
-
Empowering LLMs to Understand and Generate Complex Vector Graphics
LLM4SVG adds learnable SVG tokens and SFT data so LLMs can generate and describe scalable vector graphics much better than general-purpose LLMs.
-
VQ-SGen: A Vector Quantized Stroke Representation for Creative Sketch Generation
VQ-SGen encodes each stroke as discrete shape and location codes and generates them autoregressively, producing more coherent and diverse creative sketches than prior methods.
-
SVGDreamer++: Advancing Editability and Diversity in Text-Guided SVG Generation
SVGDreamer++ uses SAM-based hierarchical masks and adaptive path control to generate text-guided SVGs that are more editable and visually detailed.
Discussion (0). Continue with ORCID to comment.