Pith. sign in

REVIEW 2 cited by

DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15306 v3 pith:B7EFP44U submitted 2024-05-24 cs.CL cs.CV

classification cs.CLcs.CV
keywords figuresdetikzifyscientifictikzsketchesgraphicsprogramsalgorithm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Creating high-quality scientific figures can be time-consuming and challenging, even though sketching ideas on paper is relatively easy. Furthermore, recreating existing figures that are not stored in formats preserving semantic information is equally complex. To tackle this problem, we introduce DeTikZify, a novel multimodal language model that automatically synthesizes scientific figures as semantics-preserving TikZ graphics programs based on sketches and existing figures. To achieve this, we create three new datasets: DaTikZv2, the largest TikZ dataset to date, containing over 360k human-created TikZ graphics; SketchFig, a dataset that pairs hand-drawn sketches with their corresponding scientific figures; and MetaFig, a collection of diverse scientific figures and associated metadata. We train DeTikZify on MetaFig and DaTikZv2, along with synthetically generated sketches learned from SketchFig. We also introduce an MCTS-based inference algorithm that enables DeTikZify to iteratively refine its outputs without the need for additional training. Through both automatic and human evaluation, we demonstrate that DeTikZify outperforms commercial Claude 3 and GPT-4V in synthesizing TikZ programs, with the MCTS algorithm effectively boosting its performance. We make our code, models, and datasets publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AutoPresent: Designing Structured Visuals from Scratch

    cs.CV 2025-01 conditional novelty 6.0 of 10

    AutoPresent is an open 8B model trained on a new 7k-example benchmark, SlidesBench, that generates presentation slides from natural language and performs comparably to GPT-4o in one of three evaluation settings.

  2. Slow Perception: Let's Perceive Geometric Figures Step-by-step

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Decomposing geometric figures into line segments and tracing each with multiple short strokes improves LVLM geometric parsing by about 6 F1 points over direct endpoint regression.

Pith tools