REVIEW 2 cited by
DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Creating high-quality scientific figures can be time-consuming and challenging, even though sketching ideas on paper is relatively easy. Furthermore, recreating existing figures that are not stored in formats preserving semantic information is equally complex. To tackle this problem, we introduce DeTikZify, a novel multimodal language model that automatically synthesizes scientific figures as semantics-preserving TikZ graphics programs based on sketches and existing figures. To achieve this, we create three new datasets: DaTikZv2, the largest TikZ dataset to date, containing over 360k human-created TikZ graphics; SketchFig, a dataset that pairs hand-drawn sketches with their corresponding scientific figures; and MetaFig, a collection of diverse scientific figures and associated metadata. We train DeTikZify on MetaFig and DaTikZv2, along with synthetically generated sketches learned from SketchFig. We also introduce an MCTS-based inference algorithm that enables DeTikZify to iteratively refine its outputs without the need for additional training. Through both automatic and human evaluation, we demonstrate that DeTikZify outperforms commercial Claude 3 and GPT-4V in synthesizing TikZ programs, with the MCTS algorithm effectively boosting its performance. We make our code, models, and datasets publicly available.
Forward citations
Cited by 2 Pith papers
-
AutoPresent: Designing Structured Visuals from Scratch
AutoPresent is an open 8B model trained on a new 7k-example benchmark, SlidesBench, that generates presentation slides from natural language and performs comparably to GPT-4o in one of three evaluation settings.
-
Slow Perception: Let's Perceive Geometric Figures Step-by-step
Decomposing geometric figures into line segments and tracing each with multiple short strokes improves LVLM geometric parsing by about 6 F1 points over direct endpoint regression.
Discussion (0). Continue with ORCID to comment.