Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

SketchAgent: Language-Driven Sequential Sketch Generation

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A frozen multimodal LLM can draw sketches stroke-by-stroke using only a numbered grid and a few examples.

desk verdict A genuinely new sketching representation for frozen multimodal LLMs, with honest experiments—but the 'wide range' claim is overstated; the evidence supports a moderate set of simple iconic concepts. read the letter →

arxiv 2411.17673 v1 pith:KRHHZWSZ submitted 2024-11-26 cs.CV

classification cs.CV
keywords sketchgenerationmultimodallargelanguagemodelssequentialdrawingin-contextlearningchain-of-thoughtpromptinghuman-AIcollaborationchat-basededitingBéziercurvefitting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that an off-the-shelf multimodal large language model, with no training or fine-tuning, can be turned into a general-purpose sketch generator that draws stroke by stroke. The key design is a string-based sketching language on a numbered grid canvas: the model outputs coordinates and timing values as text, and a parser fits smooth Bézier curves to those points and renders them. Because strokes are produced sequentially, the same agent can edit its own sketch through chat and take turns with a human on a shared canvas. If the claim holds, the bottleneck for a general-purpose sketching agent is not sketch data or generative training but a well-designed interface on top of existing model priors. On the paper's metrics, the agent approaches human-level sketch recognizability under a CLIP classifier and is judged more human-like than direct SVG prompting in a two-alternative forced-choice study.

What carries the argument

The load-bearing object is the sketching language over a numbered grid canvas. The canvas is a 50×50 grid whose cells carry labels like x2y8, giving the text-only LLM a concrete coordinate system to reference; the paper includes a demonstration that a multimodal LLM which can name the right line to draw nevertheless fails to execute it with pixel coordinates, motivating the grid. A stroke is written as a list of grid coordinates plus a matching list of t-values, meaning the points are treated as samples along a curve rather than control points, and the system fits a cubic Bézier curve to them with a least-squares solve, splitting recursively when the fit error is large. In-context learning teaches the format and chain-of-thought prompts make the model plan stroke order in <thinking> tags; a stopping token </s{j}> allows a human to add strokes to the same canvas, which are converted back into the coordinate format by sampling the user's Bézier curve at its t-values.

What would settle it

Re-run the 500-sketch, 50-category CLIP benchmark with the identical prompts and backbone but with the grid numbers erased from the canvas image (cells left blank), and compare Top-1 accuracy against the reported 0.23; if accuracy does not drop substantially, the numbered-grid spatial scaffolding is not the mechanism that makes the method work.

Watch

Extended reading notes

Core claim

SketchAgent establishes that a frozen multimodal LLM, guided only by a system prompt, a user prompt with one worked example, and a numbered 50×50 grid canvas, can generate recognizable sketches of arbitrary textual concepts—including landmarks, scientific principles, and diagrams far beyond the 345 QuickDraw categories that bound trained sketch models. The model plans in <thinking> tags, emits strokes as coordinate sequences with t-values, and the system fits cubic Bézier curves to the sampled points by least squares, recursively splitting long curves. The same representation supports the whole interaction loop: the rendered canvas is fed back for chat-based edits, and a stopping token lets a human interleave their own strokes, which are re-sampled into the agent's coordinate format. Quantitative results on 500 sketches across 50 categories give CLIP zero-shot Top-1/Top-5 accuracy of 0.23/0.44 with the default backbone, approaching the 0.27/0.49 of human QuickDraw sketches under the same metric, and an ablation shows that removing the system prompt, chain-of-thought, or the complete in-context example each degrades accuracy.

Load-bearing premise

The load-bearing premise is that an off-the-shelf multimodal LLM can convert its semantic understanding of an object into a spatially coherent, ordered sequence of grid coordinates, given only a numbered canvas and a few in-context examples; the paper itself acknowledges that the agent often writes rich textual descriptions of parts yet struggles to turn them into effective drawing actions.

Editorial extensions

If this is right

  • A general-purpose sketching agent can be assembled without collecting human drawing data or training a generative model; the capability is inherited from the backbone LLM, so future improvements in those models should transfer directly to sketch quality.
  • Sketch editing becomes a natural conversational operation: because the canvas is part of the dialogue state, the agent can add, relocate, or annotate parts of an existing drawing in response to text, with 92% of tested editing prompts followed correctly.
  • The stroke-by-stroke output carries semantic labels assigned by the model, so sketches come with part-level annotations as a byproduct, useful for analysis and dataset construction.
  • Real-time collaborative sketching with a human partner is feasible: individual strokes take about 8 seconds in collaborative mode and a complete sketch about 20 seconds, matching the pace of human drawing.
  • The approach is largely backbone-agnostic: it works with several closed commercial models and, with lower recognition scores, with a large open-weight model, suggesting the interface is portable rather than tied to one model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the failure cases the paper lists (unicorn, human figures, letters and numbers) outline the boundary of the backbone's latent spatial competence; a natural testable prediction is that these failure modes shrink as multimodal LLMs improve, without any change to the sketching pipeline.
  • Editorial inference: the numbered-grid protocol is a general recipe for eliciting spatial output from text-only models, so the same idea could be tested for diagram layout, floor-plan drafting, or GUI wireframes, where a grid plus in-context examples might replace fine-tuning.
  • Editorial inference: the 92% editing accuracy was measured on a small hand-picked set; a broader stress test that varies object combinations and demands precise relative placement would clarify whether the agent reasons spatially or falls back on stereotyped layouts.
  • Editorial inference: because the prompts were tuned on the default backbone, the reported gap between backbone models partly reflects prompt fit rather than raw capability, so the method's portability across models is probably understated.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces SketchAgent, a training-free method for sequential sketch generation using a frozen multimodal large language model (Claude3.5-Sonnet). The agent is prompted with a grid-coordinate sketching language, in-context examples, and chain-of-thought instructions, and it outputs a sequence of strokes that are fitted to Bézier curves and rendered on a canvas. The paper claims that this approach supports text-conditioned sketch generation across a wide range of concepts, enables stroke-by-stroke sequential sketching with semantic stroke annotations, facilitates real-time human-agent collaborative sketching, and supports iterative chat-based editing. Evaluations include CLIP zero-shot classification on 500 sketches across 50 QuickDraw categories, a 2AFC human study comparing sketch human-likeness, a collaborative user study with 30 participants, an ablation study, and a small chat-editing study on 54 sketches. The supplementary provides additional qualitative results and detailed prompts.

Significance. If the central claims hold, the paper demonstrates a meaningful new capability: an off-the-shelf multimodal LLM, without any fine-tuning, can produce ordered, semantically labeled strokes that render into recognizable sketches and can participate in interactive drawing sessions. The method is simple, reproducible, and the authors commit to releasing code. The paper includes useful ablations and honest discussion of limitations, and the human studies add credibility beyond CLIP-based proxies. However, the strongest claim—"wide range of textual concepts"—is only partially supported by the evidence, and the editing and collaborative evaluations have important protocol gaps. The work is likely to be of interest to the sketch-generation and human-AI interaction communities, but the conclusions should be scaled to match the demonstrated scope.

major comments (4)
  1. [Sec. 5.1, Table 1, Supp. Figs. 21, 23] The claim in Sec. 1 that SketchAgent "can generate sketches across a wide range of textual concepts" is not established by the quantitative evidence. The average Top-1 recognition of 0.23 in Table 1 hides substantial per-category variability: the supplementary confusion matrix and top-recognized-class figures show many categories with recognition rates of 20–30% or lower, and the paper itself (Sec. 7, Fig. 14) acknowledges failures on unicorn, Frida Kahlo, and letters/numbers. The appendix also admits that some randomly selected concepts (Statue of Liberty, photosynthesis, pie chart) were unsuccessful. The claim should be narrowed to simple iconic objects, or the evaluation should be strengthened with per-category success thresholds or human recognition judgments on the full concept set.
  2. [Sec. 5.4, Chat-Based Sketch Editing] The editing evaluation reports that SketchAgent "correctly follows instructions 92% of the time" on 54 sketches, but the manuscript does not define what constitutes correct instruction-following, who made the judgment, how many raters were involved, or whether there was inter-rater agreement. Because interactive editing is one of the paper's three core contributions, this metric needs a transparent evaluation protocol to be load-bearing. The current description is insufficient for readers to assess the reliability of the 92% figure.
  3. [Sec. 5.3, Human-Agent Collaborative Sketching] The collaborative user study uses 8 concepts that were explicitly "selected based on the agent's demonstrated ability to draw them independently" (Sec. 5.3). This selection makes the study unsuitable for supporting the broader claim that SketchAgent can collaborate on arbitrary concepts. Additionally, the analysis that agent-only and user-only strokes have low CLIP recognition rates (Table 3) may conflate incompleteness with lack of contribution, since partial sketches are not necessarily expected to be recognizable. A human evaluation of partial sketches or a study with less favorably selected concepts would be needed to substantiate the collaboration claim.
  4. [Sec. 4, Method Overview] The paper repeatedly emphasizes that SketchAgent captures the "dynamic, evolving process" of sketching and "incorporates visual feedback" (Sec. 1). However, during a single sketch-generation turn, the model produces the entire coordinate sequence in one pass; the canvas is fed back only for later editing or collaborative turns. The sequential nature is primarily in the output representation and stroke ordering, not in the model's internal generation process. This distinction should be clearly stated, as the current wording overstates the mechanism.
minor comments (5)
  1. [Sec. 5, opening paragraph] There is a typo: "We demonstrate SketchAgent's capabil to generate" should read "capability".
  2. [Sec. 6, Table 2] The ablation claims that "all components contribute to the agent's full performance," but the difference between the full pipeline and the w/o System Prompt condition (0.23 vs 0.20 Top-1) is within the reported error bars. The authors should either report significance tests or temper the claim for the system-prompt component.
  3. [Sec. 5.1, Table 1] The table's last row label "Vis." is not explained in the text; a short caption indicating that it shows example sketches would improve clarity.
  4. [Sec. 5.4 and Supp. B.4] The editing prompt in the appendix says the model should "Describe the location of the added concepts first in <thinking> tags," but for the animals category the instructions (e.g., "Add a hat") contain no location; the paper should clarify whether the model is expected to infer placement in those cases and how that inference was scored.
  5. [Sec. 2, Related Work] The discussion of direct SVG prompting in Fig. 3 is useful, but there is no direct quantitative comparison to optimization-based sketch generation methods such as CLIPasso or DiffSketch on the same QuickDraw subset; adding such a comparison would help calibrate the reported numbers against prior art.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: SketchAgent's outputs are produced by an off-the-shelf LLM from prompts and evaluated with external CLIP and human judgments; no claimed result is defined by its inputs.

full rationale

SketchAgent's claimed capability is produced by an off-the-shelf multimodal LLM responding to a system prompt, a user prompt, and a numbered canvas; the only numerical post-processing, least-squares Bézier fitting in Eq. (2), is a rendering operation that does not define the semantic content of the sketch. No parameters are fitted to the CLIP recognition results or to the human 2AFC judgments, so the reported numbers are measurements rather than quantities forced by construction. The quantitative evaluation uses an external CLIP zero-shot classifier and human participants, with human QuickDraw sketches as baselines, and the recognition metric is not a re-expression of the method's inputs. The paper's self-citations (e.g., [116, 117, 126, 127] for the CLIP evaluation protocol) are methodological references rather than load-bearing justifications of the central claim. The supplement's admission that prompts were optimized for Claude3.5-Sonnet is a disclosed selection effect relevant to generalization, not a circular step, and the limitations in Sec. 7 concede unrecognizable outputs, which is inconsistent with a forced or self-fulfilling derivation. Concerns about the strength of the "wide range" claim are about evidence sufficiency, not circularity. Hence there is no circular step in the claimed derivation chain.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central claim rests on hand-chosen design choices (e.g., grid resolution) and domain assumptions about the backbone LLM's visual priors and the validity of CLIP as an evaluation proxy. No parameters are fitted to the evaluation data, and no new entities are introduced.

free parameters (1)
  • grid resolution = 50x50 cells
    Hand-chosen to balance spatial expressiveness and token length; affects stroke coordinate precision.
assumptions (3)
  • domain assumption A multimodal LLM's pretrained visual priors are sufficient to plan and execute coherent sketches from text prompts when provided a grid and in-context examples.
    Central premise of the method; the authors note in Sec. 7 that the agent sometimes produces text descriptions it cannot convert into strokes.
  • domain assumption CLIP zero-shot classification accuracy is a valid proxy for sketch recognizability and semantic fidelity.
    Used for all quantitative evaluations in Secs. 5.1-5.3; CLIP is a pretrained image-text model, not a human perceptual benchmark.
  • standard math Least-squares fitting of cubic Bezier curves to sampled grid points yields visually smooth, natural strokes.
    Standard curve fitting; the smoothness claim is supported by qualitative comparisons in Fig. 7.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SketchAgent: Language-Driven Sequential Sketch Generation." pith.science (2026). https://pith.science/paper/KRHHZWSZ

@misc{pith2026241117673,
  author       = {Pith},
  title        = {Pith review of: SketchAgent: Language-Driven Sequential Sketch Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KRHHZWSZ}},
  note         = {Machine review of arXiv:2411.17673}
}
read the original abstract

Sketching serves as a versatile tool for externalizing ideas, enabling rapid exploration and visual communication that spans various disciplines. While artificial systems have driven substantial advances in content creation and human-computer interaction, capturing the dynamic and abstract nature of human sketching remains challenging. In this work, we introduce SketchAgent, a language-driven, sequential sketch generation method that enables users to create, modify, and refine sketches through dynamic, conversational interactions. Our approach requires no training or fine-tuning. Instead, we leverage the sequential nature and rich prior knowledge of off-the-shelf multimodal large language models (LLMs). We present an intuitive sketching language, introduced to the model through in-context examples, enabling it to "draw" using string-based actions. These are processed into vector graphics and then rendered to create a sketch on a pixel canvas, which can be accessed again for further tasks. By drawing stroke by stroke, our agent captures the evolving, dynamic qualities intrinsic to sketching. We demonstrate that SketchAgent can generate sketches from diverse prompts, engage in dialogue-driven drawing, and collaborate meaningfully with human users.

Figures

Figures reproduced from arXiv: 2411.17673 by the authors.

Figure 1
Figure 1. SketchAgent leverages an off-the-shelf multimodal LLM to facilitate language-driven, sequential sketch generation through an [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Examples of sketches used across disciplines and goals. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Sketch appearance. (A) Text-to-image diffusion models [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (49 more)
Figure 4
Figure 4. Figure 4: Cubic Bezier curve. ´ Vector Graphics and Bezier ´ Curves Vector graphics allow us to create visual images directly from geometric shapes such as points, lines, curves, and polygons. Unlike raster images (represented with pixels), vector graphics are resolution-free, m…
Figure 5
Figure 5. Figure 5: Method Overview. SketchAgent (blue) receives drawing instructions and generates a string representing the intended sketch. [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 6
Figure 6. Figure 6: Although excelling in visual reasoning, multimodal [PITH_FULL_IMAGE:figures/full_fig_p004_6.png]
Figure 7
Figure 7. Figure 7: Methods for processing the agent’s coordinate sequence [PITH_FULL_IMAGE:figures/full_fig_p005_7.png]
Figure 8
Figure 8. Figure 8: Sketches produced by SketchAgent for concepts beyond [PITH_FULL_IMAGE:figures/full_fig_p005_8.png]
Figure 9
Figure 9. Figure 9: SketchAgent gradually draws stroke-by-stroke, each [PITH_FULL_IMAGE:figures/full_fig_p006_9.png]
Figure 11
Figure 11. Figure 11: Sequential sketching analysis of SketchAgent (blue) [PITH_FULL_IMAGE:figures/full_fig_p007_11.png]
Figure 10
Figure 10. Figure 10: Sequential sketching process. SVGDreamer [ [PITH_FULL_IMAGE:figures/full_fig_p007_10.png]
Figure 12
Figure 12. Figure 12: Collaborative sketching evaluation measured using [PITH_FULL_IMAGE:figures/full_fig_p008_12.png]
Figure 13
Figure 13. Figure 13: Chat-based sketch editing. We iteratively prompt [PITH_FULL_IMAGE:figures/full_fig_p008_13.png]
Figure 14
Figure 14. Figure 14: Limitations. Sketches of complex concepts (A) and [PITH_FULL_IMAGE:figures/full_fig_p008_14.png]
Figure 15
Figure 15. Figure 15: Visualization of single-stroke primitives used in the [PITH_FULL_IMAGE:figures/full_fig_p015_15.png]
Figure 16
Figure 16. Figure 16: Visualization of the simple sketch of a house provided [PITH_FULL_IMAGE:figures/full_fig_p015_16.png]
Figure 17
Figure 17. Figure 17: Sketch variability. Example of twelve different [PITH_FULL_IMAGE:figures/full_fig_p015_17.png]
Figure 18
Figure 18. Figure 18: Randomly selected sketches of scientific concepts. [PITH_FULL_IMAGE:figures/full_fig_p016_18.png]
Figure 21
Figure 21. Figure 21: Confusion matrix (showing top 10 confused classes) [PITH_FULL_IMAGE:figures/full_fig_p017_21.png]
Figure 22
Figure 22. Figure 22: Visualization of sketches from the six most confused [PITH_FULL_IMAGE:figures/full_fig_p017_22.png]
Figure 20
Figure 20. Figure 20: Randomly selected sketches of notable landmarks. [PITH_FULL_IMAGE:figures/full_fig_p017_20.png]
Figure 24
Figure 24. Figure 24: Visualization of sketches from the most recognized [PITH_FULL_IMAGE:figures/full_fig_p018_24.png]
Figure 26
Figure 26. Figure 26: Sketches generated using Llama-3.2-11B-Vision as our [PITH_FULL_IMAGE:figures/full_fig_p019_26.png]
Figure 27
Figure 27. Figure 27: Visualization of the eight top recognized classes for the [PITH_FULL_IMAGE:figures/full_fig_p019_27.png]
Figure 29
Figure 29. Figure 29: Direct prompting for SVG generation across dif [PITH_FULL_IMAGE:figures/full_fig_p020_29.png]
Figure 30
Figure 30. Figure 30: An example of our 2AFC session [PITH_FULL_IMAGE:figures/full_fig_p020_30.png]
Figure 33
Figure 33. Figure 33: Sequential four-stroke sketches of a pear, purse, and [PITH_FULL_IMAGE:figures/full_fig_p021_33.png]
Figure 34
Figure 34. Figure 34: Sequential five-stroke sketches of a pear, purse, and [PITH_FULL_IMAGE:figures/full_fig_p021_34.png]
Figure 35
Figure 35. Figure 35: Sequential six-stroke sketches of a television, bed, and [PITH_FULL_IMAGE:figures/full_fig_p022_35.png]
Figure 36
Figure 36. Figure 36: Sequential seven-stroke sketches of a backpack, fish, [PITH_FULL_IMAGE:figures/full_fig_p022_36.png]
Figure 40
Figure 40. Figure 40: Screenshot of our web interface [PITH_FULL_IMAGE:figures/full_fig_p026_40.png]
Figure 41
Figure 41. Figure 41: User instructions in the sketching interface. [PITH_FULL_IMAGE:figures/full_fig_p026_41.png]
Figure 43
Figure 43. Figure 43: Confusion matrix from CLIP classification with cate [PITH_FULL_IMAGE:figures/full_fig_p027_43.png]
Figure 44
Figure 44. Figure 44: Confusion matrix from CLIP classification with cat [PITH_FULL_IMAGE:figures/full_fig_p027_44.png]
Figure 45
Figure 45. Figure 45: Examples of sketches created in “collab” mode that [PITH_FULL_IMAGE:figures/full_fig_p028_45.png]
Figure 47
Figure 47. Figure 47: Chat-based sketch editing. We iteratively prompt [PITH_FULL_IMAGE:figures/full_fig_p028_47.png]
Figure 46
Figure 46. Figure 46: Examples of sketches from our collaborative human [PITH_FULL_IMAGE:figures/full_fig_p028_46.png]
Figure 48
Figure 48. Figure 48: Visualization of sketches produced in different cases of [PITH_FULL_IMAGE:figures/full_fig_p029_48.png]
Figure 50
Figure 50. Figure 50: ICL example ablation study. We examine the impact of [PITH_FULL_IMAGE:figures/full_fig_p029_50.png]
Figure 49
Figure 49. Figure 49: ICL example ablation study. We examine the impact [PITH_FULL_IMAGE:figures/full_fig_p029_49.png]
Figure 51
Figure 51. Figure 51: System prompt. 16 [PITH_FULL_IMAGE:figures/full_fig_p030_51.png]
Figure 52
Figure 52. Figure 52: User prompt. This prompt contains the specific sketching task as well as details about the expected format. [PITH_FULL_IMAGE:figures/full_fig_p031_52.png]
Figure 53
Figure 53. Figure 53: ICL example. This is the example of a sketch of a house we provide to the model. [PITH_FULL_IMAGE:figures/full_fig_p032_53.png]
Figure 54
Figure 54. Figure 54: ICL example. This is the example of a sketch of a house we provide to the model. [PITH_FULL_IMAGE:figures/full_fig_p033_54.png]
Figure 55
Figure 55. Figure 55: ICL example. This is the example of a sketch of a house we provide to the model. [PITH_FULL_IMAGE:figures/full_fig_p034_55.png]
Figure 57
Figure 57. Figure 57: Randomly generated sketches used in the quantitative analysis (ten sketches per category). [PITH_FULL_IMAGE:figures/full_fig_p036_57.png]
Figure 58
Figure 58. Figure 58: Sketches generated by SketchAgent for the eight categories of our human-agent collaborative study. [PITH_FULL_IMAGE:figures/full_fig_p037_58.png]
Figure 59
Figure 59. Figure 59: Sketches generated by SketchAgent for the eight categories of our human-agent collaborative study. [PITH_FULL_IMAGE:figures/full_fig_p038_59.png]
Figure 60
Figure 60. Figure 60: Sketches created collaboratively by users and SketchAgent as part of our collaborative human study. [PITH_FULL_IMAGE:figures/full_fig_p039_60.png]
Figure 61
Figure 61. Figure 61: Sketches created collaboratively by users and SketchAgent as part of our collaborative human study. [PITH_FULL_IMAGE:figures/full_fig_p040_61.png]
Figure 62
Figure 62. Figure 62: Sketches created by users in “solo” mode as part of our collaborative human study. [PITH_FULL_IMAGE:figures/full_fig_p041_62.png]
Figure 63
Figure 63. Figure 63: Sketches created by users in “solo” mode as part of our collaborative human study. [PITH_FULL_IMAGE:figures/full_fig_p042_63.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pencils to Pixels: A Systematic Study of Creative Drawings across Children, Adults and AI

    cs.HC 2025-02 conditional novelty 6.0 of 10

    A new dataset and computational framework quantify style and content in children's, adults', and DALL-E drawings, showing that expert and automated creativity scores disagree across groups.

Reference graph

Works this paper leans on

140 extracted references · 59 canonical work pages · cited by 1 Pith paper

  1. [1]

    Flamingo: a visual language model for few-shot learn- ing

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, An- toine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millicah, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Se- bastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Bink...

  2. [2]

    Drawing as a space for social-cognitive in- teraction

    Vanessa De Andrade, Sofia Freire, M ´onica Baptista, and Yael Shwartz. Drawing as a space for social-cognitive in- teraction. Education Sciences, 12(1), 2022. 3

  3. [3]

    Anthropic. Claude. https://www.anthropic.com/ claude, 2023. 2, 3, 5, 6, 1

  4. [4]

    Style and abstraction in por- trait sketching

    Itamar Berger, Ariel Shamir, Moshe Mahler, Elizabeth Carter, and Jessica Hodgins. Style and abstraction in por- trait sketching. ACM Trans. Graph., 32(4), 2013. 2

  5. [5]

    Improving image generation with better captions

    James Betker, Gabriel Goh, Li Jing, † TimBrooks, Jian- feng Wang, Linjie Li, † LongOuyang, † JuntangZhuang, † JoyceLee, † YufeiGuo, † WesamManassra, † PrafullaD- hariwal, † CaseyChu, † YunxinJiao, and Aditya Ramesh. Improving image generation with better captions. 3

  6. [6]

    Hospedales, Tao Xiang, Yu- lia Gryaditskaya, and Yi-Zhe Song

    Ayan Kumar Bhunia, Ayan Das, Umar Riaz Muhammad, Yongxin Yang, Timothy M. Hospedales, Tao Xiang, Yu- lia Gryaditskaya, and Yi-Zhe Song. Pixelor: a competitive sketching ai agent. so you think you can sketch? ACM Trans. Graph., 39:166:1–166:15, 2020. 1, 3

  7. [7]

    Doodleformer: Creative sketch drawing with transformers

    Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, Jorma Laaksonen, and Michael Felsberg. Doodleformer: Creative sketch drawing with transformers. ECCV, 2022. 2, 3

  8. [8]

    Doodleformer: Creative sketch drawing with transformers

    Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, Jorma Laaksonen, and Michael Felsberg. Doodleformer: Creative sketch drawing with transformers. In European Conference on Computer Vision, pages 338–355. Springer, 2022. 1

Show all 140 references
  1. [9]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sand- hini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Je...

  2. [10]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sand- hini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Je...

  3. [11]

    Sparks of artificial general intelligence: Early experiments with gpt-

    S ´ebastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-

  4. [12]

    arXiv preprint arXiv:2303.12712, 2023. 3

  5. [13]

    Delving into LLMs’ visual understanding ability using SVG to bridge image and text, 2024

    Mu Cai, Zeyi Huang, Yuheng Li, Haohan Wang, and Yong Jae Lee. Delving into LLMs’ visual understanding ability using SVG to bridge image and text, 2024. 2, 3

  6. [14]

    A computational approach to edge detection

    John Canny. A computational approach to edge detection. IEEE Transactions on pattern analysis and machine intelli- gence, (6):679–698, 1986. 2

  7. [15]

    Deepsvg: A hierarchical generative network for vector graphics animation, 2020

    Alexandre Carlier, Martin Danelljan, Alexandre Alahi, and Radu Timofte. Deepsvg: A hierarchical generative network for vector graphics animation, 2020. 3

  8. [16]

    On the util- ity of learning about humans for human-ai coordination

    Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan. On the util- ity of learning about humans for human-ai coordination. In Advances in Neural Information Processing Systems . Cur- ran Associates, Inc., 2019. 13

  9. [17]

    Learning to generate line drawings that convey geometry and seman- tics

    Caroline Chan, Fr ´edo Durand, and Phillip Isola. Learning to generate line drawings that convey geometry and seman- tics. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 7915–7925,

  10. [18]

    Vi- sual chain-of-thought prompting for knowledge-based vi- sual reasoning

    Zhenfang Chen, Qinhong Zhou, Yikang Shen, Yining Hong, Zhiqing Sun, Dan Gutfreund, and Chuang Gan. Vi- sual chain-of-thought prompting for knowledge-based vi- sual reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 1254–1262, 2024. 3

  11. [19]

    3doodle: Compact abstraction of objects with 3d strokes

    Changwoon Choi, Jaeah Lee, Jaesik Park, and Young Min Kim. 3doodle: Compact abstraction of objects with 3d strokes. ACM Trans. Graph., 43(4), 2024. 3

  12. [20]

    9 Visionllama: A unified llama backbone for vision tasks,

    Xiangxiang Chu, Jianlin Su, Bo Zhang, and Chunhua Shen. 9 Visionllama: A unified llama backbone for vision tasks,

  13. [21]

    B ´eziersketch: A generative model for scalable vector sketches

    Ayan Das, Yongxin Yang, Timothy Hospedales, Tao Xiang, and Yi-Zhe Song. B ´eziersketch: A generative model for scalable vector sketches. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVI 16, pages 632–647. Springer,

  14. [22]

    Drawing ap- prentice: An enactive co-creative agent for artistic collab- oration

    Nicholas Davis, Chih-PIn Hsiao, Kunwar Yashraj Singh, Lisa Li, Sanat Moningi, and Brian Magerko. Drawing ap- prentice: An enactive co-creative agent for artistic collab- oration. In Proceedings of the 2015 ACM SIGCHI Con- ference on Creativity and Cognition , page 185–186, New...

  15. [23]

    BERT: Pre-training of deep bidirectional trans- formers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional trans- formers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the As- sociation for Computational Linguistics: Human L...

  16. [24]

    Towards multimodal in-context learning for vi- sion & language models

    Sivan Doveh, Shaked Perek, M Jehanzeb Mirza, Wei Lin, Amit Alfassy, Assaf Arbelle, Shimon Ullman, and Leonid Karlinsky. Towards multimodal in-context learning for vi- sion & language models. arXiv preprint arXiv:2403.12736,

  17. [25]

    The llama 3 herd of models

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. 4

  18. [26]

    How do hu- mans sketch objects? ACM Trans

    Mathias Eitz, James Hays, and Marc Alexa. How do hu- mans sketch objects? ACM Trans. Graph. , 31(4), 2012. 2

  19. [27]

    Frank, Matthew Groh, Hope Schroeder, Amy Smith, Memo Akten, Jessica Fjeld, Hany Farid, Neil Leach, Alex Pentland, and Olga Russakovsky

    Ziv Epstein, Aaron Hertzmann, Laura Mariah Herman, Robert Mahari, Morgan R. Frank, Matthew Groh, Hope Schroeder, Amy Smith, Memo Akten, Jessica Fjeld, Hany Farid, Neil Leach, Alex Pentland, and Olga Russakovsky. Art and the science of generative ai. Science, 380:1110 – 1111, 2023. 1

  20. [28]

    clip-vit-large-patch14

    Hugging Face. clip-vit-large-patch14. https : / / huggingface . co / openai / clip - vit - large - patch14. 1

  21. [29]

    Bainbridge, Rebecca Chamberlain, and Jeffrey D

    Judy Fan, Wilma A. Bainbridge, Rebecca Chamberlain, and Jeffrey D. Wammes. Drawing as a versatile cognitive tool. Nature Reviews Psychology, 2:556 – 568, 2023. 2

  22. [30]

    Fan, Monica Dinculescu, and David Ha

    Judith E. Fan, Monica Dinculescu, and David Ha. collab- draw: An environment for collaborative sketching with an artificial agent. In Proceedings of the 2019 Conference on Creativity and Cognition , page 556–561, New York, NY , USA, 2019. Association for Computing Machinery. 3, 7

  23. [31]

    Fan, Robert D

    Judith E. Fan, Robert D. Hawkins, Mike Wu, and Noah D. Goodman. Pragmatic Inference and Visual Abstraction En- able Contextual Flexibility During Visual Communication. Computational Brain & Behavior, 3(1):86–101, 2020. 3

  24. [32]

    Drawing as a versatile cognitive tool

    Judith E Fan, Wilma A Bainbridge, Rebecca Chamberlain, and Jeffrey D Wammes. Drawing as a versatile cognitive tool. Nature Reviews Psychology, 2(9):556–568, 2023. 1

  25. [33]

    Creating drawings enhances learning by teaching

    Logan Fiorella and Shelbi Kuhlmann. Creating drawings enhances learning by teaching. Journal of Educational Psy- chology, 112(4):811, 2020. 1

  26. [34]

    Cogsketch: Sketch understanding for cognitive science research and for education

    Kenneth Forbus, Jeffrey Usher, Andrew Lovett, Kate Lock- wood, and Jon Wetzel. Cogsketch: Sketch understanding for cognitive science research and for education. Topics in Cognitive Science, 3(4):648–666, 2011. 1

  27. [35]

    Kevin Frans, L. B. Soros, and Olaf Witkowski. Clipdraw: exploring text-to-drawing synthesis through language- image encoders. In Proceedings of the 36th International Conference on Neural Information Processing Systems , Red Hook, NY , USA, 2024. Curran Associates Inc. 1, 3

  28. [36]

    Blink: Multimodal large lan- guage models can see but not perceive

    Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna. Blink: Multimodal large lan- guage models can see but not perceive. arXiv preprint arXiv:2404.12390, 2024. 4

  29. [37]

    Breath- ing life into sketches using text-to-video priors

    Rinon Gal, Yael Vinker, Yuval Alaluf, Amit Bermano, Daniel Cohen-Or, Ariel Shamir, and Gal Chechik. Breath- ing life into sketches using text-to-video priors. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4325–4336, 2024. 3

  30. [38]

    Kulkarni, Igor Babuschkin, S

    Yaroslav Ganin, Tejas D. Kulkarni, Igor Babuschkin, S. M. Ali Eslami, and Oriol Vinyals. Synthesizing programs for images using reinforced adversarial learning. ArXiv, abs/1804.01118, 2018. 3

  31. [39]

    Computer-aided design as language

    Yaroslav Ganin, Sergey Bartunov, Yujia Li, Ethan Keller, and Stefano Saliceti. Computer-aided design as language. In Neural Information Processing Systems, 2021. 3

  32. [40]

    Sketchycoco: Image genera- tion from freehand scene sketches

    Chengying Gao, Qi Liu, Qi Xu, Limin Wang, Jianzhuang Liu, and Changqing Zou. Sketchycoco: Image genera- tion from freehand scene sketches. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5174–5183, 2020. 2

  33. [41]

    Foundations of representation: Where might graphical symbol systems come from? Cognitive Science, 31(6):961–987, 2007

    Simon Garrod, Nicolas Fay, John Lee, Jon Oberlander, and Tracy MacLeod. Foundations of representation: Where might graphical symbol systems come from? Cognitive Science, 31(6):961–987, 2007. 3

  34. [42]

    Creative sketch generation

    Songwei Ge, Vedanuj Goswami, C Lawrence Zitnick, and Devi Parikh. Creative sketch generation. arXiv preprint arXiv:2011.10039, 2020. 1

  35. [43]

    Creative sketch generation

    Songwei Ge, Vedanuj Goswami, Larry Zitnick, and Devi Parikh. Creative sketch generation. In International Con- ference on Learning Representations, 2021. 2, 6

  36. [44]

    Collaborative drawing on a shared digital canvas in elementary science education: The effects of script and task awareness sup- port

    Hannie Gijlers, Armin Weinberger, Alieke Mattia van Dijk, Lars Bollen, and Wouter van Joolingen. Collaborative drawing on a shared digital canvas in elementary science education: The effects of script and task awareness sup- port. International Journal of Computer-Supported Co...

  37. [45]

    Serial sketching: visual problem solving in designing

    Gabriela Goldschmidt. Serial sketching: visual problem solving in designing. Cybernetics and System, 23(2):191– 219, 1992. 1

  38. [46]

    Grosz and Sarit Kraus

    Barbara J. Grosz and Sarit Kraus. Collaborative plans for complex group action. Artificial Intelligence, 86(2):269– 357, 1996. 13

  39. [47]

    10 Opensketch: a richly-annotated dataset of product design sketches

    Yulia Gryaditskaya, Mark Sypesteyn, Jan Willem Hofti- jzer, Sylvia Pont, Fr ´edo Durand, and Adrien Bousseau. 10 Opensketch: a richly-annotated dataset of product design sketches. ACM Trans. Graph., 38(6), 2019. 2

  40. [48]

    Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

    Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou. Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision...

  41. [49]

    A neural representation of sketch drawings

    David Ha and Douglas Eck. A neural representation of sketch drawings. CoRR, abs/1704.03477, 2017. 1, 2, 3, 6, 7

  42. [50]

    Xiaoyan Zhang, Xin Chen, Xu Yang, Zhibin Wang, Gang Yu, Bin Fu, and Hanwang Zhang

    Yucheng Han, China. Xiaoyan Zhang, Xin Chen, Xu Yang, Zhibin Wang, Gang Yu, Bin Fu, and Hanwang Zhang. Chartllama: A multimodal llm for chart understanding and generation. ArXiv, abs/2311.16483, 2023. 2, 3

  43. [51]

    Hawkins, Megumi Sano, Noah D

    Robert D. Hawkins, Megumi Sano, Noah D. Goodman, and Judith E. Fan. Visual resemblance and communicative context constrain the emergence of graphical conventions,

  44. [52]

    Visual sketchpad: Sketching as a visual chain of thought for multimodal language models

    Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Osten- dorf, Luke Zettlemoyer, Noah A Smith, and Ranjay Kr- ishna. Visual sketchpad: Sketching as a visual chain of thought for multimodal language models. arXiv preprint arXiv:2406.09403, 2024. 3

  45. [53]

    A collaborative, interactive and context-aware draw- ing agent for co-creative design

    Francisco Javier Ibarrola, Tomas Lawton, and Kazjon Grace. A collaborative, interactive and context-aware draw- ing agent for co-creative design. IEEE Transactions on Vi- sualization and Computer Graphics, 30:5525–5537, 2022. 3

  46. [54]

    Vectorfusion: Text-to-svg by abstracting pixel-based diffusion models

    Ajay Jain, Amber Xie, and Pieter Abbeel. Vectorfusion: Text-to-svg by abstracting pixel-based diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 1911–1920, 2023. 1, 3

  47. [55]

    The Quick, Draw! - A.I

    Jongejan Jonas, Rowley Henry, Kawashima Takashi, Kim Jongmin, and Fox-Gieg Nick. The Quick, Draw! - A.I. Experiment, 2016. 2, 3, 6, 7, 8

  48. [56]

    Drawing theories apart: The dispersion of Feynman diagrams in postwar physics

    David Kaiser. Drawing theories apart: The dispersion of Feynman diagrams in postwar physics . University of Chicago Press, 2019. 1

  49. [58]

    Creative sketching part- ner: an analysis of human-ai co-creativity

    Pegah Karimi, Jeba Rezwana, Safat Siddiqui, Mary Lou Maher, and Nasrin Dehbozorgi. Creative sketching part- ner: an analysis of human-ai co-creativity. In Proceedings of the 25th International Conference on Intelligent User In- terfaces, page 221–230, New York, NY , USA, 2020....

  50. [59]

    Chapter three - psychological research on joint action: The- ory and data

    G ¨unther Knoblich, Stephen Butterfill, and Natalie Sebanz. Chapter three - psychological research on joint action: The- ory and data. In Advances in Research and Theory , pages 59–101. Academic Press, 2011. 13

  51. [60]

    Cairosvg

    Kozea. Cairosvg. https://cairosvg.org/, 2023. 1

  52. [61]

    Graphic thinking for architects and designers

    Paul Laseau. Graphic thinking for architects and designers. John Wiley & Sons, 2000. 2

  53. [62]

    When is a tool a tool? user perceptions of system agency in human–ai co-creative drawing

    Tomas Lawton, Kazjon Grace, and Francisco J Ibarrola. When is a tool a tool? user perceptions of system agency in human–ai co-creative drawing. In Proceedings of the 2023 ACM Designing Interactive Systems Conference, page 1978–1996, New York, NY , USA, 2023. Association for Co...

  54. [63]

    Drawing with reframer: Emergence and control in co-creative&nbsp;ai

    Tomas Lawton, Francisco J Ibarrola, Dan Ventura, and Kazjon Grace. Drawing with reframer: Emergence and control in co-creative&nbsp;ai. In Proceedings of the 28th International Conference on Intelligent User Interfaces , page 264–277, New York, NY , USA, 2023. Association for ...

  55. [64]

    Lawrence Zitnick, and Michael F

    Yong Jae Lee, C. Lawrence Zitnick, and Michael F. Cohen. Shadowdraw: real-time user guidance for freehand draw- ing. In ACM SIGGRAPH 2011 Papers , New York, NY , USA, 2011. Association for Computing Machinery. 3

  56. [65]

    BLIP: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. BLIP: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation. In Pro- ceedings of the 39th International Conference on Machine Learning, pages 12888–12900. PMLR, 2022. 2, 3

  57. [66]

    Photo-sketching: Inferring contour draw- ings from images

    Mengtian Li, Zhe Lin, Radomir Mech, Ersin Yumer, and Deva Ramanan. Photo-sketching: Inferring contour draw- ings from images. In 2019 IEEE Winter Conference on Ap- plications of Computer Vision (WACV), pages 1403–1412. IEEE, 2019. 2

  58. [67]

    Differentiable vector graphics rasterization for editing and learning

    Tzu-Mao Li, Michal Luk ´aˇc, Gharbi Micha¨el, and Jonathan Ragan-Kelley. Differentiable vector graphics rasterization for editing and learning. ACM Trans. Graph. (Proc. SIG- GRAPH Asia), 39(6):193:1–193:15, 2020. 1, 3

  59. [68]

    Hospedales, and Shaogang Gong

    Yi Li, Yi-Zhe Song, Timothy M. Hospedales, and Shaogang Gong. Free-hand sketch synthesis with deformable stroke models. CoRR, abs/1510.02644, 2015. 1, 2

  60. [69]

    Im2pencil: Controllable pen- cil illustration from photographs

    Yijun Li, Chen Fang, Aaron Hertzmann, Eli Shechtman, and Ming-Hsuan Yang. Im2pencil: Controllable pen- cil illustration from photographs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1525–1534, 2019. 2

  61. [70]

    Hangyu Lin, Yanwei Fu, Yu-Gang Jiang, and X. Xue. Sketch-bert: Learning sketch bidirectional encoder repre- sentation from transformers by self-supervised learning of sketch gestalt. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6757–6766,

  62. [71]

    Neural strokes: Stylized line draw- ing of 3d shapes

    Difan Liu, Matthew Fisher, Aaron Hertzmann, and Evan- gelos Kalogerakis. Neural strokes: Stylized line draw- ing of 3d shapes. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision, pages 14204– 14213, 2021. 2

  63. [72]

    Sketchgan: Joint sketch completion and recognition with generative adversarial net- work

    Fang Liu, Xiaoming Deng, Yu-Kun Lai, Yong-Jin Liu, Cuixia Ma, and Hongan Wang. Sketchgan: Joint sketch completion and recognition with generative adversarial net- work. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019. 2 11

  64. [73]

    Sketchgan: Joint sketch completion and recognition with generative adversarial net- work

    Fang Liu, Xiaoming Deng, Yu-Kun Lai, Yong-Jin Liu, Cuixia Ma, and Hongan Wang. Sketchgan: Joint sketch completion and recognition with generative adversarial net- work. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 5830– 5839, 2019. 2

  65. [74]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Advances in Neural In- formation Processing Systems, pages 34892–34916. Curran Associates, Inc., 2023. 2, 3

  66. [75]

    Parallel developmental changes in chil- dren’s production and recognition of line drawings of visual concepts

    Bria Long, Judith Fan, Holly Huey, Zixian Chai, and Michael Frank. Parallel developmental changes in chil- dren’s production and recognition of line drawings of visual concepts. Nature Communications, 15, 2024. 6

  67. [76]

    Openeqa: Embodied question answering in the era of foun- dation models

    Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Sil- wal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma, Vincent Berges, Shiqi Zhang, Pulkit Agrawal, Yonatan Bisk, Dhruv Batr...

  68. [77]

    McCarthy, Robert D

    William P. McCarthy, Robert D. Hawkins, Haoliang Wang, Cameron Holdaway, and Judith E. Fan. Learning to com- municate about shared procedural abstractions, 2021. 13

  69. [78]

    McCarthy, Justin Matejka, Karl D.D

    William P. McCarthy, Justin Matejka, Karl D.D. Willis, Ju- dith E. Fan, and Yewen Pu. Communicating design intent using drawing and text. In Proceedings of the 16th Confer- ence on Creativity & Cognition, page 512–519, New York, NY , USA, 2024. Association for Computing Machinery. 3

  70. [79]

    Unsupervised doodling and painting with improved spiral

    John FJ Mellor, Eunbyung Park, Yaroslav Ganin, Igor Babuschkin, Tejas Kulkarni, Dan Rosenbaum, Andy Bal- lard, Theophane Weber, Oriol Vinyals, and SM Eslami. Unsupervised doodling and painting with improved spiral. arXiv preprint arXiv:1910.01007, 2019. 3

  71. [80]

    Learning to draw: Emer- gent communication through sketching

    Daniela Mihai and Jonathon Hare. Learning to draw: Emer- gent communication through sketching. Advances in Neu- ral Information Processing Systems , 34:7153–7166, 2021. 3

  72. [81]

    Compositional chain-of-thought prompt- ing for large multimodal models

    Chancharik Mitra, Brandon Huang, Trevor Darrell, and Roei Herzig. Compositional chain-of-thought prompt- ing for large multimodal models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14420–14431, 2024. 3

  73. [82]

    Seva: Leveraging sketches to evaluate alignment between human and machine visual abstraction

    Kushin Mukherjee, Holly Huey, Xuanchen Lu, Yael Vinker, Rio Aguina-Kang, Ariel Shamir, and Judith Fan. Seva: Leveraging sketches to evaluate alignment between human and machine visual abstraction. In Advances in Neural In- formation Processing Systems, 2023. 2

  74. [83]

    Observing by hand: sketching the nebu- lae in the nineteenth century

    Omar W Nasim. Observing by hand: sketching the nebu- lae in the nineteenth century. University of Chicago Press,

  75. [84]

    I lead, you help but only with enough details: Understanding user experience of co-creation with artificial intelligence

    Changhoon Oh, Jungwoo Song, Jinhan Choi, Seonghyeon Kim, Sungwoo Lee, and Bongwon Suh. I lead, you help but only with enough details: Understanding user experience of co-creation with artificial intelligence. In Proceedings of the 2018 CHI Conference on Human Factors in Comput...

  76. [85]

    Gpt-4 technical report, 2024

    OpenAI. Gpt-4 technical report, 2024. 2, 3, 4, 6

  77. [86]

    Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rom- bach

    Dustin Podell, Zion English, Kyle Lacey, A. Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rom- bach. Sdxl: Improving latent diffusion models for high- resolution image synthesis. ArXiv, abs/2307.01952, 2023. 2, 3

  78. [87]

    Sketchlattice: Latticed representation for sketch manipulation

    Yonggang Qi, Guoyao Su, Pinaki Nath Chowdhury, Mingkang Li, and Yi-Zhe Song. Sketchlattice: Latticed representation for sketch manipulation. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 953–961, 2021. 2

  79. [88]

    Emergent graphical conven- tions in a visual communication game

    Shuwen Qiu, Sirui Xie, Lifeng Fan, Tao Gao, Song- Chun Zhu, and Yixin Zhu. Emergent graphical conven- tions in a visual communication game. arXiv preprint arXiv:2111.14210, 2021. 3

  80. [89]

    Learning transferable vi- sual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable vi- sual models from natural language supervision. CoRR, abs/2103.000...

  81. [90]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.J. Mach. Learn. Res., 21(1), 2020. 3

  82. [91]

    Hierarchical text-conditional image generation with clip latents

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 2, 3

  83. [92]

    Col- lomosse, and Moacir Antonelli Ponti

    Leo Sampaio Ferraz Ribeiro, Tu Bui, John P. Col- lomosse, and Moacir Antonelli Ponti. Sketchformer: Transformer-based representation for sketched structure. 2020 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 14141–14150, 2020. 3

  84. [93]

    High-resolution image synthesis with latent diffusion models, 2022

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models, 2022. 1, 2, 3, 6

  85. [94]

    Visual chain of thought: bridging logical gaps with multimodal infill- ings

    Daniel Rose, Vaishnavi Himakunthala, Andy Ouyang, Ryan He, Alex Mei, Yujie Lu, Michael Saxon, Chinmay Sonar, Diba Mirza, and William Yang Wang. Visual chain of thought: bridging logical gaps with multimodal infill- ings. arXiv preprint arXiv:2305.02317, 2023. 3

  86. [95]

    Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. Photorealistic text-to- image diffusion mo...

  87. [96]

    The sketchy database: Learning to retrieve badly drawn bunnies

    Patsorn Sangkloy, Nathan Burnell, Cusuh Ham, and James Hays. The sketchy database: Learning to retrieve badly drawn bunnies. ACM Trans. Graph., 35(4), 2016. 2 12

  88. [97]

    Frida: A collaborative robot painter with a differentiable, real2sim2real planning environment

    Peter Schaldenbrand, James McCann, and Jean Oh. Frida: A collaborative robot painter with a differentiable, real2sim2real planning environment. In 2023 IEEE Inter- national Conference on Robotics and Automation (ICRA) , pages 11712–11718, 2023. 3

  89. [98]

    Cofrida: Self-supervised fine-tuning for human-robot co-painting

    Peter Schaldenbrand, Gaurav Parmar, Jun-Yan Zhu, James McCann, and Jean Oh. Cofrida: Self-supervised fine-tuning for human-robot co-painting. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE,

  90. [99]

    The reflective practitioner: How professionals think in action, 1986

    Donald A Schon and Vincent DeSanctis. The reflective practitioner: How professionals think in action, 1986. 2

  91. [100]

    Laion-5b: An open large-scale dataset for train- ing next generation image-text models, 2022

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Je- nia Jitsev. L...

  92. [101]

    A multimodal automated interpretability agent

    Tamar Rott Shaham, Sarah Schwettmann, Franklin Wang, Achyuta Rajaram, Evan Hernandez, Jacob Andreas, and Antonio Torralba. A multimodal automated interpretability agent. In Forty-first International Conference on Machine Learning, 2024. 3

  93. [102]

    Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning

    Hao Shao, Shengju Qian, Han Xiao, Guanglu Song, Zhuo- fan Zong, Letian Wang, Yu Liu, and Hongsheng Li. Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning. arXiv preprint arXiv:2403.16999, 2024. 3

  94. [103]

    A vision check-up for language models

    Pratyusha Sharma, Tamar Rott Shaham, Manel Baradad, Stephanie Fu, Adrian Rodriguez-Munoz, Shivam Duggal, Phillip Isola, and Antonio Torralba. A vision check-up for language models. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , pages 1...

  95. [104]

    Learning to sketch with shortcut cycle consistency, 2018

    Jifei Song, Kaiyue Pang, Yi-Zhe Song, Tao Xiang, and Tim- othy Hospedales. Learning to sketch with shortcut cycle consistency, 2018. 2

  96. [105]

    Ad hoc autonomous agent teams: Collaboration without pre-coordination

    Peter Stone, Gal Kaminka, Sarit Kraus, and Jeffrey Rosen- schein. Ad hoc autonomous agent teams: Collaboration without pre-coordination. Proceedings of the AAAI Con- ference on Artificial Intelligence , 24(1):1504–1509, 2010. 13

  97. [106]

    Sketchhealer: A graph-to- sequence network for recreating partial human sketches

    Guoyao Su, Yonggang Qi, Kaiyue Pang, Jie Yang, Yi-Zhe Song, and CVSSP SketchX. Sketchhealer: A graph-to- sequence network for recreating partial human sketches. In BMVC, page 5, 2020. 2

  98. [107]

    Exploring effective factors for improving vi- sual in-context learning

    Yanpeng Sun, Qiang Chen, Jian Wang, Jingdong Wang, and Zechao Li. Exploring effective factors for improving vi- sual in-context learning. arXiv preprint arXiv:2304.04748,

  99. [108]

    Sutherland

    Ivan E. Sutherland. Sketchpad—a man-machine graphi- cal communication system, page 391–408. Association for Computing Machinery, New York, NY , USA, 1998. 3

  100. [109]

    Gemini: A family of highly capable multi- modal models, 2024

    Gemini Team. Gemini: A family of highly capable multi- modal models, 2024. 2, 3

  101. [110]

    Design ideation with ai - sketching, thinking and talking with generative machine learning models

    Jakob Tholander and Martin Jonsson. Design ideation with ai - sketching, thinking and talking with generative machine learning models. In Proceedings of the 2023 ACM Design- ing Interactive Systems Conference, page 1930–1940, New York, NY , USA, 2023. Association for Computing...

  102. [111]

    Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024

    Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024. 4

  103. [112]

    Llama: Open and efficient foundation language models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth ´ee Lacroix, Bap- tiste Rozi `ere, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation la...

  104. [113]

    What do sketches say about thinking? AAAI Spring Symp

    Barbara Tversky. What do sketches say about thinking? AAAI Spring Symp. Sketch Understanding Worksh. , 2002. 2

  105. [114]

    Visualizing thought

    Barbara Tversky. Visualizing thought. In Handbook of hu- man centric visualization, pages 3–40. Springer, 2013. 1

  106. [115]

    Sketches for design and design of sketches

    Barbara Tversky, Masaki Suwa, Maneesh Agrawala, Julie Heiser, Chris Stolte, Pat Hanrahan, Doantam Phan, Jeff Klingner, Marie-Paule Daniel, Paul Lee, et al. Sketches for design and design of sketches. Human Behaviour in Design: Individuals, Teams, Tools, pages 79–86, 2003. 1

  107. [116]

    Drawing to reason and learn in science

    Russell Tytler, Vaughan Prain, George Aranda, Joseph Fer- guson, and Radhika Gorur. Drawing to reason and learn in science. Journal of Research in Science Teaching , 57(2): 209–231, 2020. 3

  108. [117]

    Bo, Ro- man Christian Bachmann, Amit Haim Bermano, Daniel Cohen-Or, Amir Zamir, and Ariel Shamir

    Yael Vinker, Ehsan Pajouheshgar, Jessica Y . Bo, Ro- man Christian Bachmann, Amit Haim Bermano, Daniel Cohen-Or, Amir Zamir, and Ariel Shamir. Clipasso: Semantically-aware object sketching. ACM Trans. Graph., 41(4), 2022. 1, 3, 6

  109. [118]

    Clipascene: Scene sketching with different types and levels of abstraction

    Yael Vinker, Yuval Alaluf, Daniel Cohen-Or, and Ariel Shamir. Clipascene: Scene sketching with different types and levels of abstraction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages 4146–4156, 2023. 3, 6

  110. [119]

    Contextseg: Sketch se- mantic segmentation by querying the context with attention

    Jiawei Wang and Changjian Li. Contextseg: Sketch se- mantic segmentation by querying the context with attention. arXiv preprint arXiv:2311.16682, 2023. 6

  111. [120]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V . Le, and Denny Zhou. Chain-of-thought prompting elicits reason- ing in large language models. Red Hook, NY , USA, 2024. Curran Associates Inc. 2, 5

  112. [121]

    Holger Winnem ¨oller, Jan Eric Kyprianidis, and Sven C. Olsen. Xdog: An extended difference-of-gaussians com- pendium including advanced image stylization. Comput. Graph., 36:740–753, 2012. 2

  113. [122]

    Scalable Vector Graphics (SVG), 1999

    World Wide Web Consortium (W3C). Scalable Vector Graphics (SVG), 1999. 3

  114. [123]

    Visual chatgpt: Talk- ing, drawing and editing with visual foundation models

    Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. Visual chatgpt: Talk- ing, drawing and editing with visual foundation models. arXiv preprint arXiv:2303.04671, 2023. 3

  115. [124]

    Iconshop: Text-guided vector icon synthesis with autoregressive trans- 13 formers

    Rong Wu, Wanchao Su, Kede Ma, and Jing Liao. Iconshop: Text-guided vector icon synthesis with autoregressive trans- 13 formers. ACM Transactions on Graphics (TOG), 42:1 – 14,

  116. [125]

    Differsketching: How differ- ently do people sketch 3d objects? ACM Transactions on Graphics (Proceedings of ACM SIGGRAPH Asia 2022), 41 (4):1–16, 2022

    Chufeng Xiao, Wanchao Su, Jing Liao, Zhouhui Lian, Yi- Zhe Song, and Hongbo Fu. Differsketching: How differ- ently do people sketch 3d objects? ACM Transactions on Graphics (Proceedings of ACM SIGGRAPH Asia 2022), 41 (4):1–16, 2022. 2

  117. [126]

    Holistically-nested edge de- tection

    Saining Xie and Zhuowen Tu. Holistically-nested edge de- tection. In Proceedings of the IEEE international confer- ence on computer vision, pages 1395–1403, 2015. 2

  118. [127]

    Diffsketcher: Text guided vec- tor sketch synthesis through latent diffusion models

    XiMing Xing, Chuang Wang, Haitao Zhou, Jing Zhang, Qian Yu, and Dong Xu. Diffsketcher: Text guided vec- tor sketch synthesis through latent diffusion models. In Advances in Neural Information Processing Systems, pages 15869–15889. Curran Associates, Inc., 2023. 3, 6

  119. [128]

    Svgdreamer: Text guided svg generation with diffusion model

    Ximing Xing, Haitao Zhou, Chuang Wang, Jing Zhang, Dong Xu, and Qian Yu. Svgdreamer: Text guided svg generation with diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4546–4555, 2024. 3, 6, 7

  120. [129]

    Hospedales, Qiyue Yin, Yi-Zhe Song, Tao Xiang, and Liang Wang

    Peng Xu, Timothy M. Hospedales, Qiyue Yin, Yi-Zhe Song, Tao Xiang, and Liang Wang. Deep learning for free- hand sketch: A survey and a toolbox, 2020. 2

  121. [130]

    The dawn of lmms: Preliminary explorations with gpt-4v(ision)

    Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. The dawn of lmms: Preliminary explorations with gpt-4v(ision). ArXiv, abs/2309.17421, 2023. 2

  122. [131]

    Idea2img: Iterative self-refinement with gpt-4v (ision) for automatic image design and generation

    Zhengyuan Yang, Jianfeng Wang, Linjie Li, Kevin Lin, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. Idea2img: Iterative self-refinement with gpt-4v (ision) for automatic image design and generation. arXiv preprint arXiv:2310.08541, 2023. 3

  123. [132]

    Apdrawinggan: Generating artistic portrait drawings from face photos with hierarchical gans

    Ran Yi, Yong-Jin Liu, Yu-Kun Lai, and Paul L Rosin. Apdrawinggan: Generating artistic portrait drawings from face photos with hierarchical gans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10743–10752, 2019. 2

  124. [133]

    Un- paired portrait drawing generation via asymmetric cycle mapping

    Ran Yi, Yong-Jin Liu, Yu-Kun Lai, and Paul L Rosin. Un- paired portrait drawing generation via asymmetric cycle mapping. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8217– 8225, 2020. 2

  125. [134]

    Zhang, Weijie Wang, Paul Pangaro, Nikolas Marte- laro, and Daragh Byrne

    C. Zhang, Weijie Wang, Paul Pangaro, Nikolas Marte- laro, and Daragh Byrne. Generative image ai using design sketches as input: Opportunities and challenges. Proceed- ings of the 15th Conference on Creativity and Cognition ,

  126. [135]

    Instruct me more! random prompt- ing for visual in-context learning

    Jiahao Zhang, Bowen Wang, Liangzhi Li, Yuta Nakashima, and Hajime Nagahara. Instruct me more! random prompt- ing for visual in-context learning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2597–2606, 2024. 3

  127. [136]

    Text-to- vector generation with neural path representation

    Peiying Zhang, Nanxuan Zhao, and Jing Liao. Text-to- vector generation with neural path representation. ACM Trans. Graph., 43(4), 2024. 3

  128. [137]

    What makes good examples for visual in-context learning? Ad- vances in Neural Information Processing Systems , 36: 17773–17794, 2023

    Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. What makes good examples for visual in-context learning? Ad- vances in Neural Information Processing Systems , 36: 17773–17794, 2023. 3

  129. [138]

    Stroke-based semantic segmentation for scene-level free-hand sketches

    Zhengming Zhang, Xiaoming Deng, Jinyao Li, Yukun Lai, Cuixia Ma, Yongjin Liu, and Hongan Wang. Stroke-based semantic segmentation for scene-level free-hand sketches. Vis. Comput., 39(12):6309–6321, 2022. 6

  130. [139]

    Multimodal chain-of- thought reasoning in language models

    Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. Multimodal chain-of- thought reasoning in language models. arXiv preprint arXiv:2302.00923, 2023. 3

  131. [140]

    Creativeseg: Semantic seg- mentation of creative sketches

    Yixiao Zheng, Kaiyue Pang, Ayan Das, Dongliang Chang, Yi-Zhe Song, and Zhanyu Ma. Creativeseg: Semantic seg- mentation of creative sketches. IEEE Transactions on Im- age Processing, 33:2266–2278, 2024. 6

  132. [141]

    shark”, which was often misclassified as a “fish

    Tao Zhou, Chen Fang, Zhaowen Wang, Jimei Yang, Byung- moon Kim, Zhili Chen, Jonathan Brandt, and Demetri Ter- zopoulos. Learning to sketch with deep q networks and demonstrated strokes. ArXiv, abs/1810.05977, 2018. 2, 3 14 SketchAgent: Language-Driven Sequential Sketch Generat...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.