REVIEW 4 major objections 5 minor 1 cited by
SketchAgent: Language-Driven Sequential Sketch Generation
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A frozen multimodal LLM can draw sketches stroke-by-stroke using only a numbered grid and a few examples.
desk verdict A genuinely new sketching representation for frozen multimodal LLMs, with honest experiments—but the 'wide range' claim is overstated; the evidence supports a moderate set of simple iconic concepts. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the sketching language over a numbered grid canvas. The canvas is a 50×50 grid whose cells carry labels like x2y8, giving the text-only LLM a concrete coordinate system to reference; the paper includes a demonstration that a multimodal LLM which can name the right line to draw nevertheless fails to execute it with pixel coordinates, motivating the grid. A stroke is written as a list of grid coordinates plus a matching list of t-values, meaning the points are treated as samples along a curve rather than control points, and the system fits a cubic Bézier curve to them with a least-squares solve, splitting recursively when the fit error is large. In-context learning teaches the format and chain-of-thought prompts make the model plan stroke order in <thinking> tags; a stopping token </s{j}> allows a human to add strokes to the same canvas, which are converted back into the coordinate format by sampling the user's Bézier curve at its t-values.
What would settle it
Re-run the 500-sketch, 50-category CLIP benchmark with the identical prompts and backbone but with the grid numbers erased from the canvas image (cells left blank), and compare Top-1 accuracy against the reported 0.23; if accuracy does not drop substantially, the numbered-grid spatial scaffolding is not the mechanism that makes the method work.
Extended reading notes
Core claim
SketchAgent establishes that a frozen multimodal LLM, guided only by a system prompt, a user prompt with one worked example, and a numbered 50×50 grid canvas, can generate recognizable sketches of arbitrary textual concepts—including landmarks, scientific principles, and diagrams far beyond the 345 QuickDraw categories that bound trained sketch models. The model plans in <thinking> tags, emits strokes as coordinate sequences with t-values, and the system fits cubic Bézier curves to the sampled points by least squares, recursively splitting long curves. The same representation supports the whole interaction loop: the rendered canvas is fed back for chat-based edits, and a stopping token lets a human interleave their own strokes, which are re-sampled into the agent's coordinate format. Quantitative results on 500 sketches across 50 categories give CLIP zero-shot Top-1/Top-5 accuracy of 0.23/0.44 with the default backbone, approaching the 0.27/0.49 of human QuickDraw sketches under the same metric, and an ablation shows that removing the system prompt, chain-of-thought, or the complete in-context example each degrades accuracy.
Load-bearing premise
The load-bearing premise is that an off-the-shelf multimodal LLM can convert its semantic understanding of an object into a spatially coherent, ordered sequence of grid coordinates, given only a numbered canvas and a few in-context examples; the paper itself acknowledges that the agent often writes rich textual descriptions of parts yet struggles to turn them into effective drawing actions.
Editorial extensions
If this is right
- A general-purpose sketching agent can be assembled without collecting human drawing data or training a generative model; the capability is inherited from the backbone LLM, so future improvements in those models should transfer directly to sketch quality.
- Sketch editing becomes a natural conversational operation: because the canvas is part of the dialogue state, the agent can add, relocate, or annotate parts of an existing drawing in response to text, with 92% of tested editing prompts followed correctly.
- The stroke-by-stroke output carries semantic labels assigned by the model, so sketches come with part-level annotations as a byproduct, useful for analysis and dataset construction.
- Real-time collaborative sketching with a human partner is feasible: individual strokes take about 8 seconds in collaborative mode and a complete sketch about 20 seconds, matching the pace of human drawing.
- The approach is largely backbone-agnostic: it works with several closed commercial models and, with lower recognition scores, with a large open-weight model, suggesting the interface is portable rather than tied to one model.
Reading between the lines
- Editorial inference: the failure cases the paper lists (unicorn, human figures, letters and numbers) outline the boundary of the backbone's latent spatial competence; a natural testable prediction is that these failure modes shrink as multimodal LLMs improve, without any change to the sketching pipeline.
- Editorial inference: the numbered-grid protocol is a general recipe for eliciting spatial output from text-only models, so the same idea could be tested for diagram layout, floor-plan drafting, or GUI wireframes, where a grid plus in-context examples might replace fine-tuning.
- Editorial inference: the 92% editing accuracy was measured on a small hand-picked set; a broader stress test that varies object combinations and demands precise relative placement would clarify whether the agent reasons spatially or falls back on stereotyped layouts.
- Editorial inference: because the prompts were tuned on the default backbone, the reported gap between backbone models partly reflects prompt fit rather than raw capability, so the method's portability across models is probably understated.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SketchAgent, a training-free method for sequential sketch generation using a frozen multimodal large language model (Claude3.5-Sonnet). The agent is prompted with a grid-coordinate sketching language, in-context examples, and chain-of-thought instructions, and it outputs a sequence of strokes that are fitted to Bézier curves and rendered on a canvas. The paper claims that this approach supports text-conditioned sketch generation across a wide range of concepts, enables stroke-by-stroke sequential sketching with semantic stroke annotations, facilitates real-time human-agent collaborative sketching, and supports iterative chat-based editing. Evaluations include CLIP zero-shot classification on 500 sketches across 50 QuickDraw categories, a 2AFC human study comparing sketch human-likeness, a collaborative user study with 30 participants, an ablation study, and a small chat-editing study on 54 sketches. The supplementary provides additional qualitative results and detailed prompts.
Significance. If the central claims hold, the paper demonstrates a meaningful new capability: an off-the-shelf multimodal LLM, without any fine-tuning, can produce ordered, semantically labeled strokes that render into recognizable sketches and can participate in interactive drawing sessions. The method is simple, reproducible, and the authors commit to releasing code. The paper includes useful ablations and honest discussion of limitations, and the human studies add credibility beyond CLIP-based proxies. However, the strongest claim—"wide range of textual concepts"—is only partially supported by the evidence, and the editing and collaborative evaluations have important protocol gaps. The work is likely to be of interest to the sketch-generation and human-AI interaction communities, but the conclusions should be scaled to match the demonstrated scope.
major comments (4)
- [Sec. 5.1, Table 1, Supp. Figs. 21, 23] The claim in Sec. 1 that SketchAgent "can generate sketches across a wide range of textual concepts" is not established by the quantitative evidence. The average Top-1 recognition of 0.23 in Table 1 hides substantial per-category variability: the supplementary confusion matrix and top-recognized-class figures show many categories with recognition rates of 20–30% or lower, and the paper itself (Sec. 7, Fig. 14) acknowledges failures on unicorn, Frida Kahlo, and letters/numbers. The appendix also admits that some randomly selected concepts (Statue of Liberty, photosynthesis, pie chart) were unsuccessful. The claim should be narrowed to simple iconic objects, or the evaluation should be strengthened with per-category success thresholds or human recognition judgments on the full concept set.
- [Sec. 5.4, Chat-Based Sketch Editing] The editing evaluation reports that SketchAgent "correctly follows instructions 92% of the time" on 54 sketches, but the manuscript does not define what constitutes correct instruction-following, who made the judgment, how many raters were involved, or whether there was inter-rater agreement. Because interactive editing is one of the paper's three core contributions, this metric needs a transparent evaluation protocol to be load-bearing. The current description is insufficient for readers to assess the reliability of the 92% figure.
- [Sec. 5.3, Human-Agent Collaborative Sketching] The collaborative user study uses 8 concepts that were explicitly "selected based on the agent's demonstrated ability to draw them independently" (Sec. 5.3). This selection makes the study unsuitable for supporting the broader claim that SketchAgent can collaborate on arbitrary concepts. Additionally, the analysis that agent-only and user-only strokes have low CLIP recognition rates (Table 3) may conflate incompleteness with lack of contribution, since partial sketches are not necessarily expected to be recognizable. A human evaluation of partial sketches or a study with less favorably selected concepts would be needed to substantiate the collaboration claim.
- [Sec. 4, Method Overview] The paper repeatedly emphasizes that SketchAgent captures the "dynamic, evolving process" of sketching and "incorporates visual feedback" (Sec. 1). However, during a single sketch-generation turn, the model produces the entire coordinate sequence in one pass; the canvas is fed back only for later editing or collaborative turns. The sequential nature is primarily in the output representation and stroke ordering, not in the model's internal generation process. This distinction should be clearly stated, as the current wording overstates the mechanism.
minor comments (5)
- [Sec. 5, opening paragraph] There is a typo: "We demonstrate SketchAgent's capabil to generate" should read "capability".
- [Sec. 6, Table 2] The ablation claims that "all components contribute to the agent's full performance," but the difference between the full pipeline and the w/o System Prompt condition (0.23 vs 0.20 Top-1) is within the reported error bars. The authors should either report significance tests or temper the claim for the system-prompt component.
- [Sec. 5.1, Table 1] The table's last row label "Vis." is not explained in the text; a short caption indicating that it shows example sketches would improve clarity.
- [Sec. 5.4 and Supp. B.4] The editing prompt in the appendix says the model should "Describe the location of the added concepts first in <thinking> tags," but for the animals category the instructions (e.g., "Add a hat") contain no location; the paper should clarify whether the model is expected to infer placement in those cases and how that inference was scored.
- [Sec. 2, Related Work] The discussion of direct SVG prompting in Fig. 3 is useful, but there is no direct quantitative comparison to optimization-based sketch generation methods such as CLIPasso or DiffSketch on the same QuickDraw subset; adding such a comparison would help calibrate the reported numbers against prior art.
Circularity Check
No circularity: SketchAgent's outputs are produced by an off-the-shelf LLM from prompts and evaluated with external CLIP and human judgments; no claimed result is defined by its inputs.
full rationale
SketchAgent's claimed capability is produced by an off-the-shelf multimodal LLM responding to a system prompt, a user prompt, and a numbered canvas; the only numerical post-processing, least-squares Bézier fitting in Eq. (2), is a rendering operation that does not define the semantic content of the sketch. No parameters are fitted to the CLIP recognition results or to the human 2AFC judgments, so the reported numbers are measurements rather than quantities forced by construction. The quantitative evaluation uses an external CLIP zero-shot classifier and human participants, with human QuickDraw sketches as baselines, and the recognition metric is not a re-expression of the method's inputs. The paper's self-citations (e.g., [116, 117, 126, 127] for the CLIP evaluation protocol) are methodological references rather than load-bearing justifications of the central claim. The supplement's admission that prompts were optimized for Claude3.5-Sonnet is a disclosed selection effect relevant to generalization, not a circular step, and the limitations in Sec. 7 concede unrecognizable outputs, which is inconsistent with a forced or self-fulfilling derivation. Concerns about the strength of the "wide range" claim are about evidence sufficiency, not circularity. Hence there is no circular step in the claimed derivation chain.
Assumptions & free parameters
free parameters (1)
- grid resolution =
50x50 cells
assumptions (3)
- domain assumption A multimodal LLM's pretrained visual priors are sufficient to plan and execute coherent sketches from text prompts when provided a grid and in-context examples.
- domain assumption CLIP zero-shot classification accuracy is a valid proxy for sketch recognizability and semantic fidelity.
- standard math Least-squares fitting of cubic Bezier curves to sampled grid points yields visually smooth, natural strokes.
Cite this review
Pith. "Pith review of SketchAgent: Language-Driven Sequential Sketch Generation." pith.science (2026). https://pith.science/paper/KRHHZWSZ
@misc{pith2026241117673,
author = {Pith},
title = {Pith review of: SketchAgent: Language-Driven Sequential Sketch Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/KRHHZWSZ}},
note = {Machine review of arXiv:2411.17673}
}
read the original abstract
Sketching serves as a versatile tool for externalizing ideas, enabling rapid exploration and visual communication that spans various disciplines. While artificial systems have driven substantial advances in content creation and human-computer interaction, capturing the dynamic and abstract nature of human sketching remains challenging. In this work, we introduce SketchAgent, a language-driven, sequential sketch generation method that enables users to create, modify, and refine sketches through dynamic, conversational interactions. Our approach requires no training or fine-tuning. Instead, we leverage the sequential nature and rich prior knowledge of off-the-shelf multimodal large language models (LLMs). We present an intuitive sketching language, introduced to the model through in-context examples, enabling it to "draw" using string-based actions. These are processed into vector graphics and then rendered to create a sketch on a pixel canvas, which can be accessed again for further tasks. By drawing stroke by stroke, our agent captures the evolving, dynamic qualities intrinsic to sketching. We demonstrate that SketchAgent can generate sketches from diverse prompts, engage in dialogue-driven drawing, and collaborate meaningfully with human users.
Figures
Figures from the paper (49 more)
Forward citations
Cited by 1 Pith paper
-
Pencils to Pixels: A Systematic Study of Creative Drawings across Children, Adults and AI
A new dataset and computational framework quantify style and content in children's, adults', and DALL-E drawings, showing that expert and automated creativity scores disagree across groups.
Reference graph
Works this paper leans on
-
[1]
Flamingo: a visual language model for few-shot learn- ing
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, An- toine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millicah, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Se- bastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sa- hand Sharifzadeh, Mikolaj Bink...
2024
-
[2]
Drawing as a space for social-cognitive in- teraction
Vanessa De Andrade, Sofia Freire, M ´onica Baptista, and Yael Shwartz. Drawing as a space for social-cognitive in- teraction. Education Sciences, 12(1), 2022. 3
2022
-
[3]
Anthropic. Claude. https://www.anthropic.com/ claude, 2023. 2, 3, 5, 6, 1
2023
-
[4]
Style and abstraction in por- trait sketching
Itamar Berger, Ariel Shamir, Moshe Mahler, Elizabeth Carter, and Jessica Hodgins. Style and abstraction in por- trait sketching. ACM Trans. Graph., 32(4), 2013. 2
2013
-
[5]
Improving image generation with better captions
James Betker, Gabriel Goh, Li Jing, † TimBrooks, Jian- feng Wang, Linjie Li, † LongOuyang, † JuntangZhuang, † JoyceLee, † YufeiGuo, † WesamManassra, † PrafullaD- hariwal, † CaseyChu, † YunxinJiao, and Aditya Ramesh. Improving image generation with better captions. 3
-
[6]
Hospedales, Tao Xiang, Yu- lia Gryaditskaya, and Yi-Zhe Song
Ayan Kumar Bhunia, Ayan Das, Umar Riaz Muhammad, Yongxin Yang, Timothy M. Hospedales, Tao Xiang, Yu- lia Gryaditskaya, and Yi-Zhe Song. Pixelor: a competitive sketching ai agent. so you think you can sketch? ACM Trans. Graph., 39:166:1–166:15, 2020. 1, 3
2020
-
[7]
Doodleformer: Creative sketch drawing with transformers
Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, Jorma Laaksonen, and Michael Felsberg. Doodleformer: Creative sketch drawing with transformers. ECCV, 2022. 2, 3
2022
-
[8]
Doodleformer: Creative sketch drawing with transformers
Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, Jorma Laaksonen, and Michael Felsberg. Doodleformer: Creative sketch drawing with transformers. In European Conference on Computer Vision, pages 338–355. Springer, 2022. 1
2022
Show all 140 references
-
[9]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sand- hini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Je...
1901
-
[10]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sand- hini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Je...
1901
-
[11]
Sparks of artificial general intelligence: Early experiments with gpt-
S ´ebastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-
-
[12]
arXiv preprint arXiv:2303.12712, 2023. 3
2023 arXiv
-
[13]
Delving into LLMs’ visual understanding ability using SVG to bridge image and text, 2024
Mu Cai, Zeyi Huang, Yuheng Li, Haohan Wang, and Yong Jae Lee. Delving into LLMs’ visual understanding ability using SVG to bridge image and text, 2024. 2, 3
2024
-
[14]
A computational approach to edge detection
John Canny. A computational approach to edge detection. IEEE Transactions on pattern analysis and machine intelli- gence, (6):679–698, 1986. 2
1986
-
[15]
Deepsvg: A hierarchical generative network for vector graphics animation, 2020
Alexandre Carlier, Martin Danelljan, Alexandre Alahi, and Radu Timofte. Deepsvg: A hierarchical generative network for vector graphics animation, 2020. 3
2020
-
[16]
On the util- ity of learning about humans for human-ai coordination
Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan. On the util- ity of learning about humans for human-ai coordination. In Advances in Neural Information Processing Systems . Cur- ran Associates, Inc., 2019. 13
2019
-
[17]
Learning to generate line drawings that convey geometry and seman- tics
Caroline Chan, Fr ´edo Durand, and Phillip Isola. Learning to generate line drawings that convey geometry and seman- tics. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 7915–7925,
-
[18]
Vi- sual chain-of-thought prompting for knowledge-based vi- sual reasoning
Zhenfang Chen, Qinhong Zhou, Yikang Shen, Yining Hong, Zhiqing Sun, Dan Gutfreund, and Chuang Gan. Vi- sual chain-of-thought prompting for knowledge-based vi- sual reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 1254–1262, 2024. 3
2024
-
[19]
3doodle: Compact abstraction of objects with 3d strokes
Changwoon Choi, Jaeah Lee, Jaesik Park, and Young Min Kim. 3doodle: Compact abstraction of objects with 3d strokes. ACM Trans. Graph., 43(4), 2024. 3
2024
-
[20]
9 Visionllama: A unified llama backbone for vision tasks,
Xiangxiang Chu, Jianlin Su, Bo Zhang, and Chunhua Shen. 9 Visionllama: A unified llama backbone for vision tasks,
-
[21]
B ´eziersketch: A generative model for scalable vector sketches
Ayan Das, Yongxin Yang, Timothy Hospedales, Tao Xiang, and Yi-Zhe Song. B ´eziersketch: A generative model for scalable vector sketches. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVI 16, pages 632–647. Springer,
2020
-
[22]
Drawing ap- prentice: An enactive co-creative agent for artistic collab- oration
Nicholas Davis, Chih-PIn Hsiao, Kunwar Yashraj Singh, Lisa Li, Sanat Moningi, and Brian Magerko. Drawing ap- prentice: An enactive co-creative agent for artistic collab- oration. In Proceedings of the 2015 ACM SIGCHI Con- ference on Creativity and Cognition , page 185–186, New...
2015
-
[23]
BERT: Pre-training of deep bidirectional trans- formers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional trans- formers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the As- sociation for Computational Linguistics: Human L...
2019
-
[24]
Towards multimodal in-context learning for vi- sion & language models
Sivan Doveh, Shaked Perek, M Jehanzeb Mirza, Wei Lin, Amit Alfassy, Assaf Arbelle, Shimon Ullman, and Leonid Karlinsky. Towards multimodal in-context learning for vi- sion & language models. arXiv preprint arXiv:2403.12736,
-
[25]
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. 4
2024 arXiv
-
[26]
How do hu- mans sketch objects? ACM Trans
Mathias Eitz, James Hays, and Marc Alexa. How do hu- mans sketch objects? ACM Trans. Graph. , 31(4), 2012. 2
2012
-
[27]
Frank, Matthew Groh, Hope Schroeder, Amy Smith, Memo Akten, Jessica Fjeld, Hany Farid, Neil Leach, Alex Pentland, and Olga Russakovsky
Ziv Epstein, Aaron Hertzmann, Laura Mariah Herman, Robert Mahari, Morgan R. Frank, Matthew Groh, Hope Schroeder, Amy Smith, Memo Akten, Jessica Fjeld, Hany Farid, Neil Leach, Alex Pentland, and Olga Russakovsky. Art and the science of generative ai. Science, 380:1110 – 1111, 2023. 1
2023
-
[28]
clip-vit-large-patch14
Hugging Face. clip-vit-large-patch14. https : / / huggingface . co / openai / clip - vit - large - patch14. 1
-
[29]
Bainbridge, Rebecca Chamberlain, and Jeffrey D
Judy Fan, Wilma A. Bainbridge, Rebecca Chamberlain, and Jeffrey D. Wammes. Drawing as a versatile cognitive tool. Nature Reviews Psychology, 2:556 – 568, 2023. 2
2023
-
[30]
Fan, Monica Dinculescu, and David Ha
Judith E. Fan, Monica Dinculescu, and David Ha. collab- draw: An environment for collaborative sketching with an artificial agent. In Proceedings of the 2019 Conference on Creativity and Cognition , page 556–561, New York, NY , USA, 2019. Association for Computing Machinery. 3, 7
2019
-
[31]
Fan, Robert D
Judith E. Fan, Robert D. Hawkins, Mike Wu, and Noah D. Goodman. Pragmatic Inference and Visual Abstraction En- able Contextual Flexibility During Visual Communication. Computational Brain & Behavior, 3(1):86–101, 2020. 3
2020
-
[32]
Drawing as a versatile cognitive tool
Judith E Fan, Wilma A Bainbridge, Rebecca Chamberlain, and Jeffrey D Wammes. Drawing as a versatile cognitive tool. Nature Reviews Psychology, 2(9):556–568, 2023. 1
2023
-
[33]
Creating drawings enhances learning by teaching
Logan Fiorella and Shelbi Kuhlmann. Creating drawings enhances learning by teaching. Journal of Educational Psy- chology, 112(4):811, 2020. 1
2020
-
[34]
Cogsketch: Sketch understanding for cognitive science research and for education
Kenneth Forbus, Jeffrey Usher, Andrew Lovett, Kate Lock- wood, and Jon Wetzel. Cogsketch: Sketch understanding for cognitive science research and for education. Topics in Cognitive Science, 3(4):648–666, 2011. 1
2011
-
[35]
Kevin Frans, L. B. Soros, and Olaf Witkowski. Clipdraw: exploring text-to-drawing synthesis through language- image encoders. In Proceedings of the 36th International Conference on Neural Information Processing Systems , Red Hook, NY , USA, 2024. Curran Associates Inc. 1, 3
2024
-
[36]
Blink: Multimodal large lan- guage models can see but not perceive
Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna. Blink: Multimodal large lan- guage models can see but not perceive. arXiv preprint arXiv:2404.12390, 2024. 4
2024 arXiv
-
[37]
Breath- ing life into sketches using text-to-video priors
Rinon Gal, Yael Vinker, Yuval Alaluf, Amit Bermano, Daniel Cohen-Or, Ariel Shamir, and Gal Chechik. Breath- ing life into sketches using text-to-video priors. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4325–4336, 2024. 3
2024
-
[38]
Kulkarni, Igor Babuschkin, S
Yaroslav Ganin, Tejas D. Kulkarni, Igor Babuschkin, S. M. Ali Eslami, and Oriol Vinyals. Synthesizing programs for images using reinforced adversarial learning. ArXiv, abs/1804.01118, 2018. 3
2018 arXiv
-
[39]
Computer-aided design as language
Yaroslav Ganin, Sergey Bartunov, Yujia Li, Ethan Keller, and Stefano Saliceti. Computer-aided design as language. In Neural Information Processing Systems, 2021. 3
2021
-
[40]
Sketchycoco: Image genera- tion from freehand scene sketches
Chengying Gao, Qi Liu, Qi Xu, Limin Wang, Jianzhuang Liu, and Changqing Zou. Sketchycoco: Image genera- tion from freehand scene sketches. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5174–5183, 2020. 2
2020
-
[41]
Foundations of representation: Where might graphical symbol systems come from? Cognitive Science, 31(6):961–987, 2007
Simon Garrod, Nicolas Fay, John Lee, Jon Oberlander, and Tracy MacLeod. Foundations of representation: Where might graphical symbol systems come from? Cognitive Science, 31(6):961–987, 2007. 3
2007
-
[42]
Creative sketch generation
Songwei Ge, Vedanuj Goswami, C Lawrence Zitnick, and Devi Parikh. Creative sketch generation. arXiv preprint arXiv:2011.10039, 2020. 1
2011 arXiv
-
[43]
Creative sketch generation
Songwei Ge, Vedanuj Goswami, Larry Zitnick, and Devi Parikh. Creative sketch generation. In International Con- ference on Learning Representations, 2021. 2, 6
2021
-
[44]
Collaborative drawing on a shared digital canvas in elementary science education: The effects of script and task awareness sup- port
Hannie Gijlers, Armin Weinberger, Alieke Mattia van Dijk, Lars Bollen, and Wouter van Joolingen. Collaborative drawing on a shared digital canvas in elementary science education: The effects of script and task awareness sup- port. International Journal of Computer-Supported Co...
2013
-
[45]
Serial sketching: visual problem solving in designing
Gabriela Goldschmidt. Serial sketching: visual problem solving in designing. Cybernetics and System, 23(2):191– 219, 1992. 1
1992
-
[46]
Grosz and Sarit Kraus
Barbara J. Grosz and Sarit Kraus. Collaborative plans for complex group action. Artificial Intelligence, 86(2):269– 357, 1996. 13
1996
-
[47]
10 Opensketch: a richly-annotated dataset of product design sketches
Yulia Gryaditskaya, Mark Sypesteyn, Jan Willem Hofti- jzer, Sylvia Pont, Fr ´edo Durand, and Adrien Bousseau. 10 Opensketch: a richly-annotated dataset of product design sketches. ACM Trans. Graph., 38(6), 2019. 2
2019
-
[48]
Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, Dinesh Manocha, and Tianyi Zhou. Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision...
2024
-
[49]
A neural representation of sketch drawings
David Ha and Douglas Eck. A neural representation of sketch drawings. CoRR, abs/1704.03477, 2017. 1, 2, 3, 6, 7
2017 arXiv
-
[50]
Xiaoyan Zhang, Xin Chen, Xu Yang, Zhibin Wang, Gang Yu, Bin Fu, and Hanwang Zhang
Yucheng Han, China. Xiaoyan Zhang, Xin Chen, Xu Yang, Zhibin Wang, Gang Yu, Bin Fu, and Hanwang Zhang. Chartllama: A multimodal llm for chart understanding and generation. ArXiv, abs/2311.16483, 2023. 2, 3
2023 arXiv
-
[51]
Hawkins, Megumi Sano, Noah D
Robert D. Hawkins, Megumi Sano, Noah D. Goodman, and Judith E. Fan. Visual resemblance and communicative context constrain the emergence of graphical conventions,
-
[52]
Visual sketchpad: Sketching as a visual chain of thought for multimodal language models
Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Osten- dorf, Luke Zettlemoyer, Noah A Smith, and Ranjay Kr- ishna. Visual sketchpad: Sketching as a visual chain of thought for multimodal language models. arXiv preprint arXiv:2406.09403, 2024. 3
2024 arXiv
-
[53]
A collaborative, interactive and context-aware draw- ing agent for co-creative design
Francisco Javier Ibarrola, Tomas Lawton, and Kazjon Grace. A collaborative, interactive and context-aware draw- ing agent for co-creative design. IEEE Transactions on Vi- sualization and Computer Graphics, 30:5525–5537, 2022. 3
2022
-
[54]
Vectorfusion: Text-to-svg by abstracting pixel-based diffusion models
Ajay Jain, Amber Xie, and Pieter Abbeel. Vectorfusion: Text-to-svg by abstracting pixel-based diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 1911–1920, 2023. 1, 3
1911
-
[55]
The Quick, Draw! - A.I
Jongejan Jonas, Rowley Henry, Kawashima Takashi, Kim Jongmin, and Fox-Gieg Nick. The Quick, Draw! - A.I. Experiment, 2016. 2, 3, 6, 7, 8
2016
-
[56]
Drawing theories apart: The dispersion of Feynman diagrams in postwar physics
David Kaiser. Drawing theories apart: The dispersion of Feynman diagrams in postwar physics . University of Chicago Press, 2019. 1
2019
-
[58]
Creative sketching part- ner: an analysis of human-ai co-creativity
Pegah Karimi, Jeba Rezwana, Safat Siddiqui, Mary Lou Maher, and Nasrin Dehbozorgi. Creative sketching part- ner: an analysis of human-ai co-creativity. In Proceedings of the 25th International Conference on Intelligent User In- terfaces, page 221–230, New York, NY , USA, 2020....
2020
-
[59]
Chapter three - psychological research on joint action: The- ory and data
G ¨unther Knoblich, Stephen Butterfill, and Natalie Sebanz. Chapter three - psychological research on joint action: The- ory and data. In Advances in Research and Theory , pages 59–101. Academic Press, 2011. 13
2011
-
[60]
Cairosvg
Kozea. Cairosvg. https://cairosvg.org/, 2023. 1
2023
-
[61]
Graphic thinking for architects and designers
Paul Laseau. Graphic thinking for architects and designers. John Wiley & Sons, 2000. 2
2000
-
[62]
When is a tool a tool? user perceptions of system agency in human–ai co-creative drawing
Tomas Lawton, Kazjon Grace, and Francisco J Ibarrola. When is a tool a tool? user perceptions of system agency in human–ai co-creative drawing. In Proceedings of the 2023 ACM Designing Interactive Systems Conference, page 1978–1996, New York, NY , USA, 2023. Association for Co...
2023
-
[63]
Drawing with reframer: Emergence and control in co-creative ai
Tomas Lawton, Francisco J Ibarrola, Dan Ventura, and Kazjon Grace. Drawing with reframer: Emergence and control in co-creative ai. In Proceedings of the 28th International Conference on Intelligent User Interfaces , page 264–277, New York, NY , USA, 2023. Association for ...
2023
-
[64]
Lawrence Zitnick, and Michael F
Yong Jae Lee, C. Lawrence Zitnick, and Michael F. Cohen. Shadowdraw: real-time user guidance for freehand draw- ing. In ACM SIGGRAPH 2011 Papers , New York, NY , USA, 2011. Association for Computing Machinery. 3
2011
-
[65]
BLIP: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. BLIP: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation. In Pro- ceedings of the 39th International Conference on Machine Learning, pages 12888–12900. PMLR, 2022. 2, 3
2022
-
[66]
Photo-sketching: Inferring contour draw- ings from images
Mengtian Li, Zhe Lin, Radomir Mech, Ersin Yumer, and Deva Ramanan. Photo-sketching: Inferring contour draw- ings from images. In 2019 IEEE Winter Conference on Ap- plications of Computer Vision (WACV), pages 1403–1412. IEEE, 2019. 2
2019
-
[67]
Differentiable vector graphics rasterization for editing and learning
Tzu-Mao Li, Michal Luk ´aˇc, Gharbi Micha¨el, and Jonathan Ragan-Kelley. Differentiable vector graphics rasterization for editing and learning. ACM Trans. Graph. (Proc. SIG- GRAPH Asia), 39(6):193:1–193:15, 2020. 1, 3
2020
-
[68]
Hospedales, and Shaogang Gong
Yi Li, Yi-Zhe Song, Timothy M. Hospedales, and Shaogang Gong. Free-hand sketch synthesis with deformable stroke models. CoRR, abs/1510.02644, 2015. 1, 2
2015 arXiv
-
[69]
Im2pencil: Controllable pen- cil illustration from photographs
Yijun Li, Chen Fang, Aaron Hertzmann, Eli Shechtman, and Ming-Hsuan Yang. Im2pencil: Controllable pen- cil illustration from photographs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1525–1534, 2019. 2
2019
-
[70]
Hangyu Lin, Yanwei Fu, Yu-Gang Jiang, and X. Xue. Sketch-bert: Learning sketch bidirectional encoder repre- sentation from transformers by self-supervised learning of sketch gestalt. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6757–6766,
2020
-
[71]
Neural strokes: Stylized line draw- ing of 3d shapes
Difan Liu, Matthew Fisher, Aaron Hertzmann, and Evan- gelos Kalogerakis. Neural strokes: Stylized line draw- ing of 3d shapes. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision, pages 14204– 14213, 2021. 2
2021
-
[72]
Sketchgan: Joint sketch completion and recognition with generative adversarial net- work
Fang Liu, Xiaoming Deng, Yu-Kun Lai, Yong-Jin Liu, Cuixia Ma, and Hongan Wang. Sketchgan: Joint sketch completion and recognition with generative adversarial net- work. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019. 2 11
2019
-
[73]
Sketchgan: Joint sketch completion and recognition with generative adversarial net- work
Fang Liu, Xiaoming Deng, Yu-Kun Lai, Yong-Jin Liu, Cuixia Ma, and Hongan Wang. Sketchgan: Joint sketch completion and recognition with generative adversarial net- work. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 5830– 5839, 2019. 2
2019
-
[74]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Advances in Neural In- formation Processing Systems, pages 34892–34916. Curran Associates, Inc., 2023. 2, 3
2023
-
[75]
Parallel developmental changes in chil- dren’s production and recognition of line drawings of visual concepts
Bria Long, Judith Fan, Holly Huey, Zixian Chai, and Michael Frank. Parallel developmental changes in chil- dren’s production and recognition of line drawings of visual concepts. Nature Communications, 15, 2024. 6
2024
-
[76]
Openeqa: Embodied question answering in the era of foun- dation models
Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Sil- wal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma, Vincent Berges, Shiqi Zhang, Pulkit Agrawal, Yonatan Bisk, Dhruv Batr...
2024
-
[77]
McCarthy, Robert D
William P. McCarthy, Robert D. Hawkins, Haoliang Wang, Cameron Holdaway, and Judith E. Fan. Learning to com- municate about shared procedural abstractions, 2021. 13
2021
-
[78]
McCarthy, Justin Matejka, Karl D.D
William P. McCarthy, Justin Matejka, Karl D.D. Willis, Ju- dith E. Fan, and Yewen Pu. Communicating design intent using drawing and text. In Proceedings of the 16th Confer- ence on Creativity & Cognition, page 512–519, New York, NY , USA, 2024. Association for Computing Machinery. 3
2024
-
[79]
Unsupervised doodling and painting with improved spiral
John FJ Mellor, Eunbyung Park, Yaroslav Ganin, Igor Babuschkin, Tejas Kulkarni, Dan Rosenbaum, Andy Bal- lard, Theophane Weber, Oriol Vinyals, and SM Eslami. Unsupervised doodling and painting with improved spiral. arXiv preprint arXiv:1910.01007, 2019. 3
1910 arXiv
-
[80]
Learning to draw: Emer- gent communication through sketching
Daniela Mihai and Jonathon Hare. Learning to draw: Emer- gent communication through sketching. Advances in Neu- ral Information Processing Systems , 34:7153–7166, 2021. 3
2021
-
[81]
Compositional chain-of-thought prompt- ing for large multimodal models
Chancharik Mitra, Brandon Huang, Trevor Darrell, and Roei Herzig. Compositional chain-of-thought prompt- ing for large multimodal models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14420–14431, 2024. 3
2024
-
[82]
Seva: Leveraging sketches to evaluate alignment between human and machine visual abstraction
Kushin Mukherjee, Holly Huey, Xuanchen Lu, Yael Vinker, Rio Aguina-Kang, Ariel Shamir, and Judith Fan. Seva: Leveraging sketches to evaluate alignment between human and machine visual abstraction. In Advances in Neural In- formation Processing Systems, 2023. 2
2023
-
[83]
Observing by hand: sketching the nebu- lae in the nineteenth century
Omar W Nasim. Observing by hand: sketching the nebu- lae in the nineteenth century. University of Chicago Press,
-
[84]
I lead, you help but only with enough details: Understanding user experience of co-creation with artificial intelligence
Changhoon Oh, Jungwoo Song, Jinhan Choi, Seonghyeon Kim, Sungwoo Lee, and Bongwon Suh. I lead, you help but only with enough details: Understanding user experience of co-creation with artificial intelligence. In Proceedings of the 2018 CHI Conference on Human Factors in Comput...
2018
-
[85]
Gpt-4 technical report, 2024
OpenAI. Gpt-4 technical report, 2024. 2, 3, 4, 6
2024
-
[86]
Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rom- bach
Dustin Podell, Zion English, Kyle Lacey, A. Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rom- bach. Sdxl: Improving latent diffusion models for high- resolution image synthesis. ArXiv, abs/2307.01952, 2023. 2, 3
2023 arXiv
-
[87]
Sketchlattice: Latticed representation for sketch manipulation
Yonggang Qi, Guoyao Su, Pinaki Nath Chowdhury, Mingkang Li, and Yi-Zhe Song. Sketchlattice: Latticed representation for sketch manipulation. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 953–961, 2021. 2
2021
-
[88]
Emergent graphical conven- tions in a visual communication game
Shuwen Qiu, Sirui Xie, Lifeng Fan, Tao Gao, Song- Chun Zhu, and Yixin Zhu. Emergent graphical conven- tions in a visual communication game. arXiv preprint arXiv:2111.14210, 2021. 3
2021 arXiv
-
[89]
Learning transferable vi- sual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable vi- sual models from natural language supervision. CoRR, abs/2103.000...
2021 arXiv
-
[90]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.J. Mach. Learn. Res., 21(1), 2020. 3
2020
-
[91]
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 2, 3
2022 arXiv
-
[92]
Col- lomosse, and Moacir Antonelli Ponti
Leo Sampaio Ferraz Ribeiro, Tu Bui, John P. Col- lomosse, and Moacir Antonelli Ponti. Sketchformer: Transformer-based representation for sketched structure. 2020 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 14141–14150, 2020. 3
2020
-
[93]
High-resolution image synthesis with latent diffusion models, 2022
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models, 2022. 1, 2, 3, 6
2022
-
[94]
Visual chain of thought: bridging logical gaps with multimodal infill- ings
Daniel Rose, Vaishnavi Himakunthala, Andy Ouyang, Ryan He, Alex Mei, Yujie Lu, Michael Saxon, Chinmay Sonar, Diba Mirza, and William Yang Wang. Visual chain of thought: bridging logical gaps with multimodal infill- ings. arXiv preprint arXiv:2305.02317, 2023. 3
2023 arXiv
-
[95]
Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. Photorealistic text-to- image diffusion mo...
-
[96]
The sketchy database: Learning to retrieve badly drawn bunnies
Patsorn Sangkloy, Nathan Burnell, Cusuh Ham, and James Hays. The sketchy database: Learning to retrieve badly drawn bunnies. ACM Trans. Graph., 35(4), 2016. 2 12
2016
-
[97]
Frida: A collaborative robot painter with a differentiable, real2sim2real planning environment
Peter Schaldenbrand, James McCann, and Jean Oh. Frida: A collaborative robot painter with a differentiable, real2sim2real planning environment. In 2023 IEEE Inter- national Conference on Robotics and Automation (ICRA) , pages 11712–11718, 2023. 3
2023
-
[98]
Cofrida: Self-supervised fine-tuning for human-robot co-painting
Peter Schaldenbrand, Gaurav Parmar, Jun-Yan Zhu, James McCann, and Jean Oh. Cofrida: Self-supervised fine-tuning for human-robot co-painting. In 2024 IEEE International Conference on Robotics and Automation (ICRA) . IEEE,
2024
-
[99]
The reflective practitioner: How professionals think in action, 1986
Donald A Schon and Vincent DeSanctis. The reflective practitioner: How professionals think in action, 1986. 2
1986
-
[100]
Laion-5b: An open large-scale dataset for train- ing next generation image-text models, 2022
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Worts- man, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Je- nia Jitsev. L...
2022
-
[101]
A multimodal automated interpretability agent
Tamar Rott Shaham, Sarah Schwettmann, Franklin Wang, Achyuta Rajaram, Evan Hernandez, Jacob Andreas, and Antonio Torralba. A multimodal automated interpretability agent. In Forty-first International Conference on Machine Learning, 2024. 3
2024
-
[102]
Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning
Hao Shao, Shengju Qian, Han Xiao, Guanglu Song, Zhuo- fan Zong, Letian Wang, Yu Liu, and Hongsheng Li. Visual cot: Advancing multi-modal language models with a com- prehensive dataset and benchmark for chain-of-thought rea- soning. arXiv preprint arXiv:2403.16999, 2024. 3
2024 arXiv
-
[103]
A vision check-up for language models
Pratyusha Sharma, Tamar Rott Shaham, Manel Baradad, Stephanie Fu, Adrian Rodriguez-Munoz, Shivam Duggal, Phillip Isola, and Antonio Torralba. A vision check-up for language models. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , pages 1...
2024
-
[104]
Learning to sketch with shortcut cycle consistency, 2018
Jifei Song, Kaiyue Pang, Yi-Zhe Song, Tao Xiang, and Tim- othy Hospedales. Learning to sketch with shortcut cycle consistency, 2018. 2
2018
-
[105]
Ad hoc autonomous agent teams: Collaboration without pre-coordination
Peter Stone, Gal Kaminka, Sarit Kraus, and Jeffrey Rosen- schein. Ad hoc autonomous agent teams: Collaboration without pre-coordination. Proceedings of the AAAI Con- ference on Artificial Intelligence , 24(1):1504–1509, 2010. 13
2010
-
[106]
Sketchhealer: A graph-to- sequence network for recreating partial human sketches
Guoyao Su, Yonggang Qi, Kaiyue Pang, Jie Yang, Yi-Zhe Song, and CVSSP SketchX. Sketchhealer: A graph-to- sequence network for recreating partial human sketches. In BMVC, page 5, 2020. 2
2020
-
[107]
Exploring effective factors for improving vi- sual in-context learning
Yanpeng Sun, Qiang Chen, Jian Wang, Jingdong Wang, and Zechao Li. Exploring effective factors for improving vi- sual in-context learning. arXiv preprint arXiv:2304.04748,
-
[108]
Sutherland
Ivan E. Sutherland. Sketchpad—a man-machine graphi- cal communication system, page 391–408. Association for Computing Machinery, New York, NY , USA, 1998. 3
1998
-
[109]
Gemini: A family of highly capable multi- modal models, 2024
Gemini Team. Gemini: A family of highly capable multi- modal models, 2024. 2, 3
2024
-
[110]
Design ideation with ai - sketching, thinking and talking with generative machine learning models
Jakob Tholander and Martin Jonsson. Design ideation with ai - sketching, thinking and talking with generative machine learning models. In Proceedings of the 2023 ACM Design- ing Interactive Systems Conference, page 1930–1940, New York, NY , USA, 2023. Association for Computing...
2023
-
[111]
Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024
Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024. 4
2024
-
[112]
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth ´ee Lacroix, Bap- tiste Rozi `ere, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation la...
2023 arXiv
-
[113]
What do sketches say about thinking? AAAI Spring Symp
Barbara Tversky. What do sketches say about thinking? AAAI Spring Symp. Sketch Understanding Worksh. , 2002. 2
2002
-
[114]
Visualizing thought
Barbara Tversky. Visualizing thought. In Handbook of hu- man centric visualization, pages 3–40. Springer, 2013. 1
2013
-
[115]
Sketches for design and design of sketches
Barbara Tversky, Masaki Suwa, Maneesh Agrawala, Julie Heiser, Chris Stolte, Pat Hanrahan, Doantam Phan, Jeff Klingner, Marie-Paule Daniel, Paul Lee, et al. Sketches for design and design of sketches. Human Behaviour in Design: Individuals, Teams, Tools, pages 79–86, 2003. 1
2003
-
[116]
Drawing to reason and learn in science
Russell Tytler, Vaughan Prain, George Aranda, Joseph Fer- guson, and Radhika Gorur. Drawing to reason and learn in science. Journal of Research in Science Teaching , 57(2): 209–231, 2020. 3
2020
-
[117]
Bo, Ro- man Christian Bachmann, Amit Haim Bermano, Daniel Cohen-Or, Amir Zamir, and Ariel Shamir
Yael Vinker, Ehsan Pajouheshgar, Jessica Y . Bo, Ro- man Christian Bachmann, Amit Haim Bermano, Daniel Cohen-Or, Amir Zamir, and Ariel Shamir. Clipasso: Semantically-aware object sketching. ACM Trans. Graph., 41(4), 2022. 1, 3, 6
2022
-
[118]
Clipascene: Scene sketching with different types and levels of abstraction
Yael Vinker, Yuval Alaluf, Daniel Cohen-Or, and Ariel Shamir. Clipascene: Scene sketching with different types and levels of abstraction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , pages 4146–4156, 2023. 3, 6
2023
-
[119]
Contextseg: Sketch se- mantic segmentation by querying the context with attention
Jiawei Wang and Changjian Li. Contextseg: Sketch se- mantic segmentation by querying the context with attention. arXiv preprint arXiv:2311.16682, 2023. 6
2023 arXiv
-
[120]
Chi, Quoc V
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V . Le, and Denny Zhou. Chain-of-thought prompting elicits reason- ing in large language models. Red Hook, NY , USA, 2024. Curran Associates Inc. 2, 5
2024
-
[121]
Holger Winnem ¨oller, Jan Eric Kyprianidis, and Sven C. Olsen. Xdog: An extended difference-of-gaussians com- pendium including advanced image stylization. Comput. Graph., 36:740–753, 2012. 2
2012
-
[122]
Scalable Vector Graphics (SVG), 1999
World Wide Web Consortium (W3C). Scalable Vector Graphics (SVG), 1999. 3
1999
-
[123]
Visual chatgpt: Talk- ing, drawing and editing with visual foundation models
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. Visual chatgpt: Talk- ing, drawing and editing with visual foundation models. arXiv preprint arXiv:2303.04671, 2023. 3
2023 arXiv
-
[124]
Iconshop: Text-guided vector icon synthesis with autoregressive trans- 13 formers
Rong Wu, Wanchao Su, Kede Ma, and Jing Liao. Iconshop: Text-guided vector icon synthesis with autoregressive trans- 13 formers. ACM Transactions on Graphics (TOG), 42:1 – 14,
-
[125]
Differsketching: How differ- ently do people sketch 3d objects? ACM Transactions on Graphics (Proceedings of ACM SIGGRAPH Asia 2022), 41 (4):1–16, 2022
Chufeng Xiao, Wanchao Su, Jing Liao, Zhouhui Lian, Yi- Zhe Song, and Hongbo Fu. Differsketching: How differ- ently do people sketch 3d objects? ACM Transactions on Graphics (Proceedings of ACM SIGGRAPH Asia 2022), 41 (4):1–16, 2022. 2
2022
-
[126]
Holistically-nested edge de- tection
Saining Xie and Zhuowen Tu. Holistically-nested edge de- tection. In Proceedings of the IEEE international confer- ence on computer vision, pages 1395–1403, 2015. 2
2015
-
[127]
Diffsketcher: Text guided vec- tor sketch synthesis through latent diffusion models
XiMing Xing, Chuang Wang, Haitao Zhou, Jing Zhang, Qian Yu, and Dong Xu. Diffsketcher: Text guided vec- tor sketch synthesis through latent diffusion models. In Advances in Neural Information Processing Systems, pages 15869–15889. Curran Associates, Inc., 2023. 3, 6
2023
-
[128]
Svgdreamer: Text guided svg generation with diffusion model
Ximing Xing, Haitao Zhou, Chuang Wang, Jing Zhang, Dong Xu, and Qian Yu. Svgdreamer: Text guided svg generation with diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4546–4555, 2024. 3, 6, 7
2024
-
[129]
Hospedales, Qiyue Yin, Yi-Zhe Song, Tao Xiang, and Liang Wang
Peng Xu, Timothy M. Hospedales, Qiyue Yin, Yi-Zhe Song, Tao Xiang, and Liang Wang. Deep learning for free- hand sketch: A survey and a toolbox, 2020. 2
2020
-
[130]
The dawn of lmms: Preliminary explorations with gpt-4v(ision)
Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. The dawn of lmms: Preliminary explorations with gpt-4v(ision). ArXiv, abs/2309.17421, 2023. 2
2023 arXiv
-
[131]
Idea2img: Iterative self-refinement with gpt-4v (ision) for automatic image design and generation
Zhengyuan Yang, Jianfeng Wang, Linjie Li, Kevin Lin, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. Idea2img: Iterative self-refinement with gpt-4v (ision) for automatic image design and generation. arXiv preprint arXiv:2310.08541, 2023. 3
-
[132]
Apdrawinggan: Generating artistic portrait drawings from face photos with hierarchical gans
Ran Yi, Yong-Jin Liu, Yu-Kun Lai, and Paul L Rosin. Apdrawinggan: Generating artistic portrait drawings from face photos with hierarchical gans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10743–10752, 2019. 2
2019
-
[133]
Un- paired portrait drawing generation via asymmetric cycle mapping
Ran Yi, Yong-Jin Liu, Yu-Kun Lai, and Paul L Rosin. Un- paired portrait drawing generation via asymmetric cycle mapping. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8217– 8225, 2020. 2
2020
-
[134]
Zhang, Weijie Wang, Paul Pangaro, Nikolas Marte- laro, and Daragh Byrne
C. Zhang, Weijie Wang, Paul Pangaro, Nikolas Marte- laro, and Daragh Byrne. Generative image ai using design sketches as input: Opportunities and challenges. Proceed- ings of the 15th Conference on Creativity and Cognition ,
-
[135]
Instruct me more! random prompt- ing for visual in-context learning
Jiahao Zhang, Bowen Wang, Liangzhi Li, Yuta Nakashima, and Hajime Nagahara. Instruct me more! random prompt- ing for visual in-context learning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2597–2606, 2024. 3
2024
-
[136]
Text-to- vector generation with neural path representation
Peiying Zhang, Nanxuan Zhao, and Jing Liao. Text-to- vector generation with neural path representation. ACM Trans. Graph., 43(4), 2024. 3
2024
-
[137]
What makes good examples for visual in-context learning? Ad- vances in Neural Information Processing Systems , 36: 17773–17794, 2023
Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. What makes good examples for visual in-context learning? Ad- vances in Neural Information Processing Systems , 36: 17773–17794, 2023. 3
2023
-
[138]
Stroke-based semantic segmentation for scene-level free-hand sketches
Zhengming Zhang, Xiaoming Deng, Jinyao Li, Yukun Lai, Cuixia Ma, Yongjin Liu, and Hongan Wang. Stroke-based semantic segmentation for scene-level free-hand sketches. Vis. Comput., 39(12):6309–6321, 2022. 6
2022
-
[139]
Multimodal chain-of- thought reasoning in language models
Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. Multimodal chain-of- thought reasoning in language models. arXiv preprint arXiv:2302.00923, 2023. 3
2023 arXiv
-
[140]
Creativeseg: Semantic seg- mentation of creative sketches
Yixiao Zheng, Kaiyue Pang, Ayan Das, Dongliang Chang, Yi-Zhe Song, and Zhanyu Ma. Creativeseg: Semantic seg- mentation of creative sketches. IEEE Transactions on Im- age Processing, 33:2266–2278, 2024. 6
2024
-
[141]
shark”, which was often misclassified as a “fish
Tao Zhou, Chen Fang, Zhaowen Wang, Jimei Yang, Byung- moon Kim, Zhili Chen, Jonathan Brandt, and Demetri Ter- zopoulos. Learning to sketch with deep q networks and demonstrated strokes. ArXiv, abs/1810.05977, 2018. 2, 3 14 SketchAgent: Language-Driven Sequential Sketch Generat...
2018 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.