REVIEW 3 major objections 6 minor 51 references
Qualitative Study for LLM-assisted Design Study Process: Strategies, Challenges, and Roles
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that LLM assistance in visualization design studies is best understood as four roles—Connector, Simulator, Programmer, Assistant—each with stage-specific strategies and challenges across the nine-stage design study process.
desk verdict Useful stage-by-stage map of LLM use in design studies, but the four-role taxonomy needs more methodological transparency and the reference list needs cleaning. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the four-role taxonomy of LLM assistance, defined by the functions the model performs for the researcher: Connector bridges knowledge and terminology gaps between visualization researchers and domain experts; Simulator stands in for users and reviewers; Programmer writes, optimizes, and debugs code; Assistant automates repetitive chores. This taxonomy organizes the interview data and is the basis of the guidelines. A second structural element is the nine-stage design study methodology framework (learn, winnow, cast, discover, design, implement, deploy, reflect, write), which supplies the stage-by-stage grid against which strategies and challenges are coded and rated.
What would settle it
Log actual LLM prompt histories from design-study project teams and have independent coders assign each logged use to one of the four roles; if a substantial fraction of uses falls outside the taxonomy or inter-coder agreement is low, the four-role claim would not survive. A complementary test: compare teams following the role-based guidelines with teams using LLMs freely; the guidelines predict fewer stage-specific failures such as hallucinated references, generic simulated feedback, and code that does not integrate into existing frameworks.
Extended reading notes
Core claim
The central discovery is that using LLMs in a design study is best understood through four recurring roles rather than as a single "ask the chatbot" activity. The Connector role uses LLMs to translate domain terminology, extract design requirements, and facilitate communication with domain experts. The Simulator role uses LLMs to impersonate questionnaire respondents, users, or reviewers, surfacing usability and reasoning problems before real humans are involved. The Programmer role covers code generation, prototyping, data cleaning, scalability testing, and usability tracking. The Assistant role covers general productivity work such as literature summarization, email drafting, transcription, and paper drafting. These roles were synthesized from interviews with 30 researchers at five expertise levels, and each role appears with stage-specific strategies and challenges across the nine stages of the design study methodology.
Load-bearing premise
The taxonomy rests on participants' accurate self-reports of how they used LLMs, and on the nine-stage design-study framework being the right structure for interpreting those reports; the paper does not observe workflows directly, and its participants may be predisposed to LLMs.
Editorial extensions
If this is right
- If the taxonomy is right, design-study guidance can be organized by role: researchers can ask which role they need at a given stage and pick the matching strategy.
- Practitioners can anticipate known failure modes per stage—hallucinated references, generic simulated feedback, code that does not integrate into existing frameworks—and apply the practices the paper collects.
- LLM tool builders can target the four roles, for example by building copilots that maintain project-wide context, since many reported challenges stem from the model's lack of global understanding.
- The questionnaire ratings on importance, difficulty, and necessity of LLM assistance can help prioritize which stages most need better tooling and training.
- Human judgment remains indispensable: the paper's participants expect LLMs to augment rather than replace researchers, especially for evaluating ideas, creative leaps, and interaction design.
Reading between the lines
- Beyond the paper: the same four roles could describe LLM use in other qualitative or co-design workflows, since translating between experts, simulating users, writing code, and doing chores are not visualization-specific; interviewing design teams in other fields would test this.
- The simulator role implies a cheap pre-screening loop the paper does not validate: if LLM-simulated users and reviewers flag the same usability and reasoning issues that real participants later raise, teams could cut pilot-study costs by pre-testing with LLMs.
- A testable extension would turn the role-based guidelines into a prompt library or checklist and randomly assign teams either to it or to unguided LLM use; the paper predicts the guided teams should meet fewer stage-specific failures such as hallucinated references and lost global context.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript presents a multi-stage qualitative study of how visualization researchers use LLMs in design studies. Based on 30 semi-structured interviews, a 36-item Likert questionnaire, and a consensus-based post-study analysis, the authors propose four roles for LLMs—Connector, Simulator, Programmer, and Assistant—and map stage-specific strategies, challenges, and practices onto the nine stages of Sedlmair et al.'s design study framework. The paper also offers role-based guidelines, future directions, and a limitations section.
Significance. The contribution is potentially useful: a shared vocabulary for LLM use in design studies and concrete stage-specific strategies and challenges would help both novice and experienced visualization researchers. Strengths include a relatively large interview sample for a qualitative study, a detailed participant profile table, the use of a structured questionnaire to complement interview data, and an unusually candid limitations section that acknowledges self-selection and the absence of empirical validation. The main weakness is that the central deliverable—the four-role taxonomy—rests on a consensus coding process that is reported only at a high level and cannot currently be audited or reproduced from the manuscript. If the authors add a codebook and a systematic evidence trace, the paper would be a solid qualitative contribution to the visualization community.
major comments (3)
- [3.3 Post-study analysis] The four roles are the paper's central contribution, but the coding process is described only as three researchers who 'independently reviewed' the data and then 'discussed... to reach a consensus.' No codebook, inclusion/exclusion criteria, initial role proposals, disagreement-resolution examples, inter-coder agreement metric, or negative-case analysis are reported. Because the role labels are broad and intuitive, a consensus procedure among authors who already know the intended framing can produce labels that fit the data without being uniquely determined by it. The questionnaire measures perceived importance and difficulty and cannot validate the role categories. Please report the coding scheme, the number and nature of disagreements, and how they were resolved, and ideally include an independent blinded coding or a member check, so readers can judge whether the taxonomy is grounded in the data rather than an interpretive overlay.
- [5 Qualitative Interview Results] The mapping from evidence to roles is presented through illustrative quotes and vague quantifiers such as 'some PhD students' and 'a proficient PhD student,' and several strategies receive dual labels such as 'connector assistant' and 'assistant connector.' This prevents readers from verifying the prevalence, consistency, and mutual exclusivity of each role across stages. Please provide a per-stage table that lists each strategy, challenge, and practice with participant IDs, the associated role label(s), and the number of participants reporting it, and explicitly discuss how multi-label cases were handled.
- [7.2.1 Trust and Usability Concerns] The authors acknowledge that participants may have a predisposition toward LLM use, but the analysis largely treats participants' retrospective self-reports as accurate descriptions of their workflows. Since no direct observation, usage logs, or artifact-based triangulation is used, statements about 'common strategies' and 'challenges' are claims about perceived practice rather than verified practice. Please qualify these claims accordingly, or add a small validation subset to strengthen the conclusions.
minor comments (6)
- [References] References [6], [11], [12], and [40] contain placeholder author names such as 'A. Author2'; these are incomplete and must be completed before publication.
- [Section 6] The heading contains a typo: 'deign studies' should be 'design studies.'
- [Fig. 3] The participant table has at least one garbled entry ('15 (10Ď)' in the publication column), and the notation '1 (0)' is not defined in the caption; please add a legend explaining the numbers.
- [Fig. 1 caption] The caption references '1 point ~ 7 points,' but Section 3.3 describes a 7-point Likert scale from 1 to 7; please harmonize the wording and state whether the bottom-section ratings are means or individual responses.
- [3.2 Data Collection] The statement that 'every participants can access and edit it anytime' is unclear: please clarify whether participants could edit their own entries only, and whether any participant-modified data were detected or excluded.
- [Section 5] Role labels such as 'connector,' 'assistant connector,' and 'simulator' are sometimes placed after punctuation without a clear visual association with the relevant strategy; consider formatting them as styled tags to improve readability.
Circularity Check
No significant circularity: the role taxonomy is an inductive qualitative coding result, not a derived prediction equivalent to its inputs.
full rationale
I walked the claimed derivation chain: interviews and questionnaires were collected from 30 participants, and a post-study analysis produced four roles through independent review and consensus discussion. The roles are not fitted to a target quantity, nor is any stated prediction forced by construction. The Sedlmair et al. nine-stage framework is an external input used to structure the interviews, not a conclusion derived from the study. The closest candidate for circularity is that the same self-reported data are used both to generate and to illustrate the roles; however, this is inherent to qualitative taxonomy construction and does not correspond to any of the enumerated circularity patterns, since there is no equation, fitted parameter, or uniqueness claim that reduces to an input. The paper's own limitations (Section 7.2.1) acknowledge the lack of empirical validation and possible participant predisposition, but these are generalizability and rigor concerns, not circularity. No load-bearing self-citation was found; the only co-authored related-work reference (Dashchat, ref. 37) is peripheral to the central derivation. Therefore the analysis is self-contained in the relevant sense and the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Sedlmair et al.'s nine-stage design study framework is a valid and comprehensive model for the design study process.
- domain assumption Participants' self-reports of their LLM usage, expertise, and experiences are accurate and representative.
- domain assumption The three researchers' consensus process reliably captures themes from the qualitative data.
invented entities (4)
-
Connector role
-
Simulator role
-
Programmer role
-
Assistant role
Cite this review
Pith. "Pith review of Qualitative Study for LLM-assisted Design Study Process: Strategies, Challenges, and Roles." pith.science (2026). https://pith.science/paper/23KNUGQG
@misc{pith2026250710024,
author = {Pith},
title = {Pith review of: Qualitative Study for LLM-assisted Design Study Process: Strategies, Challenges, and Roles},
year = {2026},
howpublished = {\url{https://pith.science/paper/23KNUGQG}},
note = {Machine review of arXiv:2507.10024}
}
read the original abstract
Design studies aim to create visualization solutions for real-world problems of different application domains. Recently, the emergence of large language models (LLMs) has introduced new opportunities to enhance the design study process, providing capabilities such as creative problem-solving, data handling, and insightful analysis. However, despite their growing popularity, there remains a lack of systematic understanding of how LLMs can effectively assist researchers in visualization-specific design studies. In this paper, we conducted a multi-stage qualitative study to fill this gap, involving 30 design study researchers from diverse backgrounds and expertise levels. Through in-depth interviews and carefully-designed questionnaires, we investigated strategies for utilizing LLMs, the challenges encountered, and the practices used to overcome them. We further compiled and summarized the roles that LLMs can play across different stages of the design study process. Our findings highlight practical implications to inform visualization practitioners, and provide a framework for leveraging LLMs to enhance the design study process in visualization research.
Figures
Reference graph
Works this paper leans on
-
[1]
G. Akbaba and N. Elmqvist. ‘Two Heads are Better than One’: Pair- Interviews for Visualization. In IEEE VIS, 2023. doi: 10.1109/VIS54172. 2023.00050 3
-
[2]
R. C. Basole and T. Major. Generative ai for visualization: Opportunities and challenges. IEEE Computer Graphics and Applications, 44(2):55–64,
-
[3]
J. Batch and N. Elmqvist. The interactive visualization gap in initial ex- ploratory data analysis. IEEE Transactions on Visualization and Computer Graphics, 24(1):104–113, 2018. doi: 10.1109/TVCG.2017.2743990 3
arXiv 2018
-
[4]
D. Cay, T. Nagel, and A. E. Yantaç. Understanding user experience of covid-19 maps through remote elicitation interviews. In 2020 IEEE Workshop on Evaluation and Beyond-Methodological Approaches to Visu- alization (BELIV), pp. 65–73. IEEE, 2020. 3
work page 2020
-
[5]
J.-F. Chen, C.-C. Ni, P.-H. Lin, and R. Lin. Designing the future: A case study on human-ai co-innovation. Creative Education, 15(3):474–494,
-
[6]
A. Chiarello and A. Author2. Generative large language models in engi- neering design: Opportunities and challenges. Design Studies, 78:101073,
-
[7]
N. E. I. . C. Eric Newburger. An interview study on the role of visualization for inferential statistics. IEEE Transactions on Visualization and Computer Graphics, 30(1):110–120, 2024. doi: 10.1109/TVCG.2023.3326521 3
arXiv 2024
-
[8]
H. Fill and A. Muff. Visualization in the era of artificial intelligence. Jusletter IT, 2023. 2
work page 2023
Show all 51 references
-
[9]
H.-G. Fill, P. Fettke, and J. Köpke. Conceptual modeling and large lan- guage models: impressions from first experiments with chatgpt.Enterprise Modelling and Information Systems Architectures (EMISAJ) , 18:1–15,
-
[10]
H.-G. Fill, F. Härer, I. Vasic, D. Borcard, B. Reitemeyer, F. Muff, S. Curty, and M. Bühlmann. Cmag: A framework for conceptual model augmented generative artificial intelligence. 2024. 2
2024
-
[11]
Gatti and A
A. Gatti and A. Author2. How chatgpt can inspire and improve serious board game design. International Journal of Game-Based Learning , 14(1):22–35, 2024. doi: 10.4018/IJGBL.2024010102 2
2024 doi
-
[12]
Gomez and A
A. Gomez and A. Author2. Large language models in complex system design. In Proceedings of the Design Society, vol. 3, pp. 123–132, 2024. doi: 10.1017/pds.2024.13 2
2024 doi
-
[13]
He, Z.-Q
J.-Y . He, Z.-Q. Cheng, C. Li, J. Sun, W. Xiang, X. Lin, X. Kang, Z. Jin, Y . Hu, and B. Luo. Wordart designer: user-driven artistic typography synthesis using large language models. arXiv preprint arXiv:2310.18332,
-
[14]
Hogan, U
T. Hogan, U. Hinrichs, and E. Hornecker. The elicitation interview tech- nique: Capturing people’s experiences of data representations. IEEE transactions on visualization and computer graphics, 22(12):2579–2593,
-
[15]
Y . Hou, M. Yang, H. Cui, L. Wang, J. Xu, and W. Zeng. C2ideas: Support- ing creative interior color design ideation with a large language model. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, pp. 1–18, 2024. 2
2024
-
[16]
Hutchinson, R
M. Hutchinson, R. Jianu, A. Slingsby, and P. Madhyastha. Llm- assisted visual analytics: Opportunities and challenges. arXiv preprint arXiv:2409.02691, 2024. 2
2024 arXiv
-
[17]
C. Y . Kim, C. P. Lee, and B. Mutlu. Understanding large-language model (llm)-powered human-robot interaction. In Proceedings of the 2024 ACM/IEEE international conference on human-robot interaction, pp. 371–380, 2024. 2
2024
-
[18]
G. Kim, H. Lee, D. Kim, H. Jung, S. Park, Y . Kim, S. Yun, T. Kil, B. Lee, and S. Park. Visually-situated natural language understanding with contrastive reading model and frozen large language models. arXiv preprint arXiv:2305.15080, 2023. 2
2023 arXiv
-
[19]
J. Kim, S. Lee, H. Jeon, K.-J. Lee, H.-J. Bae, B. Kim, and J. Seo. Phe- noflow: A human-llm driven visual analytics system for exploring large and complex stroke datasets. IEEE Transactions on Visualization and Computer Graphics, 2024. 2
2024
-
[20]
N. W. Kim, H.-K. Ko, G. Myers, and B. Bach. Chatgpt in data visualization education: A student perspective. In 2024 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), pp. 109–120. IEEE,
2024
-
[21]
Kovalerchuk, R
B. Kovalerchuk, R. Andonie, N. Datia, K. Nazemi, and E. Banissi. Visual knowledge discovery with artificial intelligence: Challenges and future directions. In Integrating artificial intelligence and visualization for visual knowledge discovery, pp. 1–27. Springer, 2022. 2
2022
-
[22]
Koziolek, S
H. Koziolek, S. Grüner, R. Hark, V . Ashiwal, S. Linsbauer, and N. Es- kandani. Llm-based and retrieval-augmented control code generation. In Proceedings of the 1st International Workshop on Large Language Models for Code, pp. 22–29, 2024. 2
2024
-
[23]
Y . Li, H. Xu, and F. Tian. From shots to stories: Llm-assisted video editing with unified language representations. arXiv preprint arXiv:2505.12237,
-
[24]
Q. V . Liao, H. Subramonyam, J. Wang, and J. Wortman Vaughan. Design- erly understanding: Information needs for model transparency to support design ideation for ai-powered user experience. In Proceedings of the 2023 CHI conference on human factors in computing systems, pp. 1–21,
2023
-
[25]
Z. Liu, X. Xie, M. He, W. Zhao, Y . Wu, L. Cheng, H. Zhang, and Y . Wu. Smartboard: Visual exploration of team tactics with llm agent. IEEE Transactions on Visualization and Computer Graphics, 2024. 2
2024
-
[26]
Y . Ma, Y . He, H. Wang, A. Wang, L. Shen, C. Qi, J. Ying, C. Cai, Z. Li, and H.-Y . Shum. Follow-your-click: Open-domain regional image animation via motion prompts. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, pp. 6018–6026, 2025. 2
2025
-
[27]
T. D. E. W. J. Mace. A qualitative interview study of distributed tracing visualisation. IEEE Transactions on Visualization and Computer Graphics, 30(7):130–140, 2024. doi: 10.1109/TVCG.2023.3241596 3
2024
-
[28]
Maddigan and M
D. Maddigan and M. Sušnjak. Chat2vis: Generating data visualizations via natural language. In Proceedings of IEEE VIS, pp. 456–467, 2022. 2
2022
-
[29]
Meyer, M
M. Meyer, M. Sedlmair, P. Quinan, and T. Munzner. Criteria for rigor in visualization design study. IEEE Transactions on Visualization and Computer Graphics, 26(1):87–97, 2019. 1
2019
-
[30]
Narechania, A
A. Narechania, A. Srinivasan, and J. Stasko. Nl4dv: A toolkit for gener- ating analytic specifications for data visualization from natural language queries. IEEE Transactions on Visualization and Computer Graphics , 27(2):369–379, 2020. 2
2020
-
[31]
Rozo-Torres, C
A. Rozo-Torres, C. J. Latorre-Rojas, and W. J. Sarmiento. Prompt engineering-based video prototyping for immersive interaction design: Limits, opportunities and perspectives. In Iberoamerican Workshop on Human-Computer Interaction, pp. 252–266. Springer, 2025. 2
2025
-
[32]
S. Saha. Human-ai collaboration: Exploring interfaces for interactive machine learning. IEEE Transactions on Visualization and Computer Graphics, 28(6):2264–2275, 2022. 2
2022
-
[33]
B. G. Schelble, C. Flathmann, N. J. McNeese, G. Freeman, and R. Mallick. Let’s think together! assessing shared mental models, performance, and trust in human-agent teams. Proceedings of the ACM on Human-Computer Interaction, 6(GROUP):1–29, 2022. 2
2022
-
[34]
Sedlmair, M
M. Sedlmair, M. Meyer, and T. Munzner. Design study methodology: Reflections from the trenches and the stacks. IEEE Transactions on Visualization and Computer Graphics, 18(12):2431–2440, 2012. 1, 2, 3, 5
2012
-
[35]
Shanbhag
P. Shanbhag. Tewen: A prompt-based system for context-aware website generation. 2025. 2
2025
-
[36]
L. Shen, H. Li, Y . Wang, and H. Qu. From data to story: Towards automatic animated data video creation with llm-based multi-agent systems. In 2024 IEEE VIS Workshop on Data Storytelling in an Era of Generative AI (GEN4DS), pp. 20–27. IEEE, 2024. 2
2024
-
[37]
S. Shen, Z. Lin, W. Liu, C. Xin, W. Dai, S. Chen, X. Wen, and X. Lan. Dashchat: Interactive authoring of industrial dashboard design proto- types through conversation with llm-powered agents. arXiv preprint arXiv:2504.12865, 2025. 2
2025 arXiv
-
[38]
S. Shin, S. Hong, and N. Elmqvist. Visualizationary: Automating de- sign feedback for visualization designers using llms. arXiv preprint arXiv:2409.13109, 2024. 2
2024 arXiv
-
[39]
F. K. Sufi. Ai-globalevents: A software for analyzing, identifying and explaining global events with artificial intelligence. Software Impacts, 11:100218, 2022. 2
2022
-
[40]
Sun and A
A. Sun and A. Author2. Llms and diffusion models in ui/ux: Advancing human-computer interaction. International Journal of Human-Computer Studies, 150:102630, 2024. doi: 10.1016/j.ijhcs.2024.102630 2
2024
-
[41]
Swanson and A
A. Swanson and A. Author2. The virtual lab: Ai agents design new sars- cov-2 nanobodies with experimental validation. Nature Biotechnology, 42:123–130, 2024. doi: 10.1038/s41587-024-01234-5 2
2024 doi
-
[42]
Walny, C
J. Walny, C. Frisson, M. West, D. Kosminsky, S. Knudsen, S. Carpendale, and W. Willett. Data changes everything: Challenges and opportunities in data visualization design handoff. IEEE Transactions on Visualization and Computer Graphics, 26(1):12–22, 2019. 2
2019
-
[43]
A. Wu, D. Deng, F. Cheng, Y . Wu, S. Liu, and H. Qu. In defence of visual analytics systems: Replies to critics. IEEE Transactions on Visualization and Computer Graphics, 29(1):1026–1036, 2022. 3
2022
-
[44]
J. Xu, W. Du, X. Liu, and X. Li. Llm4workflow: An llm-based automated workflow model generation tool. In Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering, pp. 2394– 2398, 2024. 2
2024
-
[45]
Xu and E
L. Xu and E. Wall. Exploring the capability of llms in performing low- level visual analytic tasks on svg data visualizations. IEEE Transactions on Visualization and Computer Graphics, 2023. 2
2023
-
[46]
Zamfirescu-Pereira, E
J. Zamfirescu-Pereira, E. Jun, M. Terry, Q. Yang, and B. Hartmann. Be- yond code generation: Llm-supported exploration of the program design space. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pp. 1–17, 2025. 2
2025
-
[47]
Z. Zeng, W. Watson, N. Cho, S. Rahimi, S. Reynolds, T. Balch, and M. Veloso. Flowmind: automatic workflow generation with llms. In Proceedings of the Fourth ACM International Conference on AI in Finance, pp. 73–81, 2023. 2
2023
-
[48]
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, and Z. Dong. A survey of large language models. arXiv preprint arXiv:2303.18223, 1(2), 2023. 2
2023 arXiv
-
[49]
Y . Zhao, J. Wang, L. Xiang, X. Zhang, Z. Guo, C. Turkay, Y . Zhang, and S. Chen. Lightva: Lightweight visual analytics with llm agent-based task planning and execution. IEEE Transactions on Visualization and Computer Graphics, 2024. 2
2024
-
[50]
Y . Zhao, Y . Zhang, Y . Zhang, X. Zhao, J. Wang, Z. Shao, C. Turkay, and S. Chen. Leva: Using large language models to enhance visual analytics. IEEE Transactions on Visualization and Computer Graphics, 2024. 2
2024
-
[2024]
doi: 10.1016/j.destud.2024.101073 2
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.