Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Jupybara argues that a semantic-rhetorical-pragmatic design space, encoded in LLM prompts and multi-agent critique, improves actionable data analysis and storytelling.

desk verdict Solid design-space and system paper; the abstract's effectiveness claim outruns the confounded small-n evaluation. read the letter →

arxiv 2501.16661 v1 pith:PR7VIQEQ submitted 2025-01-28 cs.HC

classification cs.HC
keywords ActionableInsightsHuman-AICollaborationMulti-AgentSystemLargeLanguageModelExploratoryDataAnalysisStorytellingSemanticsPragmatics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the messy process of turning data into decisions can be guided by a three-part design space, and that large language models can be made to follow it. The three parts are semantic precision (saying exactly what the data and results mean), rhetorical persuasion (choosing analytical methods and narrative structure that support an argument), and pragmatic relevance (connecting findings to domain knowledge and concrete actions). The authors build Jupybara, a Jupyter Notebook assistant, on two strategies that put this design space inside LLM workflows: prompts distilled from the three dimensions, and multi-agent teams in which critics and a refiner revise responses until the critics agree they are ready. Nine expert analysts gave Jupybara higher ratings than a web-based LLM tool on usability, steerability, explainability, and reparability, and rated its multi-agent mode significantly better than its single-agent mode on all three design-space dimensions. If the claim holds, AI copilots for data analysis can be steered toward actionable insight instead of merely plausible output.

What carries the argument

The load-bearing mechanism is the combination of the three-dimensional design space with two operationalization strategies. The design space is the source of the quality criteria: semantic precision, rhetorical persuasion, and pragmatic relevance. To make these criteria usable by a language model, the authors do not hand the model abstract definitions; they distill each dimension into concrete behavioral instructions, such as always interpreting statistical results and visualizations, keeping the user informed of the analysis plan, and narrating the analytical strategies used. The second strategy is multi-agent architecture: an Initial Respondent drafts a response, a set of Critics (for EDA: analysis plan, code, visualization, and interpretation/summary; for storytelling: one per design-space dimension) evaluate it, and a Refiner revises it, with the cycle repeating until all critics approve or a maximum number of discussion rounds is reached. This design means every dimension is checked by multiple agents and the final output is the product of explicit, inspectable debate rather than a single pass.

What would settle it

A controlled crossover study would settle the effectiveness claims: expert analysts analyze the same private datasets with Jupybara and with a generic web LLM tool, with order randomized, reviewers blind to which system generated each response, and processing time or cost equalized between modes. If blinded reviewers no longer rate the multi-agent output higher on semantic, rhetorical, and pragmatic quality, the architecture's claimed advantage is not established. A separate falsifier for the design space itself would be a documented recurring analyst challenge that fits none of the three dimensions.

Watch

Extended reading notes

Core claim

The central claim is that actionable exploratory data analysis and storytelling can be understood through a design space with three dimensions—semantic, rhetorical, pragmatic—and that an LLM assistant which operationalizes these dimensions produces measurably better support for analysts. Semantic work means pinning down analytical objects and results and verbalizing them accurately; rhetorical work means selecting analytical strategies and narrative moves that build a persuasive case; pragmatic work means grounding data facts in domain knowledge and translating them into recommendations. The paper reports that Jupybara, which encodes these dimensions in its prompts and in its multi-agent critique-and-refine loop, was rated by nine expert analysts as more usable, steerable, explainable, and reparable than a generic web-based LLM data tool, and that its multi-agent mode outperformed its single-agent mode on all three dimensions with significance surviving correction.

Load-bearing premise

The load-bearing premise is that the three dimensions distilled from nine expert interviews—semantic precision, rhetorical persuasion, pragmatic relevance—really do cover what makes actionable EDA and storytelling effective; if a major ingredient is missing or misweighted, Jupybara optimizes the wrong objectives and the evaluation, which rates output along those same three dimensions, cannot demonstrate true effectiveness.

Editorial extensions

If this is right

  • LLM assistants for data analysis can be built against explicit quality criteria—semantic precision, rhetorical persuasion, pragmatic relevance—rather than generic helpfulness, and those criteria double as evaluation rubrics.
  • Complex analytical questions should be routed to the multi-agent mode when response quality matters more than speed, because the multi-agent mode takes roughly five times as long as the single-agent mode but earns significantly higher quality ratings.
  • Actionable EDA and storytelling can live in one notebook workflow: analysis plans, code, interpretations, insight summaries, and editable data stories are all produced and cross-referenced inside Jupyter rather than in separate tools.
  • The design space subsumes the four analyst challenges identified in the formative study (choosing analytical strategies, tracking insights and history, finding the right language and narrative, and leveraging domain knowledge), so meeting the three dimensions addresses those challenges simultaneously.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: because the multi-agent mode also uses more time and tokens, the observed quality advantage might come from extra effort rather than from the critique architecture itself; a matched-cost comparison with equal latency and token budgets would separate those explanations.
  • Editorial extension: if the three-dimensional design space generalizes, it could serve as a shared rubric for auditing AI-generated insights in other venues, such as dashboards, business reports, and decision-support documents.
  • Editorial extension: the design space was derived from nine expert interviews, so its completeness is an open empirical question; domains with strong ethical, privacy, or fairness constraints may reveal additional dimensions beyond the pragmatic one.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a three-dimensional design space for actionable exploratory data analysis (EDA) and data storytelling, comprising semantic precision, rhetorical persuasion, and pragmatic relevance, and claims this space is grounded in theory and nine expert interviews. The authors instantiate the space in Jupybara, a Jupyter Notebook extension that uses two operationalization strategies: design-space-aware prompting and multi-agent critique-and-refine architectures. The paper reports a summative evaluation with nine expert analysts in which participants rated Jupybara against ChatGPT's data analysis plugin on usability-related dimensions and rated Jupybara's multi-agent mode against its single-agent mode on the three design-space dimensions. The central claim is that the expert evaluation supports Jupybara's usability, steerability, explainability, and reparability, and the effectiveness of the two strategies for operationalizing the design space.

Significance. If the causal claims were supported, the paper would make a useful contribution by showing that a design-space-derived prompting and multi-agent critique protocol improves actionable EDA and storytelling over single-agent LLM responses, and by providing a transferable architecture for such systems. The design space itself is a plausible synthesis of prior work and interview data, and the Jupybara implementation is broad and thoughtfully integrated with Jupyter workflows. The paper is also candid in Section 6.1 about several known limitations, which is a strength. However, the headline claims about strategy effectiveness rest on an evaluation design whose confounding factors prevent causal identification; the design space is also validated partly through rating instruments that reuse its own vocabulary. The result, as currently stated, is therefore not established, though the underlying system and framework remain potentially valuable.

major comments (4)
  1. [Section 5.2 and 5.3.1, Figure 12] The single-agent versus multi-agent comparison cannot support the claim that the multi-agent architecture is more effective. Every participant first used single-agent mode and then repeated the same analysis in multi-agent mode, for both EDA and storytelling. Mode is therefore confounded with practice, familiarity with the dataset, and carryover of earlier insights; moreover, the multi-agent pipeline produces longer, more elaborate outputs and has roughly five times the latency, a factor the authors themselves identify in Section 5.3.7 and Section 6.1 as a possible source of perceived thoroughness. The significant Wilcoxon results in Figure 12 are consistent with these alternative explanations. A counterbalanced or between-subjects design, or a condition that equalizes output length and presentation format, would be needed to attribute the ratings to the multi-agent architecture.
  2. [Section 5.2 and 5.3.1, Figure 11] The Jupybara-versus-ChatGPT comparison is not an in-session baseline: participants rated ChatGPT's data analysis plugin from prior experience and never used it during the study, as acknowledged in Section 6.1. The reported p-values are nonsignificant after Holm-Bonferroni correction, yet the Abstract states that the evaluation confirms Jupybara's superiority. The current text appropriately says 'the trend clearly indicates' within the body, but this more cautious language should also be reflected in the Abstract and the contributions list, unless a controlled comparison is conducted.
  3. [Section 4.3 and Abstract] The two strategies, design-space-aware prompting and multi-agent architectures, are never independently varied. In the evaluation, the multi-agent mode differs from the single-agent mode both in agent count and in the prompting/guidance structure, so any difference in ratings could be due to either factor or their interaction. The claimed effectiveness of 'our strategies' as a pair is therefore not separable, and the specific attribution to multi-agent collaboration in Section 5.3.7 is not identified.
  4. [Section 5.3.1, questionnaire wording] The Part 2 questionnaire asks participants to rate 'provides precise and contextually rich interpretation of results', 'generates coherent and persuasive narrative', and 'offers high-quality actionable insights', which are nearly verbatim restatements of the semantic, rhetorical, and pragmatic dimensions the system was designed to optimize. The evaluation therefore risks being circular as a validation of the design space: the outcome measures share their construct vocabulary with the intervention's design goals. A concrete test that would avoid this circularity is an independent rubric-based evaluation of the generated artifacts by blinded raters who did not see the design-space definitions, or an artifact-level analysis of whether the outputs actually satisfy the three dimensions.
minor comments (6)
  1. [Section 1, first sentence] The name 'Ben Schneiderman' should be spelled 'Ben Shneiderman'.
  2. [Section 4.4.3] The text 'Rae is able to quickly trace the the graphical representation' contains a duplicated article, which should be corrected.
  3. [Figures 11 and 12] The figures report p-values but not the corresponding Wilcoxon test statistics, sample sizes per item, or effect sizes; adding these would make the statistical claims more transparent.
  4. [Section 5.3.7] The sentence 'All but one participant preferred the multi-agent mode' would benefit from a precise count and from a direct reconciliation with the questionnaire results, especially since some participants expressed latency and verbosity concerns.
  5. [Section 6.1] The limitation paragraph is honest and welcome, but the Abstract and contributions currently state that the evaluation 'confirms' the effectiveness of the strategies; the language should be aligned so that the known confounds are not understated in the paper's front matter.
  6. [Section 3.1 and 5.1] The two participant pools both have nine experts and include two overlapping participants; it would be clearer to state explicitly that the summative evaluation is not fully independent of the formative interview sample.

Circularity Check

2 steps flagged · score 4.0 of 10

Evaluation confirms the framework by measuring the framework's own dimensions, and the framework itself is inherited from the authors' prior work.

  1. self definitional [Section 5.2 (Questionnaire), Section 4.3.2 (Multi-Agent Architectures), Abstract]
    "The second part of the questionnaire focused on evaluating the quality of responses across the three dimensions of our design space—the extent to which the systems 'provides precise and contextually rich interpretation of results,' 'generates coherent and persuasive analyses and narratives,' and 'offers high-quality actionable insights'—for the single- and multi-agent modes of Jupybara. Notably, each dimension of the design space is addressed by at least three agents, potentially enhancing the quality of the response."

    The Part 2 dependent variables are restatements of the semantic, rhetorical, and pragmatic dimensions that the system's design-space-aware prompts and dimension-specialized critics were explicitly built to optimize (Sections 4.3.1-4.3.2). The multi-agent architecture assigns at least three agents to each design-space dimension and then asks participants to rate outputs on those same three dimensions, so the reported multi-agent advantage is an internal consistency check of the architecture-to-questionnaire mapping. It cannot independently confirm the design space's claim to underpin effective EDA and storytelling; the only external comparison (ChatGPT) was non-significant after Holm-Bonferroni correction and lacked an in-session baseline.

  2. self citation load bearing [Section 2.1 (Related Work); cf. Section 3.2]
    "Their work outlined in a preliminary form a design space integrating semantics, rhetorics, and pragmatics to frame how language can better communicate actionable insights. Our current work considerably elaborates this design space and implements strategies that leverage LLMs for actionable EDA and storytelling."

    The three-dimensional design space that the whole paper operationalizes is explicitly inherited from an arXiv preprint by two of the three co-authors (Setlur and Birnbaum). The paper's own evaluation then uses those same three dimensions as the rating scales, so the framework's authority is imported from the authors' prior work rather than established by an independent benchmark. The expert interviews and the Jupybara implementation do supply independent content, which keeps the circularity partial rather than total.

full rationale

The claimed derivation chain is: prior framework plus expert interviews -> three-dimensional design space -> design-space-aware prompting and multi-agent critics -> Jupybara -> expert evaluation 'confirms' the strategies. The load-bearing circular element is at the last link: the evaluation's Part 2 instrument is built from the same three dimensions (semantic, rhetorical, pragmatic) that the prompts and critics were explicitly designed to optimize, so the significant multi-agent advantage on those scales is an internal consistency check, not an independent confirmation that the dimensions underpin effective EDA and storytelling. The design space itself also originates in the authors' own earlier preprint, making the framework's authority partially self-citational. The rest of the system (ReACT workflow, critique-refiner loop, insight DAGs, story editing, reparability features) is substantial and could in principle be tested against external benchmarks; the ChatGPT comparison was intended as an external anchor but was non-significant after correction and lacked an in-session baseline. Fixed order and latency confounds are methodological risks rather than circularity, so they do not further raise the score. Overall, the central claim retains independent content, but the evaluation's success measure is endogenous to the framework, giving a partial circularity score of 4.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper's central claims rest on the validity of the design space derived from a small interview sample, and on the assumption that subjective ratings reflect output quality. No free parameters are fitted; no new physical or conceptual entities are postulated beyond the design-space framing itself.

assumptions (4)
  • domain assumption The three-dimensional design space (semantic, rhetorical, pragmatic) captures the essential dimensions of effective actionable EDA and storytelling
    Derived from nine expert interviews and prior theory; assumed complete and generalizable. Stated in Sections 3.4 and 8.
  • domain assumption LLMs can generate coherent narratives and apply domain expertise to interpret analytical results
    Background premise for the entire system; stated in Introduction, Section 1.
  • domain assumption Expert ratings on Likert scales reflect actual quality of actionable EDA and storytelling outputs
    The evaluation relies on subjective participant ratings; assumed valid despite potential novelty and ordering biases. Sections 5.2, 6.1.
  • domain assumption ReACT prompting and multi-agent critique improve LLM response quality
    Borrowed from prior literature [112, 21, 54]; assumed to transfer to EDA and storytelling. Sections 4.2, 4.3.2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs." pith.science (2026). https://pith.science/paper/PR7VIQEQ

@misc{pith2026250116661,
  author       = {Pith},
  title        = {Pith review of: Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PR7VIQEQ}},
  note         = {Machine review of arXiv:2501.16661}
}
read the original abstract

Mining and conveying actionable insights from complex data is a key challenge of exploratory data analysis (EDA) and storytelling. To address this challenge, we present a design space for actionable EDA and storytelling. Synthesizing theory and expert interviews, we highlight how semantic precision, rhetorical persuasion, and pragmatic relevance underpin effective EDA and storytelling. We also show how this design space subsumes common challenges in actionable EDA and storytelling, such as identifying appropriate analytical strategies and leveraging relevant domain knowledge. Building on the potential of LLMs to generate coherent narratives with commonsense reasoning, we contribute Jupybara, an AI-enabled assistant for actionable EDA and storytelling implemented as a Jupyter Notebook extension. Jupybara employs two strategies -- design-space-aware prompting and multi-agent architectures -- to operationalize our design space. An expert evaluation confirms Jupybara's usability, steerability, explainability, and reparability, as well as the effectiveness of our strategies in operationalizing the design space framework with LLMs.

Figures

Figures reproduced from arXiv: 2501.16661 by the authors.

Figure 1
Figure 1. The interface of Jupybara, an AI-enabled assistant for actionable EDA and data storytelling implemented as a Jupyter Notebook extension. Jupybara operationalizes our proposed design space consisting of the semantic, rhetorical, and pragmatic dimensions. (A) For a complex user query in EDA, Jupybara identifies and presents an analysis plan before producing code. (B) In the data story generated by Jupybara, the system… view at source ↗
Figure 2
Figure 2. The multi-agent architecture for EDA in Jupybara. Given a user query, an Initial Respondent provides an initial response, which is then critiqued by four Critics. Based on the critiques, the Refiner improves the response and sends the revised version back to the Critics for review. The discussion between the Critics and the Refiner continues until all Critics agree the response is ready or a defined maximum number o… view at source ↗
Figure 3
Figure 3. The multi-agent architecture for data storytelling in [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: (A) Rae inputs her command in a Notebook cell and clicks the “Invoke AI” icon in the cell toolbar. [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: Rae tasks Jupybara with a complex query under the single-agent mode. Utilizing its agentic workflow, Jupybara produces code, executes it, and provides an interpretation. Nonetheless, this response is not immediately intuitive to Rae due to the lack of units, the limite…
Figure 6
Figure 6. Figure 6: (A) Rae switches to the multi-agent mode for the previous complex query. (See the video walkthrough in the [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: Rae wants to summarize the analysis history and insights gained so far. She navigates to the [PITH_FULL_IMAGE:figures/full_fig_p015_7.png]
Figure 8
Figure 8. Figure 8: Rae navigates to the Settings tab to activate the multi-agent mode for data storytelling (not shown here; please see video supplement). To generate a data story, Rae navigates to the Storytelling tab. She inputs her instructions and clicks “Generate Data Story.” Jupyba…
Figure 9
Figure 9. Figure 9: Rae can provide both global feedback on the entire AI-generated data story and local feedback focused on specific [PITH_FULL_IMAGE:figures/full_fig_p017_9.png]
Figure 10
Figure 10. Figure 10: Rae can also manually edit the data story by clicking “Edit,” which opens a side-by-side live HTML editor next to the [PITH_FULL_IMAGE:figures/full_fig_p018_10.png]
Figure 11
Figure 11. Figure 11: Participants separately rated ChatGPT’s data analysis plugin and [PITH_FULL_IMAGE:figures/full_fig_p018_11.png]
Figure 12
Figure 12. Figure 12: Participants separately rated the single- and multi-agent modes of [PITH_FULL_IMAGE:figures/full_fig_p019_12.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A multi-agent vision-language framework that extracts event, time, and location from public event images, evaluated with a new soft metric on VLM-augmented datasets.

Reference graph

Works this paper leans on

129 extracted references · 60 canonical work pages · cited by 1 Pith paper

  1. [1]

    Anthropic. 2024. Learn about Claude. https://docs.anthropic.com/en/docs/about- claude/models

  2. [2]

    Sriram Karthik Badam, Zhicheng Liu, and Niklas Elmqvist. 2018. Elastic Doc- uments: Coupling Text and Tables through Contextual Visualizations for En- hanced Document Reading. IEEE Transactions on Visualization and Computer Graphics 25, 1 (2018), 661–671

  3. [3]

    Konrad Banachewicz. 2024. Ireland: Gender Pay Gaps. https://www.kaggle. com/datasets/konradb/ireland-gender-pay-gaps/data

  4. [4]

    Ori Bar El, Tova Milo, and Amit Somech. 2020. Automatically Generating Data Exploration Sessions Using Deep Reinforcement Learning. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data. 1527–1537

  5. [5]

    In- sight

    Leilani Battle and Alvitta Ottley. 2023. What Do We Mean When We Say “In- sight"? A Formal Synthesis of Existing Theory.IEEE Transactions on Visualization and Computer Graphics (2023)

  6. [6]

    Alexander Bendeck, Dennis Bromley, and Vidya Setlur. 2024. SlopeSeeker: A Search Tool for Exploring a Dataset of Quantifiable Trends. In Proceedings of the 29th International Conference on Intelligent User Interfaces . 817–836

  7. [7]

    Brachman

    Philip S. Brachman. 1996. Epidemiology. In Medical Microbiology (4th ed.), Samuel Baron (Ed.). University of Texas Medical Branch at Galveston, Galveston, TX, Chapter 9. https://www.ncbi.nlm.nih.gov/books/NBK7993/

  8. [8]

    Dennis Bromley and Vidya Setlur. 2023. What Is the Difference Between a Mountain and a Molehill? Quantifying Semantic Labeling of Visual Features in Line Charts. IEEE Transactions on Visualization and Computer Graphics (2023)

Show all 129 references
  1. [9]

    Yining Cao, Jane L E, Zhutian Chen, and Haijun Xia. 2023. DataParticles: Block- Based and Language-Oriented Authoring of Animated Unit Visualizations. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–15

  2. [10]

    Olbrechts-Tyteca Chaïm Perelman

    L. Olbrechts-Tyteca Chaïm Perelman. 1969. The New Rhetoric: A Treatise on Argumentation. University of Notre Dame Press. http://www.jstor.org/stable/j. ctvpj74xx

  3. [11]

    Remco Chang, Caroline Ziemkiewicz, Tera Marie Green, and William Ribarsky

  4. [12]

    Qing Chen, Shixiong Cao, Jiazhe Wang, and Nan Cao. 2023. How Does Au- tomation Shape the Process of Narrative Visualization: A Survey of Tools. IEEE Transactions on Visualization and Computer Graphics (2023)

  5. [13]

    Liying Cheng, Xingxuan Li, and Lidong Bing. 2023. Is GPT-4 a Good Data Analyst?. In Findings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Com- putational Linguistics, Singapore, 9496–9514. https...

  6. [14]

    Eun Kyoung Choe, Bongshin Lee, et al . 2015. Characterizing Visualization Insights from Quantified Selfers’ Personal Data Presentations. IEEE Computer Graphics and Applications 35, 4 (2015), 28–37

  7. [15]

    Anamaria Crisan, Brittany Fiore-Gartland, and Melanie Tory. 2020. Passing the Data Baton: A Retrospective Analysis on Data Science Work and Workers.IEEE Transactions on Visualization and Computer Graphics 27, 2 (2020), 1860–1870

  8. [16]

    Dursun Delen and Sudha Ram. 2018. Research Challenges and Opportunities in Business Analytics. Journal of Business Analytics 1, 1 (2018), 2–12

  9. [17]

    McCurley, Ralfi Nahmias, Mukund Sundararajan, and Qiqi Yan

    Kedar Dhamdhere, Kevin S. McCurley, Ralfi Nahmias, Mukund Sundararajan, and Qiqi Yan. 2017. Analyza: Exploring Data with Conversation. In Proceedings of the 22nd International Conference on Intelligent User Interfaces (IUI 2017) . 493–504

  10. [18]

    Evanthia Dimara and John Stasko. 2021. A Critical Reflection on Visualiza- tion Research: Where do Decision Making Tasks Hide? IEEE Transactions on Visualization and Computer Graphics 28, 1 (2021), 1128–1138

  11. [19]

    It’s Like A Rubber Duck That Talks Back

    Ian Drosos, Advait Sarkar, Xiaotong (Tone) Xu, Carina Negreanu, Sean Rintel, and Lev Tankelevitch. 2024. “It’s Like A Rubber Duck That Talks Back": Under- standing Generative AI-Assisted Data Analysis Workflows through a Participa- tory Prompting Study. Proceedings of the 3rd ...

  12. [20]

    Steven Drucker, Samuel Huron, Robert Kosara, Jonathan Schwabish, and Nicholas Diakopoulos. 2018. Communicating Data to an Audience. In Data- Driven Storytelling. AK Peters/CRC Press, 211–231

  13. [21]

    Tenenbaum, and Igor Mor- datch

    Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mor- datch. 2023. Improving Factuality and Reasoning in Language Models through Multiagent Debate. ArXiv abs/2305.14325 (2023). https://api.semanticscholar. org/CorpusID:258841118

  14. [22]

    Klaus Eckelt, Kiran Gadhave, Alexander Lex, and Marc Streit. 2024. Loops: Leveraging Provenance and Visualization to Support Exploratory Data Analysis in Notebooks. In COMPUTER GRAPHICS forum , Vol. 43

  15. [23]

    Will Epperson, Vaishnavi Gorantla, Dominik Moritz, and Adam Perer. 2023. Dead or Alive: Continuous Data Profiling for Interactive Data Science. IEEE Transactions on Visualization and Computer Graphics (2023)

  16. [24]

    The Data Says Otherwise

    Yu Fu, Shunan Guo, Jane Hoffswell, Victor S. Bursztyn, Ryan Rossi, and John Stasko. 2024. " The Data Says Otherwise"—Towards Automated Fact-checking and Communication of Data Claims. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology . 1–20

  17. [25]

    Karahalios

    Tong Gao, Mira Dontcheva, Eytan Adar, Zhicheng Liu, and Karrie G. Karahalios

  18. [26]

    Samuel Gratzl, Alexander Lex, Nils Gehlenborg, Nicola Cosgrove, and Marc Streit. 2016. From Visual Exploration to Storytelling and Back Again. In Com- puter Graphics Forum, Vol. 35. Wiley Online Library, 491–500

  19. [27]

    Garrett Grolemund and Hadley Wickham. 2014. A Cognitive Interpretation of Data Analysis. International Statistical Review 82, 2 (2014), 184–204

  20. [28]

    Ken Gu, Madeleine Grunde-McLaughlin, Andrew McNutt, Jeffrey Heer, and Tim Althoff. 2024. How Do Data Analysts Respond to AI Assistance? A Wizard-of-Oz Study. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association f...

  21. [29]

    Ken Gu, Ruoxi Shang, Tim Althoff, Chenglong Wang, and Steven M. Drucker

  22. [30]

    Ken Gu, Ruoxi Shang, Tim Althoff, Chenglong Wang, and Steven M Drucker

  23. [31]

    Yi He, Shixiong Cao, Yang Shi, Qing Chen, Ke Xu, and Nan Cao. 2024. Leveraging Large Models for Crafting Narrative Visualization: A Survey. arXiv preprint arXiv:2401.14010 (2024)

  24. [32]

    Sture Holm. 1979. A Simple Sequentially Rejective Multiple Test Procedure. Scandinavian Journal of Statistics (1979), 65–70

  25. [33]

    Zezhou Huang and Eugene Wu. 2024. Cocoon: Semantic Table Profiling Using Large Language Models. In Proceedings of the 2024 Workshop on Human-In-the- Loop Data Analytics. 1–7

  26. [34]

    In Proceedings of the CHI Conference on Human Factors in Computing Systems

    How Do Analysts Understand and Verify AI-Assisted Data Analyses?. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–22

  27. [35]

    Jessica Hullman and Nicholas Diakopoulos. 2011. Visualization Rhetoric: Fram- ing Effects in Narrative Visualization. IEEE Transactions on Visualization and Computer Graphics 17 (12 2011), 2231–40. https://doi.org/10.1109/TVCG.2011. 255 CHI ’25, April 26-May 1, 2025, Yokohama,...

  28. [36]

    Joty, Md Tahmid Rahman Laskar, and Md

    Mohammed Saidul Islam, Enamul Hoque, Shafiq R. Joty, Md Tahmid Rahman Laskar, and Md. Rizwan Parvez. 2024. DataNarrative: Automated Data-Driven Storytelling with Visualizations and Texts. https://api.semanticscholar.org/ CorpusID:271855683

  29. [37]

    Peiling Jiang, Jude Rayan, Steven P Dow, and Haijun Xia. 2023. Graphologue: Exploring Large Language Model Responses with Interactive Diagrams. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–20

  30. [38]

    Jessica Hullman. 2019. The Purpose of Visualization is Insight, Not Pictures: An Interview with Visualization Pioneer Ben Shneiderman. https://medium.com/multiple-views-visualization-research-explained/the- purpose-of-visualization-is-insight-not-pictures-an-interview-with- vi...

  31. [39]

    Eunice Jun, Audrey Seo, Jeffrey Heer, and René Just. 2022. Tisane: Authoring Statistical Models via Formal Reasoning from Conceptual and Data Relationships. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. 1–16

  32. [40]

    Ju Yeon Jung, Tom Steinberger, John L King, and Mark S Ackerman. 2022. How Domain Experts Work with Data: Situating Data Science in the Practices and Settings of Craftwork. Proceedings of the ACM on Human-Computer Interaction 6, CSCW1 (2022), 1–29

  33. [41]

    Daniel Kahneman and Amos Tversky. 2013. Prospect Theory: An Analysis of Decision Under Risk. In Handbook of the Fundamentals of Financial Decision Making: Part I. World Scientific, 99–127

  34. [42]

    Eunice Jun, Maureen Daum, Jared Roesch, Sarah Chasins, Emery Berger, Rene Just, and Katharina Reinecke. 2019. Tea: A High-level Language and Runtime System for Automating Statistical Analysis. In Proceedings of the 32nd Annual ACM Symposium on User Interface Software and Techn...

  35. [43]

    Majeed Kazemitabaar, Jack Williams, Ian Drosos, Tovi Grossman, Austin Henley, Carina Negreanu, and Advait Sarkar. 2024. Improving Steering and Verifica- tion in AI-Assisted Data Analysis with Interactive Task Decomposition. arXiv preprint arXiv:2407.02651 (2024)

  36. [44]

    Ralph L. Keeney. 1992. Value-Focused Thinking: A Path to Creative Decisionmak- ing. Harvard University Press. http://www.jstor.org/stable/j.ctv322v4g7

  37. [45]

    Mary Beth Kery, Bonnie E John, Patrick O’Flaherty, Amber Horvath, and Brad A Myers. 2019. Towards Effective Foraging by Data Scientists to Find Past Anal- ysis Choices. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–13

  38. [46]

    Sean Kandel, Andreas Paepcke, Joseph M Hellerstein, and Jeffrey Heer. 2012. En- terprise Data Analysis and Visualization: An Interview Study. IEEE Transactions on Visualization and Computer Graphics 18, 12 (2012), 2917–2926

  39. [47]

    Jaeyoung Kim, Sihyeon Lee, Hyeon Jeon, Keon-Joo Lee, Hee-Joon Bae, Bohyoung Kim, and Jinwook Seo. 2024. PhenoFlow: A Human-LLM Driven Visual Analytics System for Exploring Large and Complex Stroke Datasets. IEEE Transactions on Visualization and Computer Graphics (2024)

  40. [48]

    Thomas Kluyver, Benjamin Ragan-Kelley, Fernando Pérez, Brian Granger, Matthias Bussonnier, Jonathan Frederic, Kyle Kelley, Jessica Hamrick, Jason Grout, Sylvain Corlay, et al. 2016. Jupyter Notebooks–A Publishing Format for Reproducible Computational Workflows. In Positioning ...

  41. [49]

    Qiao Lan, Dingzhu Wen, Zezhong Zhang, Qunsong Zeng, Xu Chen, Petar Popovski, and Kaibin Huang. 2021. What is Semantic Communication? A View on Conveying Meaning in the Era of Machine Intelligence. Journal of Communications and Information Networks 6 (12 2021), 336–371. https: ...

  42. [50]

    Dae Hyun Kim, Vidya Setlur, and Maneesh Agrawala. 2021. Towards Understand- ing How Readers Integrate Charts and Captions: A Case Study with Line Charts. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Sys- tems (Yokohama, Japan) (CHI ’21). Association ...

  43. [51]

    Guozheng Li, Runfei Li, Yunshan Feng, Yu Zhang, Yuyu Luo, and Chi Harold Liu. 2024. CoInsight: Visual Storytelling for Hierarchical Tables with Connected Insights. IEEE Transactions on Visualization and Computer Graphics (2024)

  44. [52]

    Haotian Li, Yun Wang, and Huamin Qu. 2023. Where Are We So Far? Understand- ing Data Storytelling Tools from the Perspective of Human-AI Collaboration. Proceedings of the CHI Conference on Human Factors in Computing Systems (2023). https://api.semanticscholar.org/CorpusID:263152607

  45. [53]

    Haotian Li, Lu Ying, Haidong Zhang, Yingcai Wu, Huamin Qu, and Yun Wang

  46. [54]

    Shahid Latif, Zheng Zhou, Yoon Kim, Fabian Beck, and Nam Wook Kim. 2021. Kori: Interactive Synthesis of Text and Charts in Data Documents. IEEE Trans- actions on Visualization and Computer Graphics 28, 1 (2021), 184–194

  47. [55]

    Xingjun Li, Yizhi Zhang, Justin Leung, Chengnian Sun, and Jian Zhao. 2023. Edassistant: Supporting Exploratory Data Analysis in Computational Notebooks with In Situ Code Search and Recommendation.ACM Transactions on Interactive Intelligent Systems 13, 1 (2023), 1–27

  48. [56]

    Yanna Lin, Haotian Li, Leni Yang, Aoyu Wu, and Huamin Qu. 2023. Inksight: Leveraging Sketch Interaction for Documenting Chart Findings in Computa- tional Notebooks. IEEE Transactions on Visualization and Computer Graphics (2023)

  49. [57]

    Yang Liu, Alex Kale, Tim Althoff, and Jeffrey Heer. 2020. Boba: Authoring and Visualizing Multiverse Analyses. IEEE Transactions on Visualization and Computer Graphics 27, 2 (2020), 1753–1763

  50. [58]

    Ziao Liu, Xiao Xie, Moqi He, Wenshuo Zhao, Yihong Wu, Liqi Cheng, Hui Zhang, and Yingcai Wu. 2024. Smartboard: Visual Exploration of Team Tactics with LLM Agent. IEEE Transactions on Visualization and Computer Graphics (2024)

  51. [59]

    Tenenbaum, Antonio Torralba, and Igor Mor- datch

    Shuang Li, Yilun Du, Joshua B. Tenenbaum, Antonio Torralba, and Igor Mor- datch. 2022. Composing Ensembles of Pre-Trained Models via Iterative Consen- sus. ArXiv abs/2210.11522 (2022). https://api.semanticscholar.org/CorpusID: 253080406

  52. [60]

    Pingchuan Ma, Rui Ding, Shuai Wang, Shi Han, and Dongmei Zhang. 2023. InsightPilot: An LLM-Empowered Automated Data Exploration System. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 346–352

  53. [61]

    Andrew M McNutt, Chenglong Wang, Robert A Deline, and Steven M Drucker

  54. [62]

    OpenAI. 2024. GPT-4o. https://platform.openai.com/docs/models/gpt-4o

  55. [63]

    Yang Ouyang, Leixian Shen, Yun Wang, and Quan Li. 2024. NotePlayer: Engaging Jupyter Notebooks for Dynamic Presentation of Analytical Processes. arXiv preprint arXiv:2408.01101 (2024)

  56. [64]

    Junhua Lu, Wei Chen, Hui Ye, Jie Wang, Honghui Mei, Yuhui Gu, Yingcai Wu, Xiaolong Luke Zhang, and Kwan-Liu Ma. 2021. Automatic Generation of Unit Visualization-Based Scrollytelling for Impromptu Data Facts Delivery. In 2021 IEEE 14th Pacific Visualization Symposium (PacificVi...

  57. [65]

    Nathalie Henry Riche, Christophe Hurter, Nicholas Diakopoulos, and Sheelagh Carpendale (Eds.). 2018. Data-Driven Storytelling. CRC Press

  58. [66]

    Cristóbal Romero and Sebastián Ventura. 2010. Educational Data Mining: A Re- view of the State of the Art. IEEE Transactions on Systems, Man, and Cybernetics, Part C (applications and reviews) 40, 6 (2010), 601–618

  59. [67]

    In Pro- ceedings of the 2023 CHI Conference on Human Factors in Computing Systems

    On the Design of AI-powered Code Assistants for Notebooks. In Pro- ceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–16

  60. [68]

    Adam Rule, Aurélien Tabard, and James D Hollan. 2018. Exploration and Expla- nation in Computational Notebooks. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems . 1–12

  61. [69]

    Gaurav Sahu, Abhay Puri, Juan Rodriguez, Alexandre Drouin, Perouz Taslakian, Valentina Zantedeschi, Alexandre Lacoste, David Vazquez, Nicolas Chapados, Christopher Pal, et al. 2024. InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation. arXiv...

  62. [70]

    Rock Yuren Pang, Ruotong Wang, Joely Nelson, and Leilani Battle. 2022. How Do Data Science Workers Communicate Intermediate Results?. In 2022 IEEE Visualization in Data Science (VDS) . IEEE, 46–54

  63. [71]

    Franz Sauer, Tyson Neuroth, Jacqueline Chu, and Kwan-Liu Ma. 2016. Audience- Targeted Design Considerations for Effective Scientific Storytelling. Computing in Science & Engineering 18, 6 (2016), 68–76

  64. [72]

    Schneider and Anne Barron

    Klaus P. Schneider and Anne Barron. 2014. Pragmatics of Discourse. De Gruyter Mouton

  65. [73]

    Aayushi Roy, Deepthi Raghunandan, Niklas Elmqvist, and Leilani Battle. 2023. How I Met Your Data Science Team: A Tale of Effective Communication. In2023 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) . 199–208. https://doi.org/10.1109/VL-HCC57772.2023.00032

  66. [74]

    Martin Schweinsberg, Michael Feldman, Nicola Staub, Olmo R van den Akker, Robbie CM van Aert, Marcel ALM Van Assen, Yang Liu, Tim Althoff, Jeffrey Heer, Alex Kale, et al. 2021. Same Data, Different Conclusions: Radical Dispersion in Empirical Results when Independent Analysts ...

  67. [75]

    Edward Segel and Jeffrey Heer. 2010. Narrative Visualization: Telling Stories with Data. IEEE Transactions on Visualization and Computer Graphics 16, 6 (2010), 1139–1148

  68. [76]

    Abhraneel Sarma, Alex Kale, Michael Jongho Moon, Nathan Taback, Fanny Chevalier, Jessica Hullman, and Matthew Kay. 2023. Multiverse: Multiplex- ing Alternative Data Analyses in R Notebooks. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–15

  69. [77]

    Vidya Setlur and Larry Birnbaum. 2024. Can Nuanced Language Lead to More Actionable Insights? Exploring the Role of Generative AI in Analytical Narrative Structure. arXiv preprint arXiv:2405.02763 (2024)

  70. [78]

    Vidya Setlur, Andriy Kanyuka, and Arjun Srinivasan. 2023. Olio: A Semantic Search Interface for Data Repositories. In Proceedings of the 36th Annual ACM Symposium on User Interface Software Technology (San Francisco, California) (UIST 2023). ACM, New York, NY, USA. https://doi...

  71. [79]

    Christoph Schröer, Felix Kruse, and Jorge Marx Gómez. 2021. A Systematic Literature Review on Applying CRISP-DM Process Model. Procedia Computer Science 181 (2021), 526–534

  72. [80]

    Vidya Setlur, Melanie Tory, and Alex Djalali. 2019. Inferencing Underspecified Natural Language Utterances in Visual Analysis. In Proceedings of the 24th International Conference on Intelligent User Interfaces (Marina del Rey, California) (IUI ’19). Association for Computing M...

  73. [81]

    Leixian Shen, Yizhi Zhang, Haidong Zhang, and Yun Wang. 2023. Data Player: Automatic Generation of Data Videos with Narration-Animation Interplay.IEEE Transactions on Visualization and Computer Graphics (2023)

  74. [82]

    Battersby, Melanie Tory, Rich Gossweiler, and Angel X

    Vidya Setlur, Sarah E. Battersby, Melanie Tory, Rich Gossweiler, and Angel X. Chang. 2016. Eviza: A Natural Language Interface for Visual Analysis. In Pro- ceedings of the 29th Annual Symposium on User Interface Software and Technology (Tokyo, Japan) (UIST 2016). ACM, New York...

  75. [83]

    Danqing Shi, Xinyue Xu, Fuling Sun, Yang Shi, and Nan Cao. 2020. Calliope: Automatic Visual Data Story Generation from a Spreadsheet. IEEE Transactions on Visualization and Computer Graphics 27, 2 (2020), 453–463

  76. [84]

    It’s a Good Idea to Put It Into Words

    Chase Stokes, Clara Hu, and Marti A Hearst. 2024. “It’s a Good Idea to Put It Into Words": Writing Rudders’ in the Initial Stages of Visualization Design. arXiv preprint arXiv:2407.15959 (2024)

  77. [85]

    Vidya Setlur and Melanie Tory. 2022. How Do You Converse With An Analytical Chatbot? Revisiting Gricean Maxims for Designing Analytical Conversational Behavior. In Proceedings of the 2022 CHI Conference on Human Factors in Com- puting Systems. 1–17

  78. [86]

    Sangho Suh, Bryan Min, Srishti Palani, and Haijun Xia. 2023. Sensecape: En- abling Multilevel Exploration and Sensemaking with Large Language Models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–18

  79. [87]

    Nicole Sultanum, Fanny Chevalier, Zoya Bylinskii, and Zhicheng Liu. 2021. Leveraging Text-Chart Links to Support Authoring of Data-Driven Articles with VizFlow. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–17

  80. [88]

    Danqing Shi, Fuling Sun, Xinyue Xu, Xingyu Lan, David Gotz, and Nan Cao

  81. [89]

    Mengdi Sun, Ligan Cai, Weiwei Cui, Yanqiu Wu, Yang Shi, and Nan Cao. 2022. Erato: Cooperative Data Story Editing via Fact Interpolation. IEEE Transactions on Visualization and Computer Graphics 29, 1 (2022), 983–993

  82. [90]

    Charles Sutton, Timothy Hobson, James Geddes, and Rich Caruana. 2018. Data Diff: Interpretable, Executable Summaries of Changes in Distributions for Data Wrangling. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 2279–2288

  83. [91]

    Tan Tang, Junxiu Tang, Jiewen Lai, Lu Ying, Yingcai Wu, Lingyun Yu, and Peiran Ren. 2022. Smartshots: An Optimization Approach for Generating Videos with Data Visualizations Embedded. ACM Transactions on Interactive Intelligent Systems (TiiS) 12, 1 (2022), 1–21

  84. [92]

    Chase Stokes, Vidya Setlur, Bridget Cogley, Arvind Satyanarayan, and Marti A Hearst. 2022. Striking a Balance: Reader Takeaways and Preferences When Integrating Text and Charts. IEEE Transactions on Visualization and Computer Graphics 29, 1 (2022), 1233–1243

  85. [93]

    Yuan Tian, Zheng Zhang, Zheng Ning, Toby Jia-Jun Li, Jonathan K Kummerfeld, and Tianyi Zhang. 2023. Interactive Text-to-SQL Generation via Editable Step- by-Step Explanations. arXiv preprint arXiv:2305.07372 (2023)

  86. [94]

    Weisz, Xuye Liu, Lingfei Wu, and Casey Dugan

    April Yi Wang, Dakuo Wang, Jaimie Drozdal, Michael Muller, Soya Park, Justin D. Weisz, Xuye Liu, Lingfei Wu, and Casey Dugan. 2022. Documentation Matters: Human-Centered AI System to Assist Data Science Code Documentation in Computational Notebooks. ACM Trans. Comput.-Hum. Int...

  87. [95]

    Nicole Sultanum and Arjun Srinivasan. 2023. DataTales: Investigating the Use of Large Language Models for Authoring Data-driven Articles. In 2023 IEEE Visualization and Visual Analytics (VIS) . IEEE, 231–235

  88. [96]

    Weisz, Michael Muller, Parikshit Ram, Werner Geyer, Casey Dugan, Yla Tausczik, Horst Samulowitz, and Alexander Gray

    Dakuo Wang, Justin D. Weisz, Michael Muller, Parikshit Ram, Werner Geyer, Casey Dugan, Yla Tausczik, Horst Samulowitz, and Alexander Gray. 2019. Human-AI Collaboration in Data Science: Exploring Data Scientists’ Perceptions of Automated AI. Proceedings of the ACM Human-Compute...

  89. [97]

    Fengjie Wang, Yanna Lin, Leni Yang, Haotian Li, Mingyang Gu, Min Zhu, and Huamin Qu. 2024. OutlineSpark: Igniting AI-powered Presentation Slides Cre- ation from Computational Notebooks through Outlines. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–16

  90. [98]

    Fengjie Wang, Xuye Liu, Oujing Liu, Ali Neshati, Tengfei Ma, Min Zhu, and Jian Zhao. 2023. Slide4N: Creating Presentation Slides from Computational Note- books with Human-AI Collaboration. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–18

  91. [99]

    Yuan Tian, Weiwei Cui, Dazhen Deng, Xinjing Yi, Yurun Yang, Haidong Zhang, and Yingcai Wu. 2024. ChartGPT: Leveraging LLMs to Generate Charts from Abstract Natural Language. IEEE Transactions on Visualization and Computer Graphics (2024)

  92. [100]

    Huichen Will Wang, Jane Hoffswell, Sao Myat Thazin Thane, Victor S Bursztyn, and Cindy Xiong Bearfield. 2024. How Aligned are Human Chart Takeaways and LLM Predictions? A Case Study on Bar Charts with Varying Layouts. arXiv preprint arXiv:2408.06837 (2024)

  93. [101]

    Xingbo Wang, Furui Cheng, Yong Wang, Ke Xu, Jiang Long, Hong Lu, and Huamin Qu. 2022. Interactive Data Analysis with Next-Step Natural Language Query Recommendation. arXiv preprint arXiv:2201.04868 (2022)

  94. [102]

    Weisz, Erick Oduor, and Casey Dugan

    Dakuo Wang, Josh Andres, Justin D. Weisz, Erick Oduor, and Casey Dugan. 2021. AutoDS: Towards Human-Centered Automation of Data Science. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–12

  95. [103]

    Le, Denny Zhou, et al

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, Denny Zhou, et al . 2022. Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models.Advances in Neural Information Processing Systems 35 (2022), 24824–24837

  96. [104]

    Luoxuan Weng, Xingbo Wang, Junyu Lu, Yingchaojie Feng, Yihan Liu, and Wei Chen. 2024. InsightLens: Discovering and Exploring Insights from Con- versational Contexts in Large-Language-Model-Powered Data Analysis. arXiv preprint arXiv:2404.01644 (2024)

  97. [105]

    Frank Wilcoxon. 1992. Individual Comparisons by Ranking Methods. In Break- throughs in Statistics: Methodology and Distribution . Springer, 196–202

  98. [106]

    Huichen Will Wang, Mitchell Gordon, Leilani Battle, and Jeffrey Heer. 2024. DracoGPT: Extracting Visualization Design Preferences from Large Language Models. arXiv preprint arXiv:2408.06845 (2024)

  99. [107]

    Kanit Wongsuphasawat, Yang Liu, and Jeffrey Heer. 2019. Goals, Process, and Challenges of Exploratory Data Analysis: An Interview Study. ArXiv abs/1911.00568 (2019). https://api.semanticscholar.org/CorpusID:207798031

  100. [108]

    Kanit Wongsuphasawat, Daniel Smilkov, James Wexler, Jimbo Wilson, Dandelion Mane, Doug Fritz, Dilip Krishnan, Fernanda B Viégas, and Martin Wattenberg

  101. [109]

    Yun Wang, Leixian Shen, Zhengxin You, Xinhuan Shu, Bongshin Lee, John Thompson, Haidong Zhang, and Dongmei Zhang. 2024. WonderFlow: Narration- Centric Design of Animated Data Videos. IEEE Transactions on Visualization and Computer Graphics (2024)

  102. [110]

    Liwenhan Xie, Chengbo Zheng, Haijun Xia, Huamin Qu, and Chen Zhu-Tian

  103. [111]

    Youfu Yan, Yu Hou, Yongkang Xiao, Rui Zhang, and Qianwen Wang. 2024. KNOWNET: Guided Health Information Seeking from LLMs via Knowledge Graph Integration. IEEE Transactions on Visualization and Computer Graphics (2024)

  104. [112]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing Reasoning and Acting in Language Models. arXiv preprint arXiv:2210.03629 (2022)

  105. [113]

    Rüdiger Wirth and Jochen Hipp. 2000. CRISP-DM: Towards a Standard Pro- cess Model for Data Mining. In Proceedings of the 4th International Conference on the Practical Applications of Knowledge Discovery and Data Mining , Vol. 1. Manchester, 29–39

  106. [114]

    Wong, Stefan Welker, Krzysztof Choromanski, Federico Tombari, Aveek Purohit, Michael S

    Andy Zeng, Adrian S. Wong, Stefan Welker, Krzysztof Choromanski, Federico Tombari, Aveek Purohit, Michael S. Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, and Peter R. Florence. 2022. Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language. ArXiv abs/22...

  107. [115]

    Xingchen Zeng, Haichuan Lin, Yilin Ye, and Wei Zeng. 2024. Advancing multi- modal large language models in chart question answering with visualization- referenced instruction tuning. IEEE Transactions on Visualization and Computer Graphics (2024)

  108. [116]

    Jian Zhao, Shenyu Xu, Senthil Chandrasegaran, Chris Bryan, Fan Du, Aditi Mishra, Xin Qian, Yiran Li, and Kwan-Liu Ma. 2021. ChartStory: Automated Partitioning, Layout, and Captioning of Charts into Comic-Style Narratives.IEEE Transactions on Visualization and Computer Graphics...

  109. [117]

    Guande Wu, Shunan Guo, Jane Hoffswell, Gromit Yeuk-Yin Chan, Ryan A Rossi, and Eunyee Koh. 2023. Socrates: Data Story Generation via Adaptive Machine- Guided Elicitation of User Feedback. IEEE Transactions on Visualization and Computer Graphics (2023)

  110. [118]

    Yuheng Zhao, Yixing Zhang, Yu Zhang, Xinyi Zhao, Junjie Wang, Zekai Shao, Cagatay Turkay, and Siming Chen. 2024. LEVA: Using large language models to enhance visual analytics. IEEE Transactions on Visualization and Computer Graphics (2024)

  111. [119]

    arXiv preprint arXiv:2408.01703 (2024)

    WaitGPT: Monitoring and Steering Conversational LLM Agent in Data Analysis with On-the-Fly Code Visualization. arXiv preprint arXiv:2408.01703 (2024)

  112. [120]

    Tongyu Zhou, Jeff Huang, and Gromit Yeuk-Yin Chan. 2024. Epigraphics: Message-Driven Infographics Authoring. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–18

  113. [122]

    Yue Yu, Leixian Shen, Fei Long, Huamin Qu, and Hao Chen. 2024. PyG- Walker: On-the-fly Assistant for Exploratory Visual Data Analysis. https: //api.semanticscholar.org/CorpusID:270559525

  114. [126]

    Yuheng Zhao, Junjie Wang, Linbin Xiang, Xiaowen Zhang, Zifei Guo, Cagatay Turkay, Yu Zhang, and Siming Chen. 2024. LightVA: Lightweight Visual Ana- lytics with LLM Agent-Based Task Planning and Execution. IEEE Transactions on Visualization and Computer Graphics (2024)

  115. [128]

    Chengbo Zheng, Dakuo Wang, April Yi Wang, and Xiaojuan Ma. 2022. Telling Stories from Computational Notebooks: AI-Assisted Presentation Slides Cre- ation for Presenting Data Science Work. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–20

  116. [2009]

    IEEE Computer Graphics and Applications 29, 2 (2009), 14–17

    Defining Insight for Visual Analytics. IEEE Computer Graphics and Applications 29, 2 (2009), 14–17

  117. [2015]

    In Proceedings of the 28th Annual ACM Symposium on User Interface Software Technology (UIST 2015)

    DataTone: Managing Ambiguity in Natural Language Interfaces for Data Visualization. In Proceedings of the 28th Annual ACM Symposium on User Interface Software Technology (UIST 2015) . ACM, New York, NY, USA, 489–500

  118. [2017]

    IEEE Transactions on Visualization and Computer Graphics 24, 1 (2017), 1–12

    Visualizing Dataflow Graphs of Deep Learning Models in Tensorflow. IEEE Transactions on Visualization and Computer Graphics 24, 1 (2017), 1–12

  119. [2021]

    In Computer Graphics Forum, Vol

    Autoclips: An Automatic Approach to Video Generation from Data Facts. In Computer Graphics Forum, Vol. 40. Wiley Online Library, 495–505

  120. [2023]

    In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems

    Notable: On-the-fly Assistant for Data Storytelling in Computational Notebooks. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–16

  121. [2024]

    In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24)

    How Do Analysts Understand and Verify AI-Assisted Data Analyses?. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 748, 22 pages. https://doi.org/10.1145/36...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.