{"id":"e51efcd1-fd68-4c28-b3b7-255751d82762","arxiv_id":"2501.16661","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Jupybara is an LLM-powered Jupyter extension that operationalizes a semantic, rhetorical, and pragmatic design space for actionable data analysis and storytelling.","lead":"Jupybara is a Jupyter Notebook plugin that uses multiple AI agents to help data analysts explore data and turn findings into persuasive stories with concrete recommendations. It also proposes a three-part design space, semantic, rhetorical, and pragmatic, for making data analysis and storytelling more actionable.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The expert evaluation cannot establish the strategy claims: multi-agent vs single-agent is confounded by fixed order and unequal output/latency, and Jupybara vs ChatGPT lacked an in-session baseline and was nonsignificant after correction.","rationale":"The reader's conditional verdict is appropriate, but my load-bearing concern is different from the reader's stated weakest assumption. The reader focused on external validity: whether the semantic/rhetorical/pragmatic design space is complete and whether the evaluation is circular because it uses the same three dimensions. My concern is internal validity: even granting the design space, the study cannot identify whether Jupybara's strategies caused the observed preferences, because the multi-agent condition is entangled with order, output length, latency, and overall system differences, and the ChatGPT comparison had no in-session control and was not significant after correction. This matters directly to the strongest claim, which asserts that an expert evaluation confirms the effectiveness of the strategies. The paper deserves credit for a real implementation, a plausible design-space derivation from interviews, and unusually honest limitation statements in Section 6.1. Those limitations are not peripheral; they are exactly the assumptions the central claim needs. A counterbalanced factorial study with blind expert rating would settle the concern. If it failed to reproduce the advantage, the paper would still be a useful design-space contribution, but the demonstrated-effectiveness claim would need to be downgraded. Since the reader already recommended conditional acceptance with similar reservations, I do not propose changing the verdict.","tokens_in":31140,"tokens_out":4525,"duration_ms":51100,"concrete_test":"Run a preregistered 2×2 factorial evaluation (design-space-aware prompt vs generic prompt × single-agent vs multi-agent) on matched tasks, counterbalancing condition order and capping or measuring output tokens and latency; have domain experts blind-rate final notebooks and stories on semantic precision, rhetorical quality, pragmatic actionability, and factual correctness. If the multi-agent advantage or the prompt advantage disappears once order and output budget are controlled, the causal claims in the Abstract are unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing weakness is the causal identifiability of the evaluation, not the completeness of the design space. Section 5.2 fixes the order: every participant first uses single-agent mode and then repeats the same analysis in multi-agent mode, and the same order is used for data storytelling. Section 5.3.1 then reports that multi-agent received higher median ratings on the three design-space dimensions, with Wilcoxon tests reported as significant. This cannot support the Abstract's claim that multi-agent architectures operationalize the design space more effectively, because mode is confounded with practice/familiarity, with response length and richness (the multi-agent pipeline produces plans, additional visualizations, and more extensive interpretations), and with latency. The paper itself notes in Section 6.1 that longer response times may lead participants to perceive outputs as more thorough, and that the fixed single-to-multi order may introduce learning bias. The ChatGPT comparison has the same identification problem: participants did not use ChatGPT during the session, and after Holm-Bonferroni correction none of the differences were significant. Finally, the two claimed strategies are never separated: there is no condition varying design-space-aware prompting while holding the agent architecture fixed, nor varying agent count while holding prompts fixed. The headline therefore overstates what the study can establish.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a three-dimensional design space for actionable exploratory data analysis (EDA) and data storytelling, comprising semantic precision, rhetorical persuasion, and pragmatic relevance, and claims this space is grounded in theory and nine expert interviews. The authors instantiate the space in Jupybara, a Jupyter Notebook extension that uses two operationalization strategies: design-space-aware prompting and multi-agent critique-and-refine architectures. The paper reports a summative evaluation with nine expert analysts in which participants rated Jupybara against ChatGPT's data analysis plugin on usability-related dimensions and rated Jupybara's multi-agent mode against its single-agent mode on the three design-space dimensions. The central claim is that the expert evaluation supports Jupybara's usability, steerability, explainability, and reparability, and the effectiveness of the two strategies for operationalizing the design space.","tokens_in":31331,"tokens_out":2579,"duration_ms":29204,"significance":"If the causal claims were supported, the paper would make a useful contribution by showing that a design-space-derived prompting and multi-agent critique protocol improves actionable EDA and storytelling over single-agent LLM responses, and by providing a transferable architecture for such systems. The design space itself is a plausible synthesis of prior work and interview data, and the Jupybara implementation is broad and thoughtfully integrated with Jupyter workflows. The paper is also candid in Section 6.1 about several known limitations, which is a strength. However, the headline claims about strategy effectiveness rest on an evaluation design whose confounding factors prevent causal identification; the design space is also validated partly through rating instruments that reuse its own vocabulary. The result, as currently stated, is therefore not established, though the underlying system and framework remain potentially valuable.","major_comments":[{"comment":"The single-agent versus multi-agent comparison cannot support the claim that the multi-agent architecture is more effective. Every participant first used single-agent mode and then repeated the same analysis in multi-agent mode, for both EDA and storytelling. Mode is therefore confounded with practice, familiarity with the dataset, and carryover of earlier insights; moreover, the multi-agent pipeline produces longer, more elaborate outputs and has roughly five times the latency, a factor the authors themselves identify in Section 5.3.7 and Section 6.1 as a possible source of perceived thoroughness. The significant Wilcoxon results in Figure 12 are consistent with these alternative explanations. A counterbalanced or between-subjects design, or a condition that equalizes output length and presentation format, would be needed to attribute the ratings to the multi-agent architecture.","section":"Section 5.2 and 5.3.1, Figure 12"},{"comment":"The Jupybara-versus-ChatGPT comparison is not an in-session baseline: participants rated ChatGPT's data analysis plugin from prior experience and never used it during the study, as acknowledged in Section 6.1. The reported p-values are nonsignificant after Holm-Bonferroni correction, yet the Abstract states that the evaluation confirms Jupybara's superiority. The current text appropriately says 'the trend clearly indicates' within the body, but this more cautious language should also be reflected in the Abstract and the contributions list, unless a controlled comparison is conducted.","section":"Section 5.2 and 5.3.1, Figure 11"},{"comment":"The two strategies, design-space-aware prompting and multi-agent architectures, are never independently varied. In the evaluation, the multi-agent mode differs from the single-agent mode both in agent count and in the prompting/guidance structure, so any difference in ratings could be due to either factor or their interaction. The claimed effectiveness of 'our strategies' as a pair is therefore not separable, and the specific attribution to multi-agent collaboration in Section 5.3.7 is not identified.","section":"Section 4.3 and Abstract"},{"comment":"The Part 2 questionnaire asks participants to rate 'provides precise and contextually rich interpretation of results', 'generates coherent and persuasive narrative', and 'offers high-quality actionable insights', which are nearly verbatim restatements of the semantic, rhetorical, and pragmatic dimensions the system was designed to optimize. The evaluation therefore risks being circular as a validation of the design space: the outcome measures share their construct vocabulary with the intervention's design goals. A concrete test that would avoid this circularity is an independent rubric-based evaluation of the generated artifacts by blinded raters who did not see the design-space definitions, or an artifact-level analysis of whether the outputs actually satisfy the three dimensions.","section":"Section 5.3.1, questionnaire wording"}],"minor_comments":[{"comment":"The name 'Ben Schneiderman' should be spelled 'Ben Shneiderman'.","section":"Section 1, first sentence"},{"comment":"The text 'Rae is able to quickly trace the the graphical representation' contains a duplicated article, which should be corrected.","section":"Section 4.4.3"},{"comment":"The figures report p-values but not the corresponding Wilcoxon test statistics, sample sizes per item, or effect sizes; adding these would make the statistical claims more transparent.","section":"Figures 11 and 12"},{"comment":"The sentence 'All but one participant preferred the multi-agent mode' would benefit from a precise count and from a direct reconciliation with the questionnaire results, especially since some participants expressed latency and verbosity concerns.","section":"Section 5.3.7"},{"comment":"The limitation paragraph is honest and welcome, but the Abstract and contributions currently state that the evaluation 'confirms' the effectiveness of the strategies; the language should be aligned so that the known confounds are not understated in the paper's front matter.","section":"Section 6.1"},{"comment":"The two participant pools both have nine experts and include two overlapping participants; it would be clearer to state explicitly that the summative evaluation is not fully independent of the formative interview sample.","section":"Section 3.1 and 5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong systems and framework contribution, and the authors are unusually transparent about limitations in Section 6.1. My main concern is that the Abstract overstates what the evaluation can establish. If the claims are softened to a usability study with exploratory comparisons, and the authors add a clear research-design caveat about the confounded single- versus multi-agent comparison, the paper could be acceptable. A full re-run with counterbalanced conditions would be ideal but may be beyond the current manuscript's scope; at minimum the causal language must be revised."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, Jupybara is a real, reasonably complete Jupyter extension for LLM-assisted EDA and data storytelling, and the paper does a genuinely useful job of turning the semantic/rhetorical/pragmatic design space from the authors' earlier position paper into concrete prompting and multi-agent patterns. Second, the evaluation evidence does not support the abstract's claim that an expert evaluation confirms the effectiveness of the strategies. The strategy comparison is under-identified: every participant used single-agent first and multi-agent second, the multi-agent outputs are longer and richer and take roughly five times as long, and the questionnaire's quality items are worded in terms of the very three dimensions the system optimizes. The authors admit the order and latency confounds in Section 6.1, which is more honest than most papers, but the abstract and contributions still state the effectiveness claim as established.\n\nWhat is actually new and valuable: the elaborated design space, the working implementation, and the critic-refiner multi-agent architecture for EDA and storytelling. The formative interview findings (fluid, integrated, iterative workflows; four recurring challenges) are useful and well-grounded. The system features are thoughtfully designed: DAG-based insight tracking, threaded clarification, direct and AI-mediated repair. The writing is clear and the usage scenario helps the reader see the tool in action. The citation practice is fair; the debt to Setlur and Birnbaum's 2024 paper is acknowledged, and concurrent work such as DataNarrative is cited.\n\nThe soft spots are concentrated in the evaluation, and I agree with the stress-test that the load-bearing weakness is causal identifiability rather than the completeness of the design space. Fixed order, output richness, and latency differences alone could account for the multi-agent preference; with nine self-selected experts, the Wilcoxon results after correction still do not establish that the architecture is what matters. The ChatGPT comparison is retrospective, and none of those differences survive Holm-Bonferroni correction. On top of this, the two proposed strategies are never separated, so there is no evidence about which component contributes what. The circularity between the design space and the rating scales is a lesser but real version of the same problem. The design space's external validity is untested, though it is plausible and well-sourced.\n\nWho gets value: visualization and human-AI-collaboration researchers. The system and framework deserve a serious referee; the evaluation deserves to be read as a usability and preference study, not a controlled effectiveness test. My recommendation: send it to peer review, but ask the authors to either temper the abstract's effectiveness claims, run a within-session randomized baseline, or clearly reframe the study as an exploratory usability assessment. Open-sourcing the code would strengthen the paper considerably. If the venue demands controlled evidence, this is not there; if it values a well-built system with a plausible framework and genuinely honest limitations, this is publishable with revision.","headline":"Solid design-space and system paper; the abstract's effectiveness claim outruns the confounded small-n evaluation.","tokens_in":31901,"tokens_out":4344,"would_cite":true,"duration_ms":40107,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Jupybara argues that a semantic-rhetorical-pragmatic design space, encoded in LLM prompts and multi-agent critique, improves actionable data analysis and storytelling.","keywords":["Actionable Insights","Human-AI Collaboration","Multi-Agent System","Large Language Model","Exploratory Data Analysis","Data Storytelling","Semantics","Pragmatics"],"falsifier":"A controlled crossover study would settle the effectiveness claims: expert analysts analyze the same private datasets with Jupybara and with a generic web LLM tool, with order randomized, reviewers blind to which system generated each response, and processing time or cost equalized between modes. If blinded reviewers no longer rate the multi-agent output higher on semantic, rhetorical, and pragmatic quality, the architecture's claimed advantage is not established. A separate falsifier for the design space itself would be a documented recurring analyst challenge that fits none of the three dimensions.","tokens_in":30901,"feed_emoji":"📊","tokens_out":7975,"duration_ms":77150,"temperature":0.7,"pith_summary":"This paper tries to establish that the messy process of turning data into decisions can be guided by a three-part design space, and that large language models can be made to follow it. The three parts are semantic precision (saying exactly what the data and results mean), rhetorical persuasion (choosing analytical methods and narrative structure that support an argument), and pragmatic relevance (connecting findings to domain knowledge and concrete actions). The authors build Jupybara, a Jupyter Notebook assistant, on two strategies that put this design space inside LLM workflows: prompts distilled from the three dimensions, and multi-agent teams in which critics and a refiner revise responses until the critics agree they are ready. Nine expert analysts gave Jupybara higher ratings than a web-based LLM tool on usability, steerability, explainability, and reparability, and rated its multi-agent mode significantly better than its single-agent mode on all three design-space dimensions. If the claim holds, AI copilots for data analysis can be steered toward actionable insight instead of merely plausible output.","feed_headline":"Critique-and-refine beats single-pass LLM analysis","feed_subtitle":"Experts rated the multi-agent Jupyter data assistant higher on all three quality checks.","key_machinery":"The load-bearing mechanism is the combination of the three-dimensional design space with two operationalization strategies. The design space is the source of the quality criteria: semantic precision, rhetorical persuasion, and pragmatic relevance. To make these criteria usable by a language model, the authors do not hand the model abstract definitions; they distill each dimension into concrete behavioral instructions, such as always interpreting statistical results and visualizations, keeping the user informed of the analysis plan, and narrating the analytical strategies used. The second strategy is multi-agent architecture: an Initial Respondent drafts a response, a set of Critics (for EDA: analysis plan, code, visualization, and interpretation/summary; for storytelling: one per design-space dimension) evaluate it, and a Refiner revises it, with the cycle repeating until all critics approve or a maximum number of discussion rounds is reached. This design means every dimension is checked by multiple agents and the final output is the product of explicit, inspectable debate rather than a single pass.","core_discovery":"The central claim is that actionable exploratory data analysis and storytelling can be understood through a design space with three dimensions—semantic, rhetorical, pragmatic—and that an LLM assistant which operationalizes these dimensions produces measurably better support for analysts. Semantic work means pinning down analytical objects and results and verbalizing them accurately; rhetorical work means selecting analytical strategies and narrative moves that build a persuasive case; pragmatic work means grounding data facts in domain knowledge and translating them into recommendations. The paper reports that Jupybara, which encodes these dimensions in its prompts and in its multi-agent critique-and-refine loop, was rated by nine expert analysts as more usable, steerable, explainable, and reparable than a generic web-based LLM data tool, and that its multi-agent mode outperformed its single-agent mode on all three dimensions with significance surviving correction.","pith_inferences":["Editorial extension: because the multi-agent mode also uses more time and tokens, the observed quality advantage might come from extra effort rather than from the critique architecture itself; a matched-cost comparison with equal latency and token budgets would separate those explanations.","Editorial extension: if the three-dimensional design space generalizes, it could serve as a shared rubric for auditing AI-generated insights in other venues, such as dashboards, business reports, and decision-support documents.","Editorial extension: the design space was derived from nine expert interviews, so its completeness is an open empirical question; domains with strong ethical, privacy, or fairness constraints may reveal additional dimensions beyond the pragmatic one."],"forward_implications":["LLM assistants for data analysis can be built against explicit quality criteria—semantic precision, rhetorical persuasion, pragmatic relevance—rather than generic helpfulness, and those criteria double as evaluation rubrics.","Complex analytical questions should be routed to the multi-agent mode when response quality matters more than speed, because the multi-agent mode takes roughly five times as long as the single-agent mode but earns significantly higher quality ratings.","Actionable EDA and storytelling can live in one notebook workflow: analysis plans, code, interpretations, insight summaries, and editable data stories are all produced and cross-referenced inside Jupyter rather than in separate tools.","The design space subsumes the four analyst challenges identified in the formative study (choosing analytical strategies, tracking insights and history, finding the right language and narrative, and leveraging domain knowledge), so meeting the three dimensions addresses those challenges simultaneously."],"supporting_citations":[{"why":"Supplies the preliminary semantic-rhetorical-pragmatic framework that this paper elaborates into the full design space.","marker":"[77]"},{"why":"Supplies the design space for AI assistants in computational notebooks that informed Jupybara's interface and integration choices.","marker":"[61]"},{"why":"Provides the ReACT reasoning-and-acting paradigm that drives Jupybara's agentic cell-by-cell EDA workflow.","marker":"[112]"},{"why":"Supplies the Chain-of-Thought prompting used in the insight-summarization and critic agents.","marker":"[103]"},{"why":"Provides the closest LLM-pipeline baseline for insight generation and tracking, which the interactive multi-agent architecture is designed to improve upon.","marker":"[104]"},{"why":"Defines insight as data facts plus provenance and domain knowledge, which shapes the pragmatic dimension and the insight-tracking graph.","marker":"[5]"},{"why":"Identified as the closest concurrent technique; Jupybara extends its agent-based storytelling approach to EDA and to design-space-guided critique.","marker":"[36]"}],"fun_headline_variants":["Multi-agent LLM loop beats single-pass in expert EDA test","Jupybara's critique-and-refine wins on all quality checks","Design space for actionable EDA powers multi-agent LLM tool","Expert-rated: Jupybara improves EDA with semantic, rhetorical, pragmatic"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the three dimensions distilled from nine expert interviews—semantic precision, rhetorical persuasion, pragmatic relevance—really do cover what makes actionable EDA and storytelling effective; if a major ingredient is missing or misweighted, Jupybara optimizes the wrong objectives and the evaluation, which rates output along those same three dimensions, cannot demonstrate true effectiveness.","fun_headline_variants_meta":{"raw":{"variants":["Multi-agent LLM loop beats single-pass in expert EDA test","Jupybara's critique-and-refine wins on all quality checks","Design space for actionable EDA powers multi-agent LLM tool","Expert-rated: Jupybara improves EDA with semantic, rhetorical, pragmatic"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000261,"raw_usage":{"total_tokens":1565,"prompt_tokens":888,"completion_tokens":677,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":504,"completion_tokens_details":{"reasoning_tokens":599}},"tokens_in":504,"tokens_out":677,"duration_ms":7420,"temperature":1.0,"reasoning_tokens":599,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T11:34:51.171525+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled crossover study would settle the effectiveness claims: expert analysts analyze the same private datasets with Jupybara and with a generic web LLM tool, with order randomized, reviewers blind to which system generated each response, and processing time or cost equalized between modes. If blinded reviewers no longer rate the multi-agent output higher on semantic, rhetorical, and pragmatic quality, the architecture's claimed advantage is not established. A separate falsifier for the design space itself would be a documented recurring analyst challenge that fits none of the three dimensions.","supporting_citations":[{"cited_title":"Can Nuanced Language Lead to More Actionable Insights? Exploring the Role of Generative AI in Analytical Narrative Structure","cited_arxiv_id":"2405.02763","evidence_quote":"Supplies the preliminary semantic-rhetorical-pragmatic framework that this paper elaborates into the full design space."},{"cited_title":"Le, Denny Zhou, et al","cited_arxiv_id":null,"evidence_quote":"Supplies the Chain-of-Thought prompting used in the insight-summarization and critic agents."}],"review_version":1}