Pith. sign in

REVIEW 7 cited by

Conversational Challenges in AI-Powered Data Science: Obstacles, Needs, and Design Opportunities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.16164 v1 pith:B26E4NR2 submitted 2023-10-24 cs.HC

classification cs.HC
keywords datachallengescontextualdesignincludingobstaclespromptsscience
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large Language Models (LLMs) are being increasingly employed in data science for tasks like data preprocessing and analytics. However, data scientists encounter substantial obstacles when conversing with LLM-powered chatbots and acting on their suggestions and answers. We conducted a mixed-methods study, including contextual observations, semi-structured interviews (n=14), and a survey (n=114), to identify these challenges. Our findings highlight key issues faced by data scientists, including contextual data retrieval, formulating prompts for complex tasks, adapting generated code to local environments, and refining prompts iteratively. Based on these insights, we propose actionable design recommendations, such as data brushing to support context selection, and inquisitive feedback loops to improve communications with AI-based assistants in data-science tools.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IntentLint: Supporting Intent Scaffolding and Prompt-time Linting in Human-AI Collaborative Data Analysis

    cs.HC 2026-08 conditional novelty 6.0 of 10

    IntentLint uses shared, editable rules to scaffold analytic intent and lint prompts, and a user study reports improved perceived collaboration awareness.

  2. Steering Semantic Data Processing With DocWrangler

    cs.HC 2025-04 conditional novelty 6.0 of 10

    An IDE for LLM-powered text data processing, with user studies showing that people convert open-ended operations into structured classifiers and use vague prompts to explore their data.

  3. Flowco: Rethinking Data Analysis in the Age of LLMs

    cs.HC 2025-04 conditional novelty 6.0 of 10

    Flowco combines visual dataflow graphs with LLM-generated code, validation checks, and unit tests to help analysts author, debug, and refine data analyses.

  4. Representing Visualization Insights as a Dense Insight Network

    cs.HC 2025-01 conditional novelty 6.0 of 10

    A five-category framework links dashboard insights into a dense network and is demonstrated in a playground and an LLM-based summarization case study.

  5. DataLab: A Unified Platform for LLM-Powered Business Intelligence

    cs.DB 2024-12 conditional novelty 6.0 of 10

    DataLab is a unified notebook-based platform for LLM-powered BI tasks that shows strong efficiency gains and competitive accuracy, but its state-of-the-art claim is not supported on several benchmarks.

  6. Effective LLM-Driven Code Generation with Pythoness

    cs.PL 2025-01 conditional novelty 5.0 of 10

    Pythoness is a test- and specification-driven embedded DSL that generates, validates, and caches LLM code, with a single example showing tests greatly improve pass rates.

  7. A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey of 80+ Deep Research systems that proposes a four-layer taxonomy (foundation models, tool use, planning, synthesis) and compares commercial and open-source implementations.

Pith tools