Pith. sign in

REVIEW 3 cited by

DataNarrative: Automated Data-Driven Storytelling with Visualizations and Texts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.05346 v3 pith:MCKNDYAD submitted 2024-08-09 cs.CL

classification cs.CL
keywords datastorieshumanstorytellingchallengescoherentdata-drivenframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data-driven storytelling is a powerful method for conveying insights by combining narrative techniques with visualizations and text. These stories integrate visual aids, such as highlighted bars and lines in charts, along with textual annotations explaining insights. However, creating such stories requires a deep understanding of the data and meticulous narrative planning, often necessitating human intervention, which can be time-consuming and mentally taxing. While Large Language Models (LLMs) excel in various NLP tasks, their ability to generate coherent and comprehensive data stories remains underexplored. In this work, we introduce a novel task for data story generation and a benchmark containing 1,449 stories from diverse sources. To address the challenges of crafting coherent data stories, we propose a multiagent framework employing two LLM agents designed to replicate the human storytelling process: one for understanding and describing the data (Reflection), generating the outline, and narration, and another for verification at each intermediary step. While our agentic framework generally outperforms non-agentic counterparts in both model-based and human evaluations, the results also reveal unique challenges in data story generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Infogen: Generating Complex Statistical Infographics from Documents

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Infogen generates complex statistical infographics from text via a metadata-then-code pipeline, and the new Infodat benchmark shows it outperforming GPT-4o and fine-tuned open LLMs on the authors' metrics.

  2. Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Text2Vis is a diverse 1,985-sample benchmark for text-to-visualization with complex queries, and a cross-modal actor-critic agent improves GPT-4o's pass rate from 26% to 42%.

  3. ChatVis: Large Language Model Agent for Generating Scientific Visualizations

    cs.HC 2025-07 conditional novelty 5.0 of 10

    A retrieval-augmented LLM assistant with iterative error correction nearly doubles the rate of generating executable ParaView visualization scripts compared with unassisted models.

Pith tools