{"id":"0c0d261c-fdca-438c-9f07-b92dfd9f740f","arxiv_id":"2501.01014","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"MDSF is an LLM-based framework for automated data insight ranking and storytelling that, by its own reported results, does not outperform GPT-4 on ranking and most narrative metrics.","lead":"The paper describes MDSF, a system that uses large language models to automatically find insights in multidimensional data and write data stories around them. It claims to outperform existing methods, but its own tables show GPT-4 beating MDSF on insight ranking and on most story-generation scores.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's central claim that MDSF outperforms existing methods in insight ranking and narrative coherence is contradicted by the paper's own Tables II and III: GPT-4 achieves lower SFD on both datasets and higher Rouge/BLEU in almost every cell. The claim is unsupported by the reported results.","rationale":"The reader's stated weakest assumption concerns unspecified baseline prompting and whether baselines received the same precomputed insight candidates. That is a legitimate reproducibility concern, but it is not the most load-bearing issue. The more decisive problem is internal inconsistency: the reported tables themselves contradict the abstract. Even if every baseline was set up perfectly fairly, the headline claim would still be false because GPT-4 outperforms MDSF on insight ranking accuracy on both datasets and on narrative quality metrics in nearly every cell. The reader's rationale does mention this contradiction, so my view partially agrees with the reader's overall rejection, but the weakest-assumption identification differs. The central claim is unsupported by the paper's own evidence, independent of any questions about baseline fairness or statistical testing. Therefore the REJECT verdict stands without modification.","tokens_in":14955,"tokens_out":2788,"duration_ms":27816,"concrete_test":"Construct a pairwise comparison table from the reported results: for each dataset and metric in Tables II and III, record whether MDSF beats GPT-4, ties, or loses, using the metrics' stated directions (lower SFD is better; higher Rouge/BLEU is better). If the signs match the printed values, GPT-4 wins both ranking comparisons and 7 of 8 story-generation comparisons. This single counting exercise settles whether the abstract's 'outperforms existing methods' claim can survive even when the reported numbers are taken at face value; it cannot, so the central claim fails unless the tables are corrected or the claim is substantially weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires the experiments to show MDSF beating existing methods on insight ranking accuracy, descriptive quality, and narrative coherence. The reported numbers do not support this. In Table II, SFD is explicitly defined as an error metric where lower is better; GPT-4 scores 5.83 on the private dataset and 7.00 on InsightBench, while MDSF scores 6.82 and 7.25. Thus GPT-4 ranks insights more accurately on both datasets. In Table III, GPT-4 has higher Rouge and BLEU than MDSF on 7 of 8 story-generation cells (private dataset, Kaggle, and InsightBench on both metrics, and Text2Analysis on Rouge); MDSF wins only Text2Analysis BLEU (0.813 vs. 0.745). Combining the two tables, GPT-4 wins 9 of 10 direct comparisons on the two headline capabilities. Only the insight-description accuracy result in Figure 3 supports MDSF. This is not a matter of disagreement over evaluation philosophy or a missing baseline prompt; if the tables are accurate, the abstract's assertion 'outperforms existing methods ... in terms of insight ranking accuracy, descriptive quality, and narrative coherence' is internally contradicted by the paper's own evidence. The framing in the results section that MDSF 'showed significant improvement over other automated models' for ranking ignores that GPT-4, an existing method, is better. The same issue affects the conclusion's claim of 'superior accuracy, operational efficiency, and user satisfaction compared to existing methods'.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MDSF, a framework for automated multidimensional data storytelling that combines data preprocessing, insight discovery, multi-criteria insight scoring, fine-tuned LLM ranking, and context-aware storytelling with an agent-based continuation mechanism. The evaluation compares MDSF with several LLM baselines on a private dataset and on Text2Analysis, InsightBench, and Kaggle data, using Spearman Footrule Distance, accuracy, ROUGE/BLEU, and a user study. The central claim is that MDSF outperforms existing methods in insight ranking accuracy, descriptive quality, and narrative coherence.","tokens_in":15292,"tokens_out":4527,"duration_ms":39720,"significance":"If its claims were supported, MDSF would be a practically useful integrated framework for automated multidimensional data storytelling. The paper includes some sound design choices: temporal train/test splits, use of public benchmarks, and a user study. The insight description accuracy result (MDSF 0.858 vs. GPT-4 0.785) is a concrete positive finding. However, the reported experiments do not support the headline claims, the scoring mechanism is underspecified, and key evaluation details such as baseline prompts and fine-tuning data size are absent. No code or data is provided, limiting independent verification.","major_comments":[{"comment":"The abstract claims MDSF outperforms existing methods in insight ranking accuracy, but Table II reports Spearman Footrule Distance (explicitly defined as lower-is-better) and shows GPT-4 at 5.83 on the private dataset and 7.00 on InsightBench, versus MDSF at 6.82 and 7.25. GPT-4 is therefore more accurate on both ranking tasks. The text's statement that MDSF 'showed significant improvement over other automated models' is contradicted by this same table, since GPT-4 is an automated model with a smaller SFD.","section":"V-B1, Table II"},{"comment":"The abstract and conclusion claim superior descriptive quality and narrative coherence, but Table III shows GPT-4 exceeding MDSF on ROUGE/BLEU in seven of eight dataset-metric cells; MDSF only wins Text2Analysis BLEU (0.813 vs. 0.745). The article itself acknowledges 'it did not surpass GPT-4 and Gemini 1.5' in the Story Generation task, so the narrative quality claim is unsupported by the reported evidence.","section":"V-B3, Table III"},{"comment":"The composite insight score is not fully specified: Section IV-A3 describes five scoring aspects, gives equations only for importance and surprise, and does not state how the aspects are combined or weighted. The fine-tuning setup omits the training-set size, annotation counts, and number of annotators, and Table II does not state which base model or configuration MDSF uses. These omissions prevent replication and make it impossible to verify the ranking results.","section":"IV-A3, V-A"},{"comment":"The user study lacks essential methodological details: no number of participants, no recruitment description, no inter-annotator reliability, and no statistical test. The rubric in Table IV is internally inconsistent (Structure is scored 0-2 while Richness uses 0-5), and Figure 4 does not clearly label its axes. The conclusion's claim of 'user satisfaction compared to existing methods' therefore goes beyond what the reported study can support.","section":"V-B4, Table IV, Figure 4"},{"comment":"The experimental comparison includes only general-purpose LLM baselines. None of the related automated storytelling systems discussed in Section II (e.g., DS-Agent, InsightPilot, Calliope) is evaluated. Since the abstract speaks of 'existing methods,' the claimed advantage over existing data storytelling frameworks is not demonstrated.","section":"II, V-A"}],"minor_comments":[{"comment":"The setup lists 'GPT-3-turbo' while Table II reports 'GPT-3.5 turbo'; please unify the model names.","section":"V-A"},{"comment":"Equation (4) defines the Spearman Footrule, but the reported values (e.g., 5.83) are not clearly normalized; specify the ranking length and whether scores are scaled.","section":"Eq. (4)"},{"comment":"The y-axis is labeled 'Accuracy' but the plotted distribution style is unclear; provide error bars or confidence intervals for the accuracy values.","section":"Figure 3"},{"comment":"Typo: 'Insigth Discovery' should be 'Insight Discovery'; also 'Mannul' in Figure 3 and 'LLMS' in Section IV-B1 should be corrected.","section":"IV-B2"},{"comment":"Several cells in Table IV are blank where a score level is undefined; consider marking these as 'N/A' for clarity.","section":"Table IV"},{"comment":"The augmented analysis methods (Prophet, SR-CNN, 3-sigma, iForest) are named but not described; include parameter choices or specific references.","section":"IV-A2"}],"recommendation":"reject","confidential_remarks":"The main issue is that the headline claims are contradicted by the paper's own tables, which is not fixable by small edits. The authors might consider a revised submission with a more modest framing, full experimental details, and a fairer baseline comparison that includes actual data storytelling systems."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper's selling point does not survive contact with its own tables. MDSF is claimed to outperform existing methods in insight ranking and narrative quality, but Table II shows GPT-4 with lower Spearman Footrule on both datasets (5.83 vs 6.82 and 7.00 vs 7.25), and Table III shows GPT-4 beating MDSF on Rouge and BLEU in seven of eight cells. The one clear win is insight-description accuracy (0.858 vs 0.785 for GPT-4). The abstract and conclusion ignore this, while the results section quietly admits that MDSF did not surpass GPT-4 or Gemini on story generation. That inconsistency is the load-bearing flaw.\n\nWhat is actually new: the framework combines subspace enumeration, anomaly detection, a five-factor scoring mechanism, a fine-tuned LLM ranker, and an agent for context-aware storytelling continuation. Each piece exists in prior work, but the integration is non-trivial and the authors provide a concrete architecture. The evaluation spans four datasets, including InsightBench and Text2Analysis, and the temporal split for training/test is a sensible choice. Fine-tuning hyperparameters are stated.\n\nSoft spots, in order of severity. The scoring equations are garbled—Eq. (2) is malformed, Eq. (3) looks like a miswritten Jensen-Shannon divergence—and the weights for the composite score are never given, so the scoring mechanism is not reproducible. No code, data, or weights are released. Baseline conditions are unspecified: we do not know whether GPT-4, Gemini, Llama, Qwen, and ChatGLM received the same precomputed insight candidates or were prompted zero-shot. The user study reports percentage scores without sample size, inter-annotator agreement, or significance tests, so its conclusions are weak.\n\nWho is it for: someone inside a company building an LLM-based augmented analytics product could use this as a starting point, and the insight-description result is promising. As a research paper, the central claim is unsupported as written. But the framework is real, the experiments are multi-dataset, and there is enough substance that a serious referee could push the authors to fix the claims and fill the gaps. I would not cite it in its current form, but I would not desk-reject it either.","headline":"The paper's central claim is contradicted by its own tables: GPT-4 beats MDSF on ranking and most story-generation metrics, so the abstract overstates the results.","tokens_in":15817,"tokens_out":2993,"would_cite":false,"duration_ms":26036,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T50"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes MDSF, a framework that automates multidimensional data storytelling by combining algorithmic insight discovery with a fine-tuned LLM ranker and a context-aware agent, reporting description accuracy close to human…","keywords":["data storytelling","large language models","augmented analysis","insight discovery","insight ranking","context-aware generation","fine-tuning","multi-dimensional data"],"falsifier":"Re-run the InsightBench ranking experiment with GPT-4 given the exact same precomputed insight candidates and cleaned inputs that MDSF's discovery stage provides. If GPT-4's Spearman-footrule distance matches or beats MDSF's 7.25, then the claimed ranking advantage does not come from MDSF's scoring and fine-tuning mechanism.","tokens_in":14733,"feed_emoji":"📊","tokens_out":6140,"duration_ms":53614,"temperature":0.7,"pith_summary":"Data analysts spend most of their effort cleaning data, hunting for patterns, and writing up findings. MDSF automates the whole path: it preprocesses multidimensional tables, discovers candidate insights with statistical algorithms, scores them by importance, significance, surprise, fatigue, and interpretability, then uses a fine-tuned large language model to rank them and turn the best into a coherent story. The paper claims this structured approach beats asking a general LLM directly, with insight-description accuracy reaching 0.858 against 1.0 for human annotation on the tasks tested. A context agent that reads the user's editing history continues the report in real time, which the authors argue lowers manual effort and interpretive bias.","feed_headline":"Framework ranks data insights close to human accuracy","feed_subtitle":"MDSF fine-tunes LLMs to score and narrate multidimensional data, beating GPT-4 on description accuracy in tests.","key_machinery":"The load-bearing object is the MDSF pipeline itself, whose central move separates insight discovery from narration. Candidate findings are extracted by augmented-analysis algorithms for time-series and non-time-series data, and each is represented as an Insight tuple carrying a scalar score. That score is the mechanism: it combines importance (weighting head versus tail subspaces), significance (fit to the insight type), surprise (Jensen-Shannon divergence between sibling and native subspaces), fatigue (suppression of repeated patterns), and interpretability (justifiability of the finding). A fine-tuned LLM ranker, trained on human-ranked examples, re-ranks these scored insights against the user's editing context, and a context agent uses the resulting order to compose and continue the story. The scoring tuple plus fine-tuned ranking is what carries the claim that MDSF knows which insights matter.","core_discovery":"On the paper's own terms, the central discovery is that inserting a scoring-and-ranking stage between insight discovery and narrative generation makes LLMs markedly more reliable for data storytelling. MDSF encodes each candidate finding as a tuple of breakdown dimensions, indicators, insight type, data model, details, and a scalar score, where the score aggregates five signals: subspace importance, statistical significance, surprise relative to sibling subspaces, fatigue from repeated similar suggestions, and interpretability. A fine-tuned LLM rank model re-ranks these scored insights against the user's edit history and profile, and a context agent then generates and continues the narrative. In experiments on a private business dataset, InsightBench, Kaggle data, and Text2Analysis, MDSF achieves the lowest Spearman-footrule ranking distance among automated systems (6.82 and 7.25 on the two ranking sets) and the highest automated description accuracy (0.858), while producing stories with Rouge and BLEU scores close to, though not above, GPT-4 and Gemini 1.5. The authors take this as evidence that algorithmic discovery plus fine-tuned ranking yields accurate, context-aware narratives with minimal manual intervention.","pith_inferences":["If the comparison setup is fair, even stronger results should appear when MDSF's scoring and ranking stage is mounted on a GPT-4- or Gemini-class base model; the paper's tables imply the ranker, rather than the base model, accounts for much of the gain.","The surprise formula based on Jensen-Shannon divergence could be lifted out and reused as a generic interestingness measure for exploratory data analysis tools beyond storytelling.","An ablation that removes the fine-tuned ranker and substitutes a zero-shot LLM would reveal exactly how much of the gain comes from the scoring mechanism; this is testable on the paper's own InsightBench setup.","The fatigue and interpretability scoring components make the framework a plausible starting point for personalized recommendation systems, where repeated suggestions must be suppressed and outputs must be explainable."],"forward_implications":["Insight ranking and description improve when the LLM selects among precomputed scored candidates rather than exploring raw tables directly.","Fine-tuning a 7-billion-parameter LLM on human-ranked insights is enough to approach manual-quality description accuracy and to beat general-purpose models on ranking distance.","A context agent that reads edit history and user profile can continue a data story in real time, making the framework usable inside an existing reporting workflow.","The same pipeline transfers across private business data, InsightBench, Kaggle datasets, and Text2Analysis, suggesting the scoring rules generalize beyond a single domain.","Using the framework reduces manual intervention and interpretive bias in turning multidimensional data into structured, conclusion-bearing reports."],"supporting_citations":[{"why":"Supplies the InsightBench benchmark used to evaluate multi-step insight generation and ranking.","marker":"[2]"},{"why":"Supplies the Text2Analysis dataset used for table-based analysis queries.","marker":"[14]"},{"why":"Represents the prior DataNarrative approach to automated data-driven storytelling with visualizations and text.","marker":"[1]"},{"why":"Provides QuickInsights, an automatic insight-discovery method motivating the discovery stage.","marker":"[67]"},{"why":"DS-Agent is the agent-based automated reporting system the paper contrasts with its own context agent.","marker":"[7]"},{"why":"InsightPilot is the LLM-empowered automated data exploration system used as reference for insight generation.","marker":"[15]"},{"why":"Benchmarks how well LLMs understand structured table data, supporting the choice of LLM bases for insight description.","marker":"[11]"},{"why":"Evaluates the statistical acumen of LLMs, informing the rationale for a dedicated ranker rather than direct generation.","marker":"[5]"}],"fun_headline_variants":["Scoring insights first makes LLM data stories more accurate","Two-stage LLM framework beats other automation in insight ranking","Fine-tuned ranking closes gap to human-level data narratives","Context-aware framework improves LLM data storytelling precision"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes the baseline models received the same cleaned data and precomputed candidate insights as MDSF; if they were run without the pipeline's preprocessing and discovery stage, the reported gains could reflect the pipeline rather than the model capability.","fun_headline_variants_meta":{"raw":{"variants":["Scoring insights first makes LLM data stories more accurate","Two-stage LLM framework beats other automation in insight ranking","Fine-tuned ranking closes gap to human-level data narratives","Context-aware framework improves LLM data storytelling precision"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000467,"raw_usage":{"total_tokens":2344,"prompt_tokens":973,"completion_tokens":1371,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":1306}},"tokens_in":589,"tokens_out":1371,"duration_ms":9585,"temperature":1.0,"reasoning_tokens":1306,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:37:03.244539+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the InsightBench ranking experiment with GPT-4 given the exact same precomputed insight candidates and cleaned inputs that MDSF's discovery stage provides. If GPT-4's Spearman-footrule distance matches or beats MDSF's 7.25, then the claimed ranking advantage does not come from MDSF's scoring and fine-tuning mechanism.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"InsightPilot is the LLM-empowered automated data exploration system used as reference for insight generation."}],"review_version":1}