{"id":"266bb962-bcb1-49e3-bda0-7d15f27439ca","arxiv_id":"2606.22977","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"StatABench is a new benchmark dataset and evaluation framework showing that even GPT-5.1 scores 68.6% on closed statistical questions and top agent frameworks score 61.86 on open tasks.","lead":"The paper introduces StatABench, a benchmark with 404 closed-format questions across 18 topics plus 30 open-ended modeling tasks to test LLMs on statistical analysis. Smart generalists might read it to gauge how close current AI systems are to performing reliable statistical work without human oversight.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Benchmark validity rests on unshown question correctness and LLM-judge calibration for open tasks","rationale":"The reader’s weakest_assumption directly identifies the same load-bearing point (question quality + judge fidelity). Because the full text was unavailable to the reader, the current UNVERDICTED / LOW verdict already reflects this uncertainty; the concrete test above would resolve it without altering the verdict category.","tokens_in":1742,"tokens_out":306,"duration_ms":12374,"concrete_test":"Sample 30 Stat-Open model outputs; have two independent statisticians score them on the same rubric used by the judge; compute Cohen’s κ and mean absolute difference; if κ < 0.7 or mean difference > 10 points on the 100-point scale, the judge protocol cannot be treated as reliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim (GPT-5.1 at 68.6 % on Stat-Closed, top agent at 61.86 on Stat-Open) requires that the 404 closed questions and 30 open modeling tasks are free of labeling errors, ambiguous wording, or domain-specific biases, and that the LLM-as-Judge protocol for Stat-Open has been calibrated against human statisticians. The abstract states the judge is “validated,” yet provides no inter-rater statistics, validation sample size, or disagreement analysis; any systematic mismatch between judge and expert would directly inflate or deflate the reported gap.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces StatABench, a benchmark for LLMs' statistical analysis capabilities consisting of Stat-Closed (404 questions across 18 topics in multiple-choice, fill-in-the-blank, decision-making, and practical formats) and Stat-Open (30 open-ended modeling tasks adapted from professional competitions). It evaluates LLMs via the LangChain MCP framework and data science agents, using a validated LLM-as-Judge protocol for open tasks, and reports that GPT-5.1 reaches only 68.6% on Stat-Closed while the best open-source model reaches 60.6%, with the top agent framework scoring 61.86 on average on Stat-Open. These results are presented as evidence of gaps in tool-grounded reasoning, methodological decision-making, and end-to-end statistical modeling.","tokens_in":1848,"tokens_out":442,"duration_ms":16461,"significance":"If the benchmark construction and judge protocol prove reliable, the work would offer a useful, multi-format evaluation resource that highlights concrete limitations in current LLMs for statistical tasks, potentially guiding improvements in agent frameworks and tool integration. The adaptation of tasks from professional competitions and the dual closed/open design are positive features that increase relevance to real statistical practice.","major_comments":[{"comment":"Abstract: The claim that the LLM-as-Judge protocol for Stat-Open is 'validated' is not accompanied by inter-rater agreement statistics, validation sample size, or disagreement analysis with human statisticians; without these, the reported 61.86 average score cannot be confidently interpreted as a reliable measure of modeling capability.","section":"Abstract"},{"comment":"Abstract (and implied dataset sections): No details are provided on question validation procedures, data exclusion rules, or checks for ambiguous wording and domain biases in the 404 closed questions and 30 open tasks; these omissions directly affect whether the performance gap (e.g., 68.6% for GPT-5.1) can be attributed to model limitations rather than benchmark artifacts.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"Thank you for the constructive feedback and the recommendation for major revision. We appreciate the emphasis on ensuring the reliability of the LLM-as-Judge protocol and the transparency of benchmark construction. We agree that additional details are needed and will revise the manuscript to incorporate them. Our point-by-point responses follow.","responses":[{"response":"We agree that the absence of quantitative validation metrics weakens the interpretation of the Stat-Open results. While the manuscript describes the LLM-as-Judge protocol and states it was validated, we did not report inter-rater agreement statistics, sample sizes, or disagreement analysis. In the revised version, we will add a new subsection in the evaluation methodology detailing: the validation sample (e.g., 10 randomly selected tasks double-annotated by human statisticians), inter-rater agreement (Cohen's kappa between LLM judge and humans), and a summary of disagreement cases with resolution process. This will directly support the 'validated' claim and allow readers to assess the reliability of the 61.86 score. The abstract will be updated to reference the added validation details if space allows.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The claim that the LLM-as-Judge protocol for Stat-Open is 'validated' is not accompanied by inter-rater agreement statistics, validation sample size, or disagreement analysis with human statisticians; without these, the reported 61.86 average score cannot be confidently interpreted as a reliable measure of modeling capability."},{"response":"We acknowledge that the current manuscript lacks explicit documentation of validation procedures for the questions and tasks, which is a valid concern for attributing performance gaps. The dataset sections describe the topics, formats, and sources (including adaptation from professional competitions) but omit the validation steps. In the revision, we will expand the 'Dataset Construction' section to include: (1) validation procedures (expert review for factual accuracy and clarity), (2) data exclusion rules (e.g., removal of questions with ambiguous interpretations or multiple correct answers), and (3) checks for ambiguous wording and domain biases (e.g., topic balance verification and bias audits across statistical subfields). These additions will strengthen the claim that observed gaps reflect model limitations. Details will appear in the main text or as supplementary material.","revision_made":"yes","referee_comment":"[Abstract] Abstract (and implied dataset sections): No details are provided on question validation procedures, data exclusion rules, or checks for ambiguous wording and domain biases in the 404 closed questions and 30 open tasks; these omissions directly affect whether the performance gap (e.g., 68.6% for GPT-5.1) can be attributed to model limitations rather than benchmark artifacts."}],"tokens_in":1441,"tokens_out":578,"duration_ms":13058,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's core contribution is StatABench itself: 404 closed questions spanning 18 topics in multiple-choice, fill-in, decision, and application formats, plus 30 open-ended modeling tasks drawn from competitions. They run these through LangChain MCP and several agents, reporting GPT-5.1 at 68.6% on the closed set, best open-source at 60.6%, and top agent at 61.86 on the open set. That setup and those numbers are new relative to the prior work cited in the abstract.\n\nThe work does a straightforward job of widening coverage beyond narrower earlier benchmarks. Multiple formats and the open component give a broader picture of where models struggle with methodological choices and end-to-end modeling.\n\nThe soft spots sit right where the stress-test note flags them. The abstract calls the LLM-as-Judge \"validated\" but supplies no inter-rater agreement numbers, sample size, or disagreement breakdown. There is also no description of how the 404 questions were checked for accuracy, ambiguity, or domain bias. Without those details the size of the claimed gap is hard to trust; any systematic error in the test items or judge would move the percentages directly. The full manuscript might contain the missing checks, but nothing in the abstract or reader's summary shows them.\n\nThis is a benchmark paper aimed at the LLM evaluation community, especially groups working on tool use and data-science agents. Readers who need a ready-made test set for statistical reasoning will find the structure useful once the validation evidence is confirmed. It is coherent on its own terms and shows honest engagement with the limits of prior benchmarks.\n\nI would send it to peer review so referees can examine the question construction and judge calibration sections.","headline":"StatABench adds a fresh set of 404 closed questions and 30 open tasks for LLM stats evaluation, but the reported performance gaps rest on unshown question validation and judge calibration.","tokens_in":2340,"tokens_out":432,"would_cite":false,"duration_ms":17564,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Current LLMs reach at most 68.6 percent on a new benchmark for statistical analysis tasks.","keywords":["statistical analysis","LLM evaluation","benchmark dataset","data science agents","tool-grounded reasoning","open-ended modeling","LLM-as-Judge","performance gap"],"falsifier":"A model or agent framework that scores above 90 percent on both Stat-Closed and Stat-Open while matching independent expert human judgments on the identical items.","tokens_in":2652,"feed_emoji":"📊","tokens_out":696,"duration_ms":21361,"temperature":0.7,"pith_summary":"The paper presents StatABench as a dataset and framework to measure LLMs on statistical analysis. Stat-Closed supplies 404 questions spanning 18 topics in multiple-choice, fill-in, decision, and application formats. Stat-Open adds 30 complex open-ended modeling tasks drawn from professional competitions. Evaluations run through the LangChain MCP framework and an LLM-as-Judge protocol show GPT-5.1 at 68.6 percent on the closed set, the strongest open-source model at 60.6 percent, and the best agent framework at 61.86 average on the open set. A reader would care because the numbers quantify how far current models remain from dependable performance on tasks that combine domain knowledge with tool use.","feed_headline":"LLMs top out at 68.6% on statistical analysis benchmark","feed_subtitle":"StatABench tests 404 questions and 30 tasks, exposing limits in reasoning and modeling for top models","key_machinery":"StatABench benchmark with its Stat-Closed and Stat-Open components, evaluated through the LangChain MCP framework and validated LLM-as-Judge protocol.","core_discovery":"StatABench comprises Stat-Closed with 404 questions across 18 statistical topics and Stat-Open with 30 open-ended tasks. When LLMs and data-science agents are tested via LangChain MCP and a validated LLM-as-Judge protocol, the highest closed-set score is 68.6 percent and the highest open-set agent average is 61.86, establishing a measurable gap between existing models and reliable statistical analysis in tool-grounded reasoning, methodological choice, and end-to-end modeling.","pith_inferences":["The benchmark could be used to steer training of specialized statistical modules inside LLMs.","Comparable test suites might expose similar limits in other fields that mix domain rules with software tools.","Higher scores on StatABench would support more automated data-analysis pipelines, though human review would likely stay necessary for critical applications.","Repeated use of the same tasks could allow tracking of progress as new models appear."],"forward_implications":["LLMs still lack reliable tool-grounded reasoning for statistical work.","Methodological decision-making remains a clear weakness.","End-to-end statistical modeling stays beyond current model reach.","Open-source models trail closed models on these tasks.","Agent frameworks provide partial gains but do not close the gap."],"fun_headline_variants":["LLMs reach 68.6% on StatABench closed questions","StatABench closed set peaks at 68.6% for LLMs","Agents score 61.86 average on StatABench open tasks","LLMs hit 68.6% on 404 StatABench statistical questions","StatABench tests 18 topics with 68.6% LLM max score"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 404 questions, 30 tasks, LangChain MCP setup, and LLM-as-Judge protocol together give an accurate, unbiased picture of statistical analysis ability.","fun_headline_variants_meta":{"raw":{"variants":["LLMs reach 68.6% on StatABench closed questions","StatABench closed set peaks at 68.6% for LLMs","Agents score 61.86 average on StatABench open tasks","LLMs hit 68.6% on 404 StatABench statistical questions","StatABench tests 18 topics with 68.6% LLM max score"]},"model":"grok-4.3","cost_usd":0.00474,"raw_usage":{"total_tokens":2351,"prompt_tokens":695,"num_sources_used":0,"completion_tokens":98,"cost_in_usd_ticks":47399500,"prompt_tokens_details":{"text_tokens":695,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1558,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":695,"tokens_out":98,"duration_ms":13097,"temperature":1.0,"reasoning_tokens":1558,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T08:23:49.401189+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A model or agent framework that scores above 90 percent on both Stat-Closed and Stat-Open while matching independent expert human judgments on the identical items.","supporting_citations":[],"review_version":1}