{"id":"05b783e6-736e-48c6-a2a9-b699902586a1","arxiv_id":"2607.28033","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"A production-grounded, sandbox-graded benchmark of 100 multi-engine data-engineering tasks shows frontier agents top out at 74.9 with no cross-engine winner.","lead":"DataClawEval is a 100-task, five-engine benchmark that tests whether LLM agents can finish real industrial data-engineering jobs end to end, not just write SQL. Top agents score only about 75 and specialize by engine, so the field still lacks a general data engineer agent.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Single-run Table 2 plus large §5.2 variance/timeouts make the engine-specialization half of the central claim statistically fragile.","rationale":"The reader correctly ACCEPTS a carefully executed benchmark paper whose core empirical finding—best agent only 74.9 on a deterministic, sandbox-graded, multi-engine suite—is well supported and leaves clear headroom. I agree that one-enterprise reconstruction and synthetic tables (§3.1) limit external generality, but that is the usual industrial-benchmark caveat and does not overturn the on-benchmark numbers. The more load-bearing internal soft spot for the abstract’s strongest wording is statistical: Table 2 is single-run while §5.2 and Appendix C document score swings and timeout deflation large enough to move engine winners and mid-pack order. That weakens “each excels on a different engine / strict domain specialization” more than it weakens “autonomous DE remains unsolved,” which survives even max@3. Hence partial agreement on the weakest link, verdict left UNCHANGED at ACCEPT: qualify specialization language after a 3-seed check, do not reject the paper. Reproducibility via released containers remains a strength if artifacts match the text.","tokens_in":24091,"tokens_out":652,"duration_ms":67600,"concrete_test":"Re-execute the full 16×100 matrix for three independent seeds under the same harness and limits; recompute per-engine argmax and overall ranking under avg@3 and majority-pass. If two or more engine crowns flip or the overall leader changes under avg@3, soften the specialization claim to match §5.2; if crowns are stable, the concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim pairs a 74.9 ceiling with “no single model dominates, as each excels on a different engine” (Abstract; §4.2; Table 2). Overall “far from solved” is robust—all 16 models sit in a ~60–75 band—but the specialization half is load-bearing for “strict domain specialization rather than omnipotent proficiency” and rests on single-run means (§4.1) over small per-engine n (12–28 tasks). The paper’s own stability study (§5.2) shows max@3−min@3 gaps of 12.6–26.8 points and pass@3−pass^3 gaps of 8–26% for four agents; Appendix C shows timeout rates up to 22% that deflate reported means by up to 13.3 points and can reorder models (GLM 5.2’s completed-run mean would lead). Engine crowns (Claude Opus PySpark 83.8, DeepSeek V4 Flash FlinkSQL 85.0, etc.) are therefore single-sample estimates that the paper’s variance evidence does not secure. Fixed CodeBuddy harness and wall-clock limits further confound model×tooling with intrinsic engine skill. Task-origin external validity (§3.1) remains a real but secondary limit for the interpretive “thus unsolved” leap.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"DataClawEval introduces a 100-task executable benchmark for autonomous data-engineering agents, derived from production code and spanning PySpark, MySQL, HiveSQL, PrestoSQL/Trino, and FlinkSQL. Tasks are reconstructed via a human-in-the-loop pipeline (intent/table synthesis, differential perturbation checks, case-specific graders) and scored in isolated Docker sandboxes by deterministic rule-based scripts that combine artifact correctness with process metrics (exploration, efficiency, self-verification; α≈0.7). Under a fixed Tencent CodeBuddy harness, 16 frontier LLMs are evaluated once per task. The strongest model reaches only 74.9 overall; MySQL is easiest and HiveSQL hardest; engine leaders differ; token/tool-call volume does not track quality. Ablations argue that LLM-as-judge scoring is inflated and unstable relative to rule-based ground truth, and multi-run checks on four models show non-trivial score and pass-rate variance. The suite, containers, and graders are released.","tokens_in":24384,"tokens_out":1438,"duration_ms":36342,"significance":"If the empirical picture holds, the paper supplies the first production-grounded, multi-engine, end-to-end harness for data-engineering agents and a clear negative result: current frontier agents are far from reliable industrial ETL/stream engineering under live execution. Strengths that should be credited explicitly include (i) case-specific deterministic graders and containerized environments rather than LLM-as-judge, (ii) differential-testing style construction to make tasks answer-identifiable, (iii) joint artifact+process scoring, (iv) a controlled 16-model comparison under one scaffold, and (v) full public release of tasks, sandboxes, and grade.py scripts. These make the benchmark immediately usable and the “unsolved” claim falsifiable by future systems.","major_comments":[{"comment":"Abstract and §4.2 treat “no single model dominates, as each excels on a different engine” as a co-equal half of the central claim with the 74.9 ceiling. Table 2 engine crowns (e.g., Claude Opus 4.8 on PySpark 83.8, DeepSeek V4 Flash on FlinkSQL 85.0) are single-run means over small per-engine n (12–28 tasks). The paper’s own §5.2 multi-run study on four agents shows max@3−min@3 gaps of 12.6–26.8 points and pass@3−pass^3 gaps of 8–26%; Appendix C shows timeout rates up to 22% that deflate means by up to 13.3 points and can reorder models (GLM 5.2). The overall “far from solved” band (~60–75) is robust; the strict specialization narrative is not yet secured. Either report multi-run engine means (or bootstrap CIs) for all 16 models, or qualify the specialization claim to match the single-run evidence.","section":"Abstract; §4.1–4.2; Table 2; §5.2; Appendix C"},{"comment":"§4.1 fixes one agent scaffold (Tencent CodeBuddy) and a wall-clock limit for all models. Engine-specific rankings and tool-call efficiency (Fig. 4) therefore confound intrinsic model skill with scaffold/tooling fit and timeout policy. Appendix C already shows timeouts can reorder the leaderboard. The manuscript should state this confound explicitly when interpreting engine winners and, where feasible, report completed-run means alongside full-run means, or a short sensitivity check under a second harness/time budget for a subset of engines.","section":"§4.1; Fig. 4; Appendix C"},{"comment":"External validity of the 74.9 ceiling and engine difficulty ordering rests on tasks reconstructed from one enterprise’s desensitized production code, with LLM-inferred intents and synthetic input tables (§3.1 Stages 2–5). Differential expert perturbations improve discriminability within this corpus, but do not establish that difficulty and dialect mix represent industrial data engineering in general. A short limitations paragraph should bound generalization (single-org provenance, synthetic tables, fixed business-domain mix in Fig. 1) so the interpretive leap “thus autonomous data engineering remains unresolved” is scoped to this harness rather than asserted universally.","section":"§3.1; Fig. 1; §7"}],"minor_comments":[{"comment":"Eq. (1) and the surrounding text set α=0.7 “in most” cases, while Appendix F case studies use α∈{0.5,0.6,0.7}. State the distribution of α across the 100 tasks and whether overall scores are sensitive to a global α sweep.","section":"§3.2 Eq. (1); Appendix F"},{"comment":"Table 1 lists “DataClawBench” and “Ours (DataClawEval)” with similar names; a one-sentence disambiguation in §1 or the table caption would reduce confusion with the related-work baseline.","section":"Table 1; §1"},{"comment":"Process sub-weights (exploration 35 / efficiency 40 / self-verification 25 in Appendix F) are free parameters not justified in the main text. Briefly motivate or note they are fixed a priori.","section":"§3.2; Appendix F"},{"comment":"Figure 1 percentages and engine counts (e.g., PrestoSQL 12%) should be checked against Table 4’s full listing for consistency in the camera-ready.","section":"Figure 1; Appendix A Table 4"},{"comment":"Typos/consistency: “Sun Yat-Sun University” on the author block; “answer-identifiable” is used well but could be defined once at first use in §3.1.","section":"Title page; §3.1"}],"recommendation":"minor_revision","confidential_remarks":"Benchmark + negative-result papers are a good fit if the venue values infrastructure. The contribution is real and the release is a strong plus. I would not block on multi-run for all 16 models if the authors clearly demote “strict domain specialization” to a provisional observation and keep the robust 74.9/unsaturated claim front-and-center; that is why I chose minor_revision rather than major_revision. Watch for over-claim in the camera-ready abstract."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is real infrastructure, not another Text-to-SQL clone. They ship 100 production-derived end-to-end tasks across PySpark, Hive, MySQL, Presto/Trino, and Flink, run agents in fresh Docker sandboxes, and grade with case-specific deterministic scripts on materialized artifacts plus a process score. Best of 16 models is 74.9. That ceiling, and the judge-inflation ablation, are the parts I’d trust first.\n\nWhat is actually new is the combination: multi-engine batch+stream engineering (not single-query SQL or CSV analysis), human-in-the-loop reconstruction with differential perturbations so inputs discriminate wrong code, and bit-exact graders instead of LLM-as-judge. Table 1 positioning is fair. The 16-model study under one CodeBuddy harness, plus token/tool-call plots, bilingual split, low-score/timeout taxonomies, and the §5.1 judge comparison, is more thorough than most leaderboard drops. Releasing containers and grade.py matters; that is the kind of work people can actually build on.\n\nSoft spots, in proportion. The abstract’s “no single model dominates / strict domain specialization” rests on single-run means over small per-engine n (12–28). Their own §5.2 stability gaps (12–27 points) and App. C timeouts (GLM 5.2 loses 13 points and would reorder) mean engine crowns are noisy estimates, not locked rankings. Fixed harness and wall-clock limits also mix tooling skill with engine skill. Task origin is one enterprise’s desensitized code with LLM-synthesized tables—standard external-validity limit for industrial benches, not a hidden circularity. α≈0.7 and process weights are author choices; they disclose them. None of that sinks the main empirical point that nobody is near saturated end-to-end DE.\n\nWho it’s for: people building or evaluating data/ETL agents, and anyone tired of SQL-only proxies. Math is scoring design, not theory; citations look appropriate; data story is transparent enough for a serious referee.\n\nI’d send it to peer review. Engage if you care about agent eval practice; skim the specialization rhetoric and lean on the ceiling, graders, and release artifacts.","headline":"Solid industrial benchmark paper: the “far from solved” ceiling is credible; the “strict engine specialization” half is thinner than the abstract sells.","tokens_in":25130,"tokens_out":567,"would_cite":true,"duration_ms":18396,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Even the best AI agent scores only 74.9 on real industrial data-engineering tasks, and no model wins across engines.","keywords":["Data Engineering","Autonomous Agent","Enterprise Data Systems","Benchmark","Multi-engine SQL","Rule-based evaluation","ETL","Streaming"],"falsifier":"Re-run the same 16 models on a second, independently sourced suite of production data-engineering tasks (different company or public multi-engine corpus) with the same sandbox-and-rule protocol; if several models clear ~90 overall or one model leads all five engines, the claimed open-challenge ceiling and specialization thesis fail.","tokens_in":24904,"feed_emoji":"⚙️","tokens_out":957,"duration_ms":19000,"temperature":0.7,"pith_summary":"DataClawEval is a new benchmark that asks whether autonomous AI agents can finish real enterprise data-engineering jobs end to end—not just write a SQL query or a short analysis script. It builds 100 tasks from production code across five engines (PySpark, MySQL, HiveSQL, PrestoSQL/Trino, and FlinkSQL), covering batch and streaming work in domains like ops, growth analytics, security, and ads. Each agent must explore live tables in an isolated sandbox, write and debug code, and materialize correct outputs; scoring uses case-specific deterministic scripts rather than an LLM judge. Across 16 frontier models the top overall score is only 74.9, HiveSQL is hardest, MySQL is easiest, and different models lead different engines. The paper’s point is that full-stack data engineering remains an open, engine-specialized challenge, and that execution-grounded grading is required to measure it honestly.","feed_headline":"Best AI data agent scores only 74.9 on real ETL tasks","feed_subtitle":"No model leads all five engines; industrial data engineering stays an open challenge","key_machinery":"DataClawEval itself: a human-in-the-loop pipeline that turns desensitized production code into answer-identifiable tasks (LLM-reconstructed intents and inputs, expert perturbations for discriminability, case-specific graders), then scores agents in fresh Docker sandboxes with a weighted mix of artifact correctness and process quality under deterministic rule-based scripts.","core_discovery":"On 100 production-derived, sandbox-executed data-engineering tasks spanning five engines, the strongest of 16 frontier agents reaches only 74.9 overall; no single model leads every engine, engine difficulty is highly uneven, and token spend does not track quality—so autonomous end-to-end data engineering is still unsolved and models show strict domain specialization rather than general proficiency.","pith_inferences":["Teams deploying a single ‘data agent’ may need engine-specialized models or routers rather than one generalist until cross-engine transfer improves.","The large process-score gap under LLM judges suggests trajectory logging and post-run verification will become first-class training signals, not just product metrics.","Bilingual and timeout analyses hint that harness limits and prompt language can quietly reorder rankings; future suites may need parallel translations and timeout-robust scoring.","If differential testing with expert perturbations is what makes tasks answer-identifiable, similar construction could transfer to neighboring ops domains (infra-as-code, ML pipeline debugging)."],"forward_implications":["Leaderboards that only test Text-to-SQL or final-answer analysis will overstate readiness for production ETL and streaming jobs.","Progress should be reported per engine (especially HiveSQL and FlinkSQL), not only as one average score.","Case-specific rule-based graders that execute outputs against live engines become the standard for this domain; generic LLM judges are shown to inflate and destabilize scores.","Released tasks, containers, and graders give a shared testbed for measuring whether future agents close the gap without changing the harness mid-comparison.","Tool-call volume and token spend are poor proxies for quality; efficient exploration matters more than retry thrash."],"fun_headline_variants":["Strongest AI data agent hits just 74.9 on real industrial ETL","No model sweeps five engines; data engineering agents top out at 74.9","16 frontier agents stall at 74.9 on production data-engineering tasks","DataClawEval: best agent 74.9; strict engine specialization, not general skill","Autonomous ETL still unsolved: peak score 74.9 across five real engines"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That tasks rebuilt from one enterprise’s cleaned production code, plus synthetic inputs and a fixed agent harness, fairly stand in for industrial data engineering in general so the 74.9 ceiling and engine specialization will hold elsewhere.","fun_headline_variants_meta":{"raw":{"variants":["Strongest AI data agent hits just 74.9 on real industrial ETL","No model sweeps five engines; data engineering agents top out at 74.9","16 frontier agents stall at 74.9 on production data-engineering tasks","DataClawEval: best agent 74.9; strict engine specialization, not general skill","Autonomous ETL still unsolved: peak score 74.9 across five real engines"]},"model":"grok-4.5","effort":"low","cost_usd":0.002086,"raw_usage":{"total_tokens":922,"prompt_tokens":807,"num_sources_used":0,"completion_tokens":95,"cost_in_usd_ticks":20864000,"prompt_tokens_details":{"text_tokens":807,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":20,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":807,"tokens_out":95,"duration_ms":2778,"temperature":1.0,"reasoning_tokens":20,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T19:51:29.030684+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same 16 models on a second, independently sourced suite of production data-engineering tasks (different company or public multi-engine corpus) with the same sandbox-and-rule protocol; if several models clear ~90 overall or one model leads all five engines, the claimed open-challenge ceiling and specialization thesis fail.","supporting_citations":[],"review_version":1}