{"id":"3f87bca2-60e0-43fc-afa5-1c8f1d300661","arxiv_id":"2605.26079","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ABA audits 168 benchmarks and flags issues in over 25.7% of tasks, with removal of flagged tasks shifting model rankings and raising scores on SWE-bench Verified and Terminal-Bench 2 by 9.9% and 9.6%.","lead":"The paper introduces ABA, an agent-based system that automatically checks AI benchmark tasks for problems like unclear instructions, missing setup details, and wrong answers. This matters because flawed tests can give misleading pictures of how capable different AI models really are.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Validation of detected defects rests on unspecified expert review and upstream PRs without reported inter-annotator agreement, precision on held-out set, or false-positive controls.","rationale":"The reader's weakest_assumption directly isolates the same validation gap that underpins the headline percentages and causal performance claims; no stronger internal inconsistency or unstated assumption appears from the supplied abstract and claim text.","tokens_in":1761,"tokens_out":312,"duration_ms":13373,"concrete_test":"Sample 100 tasks (50 flagged, 50 unflagged) stratified by domain; have two independent experts blinded to ABA output label each as defective or clean using the paper's defect taxonomy; compute Cohen's kappa and precision of ABA flags; if kappa < 0.6 or precision < 75%, the 25.7% rate and downstream deltas are unreliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that ABA's 25.7% defect rate consists of genuine benchmark flaws whose removal causally improves model rankings and scores (9.9% / 9.6%). The paper states precision is \"validated by expert review and independent third-party reports such as upstream PRs,\" yet supplies no IAA statistic, no blinded review protocol, no count of reviewed items, and no controlled false-positive measurement on tasks known to be clean. Without these, the fraction of false positives among the flagged tasks remains unknown; even modest contamination would shrink or reverse the reported performance deltas and ranking shifts.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces Auto Benchmark Audit (ABA), an agentic framework to systematically audit individual tasks in AI benchmarks for issues such as ambiguous design, environment conflicts, and incorrect ground truths. Applied to 168 benchmarks across nine domains, ABA flags critical issues in over 25.7% of tasks; precision is stated to be validated by expert review and upstream PRs. Filtering the flagged tasks is reported to shift model rankings and raise average performance by 9.9% on SWE-bench Verified and 9.6% on Terminal-Bench 2.","tokens_in":1859,"tokens_out":395,"duration_ms":13275,"significance":"If the automated detections prove reliable and the reported performance deltas are shown to be causal rather than artifacts of post-hoc selection, ABA could provide a scalable method for improving benchmark quality and the validity of capability assessments for agents and LLMs.","major_comments":[{"comment":"Abstract (validation paragraph): the central claim that ABA identifies genuine defects in 25.7% of tasks, whose removal produces 9.9% and 9.6% performance gains, rests on the assertion that 'precision ... is validated by expert review and independent third-party reports such as upstream PRs.' No inter-annotator agreement statistic, blinded review protocol, count of reviewed items, or controlled false-positive measurement on a held-out set of clean tasks is supplied. This is load-bearing for the ranking-shift and score-increase results.","section":"Abstract"},{"comment":"Results / Experimental Setup (implied): the manuscript provides no details on agent prompting, decision thresholds, or how expert validation was performed, making it impossible to assess reproducibility or the risk that post-hoc filtering effects are driven by unquantified false positives.","section":"Results section"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments emphasizing the need for transparent validation metrics and experimental details. We agree these elements are essential to substantiate the precision claims and support reproducibility. We will revise the manuscript to incorporate the requested information.","responses":[{"response":"We acknowledge that the current manuscript does not report inter-annotator agreement, a blinded protocol, or a controlled false-positive evaluation on clean tasks. The validation described relied on expert review of flagged tasks plus independent confirmation via upstream PRs. In revision we will add a new subsection under Methods that specifies: (i) the number of tasks reviewed by experts, (ii) reviewer backgrounds and any agreement statistics computed, (iii) the protocol used (including whether review was blinded), and (iv) concrete examples of upstream PRs that independently confirmed ABA flags. If additional controlled measurements on held-out clean tasks are feasible within the revision timeline, they will be included; otherwise the limitations of the current validation approach will be explicitly stated.","revision_made":"yes","referee_comment":"[Abstract] Abstract (validation paragraph): the central claim that ABA identifies genuine defects in 25.7% of tasks, whose removal produces 9.9% and 9.6% performance gains, rests on the assertion that 'precision ... is validated by expert review and independent third-party reports such as upstream PRs.' No inter-annotator agreement statistic, blinded review protocol, count of reviewed items, or controlled false-positive measurement on a held-out set of clean tasks is supplied. This is load-bearing for the ranking-shift and score-increase results."},{"response":"We agree the manuscript currently omits these implementation details. The revised version will expand the Experimental Setup and ABA Framework sections to document: the exact system and user prompts supplied to the auditing agent, the decision thresholds and aggregation rules used to flag issues, and the full expert-validation workflow (including reviewer instructions and criteria). These additions will enable independent reproduction and allow readers to evaluate the false-positive risk directly.","revision_made":"yes","referee_comment":"[Results section] Results / Experimental Setup (implied): the manuscript provides no details on agent prompting, decision thresholds, or how expert validation was performed, making it impossible to assess reproducibility or the risk that post-hoc filtering effects are driven by unquantified false positives."}],"tokens_in":1386,"tokens_out":505,"duration_ms":28868,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's core contribution is an agent-based auditing system that processes 168 benchmarks and flags problems like ambiguous instructions, environment mismatches, and wrong ground truth in 25.7% of tasks. Removing those tasks changes model orderings and lifts scores on SWE-bench Verified and Terminal-Bench 2 by roughly 10%. That scale of application and the released annotations are the genuinely new pieces.\n\nThe work is useful because it turns a known complaint about benchmarks into a repeatable process and ships the tool. Prior checks were mostly manual or narrow; this one runs across domains and produces concrete before-and-after numbers.\n\nThe soft spot is the validation step. The abstract says precision is checked by expert review and upstream PRs, yet gives no inter-annotator numbers, no count of items reviewed, and no test on clean tasks to bound false positives. If even a modest share of the flagged items are not real defects, the reported ranking shifts and score gains shrink or disappear. The method description also omits prompting details and decision thresholds, which makes it hard to judge how much the results depend on specific agent choices.\n\nThis is aimed at people who run or maintain large evaluation suites. A reader who needs a starting point for cleaning benchmarks will get concrete examples and data to build on. The central claim is plausible but not yet tightly supported, so the paper deserves a serious referee to press on the validation protocol and reproducibility of the agent runs. I would send it out rather than desk-reject.","headline":"ABA gives a practical agentic way to flag benchmark flaws at scale, but the 25.7% defect rate and ranking shifts rest on validation that lacks reported precision controls.","tokens_in":2351,"tokens_out":384,"would_cite":false,"duration_ms":14953,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"An automated agentic framework audits AI benchmarks and detects defects in more than 25 percent of tasks, altering model performance rankings when filtered.","keywords":["benchmark auditing","AI agents","LLM evaluation","automated verification","benchmark defects","agentic frameworks","task auditing","performance assessment"],"falsifier":"A controlled study measuring the false positive rate of ABA by having multiple independent experts review a random sample of flagged and unflagged tasks and comparing agreement rates.","tokens_in":2644,"feed_emoji":"🔍","tokens_out":650,"duration_ms":17146,"temperature":0.7,"pith_summary":"Modern AI benchmarks have grown too complex for reliable human verification, often containing hidden flaws like ambiguous instructions or wrong ground truths. The paper introduces Auto Benchmark Audit (ABA), an agentic system that automatically checks each task for issues such as environment conflicts and incomplete specifications. Running ABA across 168 benchmarks in nine domains reveals problems in over 25.7 percent of tasks. Expert validation supports the findings, and removing the flawed tasks changes how models rank and raises average scores on benchmarks like SWE-bench Verified by 9.9 percent. This suggests that current capability assessments for agents and LLMs are distorted by these defects.","feed_headline":"Agent audit finds defects in 25.7% of AI benchmark tasks","feed_subtitle":"Removing them shifts model rankings and lifts average scores by up to 9.9% on verified suites.","key_machinery":"The Auto Benchmark Audit (ABA) agentic framework that deploys agents to check individual benchmark tasks for hidden dependencies, specification gaps, and grading logic problems.","core_discovery":"The ABA framework systematically audits benchmark tasks and identifies critical issues including ambiguous task design, execution environment conflicts, and incorrect ground truths in over 25.7% of the evaluated tasks. Validation through expert review and upstream PRs confirms the detections, and filtering these problematic tasks shifts model rankings while increasing average performance on SWE-bench Verified and Terminal-Bench 2 by 9.9% and 9.6%, respectively.","pith_inferences":["Future benchmarks may need to incorporate automated auditing as a standard pre-release step.","Model comparisons in existing papers could shift if similar undetected flaws exist in other common suites.","Agentic auditing tools could be extended to suggest or apply fixes for the identified issues.","Rates of defects might differ across domains not covered in the nine studied here."],"forward_implications":["Filtering flawed tasks changes model rankings on the audited benchmarks.","Average performance increases by 9.9% on SWE-bench Verified after removal.","Average performance increases by 9.6% on Terminal-Bench 2 after removal.","The audits cover 168 benchmarks across nine domains from frontier LLM evaluations and prior NeurIPS papers."],"fun_headline_variants":["ABA flags 25.7% benchmark tasks with issues","25.7% of benchmark tasks defective per ABA","ABA finds 25.7% AI benchmark tasks flawed","Automated audit detects 25.7% faulty tasks"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The automated audits accurately detect real defects rather than generating many false positives, with validation relying on expert review without quantified agreement metrics.","fun_headline_variants_meta":{"raw":{"variants":["ABA flags 25.7% benchmark tasks with issues","25.7% of benchmark tasks defective per ABA","ABA finds 25.7% AI benchmark tasks flawed","Automated audit detects 25.7% faulty tasks"]},"model":"grok-4.3","cost_usd":0.006704,"raw_usage":{"total_tokens":3120,"prompt_tokens":662,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":67037000,"prompt_tokens_details":{"text_tokens":662,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2395,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":662,"tokens_out":63,"duration_ms":19413,"temperature":1.0,"reasoning_tokens":2395,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T21:38:32.674850+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled study measuring the false positive rate of ABA by having multiple independent experts review a random sample of flagged and unflagged tasks and comparing agreement rates.","supporting_citations":[],"review_version":1}