{"id":"79ba19a0-5b99-4063-99db-0204ac9a2972","arxiv_id":"1909.00621","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"ICCMA'17 was won by pyglaf in three tracks, argmat-sat in two, and ArgSemSAT, CoQuiAAS, and argmat-dvisat in one each.","lead":"This paper reports the design and results of the 2017 International Competition on Computational Models of Argumentation, where 16 solvers competed on abstract argumentation problems across seven semantics and a special triathlon track. It documents novel design choices, including new semantics, a correctness-focused scoring scheme, and a benchmark selection stage based on empirical hardness.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"D3 prize hinges on unverified answers: argmat-dvisat's 276 points top pyglaf's 275 while carrying five unchecked USC, so the reported D3 winner is not yet robust.","rationale":"The paper is a transparent competition report and most reported winners are separated by comfortable margins. The benchmark-classification assumption identified by the reader is a legitimate representativeness concern, but it does not invalidate the stated winners under the competition's own rules. The D3 track is different: the margin between first and second place is one point, and the winner has five unverified unique contributions counted as correct. Exactly the class of uncertainty the paper acknowledges in Section 3.4 is decisive here. This does not imply any solver cheated or that organizers acted improperly; it means the publicly reported data do not yet demonstrate that argmat-dvisat outperformed pyglaf on D3. A focused re-verification or a score recomputation would settle it. Since all other award claims appear robust, the appropriate disposition is conditional acceptance rather than rejection.","tokens_in":43646,"tokens_out":7999,"duration_ms":172756,"concrete_test":"Re-verify the five unverified unique solutions of argmat-dvisat and the one of pyglaf in the D3 track with an independent checker for grounded, stable, and preferred extension enumeration, or recompute D3 scores with all unchecked solutions scored as 0 points. If argmat-dvisat's score falls below pyglaf's, the D3 winner changes and the paper's central result list must be corrected.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 3.4 states that solutions that could not be verified were 'rated with 1 point' and claims 'in none of the tracks these had an influence on the ranking.' The D3 results in Section 7.1 contradict that assurance. argmat-dvisat wins D3 with 276 points against pyglaf's 275, a one-point margin, and the table lists 5 unchecked USC for argmat-dvisat versus 1 unchecked USC for pyglaf. Since each unchecked answer contributes +1 if treated as correct but would contribute -5 if wrong, even one false-positive among argmat-dvisat's unverified unique solutions (while pyglaf's unverified answer is correct) reverses the order: 276 - 6 = 270 < 275. Excluding unverified answers entirely also gives pyglaf the win. The reported claim 'argmat-dvisat won D3' is therefore not established by the data as presented, and the paper's own verification note is inaccurate for this track.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports the design and results of the Second International Competition on Computational Models of Argumentation (ICCMA'17), covering the task/track structure, the benchmark suite and call for benchmarks, the hardness-based instance selection, the scoring scheme with a -5 penalty for wrong answers, the answer verification procedure, the 16 participating solvers, and the per-track results and awards. The main claims are the listed winners: pyglaf for CO, ST, and ID; argmat-sat for SST and STG; ArgSemSAT for PR; CoQuiAAS for GR; and argmat-dvisat for the Dung's Triathlon (D3). The paper also compares the 2017 winners with the best ICCMA'15 solvers on common tasks.","tokens_in":43865,"tokens_out":7672,"duration_ms":65889,"significance":"If the reported results are accurate, this paper is a useful archival record for the argumentation-solver community and a methodological reference for future competitions. Its strengths are transparency and reproducibility: the verification chain is described in detail (ASPARTIX-D reference solutions, publicly available ASP encodings for checking SE/EE answers, and a majority-vote fallback for the small fraction of unverifiable cases), and the benchmark selection procedure is specified with enough detail to be re-run. The paper also reports per-solver correct/wrong/timeout/unverified counts, which is the right level of detail for a competition report. However, the archival value depends on the reliability of the awards; the D3 award is currently not supported by the verified data as presented, so the winners list needs correction before the paper can be accepted.","major_comments":[{"comment":"Section 3.4 states that the approximately 0.1% of solutions that could not be verified were 'rated with 1 point' and that 'in none of the tracks these had an influence on the ranking of the solvers.' The D3 results in Figure 9 contradict this claim. argmat-dvisat is reported as the D3 winner with 276 points against pyglaf's 275, while the table lists 5 unchecked USC for argmat-dvisat and 1 unchecked USC for pyglaf. Because unverified solutions were counted as 1 point each, excluding them gives pyglaf 274 verified points and argmat-dvisat 271 verified points; if any single one of argmat-dvisat's five unverified answers were in fact wrong, the -5 penalty would reduce its score below pyglaf's. The reported D3 winner is therefore not robustly established by the data as presented. Please verify the five unchecked USC of argmat-dvisat, or withdraw/qualify the D3 award, and correct the Section 3.4 claim that no track was affected by unverified solutions.","section":"Section 3.4 and Section 7.1 (D3 results, Figure 9)"},{"comment":"The benchmark sets for the new SST, STG, and ID tasks are inherited from task group A, whose hardness classification was performed only for the representative enumeration task EE-PR. The paper asserts that enumeration tasks are representative because decision-task difficulty depends on the queried argument, but it provides no evidence that EE-PR hardness predicts difficulty for the newly introduced SST/STG/ID tasks, which have different complexity profiles and were not present in ICCMA'15. Since the reported winners for these semantics (argmat-sat in SST and STG, pyglaf in ID) are central claims, I ask the authors to either state explicitly and prominently that these rankings are conditional on benchmarks selected by EE-PR hardness, or provide a robustness check (for example, re-ranking the top solvers on benchmarks selected under an alternative classification).","section":"Section 5.1-5.2 (benchmark selection for groups D and E)"}],"minor_comments":[{"comment":"Please clarify whether the 'Correct' counts in the result tables include the unverified solutions that were 'rated with 1 point'; this ambiguity is what makes the D3 margin difficult to interpret from the tables alone.","section":"Section 3.4 / 7.1"},{"comment":"The sentence 'for 22 instances no answer is provided by ASPARTIX-D' should be tied to the subsequent count: please state explicitly that these 22 instances are not included in the 114 instances classified as having no stable extensions, or explain how they were treated.","section":"Section 5.2 (No stable extensions)"},{"comment":"There is a typo: 'usefull comments' should be 'useful comments.'","section":"Acknowledgments"},{"comment":"The phrase 'It is also worth to be noted' is ungrammatical; consider 'It is also worth noting.'","section":"Section 7.1"}],"recommendation":"major_revision","confidential_remarks":"The D3 issue is the main obstacle to acceptance. Since the organizers presumably still have the raw outputs for the five unverified USC of argmat-dvisat, verifying them or declaring the D3 award void is a feasible fix rather than a reason for rejection. The benchmark-selection assumption for groups D and E is a fairness concern that the authors should address explicitly, but it is not, by itself, a blocking error."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know two things about this paper before reading it. It is the definitive write-up of ICCMA'17: design, benchmark selection, verification, full results, and a fair comparison against ICCMA'15 solvers. It deserves to be the standard citation for that event. The second thing is the D3 result. The paper awards the Dung's Triathlon to argmat-dvisat with 276 points against pyglaf's 275. Section 3.4 says unverified answers were scored as correct and that 'in none of the tracks these had an influence on the ranking.' That claim is wrong for D3. argmat-dvisat has 5 unchecked USC, pyglaf has 1. One false positive among argmat-dvisat's unverified answers, while pyglaf's is correct, flips the margin: 276 becomes 270, and pyglaf wins. The paper's own data do not establish that argmat-dvisat won D3.\n\nWhat the paper does well: the design novelties are real and well motivated—new semantics, the correctness-focused scoring with the −5 penalty, the multi-semantics D3 track, and a benchmark-hardness classification that avoids the easy-instance problem of ICCMA'15. The verification pipeline is described honestly: reference answers from ASPARTIX-D, dedicated ASP checks, majority vote for only about 0.1% of 105,350 solutions, with USC counts per solver in every result table. The solver descriptions and the retrospective comparison with ICCMA'15 are useful for anyone tracking progress in argumentation solving. The classification assumption—using EE-PR, EE-ST, SE-GR as representatives and reusing group A's set for SST/STG/ID—is a defensible practical choice, though it means the new-semantics rankings are partly built on an extrapolated difficulty measure.\n\nThe soft spots are proportional. The D3 issue is real and should be fixed: either re-verify those six unchecked USC, or report the D3 ranking as provisional and note that pyglaf may in fact be the winner. Section 3.4's blanket claim that unverified answers never affected rankings is inaccurate as written. The benchmark-representativity point is a minor caveat, not a flaw; competition design always involves such compromises.\n\nFor whom: anyone working on abstract argumentation solvers, competitions, or empirical evaluation of AI systems. I would bring it to a reading group and would cite it for the ICCMA'17 results and benchmark-selection methodology. My own verdict is skeptical only on the D3 award, not on the rest of the paper.\n\nRecommendation: yes, send it to peer review. A good referee will catch the D3 issue; the rest of the paper holds up and the community needs this record.","headline":"A thorough, transparent ICCMA'17 report whose D3 prize call is one unchecked answer away from being wrong—worth refereeing, but the organizers should re-examine their own verification claim.","tokens_in":44377,"tokens_out":2263,"would_cite":true,"duration_ms":20022,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports the design and results of the second International Competition on Computational Models of Argumentation, stating that under its scoring and verification procedure pyglaf won the complete, stable, and ideal semantics…","keywords":["abstract argumentation","solver competition","ICCMA'17","argumentation semantics","benchmark selection","solver evaluation","computational logic","Dung's triathlon"],"falsifier":"Re-score the competition after replacing each unverifiable answer with 0 points instead of 1 point and independently checking the instances for which the reference solver produced no answer; if any track's top solver changes under either operation, the reported winner list is not robust to the verification rule.","tokens_in":43468,"feed_emoji":"🏆","tokens_out":12970,"duration_ms":104490,"temperature":0.7,"pith_summary":"The paper reports on the second International Competition on Computational Models of Argumentation, a benchmark contest for solvers that compute with abstract argumentation frameworks. It establishes who won each track under the competition's own rules and argues that those results are meaningful because of three design features: a call for benchmarks, a hardness-based instance selection stage, and a scoring scheme that penalizes wrong answers. The winners are pyglaf for complete, stable, and ideal semantics; argmat-sat for semi-stable and stage semantics; ArgSemSAT for preferred; CoQuiAAS for grounded; and argmat-dvisat for Dung's Triathlon. The reason to care is that the competition is the main instrument for measuring progress in this area of AI, and the paper's comparison with the 2015 edition shows where solver technology has improved and where it has not.","feed_headline":"Newcomers win six of eight tracks at the 2017 argumentation competition","feed_subtitle":"pyglaf wins three tracks, argmat-sat two; ArgSemSAT, CoQuiAAS, argmat-dvisat one each.","key_machinery":"The central object is the abstract argumentation framework, a directed graph in which nodes are arguments and edges are attacks, together with the seven semantics under evaluation: complete, preferred, stable, semi-stable, stage, grounded, and ideal. The machinery that carries the argument is the competition protocol: 24 tasks grouped into five hardness classes (A to E), a benchmark suite of 3990 instances across 11 domains, a selection stage that classifies instances from very easy to too hard using reference solvers from the 2015 edition, a scoring rule of +1 per correct answer and -5 per wrong answer with ties broken by cumulative runtime, and an answer-verification pipeline based on ASPARTIX-D reference solutions, extension-checking encodings, and majority agreement.","core_discovery":"The central claim is that the second International Competition on Computational Models of Argumentation was run as designed and that its recorded winners are correct under that design. The paper describes the verification pipeline: reference solutions were generated with ASPARTIX-D, single extensions were checked with dedicated ASP encodings, and the few answers that could not be checked were resolved by majority agreement; only about 0.1% of the 105,350 submitted solutions could not be verified, and the paper states that these did not influence any track ranking. The reported winners are pyglaf for complete, stable, and ideal semantics; argmat-sat for semi-stable and stage semantics; ArgSemSAT for preferred; CoQuiAAS for grounded; and argmat-dvisat for Dung's Triathlon. On the competition's common tracks with 2015, the comparison shows mixed progress: some 2017 winners outperform their predecessors, while others remain comparable or even behind.","pith_inferences":["Beyond the paper's claims, the transferability of the benchmark classification can be tested directly by running a dedicated hardness classification for the semi-stable and stage tasks and comparing it with the group A classification; if the distributions differ, the rankings for the new semantics are shaped by the benchmark choice rather than by solver strength.","Beyond the paper's claims, the verification rule can be stress-tested by re-scoring all submitted answers with 0 points for the unverifiable cases instead of 1 point; a changed ordering in any track would show that the winner list is sensitive to the verification rule.","Beyond the paper's claims, the 2015-versus-2017 comparison can be extended to the new semantics by running the 2015 winners on the semi-stable, stage, and ideal benchmarks, yielding a complete two-year progress picture.","Beyond the paper's claims, the Dung's Triathlon idea can be generalized to other triples of semantics at different complexity levels, such as semi-stable, stage, and ideal, to see whether exploiting inter-semantics relationships transfers beyond the original trio."],"forward_implications":["The winners become the reference points that future solvers must beat on their respective tracks.","The scoring scheme with a -5 penalty for wrong answers pushes unreliable solvers to the bottom of the rankings, so the rule can be retained to discourage guessing.","The hardness-based instance selection gives a reproducible way to build a balanced benchmark set, so future editions can apply the same classification procedure to new domains.","The common-track comparison shows that progress between 2015 and 2017 is uneven: some 2017 winners clearly outperform their predecessors, while others are comparable or slightly worse."],"supporting_citations":[{"why":"Defines abstract argumentation frameworks and the grounded, preferred, and stable semantics; the D3 track is built on these three semantics.","marker":"(Dung, 1995)"},{"why":"Introduces semi-stable semantics, one of the new tracks added in this edition.","marker":"(Caminada et al., 2012)"},{"why":"Introduces stage semantics, another new track in the competition.","marker":"(Verheij, 1996)"},{"why":"Introduces ideal semantics and the ideal extension computation used in the ID track.","marker":"(Dung et al., 2007)"},{"why":"Reports the design and results of ICCMA'15, providing the baseline, the benchmark generators' context, and the reference solvers used for hardness classification.","marker":"(Thimm and Villata, 2017)"},{"why":"Provides the implementation of the three ICCMA'15 benchmark generators that form part of the benchmark suite.","marker":"(Cerutti et al., 2014b)"},{"why":"Presents the ASPARTIX encodings underlying the reference solution generator used in the verification pipeline.","marker":"(Egly et al., 2010)"},{"why":"Describes the improved ASPARTIX-D system used to generate reference solutions and to check stable-extension existence.","marker":"(Gaggl et al., 2015)"}],"fun_headline_variants":["ICCMA'17: six track wins for newcomer solvers","Argumentation contest 2017: new solvers take six of eight","Second argumentation competition: newcomers win six of eight","Newcomer solvers win six of eight tracks at 2017 argumentation event","Report: newcomers win six of eight at argumentation contest"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole ranking depends on the assumption that how hard an instance is for the representative enumeration tasks (enumerating preferred extensions, enumerating stable extensions, and finding one grounded extension) predicts how hard it is for every other task in the same group, including the newly added semi-stable, stage, and ideal tasks that were assigned to the same benchmark set without their own difficulty classification.","fun_headline_variants_meta":{"raw":{"variants":["ICCMA'17: six track wins for newcomer solvers","Argumentation contest 2017: new solvers take six of eight","Second argumentation competition: newcomers win six of eight","Newcomer solvers win six of eight tracks at 2017 argumentation event","Report: newcomers win six of eight at argumentation contest"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001721,"raw_usage":{"total_tokens":6787,"prompt_tokens":903,"completion_tokens":5884,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":519,"completion_tokens_details":{"reasoning_tokens":5794}},"tokens_in":519,"tokens_out":5884,"duration_ms":35572,"temperature":1.0,"reasoning_tokens":5794,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:41:42.830595+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score the competition after replacing each unverifiable answer with 0 points instead of 1 point and independently checking the instances for which the reference solver produced no answer; if any track's top solver changes under either operation, the reported winner list is not robust to the verification rule.","supporting_citations":[{"cited_title":"Heureka: A general heuristic backtracking solver for abstract argumentation","cited_arxiv_id":null,"evidence_quote":"Reports the design and results of ICCMA'15, providing the baseline, the benchmark generators' context, and the reference solvers used for hardness classification."}],"review_version":1}