REVIEW 3 major objections 6 minor 22 references
Decision Making with Argumentation Graphs
T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A structured online debate and voting system allocated a real 20,000-euro budget with broad acceptance in the reported experiment.
desk verdict A real participatory budgeting case study with open data and honest self-criticism, whose headline acceptance claim is weakened by a familiar outcome-satisfaction confound. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the argumentation graph of D-BAS, the Dialog-Based Argumentation System: a directed acyclic graph whose nodes are the issue, positions (proposals), and atomic statements, and whose directed edges are arguments marked as support or attack; it lets a participant argue with the system rather than with every other participant, preserving structure as the discussion grows. The decision machinery is a deliberately transparent scoring rule: proposals carry a fixed cost, voters approve and rank them, first preference earns $N$ points, second earns $N-1$, down to $0$ for unranked, and approval counts break ties. Winners are then chosen greedily by score while staying inside the budget. This combination is what the paper claims makes the result understandable and acceptable to untrained participants.
What would settle it
Repeat the same procedure with a budget that can fund only one or two of many competing proposals, and check whether fairness and acceptance ratings stay high among participants whose proposals lose. If acceptance drops sharply when most people lose, the original satisfaction was mostly about outcome, not about the argumentation-and-voting procedure.
Extended reading notes
Core claim
Central claim: a decision procedure built on dialog-based argumentation can distribute real funds and be accepted by participants. In the experiment, students could submit proposals with a price tag, discuss them in D-BAS, and then approve and rank proposals in decide. Rankings were scored with a Borda-style count in which a voter's first preference gets $N$ points, the second gets $N-1$, and unranked proposals get $0$, with the number of approvals as tiebreaker; the winning set was built greedily by taking top-scoring proposals that fit the remaining budget. Eight final proposals went to a vote and five won; the participant survey returned means above 5 on the 7-point scale for most fairness and acceptance items, including agreement that the decision should be made this way in the future. The paper's own framing is that the exercise was 'well accepted and thus successful'.
Load-bearing premise
The acceptance conclusion rests on the assumption that survey answers from a self-selected group measure the quality of the procedure itself, rather than their happiness that five of eight proposals they liked won; the paper itself flags that high satisfaction may have been caused by the large number of winning proposals.
Editorial extensions
If this is right
- A structured argumentation graph can feed a real resource-allocation vote, not just a discussion forum.
- An unmonitored process is not enough: rule-breaking and vague proposals will occur, so an editorial review phase is needed before voting.
- When the vote is separated from the argumentation, most participants choose to vote rather than argue; only 10 arguments were added during the voting phase.
- In this case the winning set was robust to the scoring method: Borda, approval, and single-vote rankings agreed on the top proposals, though a Top-2 approval rule would have changed one winner.
- Future runs should limit the number of winners if the goal is to test acceptance of the procedure independently of outcome.
Reading between the lines
- Editorial inference: the D-BAS-plus-decide pattern is a reusable template for small-group budget allocation wherever single-sign-on authentication exists; the same phases could serve departmental, neighbourhood, or campus funds.
- Editorial inference: a natural next experiment is to vary the win rate deliberately to separate outcome effects from process effects, since the paper itself suspects high satisfaction came from many proposals winning.
- Editorial inference: if argumentation participation stays lower than voting, an automated decision agent that infers preferences from argumentation alone would not yet be feasible; the paper itself notes this as a future scenario.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The thesis extends the Dialog-Based Argumentation System (D-BAS) with cost-bearing proposals and a companion voting service called "decide," then deploys both in a real participatory budgeting process at Heinrich-Heine-University Düsseldorf, where computer science students allocated €20,000 of quality-enhancement funds. The paper describes the argumentation graph, the preference aggregation scheme (a truncated Borda count with approval voting as tiebreaker), the three-phase experiment (proposals and argumentation, an unplanned review/correction phase, and voting), and the results: 142 voters chose five of eight proposals, spending €18,650 of the budget. A post-experiment survey (N between 22 and 47 per item) is used to support the central claim that "the procedure is well accepted and thus successful." The thesis also discusses future improvements and compares the approach with participatory budgeting in Porto Alegre, Wuppertal, and Reykjavík.
Significance. The work is a useful case study at the intersection of structured argumentation and participatory budgeting. Its strengths include a real deployment with actual funds, an open-source implementation, raw preference data in the appendix, and an unusually candid account of design changes and limitations. The paper also gives a clear, accessible description of the scoring method and compares it with several real-world participatory budgeting processes. However, the headline empirical claim—that the procedure is well accepted and therefore successful—rests on survey evidence that is confounded with outcome satisfaction and collected from a self-selected subsample after a favorable result. As a single-case pilot with mid-experiment redesign, the study can support a qualified claim about feasibility and reported satisfaction, but not the strong acceptance claim made in the abstract and conclusion. The central derivation of the decision procedure is sound, and the limitations are acknowledged in the future-work section, so the weaknesses are addressable through reframing and additional analysis rather than being irreparable.
major comments (3)
- [Abstract; Ch. 7; §5.6] The central claim that "the procedure is well accepted and thus successful" is not established by the survey data, because the survey was administered only after the winners were announced, when five of eight proposals won and 93% of the budget was spent. The respondents were self-selected (item Ns of 22–47 versus 142 voters, with no response rate reported), and the author explicitly concedes the confound in Chapter 7: "the satisfaction in the outcome is high because most of the proposals were able to be included in the winning set of proposals." The pattern in Figure 5.7 is consistent with this: decision-acceptance items (e.g., "I accept the decision," mean 6.51) are notably higher than procedural-justice items (means 5.13–5.76). The abstract and conclusion should be revised to claim, at most, that participants reported high satisfaction with the outcome and expressed desire for future procedures, unless the author can provide analysis that separates outcome favorability from procedure acceptance, such as responses from participants whose proposals lost or items asked before the outcome was known.
- [§5.3, §5.4, Ch. 7] The experiment was redesigned mid-course in ways that undermine the claim that the tested procedure is the procedure described in the design chapters. Phase 2 was added after the first day when many proposals violated the rules, the proposal-submission window was closed early, and the voting phase was shortened from one week to five days. The author himself observes in Chapter 7 that the voting phase "actually became just a traditional vote." Thus the acceptance evidence pertains to an ad hoc, supervised process rather than to the proposed unsupervised D-BAS-plus-decide workflow. This should be treated as a fundamental limitation of the experiment as a test of the original design, and the paper should state that the experiment was a pilot that iterated on the procedure, not a confirmatory test of the designed process.
- [§5.6] The survey analysis needs a discussion of non-response and self-selection to assess the representativeness of the acceptance figures. The paper reports item Ns (22–47) and the total number of voters (142), but never reports the number of students invited to the survey, the response rate, or any comparison between survey respondents and the full voter population. Without this information, it is impossible to know whether the positive averages reflect the views of typical participants or of a self-selected subset, particularly those who were satisfied with the outcome. At a minimum, the response rate should be reported, and the limitations of the survey as a non-probability sample should be acknowledged directly in Section 5.6.
minor comments (6)
- [Abstract] The phrase "The results indicate that the procedure is well accepted and thus successful" conflates outcome satisfaction with procedural acceptance; consider rewording to "participants reported high satisfaction with the outcome and expressed interest in future procedures."
- [§4.4] The description of the aggregation method would be clearer if the truncated Borda parameter N were explicitly defined as the maximum number of preferences cast by any single participant, and if an example showed how unranked proposals receive the implicit score of 0; the example in §4.5 covers this well, so a forward reference there would help.
- [Table 5.1] The table lists both Borda and approval scores without a common scale, which makes the statement that the outcome would be the same under approval voting easier to verify if the two columns were plotted together or normalized; consider adding a short note clarifying that the ranking, not the absolute values, is the basis for the comparison.
- [Figure 5.7] The boxplot figure is difficult to parse because the extracted statistics table duplicates rows and the N annotations are garbled in the text; please ensure the figure itself clearly labels each item N and quartile values, and check the rendering of the German-language table.
- [§5.2] The paper reports 52 participants in the argumentation and 142 voters, and later notes that 18 of the 52 arguers did not vote. It would be useful to state explicitly that participation in argumentation and voting were measured on different bases (registered versus voted), since the text could be misread as a drop in total participation.
- [§5.8.1] The typo "as+ it could undermine" should be corrected, and the brief discussion of structured argumentation violations could be tightened by adding examples for each of the three violation types.
Circularity Check
No significant circularity: the acceptance claim rests on independent survey data and experimental records, not on the cited D-BAS results.
full rationale
The thesis's central empirical claim—that the procedure was well accepted and thus successful—is supported by the experiment's own survey responses, participation counts, and vote data, not by the D-BAS publications cited from the same laboratory. The scoring procedure is a transparent mix of approval and Borda scores described in the thesis itself, and the simulated alternative scoring methods in Section 5.5 are computed from the same raw preference data rather than being presented as independent predictions. The author's Future Work caveat that satisfaction with the outcome may be high because most proposals won is an acknowledged validity limitation of the acceptance inference, not a definitional reduction of the procedure's acceptance to its outcome; it does not make the survey evidence circular in the technical sense used here. Self-citations describe the prior D-BAS system and earlier field experiments, but the load-bearing evidence for the acceptance conclusion is the new experiment's own measurements. Therefore no circular step can be exhibited, and the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (4)
- N for Borda scoring =
8 (maximum number of preferences submitted by any voter)
- Minimum proposal cost =
100 EUR
- Maximum proposal cost =
20000 EUR
- Tiebreaker order =
approval score, then proposal creation order
assumptions (4)
- domain assumption D-BAS argumentation graph structure with issues, positions, statements, and attack/support relations adequately represents the reasoning exchanged.
- domain assumption Participants' self-assessed satisfaction is a valid measure of procedural success.
- domain assumption LDAP enrollment filtering correctly identifies eligible computer science students.
- standard math Voter preferences can be meaningfully aggregated by summing Borda scores with implicit zero for unranked proposals.
Cite this review
Pith. "Pith review of Decision Making with Argumentation Graphs." pith.science (2026). https://pith.science/paper/77QLCP72
@misc{pith2026190803357,
author = {Pith},
title = {Pith review of: Decision Making with Argumentation Graphs},
year = {2026},
howpublished = {\url{https://pith.science/paper/77QLCP72}},
note = {Machine review of arXiv:1908.03357}
}
read the original abstract
This work is about making decisions by digital means. Funds should be distributed by the students of Heinrich-Heine-University. The proposals were made by the students themselves without further influence. For this purpose, dialog-based argumentation is used to give the participants a better understanding of various arguments. In addition, a software service has been developed which allows the students to express their preferences for various proposals. An experiment was carried out at the university, which should prove whether students are satisfied with this type of participation. The results indicate that the procedure is well accepted and thus successful. However, improvements to the process itself were necessary during the experiment and should be considered for future procedures. Further procedures are desired.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Abers , Rebecca; Brandão , Igor; King , Robin; Votto , Daniely: Porto Alegre: Participatory Budgeting and the Challenge of Sustaining Transformative Change. (2018), June
work page 2018
-
[2]
https://web.archive.org/web/20140809032935/http://eutopiamagazine.eu/en/r \,Version:\,07 2014
Bjarnason , Róbert: 'Your Priorities': An icelandic story of e-democracy. https://web.archive.org/web/20140809032935/http://eutopiamagazine.eu/en/r \,Version:\,07 2014
-
[3]
Bjarnason , Róbert: Citizen participation and digital tools for upgrading democracy in Iceland and beyond. https://oecd-opsi.org/wp-content/uploads/2019/04/Citizens-Foundation-Citizen-participation-and-digital-tools-V4.pdf. \,Version:\,April 2019
work page 2019
- [4]
-
[5]
http://siteresources.worldbank.org/INTEMPOWERMENT/Resources/14657_Partic-Budg-Brazil-web.pdf
Bhatnagar , Deepti; Rathore , Animesh; Moreno Torres , Magüi; Kanungo , Parameeta: Participatory Budgeting in Brazil. http://siteresources.worldbank.org/INTEMPOWERMENT/Resources/14657_Partic-Budg-Brazil-web.pdf
-
[6]
https://citizens.is/portfolio_page/my-neighbourhood/
Citizens Foundation : My Neighbourhood. https://citizens.is/portfolio_page/my-neighbourhood/. \,Version:\,December 2011
work page 2011
-
[7]
Dummett , Michael: Voting procedures. (1984)
work page 1984
-
[8]
Dummett , Michael A.: Principles of electoral reform. Oxford University Press, 1997
work page 1997
Show all 22 references
-
[9]
In: Social Choice and Welfare 40 (2013), Feb, Nr
Emerson , Peter: The original Borda count and partial voting. In: Social Choice and Welfare 40 (2013), Feb, Nr. 2, 353--358. http://dx.doi.org/10.1007/s00355-011-0603-9. DOI 10.1007/s00355--011--0603--9. ISSN 1432--217X
2013 doi
-
[10]
In: Australian Journal of Political Science 49 (2014), 04
Fraenkel , Jon; Grofman , Bernard: The Borda Count and its real-world alternatives: Comparing scoring rules in Nauru and Slovenia. In: Australian Journal of Political Science 49 (2014), 04. http://dx.doi.org/10.1080/10361146.2014.900530. DOI 10.1080/10361146.2014.900530
2014
-
[11]
In: Computational Models of Argument, Potsdam, 2016, S
Krauthoff , Tobias; Betz , Gregor; Baurmann , Michael; Mauve , Martin: Dialog-Based Online Argumentation . In: Computational Models of Argument, Potsdam, 2016, S. 33--40
2016
-
[12]
In: International Journal of Electronic Government Research 13 (2017), 04, S
Kaur Kapoor , Kawaljeet; Omar , Amizan; Sivarajah , Uthayasankar: Enabling Multichannel Participation Through ICT Adaptation. In: International Journal of Electronic Government Research 13 (2017), 04, S. 66--80. http://dx.doi.org/10.4018/IJEGR.2017040104. DOI 10.4018/IJEGR.2017040104
2017 doi
-
[13]
In: Computational Models of Argument: Proceedings of COMMA 2018 305 (2018), S
KRAUTHOFF , Tobias; METER , Christian; BAURMANN , Michael; BETZ , Gregor; MAUVE , Martin: D-BAS-A Dialog-Based Online Argumentation System. In: Computational Models of Argument: Proceedings of COMMA 2018 305 (2018), S. 325
2018
-
[14]
In: Proceedings of the 1st Workshop on Advances in Argumentation in Artificial Intelligence, Bari, 2017, S
Krauthoff , Tobias; Meter , Christian; Mauve , Martin: Dialog-Based Online Argumentation: Findings from a Field Experiment . In: Proceedings of the 1st Workshop on Advances in Argumentation in Artificial Intelligence, Bari, 2017, S. 85--99
2017
-
[15]
http://nbn-resolving.de/urn/resolver.pl?urn=urn:nbn:de:hbz:061-20180628-094005-1
Krauthoff , Tobias: Dialog-Based Online Argumentation, Heinrich-Heine-University, dissertation, June 2018. http://nbn-resolving.de/urn/resolver.pl?urn=urn:nbn:de:hbz:061-20180628-094005-1. URN urn:nbn:de:hbz:061--20180628--094005--1
2018
-
[16]
Springer Science & Business Media, 2012
Lueg , Christopher; Fisher , Danyel: From Usenet to CoWebs: interacting with social information spaces. Springer Science & Business Media, 2012
2012
-
[17]
In: Computational Models of Argument, Warsaw, 2018, S
Meter , Christian; Schneider , Alexander; Mauve , Martin: EDEN: Extensible Discussion Entity Network . In: Computational Models of Argument, Warsaw, 2018, S. 257--268
2018
-
[18]
http://dip21.bundestag.de/dip21/btd/18/116/1811614.pdf
Entwurf eines Ersten Gesetzes zur Änderung des E-Government-Gesetzes. http://dip21.bundestag.de/dip21/btd/18/116/1811614.pdf. \,Version:\,March 2017. Drucksache 18/11614
2017
-
[19]
In: International Political Science Review 23 (2002), Nr
Reilly , Benjamin: Social choice in the south seas: Electoral innovation and the borda count in the pacific island countries. In: International Political Science Review 23 (2002), Nr. 4, S. 355--372
2002
-
[20]
(2018), 02, 21--66
Ruesch , Michelle: D3.2 Pilots implementation - final. (2018), 02, 21--66. https://empatia-project.eu/wp-content/uploads/2018/07/EMPATIA_Deliverable_3.2_13FINAL.pdf
2018
-
[21]
(2000), October
Wampler , Brian: A Guide to Participatory Budgeting. (2000), October. https://www.partizipation.at/fileadmin/media_data/Downloads/themen/A_guide_to_PB.pdf
2000
-
[22]
write newline
" write newline "" before.all 'output.state := FUNCTION fin.entry write newline FUNCTION set.period.dash output.state before.all = 'skip period.dash 'output.state := if FUNCTION set.period output.state before.all = 'skip period.dash 'output.state := if FUNCTION set.period.dash...
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.