Pith. sign in

REVIEW 3 major objections 6 minor 22 references

Decision Making with Argumentation Graphs

T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A structured online debate and voting system allocated a real 20,000-euro budget with broad acceptance in the reported experiment.

desk verdict A real participatory budgeting case study with open data and honest self-criticism, whose headline acceptance claim is weakened by a familiar outcome-satisfaction confound. read the letter →

arxiv 1908.03357 v1 pith:77QLCP72 submitted 2019-08-09 cs.CY cs.GT

classification cs.CYcs.GT
keywords participatorybudgetingargumentationgraphD-BASonlinedeliberationBordacountapprovalvotinge-participationcasestudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that an online participatory budgeting process can work: a community of students decides how to spend a real 20,000-euro fund by proposing projects, arguing for and against them in a structured argumentation graph, then voting on the final proposals through a simple ranking interface. The reported experiment drew 142 voters and produced five winning proposals that together used 18,650 of the 20,000 euros. Survey responses, mostly above 5 on a 1-to-7 scale, led the author to conclude that the procedure was well accepted and therefore successful. The thesis also acknowledges that the process had to be changed mid-course: rule-violating or underspecified proposals forced an unplanned review phase, so the final vote was preceded by editorial filtering rather than continued discussion.

What carries the argument

The load-bearing object is the argumentation graph of D-BAS, the Dialog-Based Argumentation System: a directed acyclic graph whose nodes are the issue, positions (proposals), and atomic statements, and whose directed edges are arguments marked as support or attack; it lets a participant argue with the system rather than with every other participant, preserving structure as the discussion grows. The decision machinery is a deliberately transparent scoring rule: proposals carry a fixed cost, voters approve and rank them, first preference earns $N$ points, second earns $N-1$, down to $0$ for unranked, and approval counts break ties. Winners are then chosen greedily by score while staying inside the budget. This combination is what the paper claims makes the result understandable and acceptable to untrained participants.

What would settle it

Repeat the same procedure with a budget that can fund only one or two of many competing proposals, and check whether fairness and acceptance ratings stay high among participants whose proposals lose. If acceptance drops sharply when most people lose, the original satisfaction was mostly about outcome, not about the argumentation-and-voting procedure.

Watch

Extended reading notes

Core claim

Central claim: a decision procedure built on dialog-based argumentation can distribute real funds and be accepted by participants. In the experiment, students could submit proposals with a price tag, discuss them in D-BAS, and then approve and rank proposals in decide. Rankings were scored with a Borda-style count in which a voter's first preference gets $N$ points, the second gets $N-1$, and unranked proposals get $0$, with the number of approvals as tiebreaker; the winning set was built greedily by taking top-scoring proposals that fit the remaining budget. Eight final proposals went to a vote and five won; the participant survey returned means above 5 on the 7-point scale for most fairness and acceptance items, including agreement that the decision should be made this way in the future. The paper's own framing is that the exercise was 'well accepted and thus successful'.

Load-bearing premise

The acceptance conclusion rests on the assumption that survey answers from a self-selected group measure the quality of the procedure itself, rather than their happiness that five of eight proposals they liked won; the paper itself flags that high satisfaction may have been caused by the large number of winning proposals.

Editorial extensions

If this is right

  • A structured argumentation graph can feed a real resource-allocation vote, not just a discussion forum.
  • An unmonitored process is not enough: rule-breaking and vague proposals will occur, so an editorial review phase is needed before voting.
  • When the vote is separated from the argumentation, most participants choose to vote rather than argue; only 10 arguments were added during the voting phase.
  • In this case the winning set was robust to the scoring method: Borda, approval, and single-vote rankings agreed on the top proposals, though a Top-2 approval rule would have changed one winner.
  • Future runs should limit the number of winners if the goal is to test acceptance of the procedure independently of outcome.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the D-BAS-plus-decide pattern is a reusable template for small-group budget allocation wherever single-sign-on authentication exists; the same phases could serve departmental, neighbourhood, or campus funds.
  • Editorial inference: a natural next experiment is to vary the win rate deliberately to separate outcome effects from process effects, since the paper itself suspects high satisfaction came from many proposals winning.
  • Editorial inference: if argumentation participation stays lower than voting, an automated decision agent that infers preferences from argumentation alone would not yet be feasible; the paper itself notes this as a future scenario.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The thesis extends the Dialog-Based Argumentation System (D-BAS) with cost-bearing proposals and a companion voting service called "decide," then deploys both in a real participatory budgeting process at Heinrich-Heine-University Düsseldorf, where computer science students allocated €20,000 of quality-enhancement funds. The paper describes the argumentation graph, the preference aggregation scheme (a truncated Borda count with approval voting as tiebreaker), the three-phase experiment (proposals and argumentation, an unplanned review/correction phase, and voting), and the results: 142 voters chose five of eight proposals, spending €18,650 of the budget. A post-experiment survey (N between 22 and 47 per item) is used to support the central claim that "the procedure is well accepted and thus successful." The thesis also discusses future improvements and compares the approach with participatory budgeting in Porto Alegre, Wuppertal, and Reykjavík.

Significance. The work is a useful case study at the intersection of structured argumentation and participatory budgeting. Its strengths include a real deployment with actual funds, an open-source implementation, raw preference data in the appendix, and an unusually candid account of design changes and limitations. The paper also gives a clear, accessible description of the scoring method and compares it with several real-world participatory budgeting processes. However, the headline empirical claim—that the procedure is well accepted and therefore successful—rests on survey evidence that is confounded with outcome satisfaction and collected from a self-selected subsample after a favorable result. As a single-case pilot with mid-experiment redesign, the study can support a qualified claim about feasibility and reported satisfaction, but not the strong acceptance claim made in the abstract and conclusion. The central derivation of the decision procedure is sound, and the limitations are acknowledged in the future-work section, so the weaknesses are addressable through reframing and additional analysis rather than being irreparable.

major comments (3)
  1. [Abstract; Ch. 7; §5.6] The central claim that "the procedure is well accepted and thus successful" is not established by the survey data, because the survey was administered only after the winners were announced, when five of eight proposals won and 93% of the budget was spent. The respondents were self-selected (item Ns of 22–47 versus 142 voters, with no response rate reported), and the author explicitly concedes the confound in Chapter 7: "the satisfaction in the outcome is high because most of the proposals were able to be included in the winning set of proposals." The pattern in Figure 5.7 is consistent with this: decision-acceptance items (e.g., "I accept the decision," mean 6.51) are notably higher than procedural-justice items (means 5.13–5.76). The abstract and conclusion should be revised to claim, at most, that participants reported high satisfaction with the outcome and expressed desire for future procedures, unless the author can provide analysis that separates outcome favorability from procedure acceptance, such as responses from participants whose proposals lost or items asked before the outcome was known.
  2. [§5.3, §5.4, Ch. 7] The experiment was redesigned mid-course in ways that undermine the claim that the tested procedure is the procedure described in the design chapters. Phase 2 was added after the first day when many proposals violated the rules, the proposal-submission window was closed early, and the voting phase was shortened from one week to five days. The author himself observes in Chapter 7 that the voting phase "actually became just a traditional vote." Thus the acceptance evidence pertains to an ad hoc, supervised process rather than to the proposed unsupervised D-BAS-plus-decide workflow. This should be treated as a fundamental limitation of the experiment as a test of the original design, and the paper should state that the experiment was a pilot that iterated on the procedure, not a confirmatory test of the designed process.
  3. [§5.6] The survey analysis needs a discussion of non-response and self-selection to assess the representativeness of the acceptance figures. The paper reports item Ns (22–47) and the total number of voters (142), but never reports the number of students invited to the survey, the response rate, or any comparison between survey respondents and the full voter population. Without this information, it is impossible to know whether the positive averages reflect the views of typical participants or of a self-selected subset, particularly those who were satisfied with the outcome. At a minimum, the response rate should be reported, and the limitations of the survey as a non-probability sample should be acknowledged directly in Section 5.6.
minor comments (6)
  1. [Abstract] The phrase "The results indicate that the procedure is well accepted and thus successful" conflates outcome satisfaction with procedural acceptance; consider rewording to "participants reported high satisfaction with the outcome and expressed interest in future procedures."
  2. [§4.4] The description of the aggregation method would be clearer if the truncated Borda parameter N were explicitly defined as the maximum number of preferences cast by any single participant, and if an example showed how unranked proposals receive the implicit score of 0; the example in §4.5 covers this well, so a forward reference there would help.
  3. [Table 5.1] The table lists both Borda and approval scores without a common scale, which makes the statement that the outcome would be the same under approval voting easier to verify if the two columns were plotted together or normalized; consider adding a short note clarifying that the ranking, not the absolute values, is the basis for the comparison.
  4. [Figure 5.7] The boxplot figure is difficult to parse because the extracted statistics table duplicates rows and the N annotations are garbled in the text; please ensure the figure itself clearly labels each item N and quartile values, and check the rendering of the German-language table.
  5. [§5.2] The paper reports 52 participants in the argumentation and 142 voters, and later notes that 18 of the 52 arguers did not vote. It would be useful to state explicitly that participation in argumentation and voting were measured on different bases (registered versus voted), since the text could be misread as a drop in total participation.
  6. [§5.8.1] The typo "as+ it could undermine" should be corrected, and the brief discussion of structured argumentation violations could be tightened by adding examples for each of the three violation types.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the acceptance claim rests on independent survey data and experimental records, not on the cited D-BAS results.

full rationale

The thesis's central empirical claim—that the procedure was well accepted and thus successful—is supported by the experiment's own survey responses, participation counts, and vote data, not by the D-BAS publications cited from the same laboratory. The scoring procedure is a transparent mix of approval and Borda scores described in the thesis itself, and the simulated alternative scoring methods in Section 5.5 are computed from the same raw preference data rather than being presented as independent predictions. The author's Future Work caveat that satisfaction with the outcome may be high because most proposals won is an acknowledged validity limitation of the acceptance inference, not a definitional reduction of the procedure's acceptance to its outcome; it does not make the survey evidence circular in the technical sense used here. Self-citations describe the prior D-BAS system and earlier field experiments, but the load-bearing evidence for the acceptance conclusion is the new experiment's own measurements. Therefore no circular step can be exhibited, and the appropriate finding is no significant circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The thesis introduces no new theoretical entities. The free parameters are procedural design choices rather than fitted values, and the axioms are assumptions about the representational power of D-BAS, the validity of self-reported satisfaction, the reliability of LDAP authentication, and the chosen aggregation rule. The central claim depends most heavily on the second axiom, which is directly challenged by the outcome-favorability confound the author himself notes in Chapter 7.

free parameters (4)
  • N for Borda scoring = 8 (maximum number of preferences submitted by any voter)
    Section 4.4: the Borda maximum score N is set to the highest number of preferences a participant has chosen, making the scoring scale data-dependent rather than fixed.
  • Minimum proposal cost = 100 EUR
    Experimenter-chosen threshold in Section 5.1 to exclude trivial proposals; affects which proposals enter the vote.
  • Maximum proposal cost = 20000 EUR
    Section 5.1: equal to the total budget, a procedural constraint rather than a fitted constant.
  • Tiebreaker order = approval score, then proposal creation order
    Section 4.4: chosen for transparency; it determines the winner in the artificial four-participant example but not the real outcome because Borda and approval rankings coincided (Section 5.5).
assumptions (4)
  • domain assumption D-BAS argumentation graph structure with issues, positions, statements, and attack/support relations adequately represents the reasoning exchanged.
    Sections 2.5 and 4 rely on D-BAS's representational adequacy to claim educational benefits; the experiment itself found users violated the statement format (Section 5.8.1).
  • domain assumption Participants' self-assessed satisfaction is a valid measure of procedural success.
    The central acceptance claim rests on the Likert survey (Section 5.6) without validation against behavioral or objective outcome measures.
  • domain assumption LDAP enrollment filtering correctly identifies eligible computer science students.
    Section 5.1 admits 'the trust in the correctness of this directory was not absolute', and voting eligibility depends on it.
  • standard math Voter preferences can be meaningfully aggregated by summing Borda scores with implicit zero for unranked proposals.
    Section 4.4 defines the scoring; unranked proposals receive 0, a truncation assumption the paper adopted for simplicity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Decision Making with Argumentation Graphs." pith.science (2026). https://pith.science/paper/77QLCP72

@misc{pith2026190803357,
  author       = {Pith},
  title        = {Pith review of: Decision Making with Argumentation Graphs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/77QLCP72}},
  note         = {Machine review of arXiv:1908.03357}
}
read the original abstract

This work is about making decisions by digital means. Funds should be distributed by the students of Heinrich-Heine-University. The proposals were made by the students themselves without further influence. For this purpose, dialog-based argumentation is used to give the participants a better understanding of various arguments. In addition, a software service has been developed which allows the students to express their preferences for various proposals. An experiment was carried out at the university, which should prove whether students are satisfied with this type of participation. The results indicate that the procedure is well accepted and thus successful. However, improvements to the process itself were necessary during the experiment and should be considered for future procedures. Further procedures are desired.

Figures

Figures reproduced from arXiv: 1908.03357 by the authors.

Figure 2.1
Figure 2.1. Exploding complexity with a growing number of users, as each user has to argue with n−1 other users [PITH_FULL_IMAGE:figures/full_fig_p012_2_1.png] view at source ↗
Figure 2.3
Figure 2.3. Horizontal scaling of argumentation with via an automated intermediary. Each user just has to argue with the system. of the user interface and are therefore not considered in the rules of the dialog game execution platform. The dialog game execution platform is the logic behind D-BAS, that chooses on how to proceed in an argumentation by consulting the given argumentation graph and the user’s current location in it.… view at source ↗
Figure 2.4
Figure 2.4. Example of an argumentation graph. I denotes an issue, P a position, A-D are arguments. The targets of the edges are called conclusion, the sources premise. A supports P. B rebuts A on P, by attacking the conclusion of the supported P. C undermines A, by attacking its premise. D undercuts C, by directly attacking C. 10 [PITH_FULL_IMAGE:figures/full_fig_p017_2_4.png] view at source ↗
Figures from the paper (10 more)
Figure 3.1
Figure 3.1. Figure 3.1: The ballots for the voting methods. Left: Approval Voting, Right: Borda Count Furthermore, the voting procedure is relatively simple, so that no deep knowledge is required from the voters and the procedure takes place with regular ballots. Like stated in the section …
Figure 4.1
Figure 4.1. Figure 4.1: The voting interface. Participants could approve proposals from below. After￾wards these proposals could be ranked to express priority. (This figure is translated from German). 22 [PITH_FULL_IMAGE:figures/full_fig_p029_4_1.png]
Figure 5
Figure 5. Figure 5: shows the share each user has in the total number of arguments. Even more than in [PITH_FULL_IMAGE:figures/full_fig_p040_5.png]
Figure 5.1
Figure 5.1. Figure 5.1: This is the timeline of the procedure. The participation process started on the 23rd of April 2019 at 12:36 and ended on the 13th of May 2019 at 12:00. Each vertical line represents the moment an e-mail was sent to the participants to in￾form them about the ongoing p…
Figure 5.2
Figure 5.2. Figure 5.2: Proportion of participants of the total number of arguments. The full circle rep￾resents 197 arguments. The Need to have a Review of the First Phase After the first day, it was found that participants did not read or straight ignored the rules the proposals had to ad…
Figure 5.3
Figure 5.3. Figure 5.3: The final argumentation graph, after the filtering of phase 2. The grey point is the issue, blue points are proposals (positions), yellow points are statements. Green and red arrows represent arguments pro or contra the argument, they are pointing at. 36 [PITH_FULL_…
Figure 5.4
Figure 5.4. Figure 5.4: Proportion of participants in all levels of participation. Proposed has to be a subset of Argued. Voted: 142, Argued: 52, Proposed: 19, Argued & Voted: 39, Proposed & Voted: 14 5.5. Results The results were published twenty days after the start of the experiment. The…
Figure 5.5
Figure 5.5. Figure 5.5: Score distribution of the proposals. The two other scores chosen are: Single-Vote that simulates the case where every participant can just vote for one proposal. The proposal gets a score of 1. This system is included because of the frequency this voting system gets …
Figure 5.6
Figure 5.6. Figure 5.6: Priority Distribution ID Costs 1st 2nd 3rd 4th 5th Borda Approval Single Top 3 790 1000,00 € 40 27 7 3 1 336 78 40 74 821 20000,00 € 17 0 0 0 0 85 17 17 17 823 12000,00 € 25 16 7 7 3 227 58 25 48 774 150,00 € 14 11 18 8 3 187 54 14 43 746 4000,00 € 12 13 10 10 1 163 …
Figure 5.7
Figure 5.7. Figure 5.7: Questions and their results. 1 is equal to not correct at all and 7 is In any case, correct. The boxes are showing the range of the median 50% of scores, while the whiskers are showing the minimum and maximum scores. The thick bar in the diagram denotes the median. T…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

22 extracted references · 20 canonical work pages

  1. [1]

    (2018), June

    Abers , Rebecca; Brandão , Igor; King , Robin; Votto , Daniely: Porto Alegre: Participatory Budgeting and the Challenge of Sustaining Transformative Change. (2018), June

  2. [2]

    https://web.archive.org/web/20140809032935/http://eutopiamagazine.eu/en/r \,Version:\,07 2014

    Bjarnason , Róbert: 'Your Priorities': An icelandic story of e-democracy. https://web.archive.org/web/20140809032935/http://eutopiamagazine.eu/en/r \,Version:\,07 2014

  3. [3]

    https://oecd-opsi.org/wp-content/uploads/2019/04/Citizens-Foundation-Citizen-participation-and-digital-tools-V4.pdf

    Bjarnason , Róbert: Citizen participation and digital tools for upgrading democracy in Iceland and beyond. https://oecd-opsi.org/wp-content/uploads/2019/04/Citizens-Foundation-Citizen-participation-and-digital-tools-V4.pdf. \,Version:\,April 2019

  4. [4]

    55--66 S

    Black , Duncan: Which Candidate Ought to be Elected? Springer Netherlands, 1987. 55--66 S

  5. [5]

    http://siteresources.worldbank.org/INTEMPOWERMENT/Resources/14657_Partic-Budg-Brazil-web.pdf

    Bhatnagar , Deepti; Rathore , Animesh; Moreno Torres , Magüi; Kanungo , Parameeta: Participatory Budgeting in Brazil. http://siteresources.worldbank.org/INTEMPOWERMENT/Resources/14657_Partic-Budg-Brazil-web.pdf

  6. [6]

    https://citizens.is/portfolio_page/my-neighbourhood/

    Citizens Foundation : My Neighbourhood. https://citizens.is/portfolio_page/my-neighbourhood/. \,Version:\,December 2011

  7. [7]

    Dummett , Michael: Voting procedures. (1984)

  8. [8]

    Oxford University Press, 1997

    Dummett , Michael A.: Principles of electoral reform. Oxford University Press, 1997

Show all 22 references
  1. [9]

    In: Social Choice and Welfare 40 (2013), Feb, Nr

    Emerson , Peter: The original Borda count and partial voting. In: Social Choice and Welfare 40 (2013), Feb, Nr. 2, 353--358. http://dx.doi.org/10.1007/s00355-011-0603-9. DOI 10.1007/s00355--011--0603--9. ISSN 1432--217X

  2. [10]

    In: Australian Journal of Political Science 49 (2014), 04

    Fraenkel , Jon; Grofman , Bernard: The Borda Count and its real-world alternatives: Comparing scoring rules in Nauru and Slovenia. In: Australian Journal of Political Science 49 (2014), 04. http://dx.doi.org/10.1080/10361146.2014.900530. DOI 10.1080/10361146.2014.900530

  3. [11]

    In: Computational Models of Argument, Potsdam, 2016, S

    Krauthoff , Tobias; Betz , Gregor; Baurmann , Michael; Mauve , Martin: Dialog-Based Online Argumentation . In: Computational Models of Argument, Potsdam, 2016, S. 33--40

  4. [12]

    In: International Journal of Electronic Government Research 13 (2017), 04, S

    Kaur Kapoor , Kawaljeet; Omar , Amizan; Sivarajah , Uthayasankar: Enabling Multichannel Participation Through ICT Adaptation. In: International Journal of Electronic Government Research 13 (2017), 04, S. 66--80. http://dx.doi.org/10.4018/IJEGR.2017040104. DOI 10.4018/IJEGR.2017040104

  5. [13]

    In: Computational Models of Argument: Proceedings of COMMA 2018 305 (2018), S

    KRAUTHOFF , Tobias; METER , Christian; BAURMANN , Michael; BETZ , Gregor; MAUVE , Martin: D-BAS-A Dialog-Based Online Argumentation System. In: Computational Models of Argument: Proceedings of COMMA 2018 305 (2018), S. 325

  6. [14]

    In: Proceedings of the 1st Workshop on Advances in Argumentation in Artificial Intelligence, Bari, 2017, S

    Krauthoff , Tobias; Meter , Christian; Mauve , Martin: Dialog-Based Online Argumentation: Findings from a Field Experiment . In: Proceedings of the 1st Workshop on Advances in Argumentation in Artificial Intelligence, Bari, 2017, S. 85--99

  7. [15]

    http://nbn-resolving.de/urn/resolver.pl?urn=urn:nbn:de:hbz:061-20180628-094005-1

    Krauthoff , Tobias: Dialog-Based Online Argumentation, Heinrich-Heine-University, dissertation, June 2018. http://nbn-resolving.de/urn/resolver.pl?urn=urn:nbn:de:hbz:061-20180628-094005-1. URN urn:nbn:de:hbz:061--20180628--094005--1

  8. [16]

    Springer Science & Business Media, 2012

    Lueg , Christopher; Fisher , Danyel: From Usenet to CoWebs: interacting with social information spaces. Springer Science & Business Media, 2012

  9. [17]

    In: Computational Models of Argument, Warsaw, 2018, S

    Meter , Christian; Schneider , Alexander; Mauve , Martin: EDEN: Extensible Discussion Entity Network . In: Computational Models of Argument, Warsaw, 2018, S. 257--268

  10. [18]

    http://dip21.bundestag.de/dip21/btd/18/116/1811614.pdf

    Entwurf eines Ersten Gesetzes zur Änderung des E-Government-Gesetzes. http://dip21.bundestag.de/dip21/btd/18/116/1811614.pdf. \,Version:\,March 2017. Drucksache 18/11614

  11. [19]

    In: International Political Science Review 23 (2002), Nr

    Reilly , Benjamin: Social choice in the south seas: Electoral innovation and the borda count in the pacific island countries. In: International Political Science Review 23 (2002), Nr. 4, S. 355--372

  12. [20]

    (2018), 02, 21--66

    Ruesch , Michelle: D3.2 Pilots implementation - final. (2018), 02, 21--66. https://empatia-project.eu/wp-content/uploads/2018/07/EMPATIA_Deliverable_3.2_13FINAL.pdf

  13. [21]

    (2000), October

    Wampler , Brian: A Guide to Participatory Budgeting. (2000), October. https://www.partizipation.at/fileadmin/media_data/Downloads/themen/A_guide_to_PB.pdf

  14. [22]

    write newline

    " write newline "" before.all 'output.state := FUNCTION fin.entry write newline FUNCTION set.period.dash output.state before.all = 'skip period.dash 'output.state := if FUNCTION set.period output.state before.all = 'skip period.dash 'output.state := if FUNCTION set.period.dash...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.