{"id":"574e9867-609a-498a-ab7f-cd6cd6f7479c","arxiv_id":"1908.03357","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An online participatory budgeting process combining D-BAS argumentation with approval-plus-Borda voting was accepted by its participants in a real university funding decision.","lead":"This master's thesis reports a participatory budgeting experiment at a German university, where students discussed spending proposals in an argumentation graph and then voted with approval and Borda scores. It finds acceptable satisfaction: 142 students voted, five of eight proposals were funded, and survey ratings were mostly positive.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Outcome satisfaction cannot be separated from procedure acceptance in the evidence for the headline claim.","rationale":"The reader's weakest assumption is the same one I would flag, so this pass does not move the verdict. The thesis has independent strengths: open-source code, raw preference data, a real budget, and a field survey; the abstract is appropriately hedged with 'indicate.' The load-bearing defect is internal to the acceptance inference: the survey was delivered after the outcome was known, and the author himself identifies the favorable outcome as the likely driver of satisfaction. The highest-scoring items are outcome and decision-acceptance items; the procedural-justice items are lower, which is the pattern the confound predicts. The unplanned third phase and the author's admission that the vote became 'just a traditional vote' also mean that what was accepted is a patched single instance, not the originally specified procedure. A mediation analysis on the individual-level survey data is the minimal check that could separate procedure acceptance from outcome satisfaction; the aggregate statistics in Fig. 5.7 cannot do so. If the requested data are unavailable, a pre-registered replication with a scarce budget would be the appropriate follow-up. Because the paper's own Future Work section already concedes the core limitation, the CONDITIONAL verdict stands; no adjustment is warranted.","tokens_in":33627,"tokens_out":8108,"duration_ms":93228,"concrete_test":"Obtain the individual-level survey responses from the sociological partners (the thesis states full results will be published, §5.6) and fit a mediation model for 'I accept the decision' with two composites: procedural justice (rules applied consistently, impartial, informed) and distributive justice/outcome (fair balance, no unreasonable disadvantage, broad agreement, needs met). Test whether the procedural-justice path remains significant after controlling for the distributive-justice composite. If the outcome composite fully mediates acceptance, the headline 'procedure well accepted' is unsupported; if procedural justice remains predictive, the concern is refuted. The aggregate boxplots in Fig. 5.7 cannot settle this because they are marginal distributions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim—'the procedure is well accepted and thus successful' (Abstract; Ch. 7)—depends on treating post-outcome survey scores as evidence about the procedure. That inference is the least secure step. The survey was administered only after the winners were announced, when five of eight proposals had won and 93% of the budget was spent; respondents were self-selected (item Ns 22–47 versus 142 voters, no response rate reported in §5.6); and the realized procedure included an unplanned correction phase (§5.3) and a shortened voting phase (§5.4) that, in the author's own words, made the vote 'just a traditional vote' (Ch. 7). More importantly, the author concedes the confound in Future Work: 'the satisfaction in the outcome is high because most of the proposals were able to be included in the winning set of proposals.' Consistent with that, the outcome/decision-acceptance items ('I accept the decision': mean 6.51, median 7.0) are higher than the procedural-justice items (means 5.13–5.76 in Fig. 5.7). If these scores track the favorable result rather than the decision process, the central claim is not established. This is not a generalization concern; it is a question of what this experiment's own data can support.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The thesis extends the Dialog-Based Argumentation System (D-BAS) with cost-bearing proposals and a companion voting service called \"decide,\" then deploys both in a real participatory budgeting process at Heinrich-Heine-University Düsseldorf, where computer science students allocated €20,000 of quality-enhancement funds. The paper describes the argumentation graph, the preference aggregation scheme (a truncated Borda count with approval voting as tiebreaker), the three-phase experiment (proposals and argumentation, an unplanned review/correction phase, and voting), and the results: 142 voters chose five of eight proposals, spending €18,650 of the budget. A post-experiment survey (N between 22 and 47 per item) is used to support the central claim that \"the procedure is well accepted and thus successful.\" The thesis also discusses future improvements and compares the approach with participatory budgeting in Porto Alegre, Wuppertal, and Reykjavík.","tokens_in":33815,"tokens_out":3368,"duration_ms":37994,"significance":"The work is a useful case study at the intersection of structured argumentation and participatory budgeting. Its strengths include a real deployment with actual funds, an open-source implementation, raw preference data in the appendix, and an unusually candid account of design changes and limitations. The paper also gives a clear, accessible description of the scoring method and compares it with several real-world participatory budgeting processes. However, the headline empirical claim—that the procedure is well accepted and therefore successful—rests on survey evidence that is confounded with outcome satisfaction and collected from a self-selected subsample after a favorable result. As a single-case pilot with mid-experiment redesign, the study can support a qualified claim about feasibility and reported satisfaction, but not the strong acceptance claim made in the abstract and conclusion. The central derivation of the decision procedure is sound, and the limitations are acknowledged in the future-work section, so the weaknesses are addressable through reframing and additional analysis rather than being irreparable.","major_comments":[{"comment":"The central claim that \"the procedure is well accepted and thus successful\" is not established by the survey data, because the survey was administered only after the winners were announced, when five of eight proposals won and 93% of the budget was spent. The respondents were self-selected (item Ns of 22–47 versus 142 voters, with no response rate reported), and the author explicitly concedes the confound in Chapter 7: \"the satisfaction in the outcome is high because most of the proposals were able to be included in the winning set of proposals.\" The pattern in Figure 5.7 is consistent with this: decision-acceptance items (e.g., \"I accept the decision,\" mean 6.51) are notably higher than procedural-justice items (means 5.13–5.76). The abstract and conclusion should be revised to claim, at most, that participants reported high satisfaction with the outcome and expressed desire for future procedures, unless the author can provide analysis that separates outcome favorability from procedure acceptance, such as responses from participants whose proposals lost or items asked before the outcome was known.","section":"Abstract; Ch. 7; §5.6"},{"comment":"The experiment was redesigned mid-course in ways that undermine the claim that the tested procedure is the procedure described in the design chapters. Phase 2 was added after the first day when many proposals violated the rules, the proposal-submission window was closed early, and the voting phase was shortened from one week to five days. The author himself observes in Chapter 7 that the voting phase \"actually became just a traditional vote.\" Thus the acceptance evidence pertains to an ad hoc, supervised process rather than to the proposed unsupervised D-BAS-plus-decide workflow. This should be treated as a fundamental limitation of the experiment as a test of the original design, and the paper should state that the experiment was a pilot that iterated on the procedure, not a confirmatory test of the designed process.","section":"§5.3, §5.4, Ch. 7"},{"comment":"The survey analysis needs a discussion of non-response and self-selection to assess the representativeness of the acceptance figures. The paper reports item Ns (22–47) and the total number of voters (142), but never reports the number of students invited to the survey, the response rate, or any comparison between survey respondents and the full voter population. Without this information, it is impossible to know whether the positive averages reflect the views of typical participants or of a self-selected subset, particularly those who were satisfied with the outcome. At a minimum, the response rate should be reported, and the limitations of the survey as a non-probability sample should be acknowledged directly in Section 5.6.","section":"§5.6"}],"minor_comments":[{"comment":"The phrase \"The results indicate that the procedure is well accepted and thus successful\" conflates outcome satisfaction with procedural acceptance; consider rewording to \"participants reported high satisfaction with the outcome and expressed interest in future procedures.\"","section":"Abstract"},{"comment":"The description of the aggregation method would be clearer if the truncated Borda parameter N were explicitly defined as the maximum number of preferences cast by any single participant, and if an example showed how unranked proposals receive the implicit score of 0; the example in §4.5 covers this well, so a forward reference there would help.","section":"§4.4"},{"comment":"The table lists both Borda and approval scores without a common scale, which makes the statement that the outcome would be the same under approval voting easier to verify if the two columns were plotted together or normalized; consider adding a short note clarifying that the ranking, not the absolute values, is the basis for the comparison.","section":"Table 5.1"},{"comment":"The boxplot figure is difficult to parse because the extracted statistics table duplicates rows and the N annotations are garbled in the text; please ensure the figure itself clearly labels each item N and quartile values, and check the rendering of the German-language table.","section":"Figure 5.7"},{"comment":"The paper reports 52 participants in the argumentation and 142 voters, and later notes that 18 of the 52 arguers did not vote. It would be useful to state explicitly that participation in argumentation and voting were measured on different bases (registered versus voted), since the text could be misread as a drop in total participation.","section":"§5.2"},{"comment":"The typo \"as+ it could undermine\" should be corrected, and the brief discussion of structured argumentation violations could be tightened by adding examples for each of the three violation types.","section":"§5.8.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a Master's thesis and reads as an honest, detailed case report rather than a full scientific study. The main empirical claim is overstated relative to the evidence, but the underlying system is described clearly and the raw data are provided. I recommend major revision with a request to reframe the acceptance claim, add a threats-to-validity section covering the mid-experiment redesign and self-selected survey, and report response rates. The paper could then be a valuable practical contribution to the participatory budgeting and argumentation literature. I saw no evidence of problematic citation practices; the self-citations to D-BAS are appropriate given the system under study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a Master's thesis that actually ran a participatory budgeting process with real money (€20,000), real students, and real consequences. The components are not new—D-BAS argumentation, Borda and approval scoring, participatory budgeting—but the integration and the deployment are. The author also ships the software (open source on GitHub), includes raw preference data in the appendix, and reports a survey run with a sociology department. That is more than most case studies in this area provide, and it deserves credit. The writing is also unusually honest: the author explicitly says the voting phase became \"just a traditional vote\" after an unplanned review phase, and the Future Work section concedes that satisfaction with the outcome is probably high because five of eight proposals won. These admissions are not buried; they are right there in the text.\n\nThe central claim—\"the procedure is well accepted and thus successful\"—is only partially supported. The survey was administered after the winners were announced, with self-selected respondents (item Ns of 22–47 versus 142 voters, no response rate given). The decision-acceptance items are higher (median 7) than the procedural-justice items (medians 5–6), which is exactly the pattern you would expect if the scores track a favorable outcome rather than the procedure itself. The author's own future-work statement acknowledges that confound, but the abstract and conclusion still lean on it. A second, smaller issue: the paper notes that approval-only ranking would have produced the same winning set, so the combined Borda-plus-approval scoring is not actually shown to matter in this case. That is not a fatal flaw, but it weakens the claimed contribution of the scoring method.\n\nI am not worried about the citation pattern or the dependence on prior D-BAS work; the empirical claim is grounded in this experiment's own data, and the D-BAS background is appropriately cited. The paper is a single case with no baseline and a mid-process redesign, so the generalizability is minimal. But it is a useful, well-documented case study of a real e-participation deployment, and the author has been clear about what can and cannot be concluded.\n\nA serious referee should see this. It will benefit from a revision that rephrases the success claim as \"participants were satisfied under favorable conditions\" and that engages with the outcome confound more directly. For someone working on participatory budgeting or argumentation-based decision support, this is worth reading and possibly citing for its empirical data; I would not cite it for the scoring method.\n\nRecommendation: send it to peer review as a case study, not as a generalizable result.","headline":"A real participatory budgeting case study with open data and honest self-criticism, whose headline acceptance claim is weakened by a familiar outcome-satisfaction confound.","tokens_in":34376,"tokens_out":1402,"would_cite":false,"duration_ms":18014,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A structured online debate and voting system allocated a real 20,000-euro budget with broad acceptance in the reported experiment.","keywords":["participatory budgeting","argumentation graph","D-BAS","online deliberation","Borda count","approval voting","e-participation","case study"],"falsifier":"Repeat the same procedure with a budget that can fund only one or two of many competing proposals, and check whether fairness and acceptance ratings stay high among participants whose proposals lose. If acceptance drops sharply when most people lose, the original satisfaction was mostly about outcome, not about the argumentation-and-voting procedure.","tokens_in":33370,"feed_emoji":"🗳️","tokens_out":7243,"duration_ms":75846,"temperature":0.7,"pith_summary":"This paper tries to show that an online participatory budgeting process can work: a community of students decides how to spend a real 20,000-euro fund by proposing projects, arguing for and against them in a structured argumentation graph, then voting on the final proposals through a simple ranking interface. The reported experiment drew 142 voters and produced five winning proposals that together used 18,650 of the 20,000 euros. Survey responses, mostly above 5 on a 1-to-7 scale, led the author to conclude that the procedure was well accepted and therefore successful. The thesis also acknowledges that the process had to be changed mid-course: rule-violating or underspecified proposals forced an unplanned review phase, so the final vote was preceded by editorial filtering rather than continued discussion.","feed_headline":"Structured online debate and voting split a 20,000-euro fund","feed_subtitle":"A real participatory-budgeting trial found broad acceptance, yet most engagement came from voting, not debate.","key_machinery":"The load-bearing object is the argumentation graph of D-BAS, the Dialog-Based Argumentation System: a directed acyclic graph whose nodes are the issue, positions (proposals), and atomic statements, and whose directed edges are arguments marked as support or attack; it lets a participant argue with the system rather than with every other participant, preserving structure as the discussion grows. The decision machinery is a deliberately transparent scoring rule: proposals carry a fixed cost, voters approve and rank them, first preference earns $N$ points, second earns $N-1$, down to $0$ for unranked, and approval counts break ties. Winners are then chosen greedily by score while staying inside the budget. This combination is what the paper claims makes the result understandable and acceptable to untrained participants.","core_discovery":"Central claim: a decision procedure built on dialog-based argumentation can distribute real funds and be accepted by participants. In the experiment, students could submit proposals with a price tag, discuss them in D-BAS, and then approve and rank proposals in decide. Rankings were scored with a Borda-style count in which a voter's first preference gets $N$ points, the second gets $N-1$, and unranked proposals get $0$, with the number of approvals as tiebreaker; the winning set was built greedily by taking top-scoring proposals that fit the remaining budget. Eight final proposals went to a vote and five won; the participant survey returned means above 5 on the 7-point scale for most fairness and acceptance items, including agreement that the decision should be made this way in the future. The paper's own framing is that the exercise was 'well accepted and thus successful'.","pith_inferences":["Editorial inference: the D-BAS-plus-decide pattern is a reusable template for small-group budget allocation wherever single-sign-on authentication exists; the same phases could serve departmental, neighbourhood, or campus funds.","Editorial inference: a natural next experiment is to vary the win rate deliberately to separate outcome effects from process effects, since the paper itself suspects high satisfaction came from many proposals winning.","Editorial inference: if argumentation participation stays lower than voting, an automated decision agent that infers preferences from argumentation alone would not yet be feasible; the paper itself notes this as a future scenario."],"forward_implications":["A structured argumentation graph can feed a real resource-allocation vote, not just a discussion forum.","An unmonitored process is not enough: rule-breaking and vague proposals will occur, so an editorial review phase is needed before voting.","When the vote is separated from the argumentation, most participants choose to vote rather than argue; only 10 arguments were added during the voting phase.","In this case the winning set was robust to the scoring method: Borda, approval, and single-vote rankings agreed on the top proposals, though a Top-2 approval rule would have changed one winner.","Future runs should limit the number of winners if the goal is to test acceptance of the procedure independently of outcome."],"supporting_citations":[{"why":"It introduces D-BAS and the dialog-based argumentation graph that this work extends with costs and voting.","marker":"[KBBM16]"},{"why":"It documents the D-BAS system that serves as the argumentation and authentication backend for the experiment.","marker":"[KMB+18]"},{"why":"It reports the earlier field experiment with D-BAS whose participation patterns and attack-heavy argumentation inform the design and expectations.","marker":"[KMM17]"},{"why":"It provides the D-BAS codebase and doctoral thesis from which the decide extension and experiment are built.","marker":"[Kra18]"},{"why":"It defines Borda count, the scoring rule the thesis adapts into truncated preference scoring with an approval tiebreaker.","marker":"[Bla87]"},{"why":"It documents a comparable city-level participatory budgeting pilot, used as a benchmark and source of design differences.","marker":"[Rue18]"},{"why":"It describes the Reykjavík neighbourhood participatory budgeting process, compared for its continuous voting and externally set costs.","marker":"[Cit11]"}],"fun_headline_variants":["Argumentation graphs split student fund","Debate plus voting allocates €20k","Borda counting decides student proposals","Students accept dialog-based fund tool","Voting dominates debate in fund trial"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The acceptance conclusion rests on the assumption that survey answers from a self-selected group measure the quality of the procedure itself, rather than their happiness that five of eight proposals they liked won; the paper itself flags that high satisfaction may have been caused by the large number of winning proposals.","fun_headline_variants_meta":{"raw":{"variants":["Argumentation graphs split student fund","Debate plus voting allocates €20k","Borda counting decides student proposals","Students accept dialog-based fund tool","Voting dominates debate in fund trial"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000238,"raw_usage":{"total_tokens":1450,"prompt_tokens":826,"completion_tokens":624,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":442,"completion_tokens_details":{"reasoning_tokens":564}},"tokens_in":442,"tokens_out":624,"duration_ms":7232,"temperature":1.0,"reasoning_tokens":564,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:15:04.786216+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the same procedure with a budget that can fund only one or two of many competing proposals, and check whether fairness and acceptance ratings stay high among participants whose proposals lose. If acceptance drops sharply when most people lose, the original satisfaction was mostly about outcome, not about the argumentation-and-voting procedure.","supporting_citations":[],"review_version":1}