{"id":"50dd32cf-fd34-463e-81ee-9a7a68d5009f","arxiv_id":"2502.05017","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The paper introduces and field-tests three methods that use voting data to structure deliberation groups, let participants adjust an algorithmic budget allocation, and map opinion shifts before and after discussion.","lead":"A field study reports three algorithm-assisted designs that combine online voting with face-to-face deliberation, tested in a Swiss cultural funding process and a vTaiwan AI-policy workshop. The results suggest such hybrids are practical and acceptable to participants, though the evidence is observational and small-scale.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PCD relies on radial clustering in 2D PCA space with silhouette 0.238; if these clusters do not reflect preference similarity, every PCD contrast loses construct validity.","rationale":"The reader's verdict (CONDITIONAL) is well-calibrated: the paper is a transparent field report of three algorithmic methods, with honest limitations and public data. My stress test focused on the single most load-bearing assumption: the PCD framework's independent variable—homogeneous vs heterogeneous group composition—must be constructed validly for any PCD comparison to be interpretable. The radial clustering implementation has two concrete weaknesses: it uses only angular position, ignoring preference magnitude, and its silhouette score (0.238) is in the 'weak structure' range. Section 6.2's admission that PCA flattens preference complexity compounds this. If the clusters are not preference-homogeneous, then the headline findings (Section 5.1.1 and 5.1.3) cannot support the paper's central claim about preference-based group formation. I agree with the reader's weakest_assumption on this point. I considered two other candidate concerns: (1) the fixed round order (homogeneous always first) confounds all PCD comparisons, which is serious but cannot be resolved post hoc without additional data; and (2) an internal inconsistency in Table 2—the text claims all three statements with BC ≥ 0.555 fell below threshold after deliberation, but Statement 4's after-BC is 0.619, still above 0.555. This is a factual error in the ReadTheRoom section, but it is localized and does not undermine the universal BC decreases. The clustering validity, by contrast, determines whether the PCD findings have any causal interpretation. The proposed test—computing within- vs between-group preference similarity on the raw approval vectors—would directly settle whether the radial sectors are preference-homogeneous. The order confound should be acknowledged as a limiting factor in any future replication, but the clustering test is the more fundamental check. Since this concern does not move the overall verdict (the methods remain worthy of conditional acceptance and the limitations are disclosed), I keep the reader's CONDITIONAL verdict unchanged.","tokens_in":19170,"tokens_out":8648,"duration_ms":92503,"concrete_test":"Using the KK24 public dataset, compute the mean pairwise Jaccard (or cosine) similarity of approval vectors within each of the six Radial Clustering groups and between groups, under a permutation null (randomly permuting group labels and recomputing the within-minus-between difference). If the observed difference is not significant, the 'homogeneous' groups are not truly preference-similar. As a secondary check, re-estimate the main outcomes (round-level voting-deliberation correlation, ease/alignment ratings) using groups formed by high-dimensional balanced K-Means (silhouette 0.429 in Table 3) instead of Radial Clustering; if the substantive PCD effects disappear, they are artifacts of the angular partitioning rather than of preference homogeneity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that Radial Clustering on the two-dimensional PCA projection (Section 3.1.3) creates genuinely homogeneous preference groups. The algorithm assigns participants to sectors using only their angular position θ = arctan((PC2_i − mean)/(PC1_i − mean)), discarding the radial distance from the center. Consequently, a participant near the origin with weak or mixed preferences and a participant far out on the same ray with strong, focused preferences fall into the same 'homogeneous' group, while two participants with nearly identical preference vectors but different overall intensity can be split across sectors. The reported silhouette score of 0.238 (Table 3) indicates weak cluster structure, and Section 6.2 concedes that the projection 'inevitably flattens the complexity of participant preferences.' If the groups are not actually preference-homogeneous, then every PCD result—the significant voting-deliberation correlation in heterogeneous rounds (r=0.678, p=0.000527), the ease/alignment self-reports (83%/76% vs 65%/55%), and the project-cost differences in Table 1—cannot be attributed to preference homogeneity; the contrasts could instead reflect arbitrary group composition, facilitator behaviour, or the fixed ordering in which homogeneous deliberation always preceded heterogeneous deliberation. This threatens the paper's first contribution, which is its most novel and extensively evaluated component.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents three algorithmic methods to bridge online voting and face-to-face deliberation: Preference-based Clustering for Deliberation (PCD), Human-in-the-loop Method of Equal Shares (MES), and the ReadTheRoom deliberation method. These are evaluated in two real-world case studies: the Kultur Komitee 2024 (KK24) budgeting assembly (N=35) and a vTaiwan AI-regulation workshop (N=44). The reported results include: heterogeneous deliberation outcomes correlating with pre-deliberation voting (r=0.678, p=0.000527), homogeneous deliberation being rated easier and more preference-aligned (83%/76% vs 65%/55%), MES distributing a budget more fairly than a Greedy baseline, and ReadTheRoom deliberation decreasing the Bimodality Coefficient across all five debated statements. The authors provide public datasets and code and candidly discuss limitations in Section 6.2.","tokens_in":19388,"tokens_out":5525,"duration_ms":53590,"significance":"If the causal claims held, the paper would offer practical recipes for integrating voting and deliberation, with implications for participatory budgeting and deliberative mini-publics. The strongest assets are the genuine field deployment, the public datasets/code, and the explicit acknowledgment of threats to validity in Section 6.2. However, the current evidence is predominantly descriptive or correlational; the causal language in Sections 5 and 6.1 goes beyond what the study design can support. The contribution is best framed as context-rich design insights rather than a validated test of the methods' effects.","major_comments":[{"comment":"The homogeneous-versus-heterogeneous contrast is the load-bearing evidence for PCD, but the manuscript does not establish that Radial Clustering actually produces preference-homogeneous groups. The reported silhouette score of 0.238 indicates weak cluster structure, and the assignment rule uses only the angular coordinate theta = arctan((PC2_i - mean)/(PC1_i - mean)), discarding radial distance. Consequently, a participant with weak, mixed preferences near the center can be grouped with a participant on the same ray with strong, focused preferences, while two participants with nearly identical preference profiles but different intensity can be split across sectors. Section 6.2 concedes that the two-dimensional PCA projection 'inevitably flattens the complexity of participant preferences,' yet no validation is provided that the resulting groups are similar in the original high-dimensional voting space. Without such validation, the alignment, ease, and cost differences reported in Sections 5.1.1-5.1.4 cannot be attributed to preference homogeneity.","section":"Section 3.1.3, Table 3, Figure 1"},{"comment":"The two deliberation rounds were not counterbalanced: homogeneous groups always met first and heterogeneous groups second (step 3 of Section 3.1.3). Any observed difference between the rounds—the voting-deliberation correlation (r=0.678 vs r=0.366), the ease and alignment ratings (83%/76% vs 65%/55%), or the project-cost patterns in Table 1—could be due to order, fatigue, facilitator behavior, or learning effects rather than group composition. There is no control condition or washout. The causal phrasing in Sections 5.1.1-5.1.4 and 6.1 ('suggests that using Radial clustering... creates a perceivable difference', 'heterogeneous deliberation funds more costly projects') should be tempered to associational claims, or the authors should provide additional evidence, such as a reversed-order implementation or a within-subject design with balanced order.","section":"Section 3.1.3, Section 5.1"},{"comment":"The paper reports a large number of significance tests without any multiple-comparison correction. In Figure 3, only one of the five demographic-group comparisons reaches p<0.05 (Age ≤33, p=0.047), and it would not survive a simple Bonferroni correction given the number of tests. In Table 2, only one of the five mean opinion changes is statistically significant (Statement 4), and the central polarisation-reduction claim rests on Bimodality Coefficient and Consensus Index changes for which no standard errors, confidence intervals, or significance tests are reported. With N=44 and five-point Likert items, the BC metric is highly sensitive to response distributions; the universal decrease in BC is suggestive but not by itself evidence of a robust effect. I request effect sizes and confidence intervals for the BC/CI changes and either explicit multiple-comparison control or an explicit statement that the analyses are exploratory.","section":"Figure 3, Table 2, Section 5.3.1"},{"comment":"The Human-in-the-loop MES evaluation does not support the causal claim that the method 'builds algorithmic trust' (Section 3.2.1). The MES-versus-Greedy comparison in Figure 8 is a retrospective computational baseline, not a field experiment, and the participant responses (62% 'Very fair', 81% supporting a 50:50 ratio) are single-group post-hoc ratings with no control or pre-measurement. The paper also notes in Section 6.2 that participants modified budgets after the MES calculation, potentially compromising the proportionality that the fairness claim is based on. Please distinguish clearly between (a) feasibility and participant acceptability, which the data support, and (b) causal effects on trust or perceived fairness, which the design cannot establish.","section":"Section 5.2, Figure 8"}],"minor_comments":[{"comment":"The roles of the authors are described inconsistently: Section 1.1.2 says 'The corresponding author offered the ReadTheRoom Deliberation method,' while Section 4 says 'the first author participated in the public deliberation workshops.' Please clarify which author held which role.","section":"Section 1.1.2"},{"comment":"The Bimodality Coefficient formula is written as 'BC = skewness2+1 / kurtosis+3'; the standard formula is (skewness^2 + 1) / (kurtosis + 3). Please add the parentheses and verify the reference to Knapp (2007).","section":"Section 3.3.4"},{"comment":"The sentence 'We also observed more groups using sticky notes as votes' references Figure 4, but the main text does not explain what the figure shows beyond that clause. Please expand the caption or add an explicit description in the text.","section":"Section 5.1.1, Figure 4"},{"comment":"The column headers 'SS', 'BG', and 'OC' are not expanded in the table itself; the caption gives the meaning, but it would be clearer to add them directly to the header (e.g., 'Silhouette Score (SS)', 'Balanced Groups (BG)', 'Overlapping Clusters (OC)').","section":"Table 3"},{"comment":"References [39] and [40] appear to be the same paper (one entry contains a typo in the author name 'Hnggli'). Please deduplicate and provide a single correct citation.","section":"References"},{"comment":"Section 6.2 states that post-selection budget modifications 'potentially compromised MES's mathematical proportionality guarantees,' yet Section 5.2.2 presents the same modification step as a positive feature. The tension is real and should be discussed in the results, not only in the limitations.","section":"Section 6.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a candid field report with useful public datasets, and the limitations section is unusually transparent. However, the gap between the claims and the evidence is substantial, particularly for PCD: the clustering validity (silhouette 0.238) and the ordering confound are unresolved load-bearing issues. The statistical analyses also need multiple-comparison control or a clear exploratory label. I believe the paper can be made publishable after major revision, but in its current form the central causal claims are not established. The editor may also consider whether the paper is better positioned as a design-and-deployment case study rather than as a demonstration of algorithmic effects."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real field report from two actual participatory processes, and the authors are transparent about their roles and limits. The three methods—PCD, Human-in-the-loop MES, ReadTheRoom—are not new algorithms in a strict sense; they are thoughtful recombinations of existing tools (PCA, k-means, MES, Polis, before/after voting). That is still a useful contribution: civic-tech practitioners get a concrete playbook. KK24 gives a real budget, real projects, real participants; vTaiwan gives an actual policy deliberation on AI regulation. They share data, acknowledge their dual roles, and flag their own limitations in Section 6.2. That honesty is earned.\n\nThe soft spot is the PCD evidence. Radial clustering in 2D PCA space uses only angular position and discards distance from the center, so preference intensity is ignored. The reported silhouette score (0.238) is weak, and Section 6.2 concedes the projection flattens preference complexity. That means the homogeneous-vs-heterogeneous contrasts—the r=0.678 correlation, the ease/alignment self-reports, the project cost differences—cannot be cleanly attributed to preference homogeneity. Homogeneous always came first, so ordering and facilitator effects are confounds. Samples are small (N=35 for KK24, 44 for vTaiwan), there is no control group, and many significance tests run without correction. ReadTheRoom has the cleanest before/after design, but even there only one mean change reaches p<0.05, and the BC decreases are consistent but on small n.\n\nNone of this is fatal. The paper does not overclaim; it explicitly presents results as preliminary. The real value is in the documented practice: the procedures are concrete, the data are public, and the authors candidly discuss what could be driving their observations. The right framing is \"here is what we did and what we saw,\" not \"we proved algorithms improve deliberation.\" A serious referee should engage because these methods are already being used in real democratic processes, and the design details matter for anyone replicating or adapting them.\n\nRecommendation: send it out. It deserves referee time, with the expectation of major revisions—especially around clustering validity, ordering effects, and multiple comparisons. I would not cite it for causal claims, but I would cite it as a documented field application.","headline":"Genuine field report with honest limitations, but the PCD evidence rests on a clustering step with weak construct validity.","tokens_in":19935,"tokens_out":1497,"would_cite":true,"duration_ms":17971,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that algorithms that bridge online voting with face-to-face deliberation work in real-world settings, producing measurable changes in how groups decide and how polarised they become.","keywords":["Participatory budgeting","Citizens' assemblies","Method of Equal Shares","Group deliberation","Democratic innovation","Opinion clustering","Social computing","Hybrid democracy"],"falsifier":"Compute silhouette scores for the Radial Clustering assignments in the original high-dimensional preference space rather than the two-dimensional projection; if the assignments show near-zero or negative silhouette values there, or if a placebo re-randomisation of participants into 'homogeneous' and 'heterogeneous' groups reproduces the same ease/alignment differences, the PCD effect would be indistinguishable from group-size or facilitation artefacts. Similarly, a delayed re-administered survey of vTaiwan participants weeks after the workshop would settle whether the reduced Bimodality Coefficients persist.","tokens_in":18930,"feed_emoji":"🗳️","tokens_out":5347,"duration_ms":46630,"temperature":0.7,"pith_summary":"The paper claims that algorithms can productively bridge large-scale online voting with face-to-face deliberation, and that doing so changes both process and outcomes in measurable ways. Three methods are tested in two real democratic settings: grouping deliberators by voting similarity (PCD), letting participants steer how much of a budget is decided by a fair-voting algorithm (Human-in-the-loop MES), and using opinion-space maps plus before/after voting to guide discussion (ReadTheRoom). Field results show heterogeneous deliberation tracks individual voting more closely, homogeneous deliberation feels easier and more preference-aligned, equal-shares budgeting is seen as fair while preserving a proportionality guarantee, and structured deliberation lowers polarisation on divisive statements. A sympathetic reader would take the paper as evidence that voting data can be reused, not just aggregated, to make participatory decision-making more representative and trusted.","feed_headline":"Voting data can steer deliberation groups — and it changes outcomes","feed_subtitle":"Field tests show fairer budgets, less polarisation, and higher trust when algorithms set the stage.","key_machinery":"The load-bearing machinery is threefold: (1) PCD's Radial Clustering, which projects participants' approval votes into a two-dimensional PCA opinion space and divides the angular space into balanced 'pizza-slice' sectors to create preference-homogeneous and later heterogeneous groups; (2) Human-in-the-loop MES, the Method of Equal Shares algorithm (each voter gets an equal budget share, projects are selected when affordable by supporters' capped shares) wrapped in an interactive interface where participants adjust the algorithmic budget share in real time and may subsequently adjust project budgets; (3) ReadTheRoom, which uses Polis's opinion mapping to identify divisive statements, builds a visual decision tree from opinion groups, and runs paired before/after Likert voting on those statements to quantify polarisation via the Bimodality Coefficient and consensus via 1/(1+SD). These three mechanisms carry the argument that voting data can structure deliberation, that algorithmic fairness can be made human-controllable, and that opinion shifts can be tracked and guided.","core_discovery":"The central claim is that voting and deliberation are complementary, and that computational methods can join them: PCD uses pre-deliberation votes to form balanced homogeneous and heterogeneous groups, producing different and measurable deliberation dynamics; Human-in-the-loop MES extends the Method of Equal Shares so participants decide how much of the budget the algorithm decides, preserving MES's proportionality while building trust; ReadTheRoom maps the online opinion space onto a decision tree and uses spectrum-based before/after voting to make opinion shifts visible and reduce polarisation. In KK24, heterogeneous group decisions closely mirrored individual online votes (Pearson r=0.678, p=0.000527) while homogeneous groups were rated easier by 83% of participants and more preference-aligned by 76%; in vTaiwan, Bimodality Coefficients fell below the 0.555 polarisation threshold for all three initially-divisive statements after deliberation. The paper argues these structured integrations make deliberation more inclusive of niche interests while keeping the breadth of voting.","pith_inferences":["A direct implication the authors leave implicit: PCD's homogeneous-first, heterogeneous-second sequence could be reordered or iterated, and the framework predicts that each reordering produces a different trade-off between niche representation and outcome alignment; this is testable with the same clustering pipeline.","The ReadTheRoom effect on polarisation is only measured immediately after a single session; an obvious extension is a follow-up survey weeks later to test whether the convergence is durable or a temporary conformity effect — a concern the authors themselves flag.","The human-in-the-loop interface could be extended beyond budget share to let participants steer which fairness criterion (e.g., Greedy vs MES) is used, connecting to the broader question of algorithmic legitimacy as a function of control, not just outcome.","One unresolved tension: if post-selection budget adjustments compromise MES's proportionality guarantee, as the authors concede, the method is more accurately described as a deliberation-supported heuristic than a strict fairness mechanism; formalising those adjustments as preference updates is a natural next step."],"forward_implications":["In processes like KK24, using voting data to form homogeneous-then-heterogeneous deliberation groups can surface niche projects that simple voting would miss, while later heterogeneous rounds keep outcomes broadly aligned with voter preferences.","Offering participants a visible, adjustable budget share for MES can produce near-unanimous endorsement of the voting-to-deliberation ratio and of the algorithm's fairness, without vetoing algorithmic selections.","Structured deliberations built on opinion-space maps can move divisive statements below the bimodality polarisation threshold and increase consensus indices in a single session, even with modest participation.","Fair voting methods like MES can fund more projects per voter and lower the Gini coefficient of budget allocation relative to the common Greedy method under the same budget.","These bridging methods transfer across settings: the same algorithmic ideas were implemented in a Swiss cultural budgeting assembly and a Taiwanese AI-regulation roundtable, suggesting generalizability to other participatory fora."],"supporting_citations":[{"why":"Peters, Pierczynski, and Skowron 2021 — supplies the Method of Equal Shares algorithm whose fairness and proportionality the Human-in-the-loop extension builds on.","marker":"[31]"},{"why":"Small et al. 2021 — supplies the Polis platform's opinion-space mapping and consensus/divisive statement detection used in ReadTheRoom.","marker":"[33]"},{"why":"Sunstein 2017 — provides the enclave-deliberation concept that motivates homogeneous-first PCD group formation.","marker":"[34]"},{"why":"Knapp 2007 — supplies the Bimodality Coefficient and the 0.555 polarisation threshold used to measure ReadTheRoom's polarisation reduction.","marker":"[22]"},{"why":"Chambers and Warren 2023 — theorises the complementary roles of voting and deliberation that the paper's bridging argument relies on.","marker":"[7]"},{"why":"Hendriks and Michels 2024 — provides the hybrid democratic innovation frame and participatory budgeting new style as the paper's target setting.","marker":"[17]"},{"why":"Yang et al. 2024 — prior experiment showing participants perceive MES as fairer than Greedy, the baseline the KK24 fairness findings extend.","marker":"[40]"},{"why":"Mansbridge 1994 — contribution of protected enclaves for marginalised perspectives, supporting PCD's homogeneous phase.","marker":"[27]"}],"fun_headline_variants":["Algorithms bridge voting and deliberation to reduce polarisation","Voting-guided deliberation: less polarisation, fairer budgets","New algorithms make deliberation more inclusive and trusted","Field-tested: algorithms that link votes to discussion","PCD, MES, ReadTheRoom: algorithms for better deliberation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The PCD results rest on the assumption that Radial Clustering on a two-dimensional PCA projection actually groups participants by genuine preference similarity; with a silhouette score of 0.238 and the authors' own concession that the projection flattens preference complexity, the observed contrasts between homogeneous and heterogeneous rounds could partly reflect noise or facilitator behaviour rather than true preference homogeneity.","fun_headline_variants_meta":{"raw":{"variants":["Algorithms bridge voting and deliberation to reduce polarisation","Voting-guided deliberation: less polarisation, fairer budgets","New algorithms make deliberation more inclusive and trusted","Field-tested: algorithms that link votes to discussion","PCD, MES, ReadTheRoom: algorithms for better deliberation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000274,"raw_usage":{"total_tokens":1669,"prompt_tokens":1001,"completion_tokens":668,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":588}},"tokens_in":617,"tokens_out":668,"duration_ms":7037,"temperature":1.0,"reasoning_tokens":588,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T20:34:08.593796+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute silhouette scores for the Radial Clustering assignments in the original high-dimensional preference space rather than the two-dimensional projection; if the assignments show near-zero or negative silhouette values there, or if a placebo re-randomisation of participants into 'homogeneous' and 'heterogeneous' groups reproduces the same ease/alignment differences, the PCD effect would be indistinguishable from group-size or facilitation artefacts. Similarly, a delayed re-administered survey of vTaiwan participants weeks after the workshop would settle whether the reduced Bimodality Coefficients persist.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Peters, Pierczynski, and Skowron 2021 — supplies the Method of Equal Shares algorithm whose fairness and proportionality the Human-in-the-loop extension builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Small et al. 2021 — supplies the Polis platform's opinion-space mapping and consensus/divisive statement detection used in ReadTheRoom."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Sunstein 2017 — provides the enclave-deliberation concept that motivates homogeneous-first PCD group formation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Knapp 2007 — supplies the Bimodality Coefficient and the 0.555 polarisation threshold used to measure ReadTheRoom's polarisation reduction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Chambers and Warren 2023 — theorises the complementary roles of voting and deliberation that the paper's bridging argument relies on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Hendriks and Michels 2024 — provides the hybrid democratic innovation frame and participatory budgeting new style as the paper's target setting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Mansbridge 1994 — contribution of protected enclaves for marginalised perspectives, supporting PCD's homogeneous phase."}],"review_version":1}