Pith. sign in

REVIEW 4 major objections 6 minor 127 references

Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist Data

T0 review · 4 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper claims that a large language model instructed to run a Socratic dialogue can stand in for synchronous human deliberation in crowd annotation, improving accuracy and confidence while avoiding the coordination costs of real-time di

desk verdict Useful new system, but the accuracy win over synchronous deliberation is not yet proven—the paper's own failure analysis may explain the effect. read the letter →

arxiv 2508.09911 v1 pith:7DVLTFCS submitted 2025-08-13 cs.HC

classification cs.HC
keywords dataannotationperspectivismSocraticdialoguelargelanguagemodelscrowdsourcingdeliberationconfidencelabeldisagreements
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a large language model instructed to run a Socratic dialogue can take the place of other crowdworkers in the deliberation step of data annotation, making perspectivist data collection—where disagreement and diverse perspectives are preserved rather than averaged away—practical at scale. The paper builds an asynchronous chat system in which annotators defend their labels to an LLM that asks probing questions, then re-annotate. Compared against a published synchronous human-deliberation benchmark on two tasks, the system produced more accurate labels on the Relation task where ground truth exists (64.79% post-deliberation accuracy versus 48.86%), moved more annotations to higher confidence, and drew longer, more numerous messages from annotators. The central claim is that the benefits of human deliberation for preserving perspectives can be obtained without coordinating human schedules, and that LLMs can serve this role responsibly when constrained to a Socratic persona.

What carries the argument

The engine is the Socratic dialogue loop, specifically the elenchus stage: after the annotator asserts a claim, the LLM asks one question at a time, probes for counter-evidence and hypothetical boundary cases, and only approves the claim if the reasoning holds. The system prompt encodes five steps of the Socratic method, a temperament (humility, respect, joy, mutual understanding), and guardrails that forbid outside knowledge, limit responses to three sentences, and require a minimum of two rounds. The wording of the prompt matters: it is the entire mechanism that turns a general-purpose chat model (Claude 3 Haiku) into a deliberation partner, and the paper reports that six prompt iterations

What would settle it

Take the same 40 datapoints, recruit one pool of annotators, and randomly assign each datapoint to either Socratic-LLM deliberation or synchronous human deliberation, with identical instructions, the same number of annotators per datapoint, and identical pre/post confidence scales. If the Relation post-deliberation accuracy gap (64.79% vs 48.86%) does not reproduce in that matched design, the paper's central comparison is confounded; additionally, auditing conversation logs for any use of outside factual knowledge would test whether the accuracy gain came from the LLM leaking answers rather th

Watch

Extended reading notes

Core claim

The paper's central discovery is that a constrained Socratic LLM outperforms a synchronous human-deliberation benchmark on key annotation metrics. In the Relation task, participants changed labels more often (23.85% annotation-level flips versus 7.63%), and those changes moved overall accuracy from 52.52% to 64.79%, whereas the benchmark moved only from 47.16% to 48.86%; the improvement came mostly from correcting over-eager 'Expressed' relations. Confidence also increased: post-deliberation, 85.34% of annotations were marked high confidence versus 66.39% in the benchmark, with 28% of annotations moving from medium to high confidence. Qualitatively, the LLM took on four roles—argument evalua

Load-bearing premise

The headline results stand only if the asynchronous Socratic study and the synchronous human benchmark are genuinely comparable in participant pools, task instructions, and annotation counts; if they are not, the accuracy and confidence differences are confounded.

Editorial extensions

If this is right

  • Asynchronous Socratic deliberation can replace synchronous human deliberation for at least some annotation tasks, cutting coordination costs while improving or matching outcomes.
  • On the Relation task, deliberation with the Socratic LLM raised post-deliberation accuracy to 64.79% from 52.52%, beating the benchmark's 48.86% and mostly by correcting false positives.
  • Annotator confidence rose substantially: 28% of annotations moved from medium to high confidence, versus no net change in the benchmark.
  • Annotators engaged more (7.6 vs 5.4 messages; 104.7 vs 75.3 characters per message), suggesting the format sustains deliberation depth.
  • The system's Socratic roles—argument evaluator, boundary negotiator, cognitive support tool, and validator—give dataset curators a way to inspect reasoning and class boundaries in the resulting conversation logs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An untested extension of the role analysis: deliberately switching the LLM's persona (evaluator, boundary negotiator, cognitive support, validator) across dialogue states could improve deliberation, but would place new demands on guardrails.
  • The confidence metric could be repurposed: on subjective tasks without ground truth, a 28% medium-to-high confidence shift may serve as a quality proxy where accuracy is unavailable.
  • The asymmetry in Relation flips (18.46% Expressed to Not Expressed vs 4.41% in the benchmark) raises a question the paper does not settle: whether Socratic questioning biases annotators toward a more conservative reading of relations.
  • The paper's comparison suggests a direct test: a matched between-subjects study with identical participant pools and annotation counts would reveal whether the accuracy gain is attributable to the Socratic mechanism itself rather than to differences between the two studies.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a Socratic LLM dialogue system that replaces a human deliberation partner during crowd annotation, allowing asynchronous reflection on binary labels. The authors benchmark against Schaekermann et al. (2018), using the same Sarcasm and Relation datasets, and report that their intervention produces more annotation flips on the Relation task, higher post-deliberation accuracy on 21 ground-truth Relation datapoints (64.79% vs. 48.86%), increased annotator confidence, and longer/more numerous discussion messages. They also present a qualitative analysis of LLM roles (argument evaluator, boundary negotiator, cognitive support tool, validator) and a candid discussion of failures, including LLM misrepresentations of the Relation task.

Significance. If its central claims hold, the paper makes a useful contribution to perspectivist data annotation by showing that a constrained Socratic LLM can partially substitute for costly synchronous deliberation, with the release of the full system prompt and the use of the benchmark's public dataset enabling replication. The qualitative analysis of how the LLM supports argumentation and the explicit treatment of failures are strengths, as is the authors' willingness to publish a complete, human-readable prompt (Appendix A). However, the main quantitative evidence is currently not strong enough to support the headline accuracy and confidence claims: the comparison is to a prior study rather than a matched control, the decisive accuracy effect is small (roughly 9 annotations), and a known LLM error is aligned with the observed flip direction. The paper's value is more persuasive as a design exploration and qualitative study than as a demonstration of improved annotation accuracy.

major comments (4)
  1. [§5.2, Table 2] The headline accuracy improvement is not established statistically. The increase from 52.52% to 64.79% (a 12.27 pp gain on n=71 annotations, i.e., about 8.7 flips) is reported without confidence intervals, a significance test, or an effect size. The comparison to the benchmark's n=1003 annotations is a cross-study comparison with different participant pools, task presentation, and annotation counts per datapoint, so the apparent advantage could reflect population or design differences. The text also contains an internal inconsistency: pre-deliberation accuracy is given as 53.52% in one sentence and 52.52% in the next. Please report a binomial confidence interval for the post-deliberation accuracy, test the pre/post change, and clearly state the limitations of the cross-study comparison.
  2. [§5.7.1, Table 2, Fig. 3] The accuracy gain may be attributable to the LLM's task misrepresentation rather than to Socratic deliberation. Section 5.7.1 states that the LLM 'occasionally' misrepresented the Relation task, enforcing the incorrect rule that relations must be explicitly stated rather than implied, and that this 'led to some annotation flips.' The dominant flip direction in Fig. 3 (Expressed→Not Expressed: 18.46% vs. 4.41% in the benchmark) and the drop in false positives in Table 2 (35.21%→22.54%) account for essentially the entire accuracy improvement. Given that only about 9 decisive flips are needed for the headline gain, even a handful of misrepresentation-driven flips could explain the result. The paper reports neither the number of such flips nor a sensitivity analysis excluding or reclassifying them. This is a load-bearing gap.
  3. [§4.2.3, §6.2, footnote 12] The re-annotation phase introduces a 'Not Sure' option that was absent in the first annotation phase, and the analysis excludes flips to 'Not Sure' in both conditions. This is a deliberate design choice with a perspectivist rationale, but it changes the response surface between pre- and post-deliberation and may inflate flip rates or confidence changes. The benchmark's 'Irresolvable' category is not identical to 'Not Sure,' so the exclusion is not symmetric. Please report how many annotations were affected, and show that the main flip-rate and accuracy conclusions are robust to including these responses or treating them as missing rather than excluded.
  4. [§5.5.4, Appendix A] The 'Validator' role is described as emergent, but it is directly instructed in the system prompt. Appendix A, step 5, tells the LLM: 'If the reasoning is sound based on the discussion, you should encourage them to continue on to re-annotate the item below their chat.' Similarly, the 'Argument Evaluator' and 'Classification Boundary Negotiator' behaviors closely follow prompt steps 2–4. The qualitative finding is better framed as 'prompt-induced behaviors' rather than emergent roles. This does not invalidate the qualitative analysis, but the current framing overstates the serendipity of the design.
minor comments (6)
  1. [Throughout] Statistical notation: p-values are printed as 'p « 0.001' (e.g., §5.1.1, §5.1.2, §5.3). This should be 'p < 0.001.'
  2. [§5.2] Inconsistent accuracy values: the text says 'Pre-deliberation, our participants had slightly higher accuracy in labeling than the benchmark (53.52% vs. 47.16%),' then later says 'Our overall accuracy increased substantially to 64.79% (from 52.52%).' Please correct the pre-deliberation value.
  3. [§5.3] The sentence 'The change in confidence (pre-intervention confidence minus post-intervention confidence) was significant...' appears to have the sign reversed for the reported increase. Clarify whether the t-test was computed on post−pre.
  4. [§5.1.1, Fig. 3] The Sankey diagrams are informative, but the figure does not show marginal totals. Adding per-arrow counts or percentages would make the comparison of flip asymmetry easier to verify.
  5. [§4.3] The authors report excluding participants who used outside LLMs or completed tasks too quickly, but do not state how many participants were excluded for each reason. Reporting these numbers would help assess data quality.
  6. [§5.2] The comparison group (n=1003 annotation-level observations from the benchmark) is much larger than the treatment group (n=71). While this is a consequence of the benchmark design, the authors should note that the large n gives the benchmark narrow confidence intervals and that the comparison is underpowered on the treatment side.

Circularity Check

1 steps flagged · score 2.0 of 10

Quantitative accuracy/confidence comparisons are measured and self-contained; the qualitative 'emergent roles' are partly prompt instructions relabeled as discoveries.

  1. self definitional [Section 5.5 ('LLM Roles for Supporting Data Annotation') and Appendix A prompt steps 2/5; contradicted by Section 7 limitation sentence]
    "Through our qualitative analysis of the conversation logs ... we identified that the Socratic LLM’s adherence to the prompt instructions (provided in Appendix A) cause distinct patterns of LLM behavior to emerge. We present these patterns as emergent roles ... Validator, approving the annotator’s label once sufficient deliberation has occurred. ... [Appendix A step 5:] If the reasoning is sound based on the discussion, you should encourage them to continue on to re-annotate the item below their chat. ... [Section 7:] These roles were largely effective and appreciated by participants, despite n"

    The claimed 'emergent roles' are textually equivalent to the prompt steps that prescribe them. The Validator role (approving a label once reasoning is sound) is exactly Appendix A step 5; the Classification Boundary Negotiator (using counter-examples to find boundaries) is exactly Appendix A step 2 ('ask about how they might categorize counter-examples to help identify boundaries in their logic'). Section 7 asserts the roles were 'not explicitly designed in the prompt instructions,' but they appear verbatim there. Thus the qualitative role taxonomy reduces to a summary of the system prompt by construction, rather than an independent empirical discovery. This is presentational circularity only; the accuracy and confidence outcomes are measured separately and do not depend on this taxonomy.

full rationale

The paper's headline numerical claims are not circular: Relation-task accuracy (52.52% pre vs. 64.79% post, vs. benchmark 47.16% to 48.86%) and confidence changes (28% medium-to-high) come from collected annotations compared against Schaekermann et al.'s public benchmark. No parameter is fitted and renamed a prediction, no prior result by these authors is invoked as a load-bearing uniqueness argument, and the derivation chain for the quantitative findings is external and measured. The only circularity found is presentational: the 'emergent roles' of Section 5.5 are relabeled versions of the system prompt's own Socratic steps (Validator = step 5; Boundary Negotiator = step 2), and the paper's own limitation statement disclaims that they were 'not explicitly designed' when the appendix shows they were. This affects a qualitative framing contribution, not the central accuracy/confidence comparisons. The skeptic concern about LLM task misrepresentation driving flips (§5.7.1) is a serious internal-validity threat, but it is not circularity: the accuracy outcome is still measured, not derived from the prompt by construction. Overall score reflects one minor, non-load-bearing circular framing.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The central claims rest on the comparability of two studies, on the validity of self-reported confidence, and on the subset of ground-truth datapoints. No fitted model parameters are involved; the only numeric design parameter affecting a headline metric is the minimum required message count. The 'Not Sure' re-annotation option and the supportive LLM temperament are design choices that interact with the measurement of flips and confidence.

free parameters (1)
  • Minimum required discussion messages per datapoint = 2
    Set in the study protocol (Section 4.2.2) to mirror the benchmark's minimum participation; directly influences the engagement metric (average message count).
assumptions (5)
  • domain assumption The prior synchronous-deliberation study by Schaekermann et al. [102] is a valid baseline for comparison despite differences in participant pool, task conditions, and data size.
    Section 4 states the design mirrors the benchmark; Section 7 acknowledges 'difference in design between the synchronous deliberation approach of the benchmark paper and our asynchronous implementation', making comparability load-bearing for all quantitative claims.
  • domain assumption Self-reported confidence measures genuine certainty, not social desirability or the effect of the LLM's supportive temperament.
    Section 5.3 interprets confidence increases as stemming from the Socratic process; the prompt (Appendix A) instructs traits like humility, respect, and mutual understanding, which could inflate confidence.
  • domain assumption The 21 Relation datapoints with ground truth are representative and their ground-truth labels are correct.
    Section 5.2 restricts accuracy analysis to 21 datapoints, noting that 4 of 25 grounded items were not deliberated in the benchmark; this subset selection affects the accuracy claim.
  • standard math Statistical tests treat individual annotations as independent observations, although two annotations come from the same participant.
    Sections 5.1 and 5.3 use z-tests and t-tests on annotations without accounting for participant clustering.
  • ad hoc to paper The 'Not Sure' re-annotation option added only in the post-deliberation phase is a valid design, and flips to this option can be excluded from analysis.
    Section 4.2.3 introduces the 'Not Sure' option only in re-annotation; Section 5.1 excludes such flips, an asymmetric design choice that affects measured flip rates.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist Data." pith.science (2026). https://pith.science/paper/7DVLTFCS

@misc{pith2026250809911,
  author       = {Pith},
  title        = {Pith review of: Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist Data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7DVLTFCS}},
  note         = {Machine review of arXiv:2508.09911}
}
read the original abstract

Data annotation underpins the success of modern AI, but the aggregation of crowd-collected datasets can harm the preservation of diverse perspectives in data. Difficult and ambiguous tasks cannot easily be collapsed into unitary labels. Prior work has shown that deliberation and discussion improve data quality and preserve diverse perspectives -- however, synchronous deliberation through crowdsourcing platforms is time-intensive and costly. In this work, we create a Socratic dialog system using Large Language Models (LLMs) to act as a deliberation partner in place of other crowdworkers. Against a benchmark of synchronous deliberation on two tasks (Sarcasm and Relation detection), our Socratic LLM encouraged participants to consider alternate annotation perspectives, update their labels as needed (with higher confidence), and resulted in higher annotation accuracy (for the Relation task where ground truth is available). Qualitative findings show that our agent's Socratic approach was effective at encouraging reasoned arguments from our participants, and that the intervention was well-received. Our methodology lays the groundwork for building scalable systems that preserve individual perspectives in generating more representative datasets.

Figures

Figures reproduced from arXiv: 2508.09911 by the authors.

Figure 1
Figure 1. Socratic Method in practice. Participants progress through the phases of the Socratic Method as they complete annotation tasks. (A) Wonder, the participant considers the datapoint provided. (B) Hypothe￾sis, participants generate a hypothesis based on the data and respond to questions to record their annotation. (C) Elenchus, engagement with the Socratic LLM encourages participants to reflect on their hypothesis and … view at source ↗
Figure 2
Figure 2. Workflow diagram of the Socratic LLM-assisted annotation process. Crowdworkers (N=133) first per [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Annotation-level flips from both our results (left) and Schaekermann et al. [102] (right). Results are [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Changes in confidence between initial annotations and post-deliberation re-annotations, both for us [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]
Figure 5
Figure 5. Figure 5: Example screenshot from our system during the pre-deliberation annotation phase. Datapoints [PITH_FULL_IMAGE:figures/full_fig_p034_5.png]
Figure 6
Figure 6. Figure 6: Example screenshot from our system at the point of the post-deliberation annotation phase. The [PITH_FULL_IMAGE:figures/full_fig_p035_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

127 extracted references · 44 canonical work pages

  1. [1]

    [n. d.]. Socratic Methods - Wikiversity — en.wikiversity.org. https://en.wikiversity.org/wiki/Socratic_Methods. [Ac- cessed 2024-10-25]

  2. [2]

    Jamie R. Abrams. 2015. Reframing the Socratic Method. 64, 4 (2015), 562–585. jstor:24716713 https://www.jstor.org/ stable/24716713

  3. [3]

    Erfan Al-Hossami, Razvan Bunescu, Justin Smith, and Ryan Teehan. 2024. Can Language Models Employ the Socratic Method? Experiments with Code Debugging. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE 2024) . Association for Computing Machinery, New York, NY, USA, 53–59. doi:10. 1145/3626252.3630799

  4. [4]

    Reham Al Tamime, Joni Salminen, Soon-Gyo Jung, and Bernard Jansen. 2024. Evaluating LLM-Generated Topics from Survey Responses: Identifying Challenges in Recruiting Participants through Crowdsourcing. In 2024 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) . IEEE, 412–416

  5. [5]

    Maya Aloni and Christine Harrington. 2018. Research Based Practices for Improving the Effectiveness of Asyn- chronous Online Discussion Boards. Scholarship of Teaching and Learning in Psychology 4, 4 (Dec. 2018), 271–289. doi:10.1037/stl0000121

  6. [6]

    Omar Alonso and Stefano Mizzaro. 2012. Using crowdsourcing for TREC relevance assessment. Information process- ing & management 48, 6 (2012), 1053–1066

  7. [7]

    Paul André, Aniket Kittur, and Steven P Dow. 2014. Crowd synthesis: Extracting categories and clusters from complex data. In Proceedings of the 17th ACM conference on Computer supported cooperative work & social computing. 989–998

  8. [8]

    Lora Aroyo and Chris Welty. 2015. Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation. AI Magazine 36, 1 (March 2015), 15–24. doi:10.1609/aimag.v36i1.2564

Show all 127 references
  1. [9]

    Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of sto- chastic parrots: Can language models be too big?. InProceedings of the 2021 ACM conference on fairness, accountability, and transparency. 610–623

  2. [10]

    Karim Benharrak, Tim Zindulka, Florian Lehmann, Hendrik Heuer, and Daniel Buschek. 2024. Writer-Defined AI Personas for On-Demand Feedback Generation. In Proceedings of the 2024 CHI Conference on Human Factors in Com- puting Systems (Honolulu, HI, USA) (CHI ’24). Association f...

  3. [11]

    Reuben Binns, Michael Veale, Max Van Kleek, and Nigel Shadbolt. 2017. Like Trainer, Like Bot? Inheritance of Bias in Algorithmic Content Moderation. In Social Informatics (Cham, 2017), Giovanni Luca Ciampaglia, Afra Mashhadi, and Taha Yasseri (Eds.). Springer International Pub...

  4. [12]

    Eugenia Arazo Boa, Amornrat Wattanatorn, and Kanchit Tagong. 2018. The Development and Validation of the Blended Socratic Method of Teaching (BSMT): An Instructional Model to Enhance Critical Thinking Skills of Under- graduate Business Students. 39, 1 (2018), 81–89. doi:10.101...

  5. [13]

    Jonathan Bragg, Mausam, and Daniel S. Weld. 2018. Sprout: Crowd-Powered Task Design for Crowdsourcing. In Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology (UIST ’18) . Association for Computing Machinery, New York, NY, USA, 165–176. doi:10...

  6. [14]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101

  7. [15]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. 5 (2021), 188:1–188:21. Issue CSCW1. doi:10.1145/ 3449287

  8. [16]

    Federico Cabitza, Andrea Campagner, and Valerio Basile. 2023. Toward a Perspectivist Turn in Ground Truthing for Predictive Computing. Proceedings of the AAAI Conference on Artificial Intelligence 37, 6 (June 2023), 6860–6868. doi:10.1609/aaai.v37i6.25840

  9. [17]

    Carrie J Cai, Shamsi T Iqbal, and Jaime Teevan. 2016. Chain reactions: The impact of order on microtask chains. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems . 3143–3154

  10. [18]

    Scott Allen Cambo and Darren Gergle. 2022. Model Positionality and Computational Reflexivity: Promoting Reflex- ivity in Data Science. In CHI Conference on Human Factors in Computing Systems . ACM, New Orleans LA USA, 1–19. doi:10.1145/3491102.3501998 Proc. ACM Hum.-Comput. In...

  11. [19]

    Joseph Chee Chang, Saleema Amershi, and Ece Kamar. 2017. Revolt: Collaborative Crowdsourcing for Labeling Machine Learning Datasets. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (CHI ’17). Association for Computing Machinery, New York, NY, US...

  12. [20]

    Adriane Chapman, Philip Grylls, Pamela Ugwudike, David Gammack, and Jacqui Ayling. 2022. A Data-driven Anal- ysis of the Interplay between Criminological Theory and Predictive Policing Algorithms. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Trans...

  13. [21]

    Ana Paula Chaves and Marco Aurelio Gerosa. 2021. How should my chatbot interact? A survey on social charac- teristics in human–chatbot interaction design. International Journal of Human–Computer Interaction 37, 8 (2021), 729–758

  14. [22]

    Quanze Chen, Jonathan Bragg, Lydia B Chilton, and Dan S Weld. 2019. Cicero: Multi-turn, contextual argumentation for accurate crowdsourcing. In Proceedings of the 2019 chi conference on human factors in computing systems . 1–14

  15. [23]

    Weld, and Amy X

    Quan Ze Chen, Daniel S. Weld, and Amy X. Zhang. 2021. Goldilocks: Consistent Crowdsourced Scalar Annotations with Relative Uncertainty. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (Oct. 2021), 335:1– 335:25. doi:10.1145/3476076

  16. [24]

    Quan Ze Chen and Amy X. Zhang. 2023. Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus Disagreement. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2 (Oct. 2023), 283:1–283:26. doi:10.1145/3610074

  17. [25]

    Julian Chingoma and Adrian Haret. 2023. Deliberation as Evidence Disclosure: A Tale of Two Protocol Types. In Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems (AAMAS ’23) . Inter- national Foundation for Autonomous Agents and Multiag...

  18. [26]

    Florian Daniel, Pavel Kucherbaev, Cinzia Cappiello, Boualem Benatallah, and Mohammad Allahbakhsh. 2018. Quality Control in Crowdsourcing: A Survey of Quality Attributes, Assessment Techniques, and Assurance Actions. 51, 1 (2018), 7:1–7:40. doi:10.1145/3148148

  19. [27]

    Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022. Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations. Transactions of the Association for Computational Linguistics 10 (Jan. 2022), 92–110. doi:10.1162/tacl_a_00449

  20. [28]

    Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated hate speech detection and the problem of offensive language. InProceedings of the international AAAI conference on web and social media, Vol. 11. 512–515

  21. [29]

    Valerio De Stefano. 2015. The Rise of the ’Just-in-Time Workforce’: On-Demand Work, Crowd Work and Labour Protec- tion in the ’Gig-Economy’ . Social Science Research Network:2682602 doi:10.2139/ssrn.2682602

  22. [30]

    Amanda Delaney, Bella Lough, Michelle Whelan, Max Cameron, et al. 2004. A review of mass media campaigns in road safety. Monash University Accident Research Centre Reports 220 (2004), 85

  23. [31]

    Haris Delić and Senad Bećirović. 2016. Socratic Method as an Approach to Teaching. European Researcher 111, 10 (Oct. 2016). doi:10.13187/er.2016.111.511

  24. [32]

    Yuyang Ding, Hanglei Hu, Jie Zhou, Qin Chen, Bo Jiang, and Liang He. 2024. Boosting Large Language Models with Socratic Method for Conversational Mathematics Teaching. arXiv:2407.17349 [cs]

  25. [33]

    Carl DiSalvo. 2012. Adversarial Design. The MIT Press. doi:10.7551/mitpress/8732.001.0001

  26. [34]

    Ryan Drapeau, Lydia Chilton, Jonathan Bragg, and Daniel Weld. 2016. Microtalk: Using argumentation to improve crowdsourcing accuracy. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , Vol. 4. 32–41

  27. [35]

    Such, Mark Coté, and Natalia Criado

    Xavier Ferrer, prefix=van useprefix=false family=Nuenen, given=Tom, Jose M. Such, Mark Coté, and Natalia Criado

  28. [36]

    Elena Filatova. 2012. Irony and Sarcasm: Corpus Generation and Analysis Using Crowdsourcing.. In Lrec. Citeseer, 392–398

  29. [37]

    Eve Fleisig, Su Lin Blodgett, Dan Klein, and Zeerak Talat. 2024. The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels. (May 2024). doi:10.48550/ARXIV.2405.05860

  30. [38]

    Soon Yen Foo and Choon Lang Quek. 2019. Developing Students’ Critical Thinking through Asynchronous Online Discussions: A Literature Review. Malaysian Online Journal of Educational Technology 7, 2 (2019), 37–58. doi:10. 17220/mojet.2019.02.003

  31. [39]

    Giles Foody, Linda See, Steffen Fritz, Inian Moorthy, Christoph Perger, Christian Schill, and Doreen Boyd. 2018. Increasing the accuracy of crowdsourced information on land cover via a voting procedure weighted by information inferred from the contributed data. ISPRS Internati...

  32. [40]

    Benoît Frénay and Michel Verleysen. 2014. Classification in the Presence of Label Noise: A Survey.IEEE Transactions on Neural Networks and Learning Systems 25, 5 (May 2014), 845–869. doi:10.1109/TNNLS.2013.2292894 Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW52...

  33. [41]

    Simona Frenda, Gavin Abercrombie, Valerio Basile, Alessandro Pedrani, Raffaella Panizzon, Alessandra Teresa Cignarella, Cristina Marco, and Davide Bernardi. 2024. Perspectivist approaches to natural language processing: a survey. Language Resources and Evaluation (2024), 1–28

  34. [42]

    Gordon, Michelle S

    Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S. Bernstein. 2022. Jury Learning: Integrating Dissenting Voices into Machine Learning Models. In CHI Conference on Human Factors in Computing Systems . ACM, New Or...

  35. [43]

    Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S

    Mitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S. Bernstein. 2021. The Disagree- ment Deconvolution: Bringing Machine Learning Performance Metrics In Line With Reality. InProceedings of the 2021 CHI Conference on Human Factors in Computing Syst...

  36. [44]

    Tanya Goyal, Tyler McDonnell, Mucahid Kutlu, Tamer Elsayed, and Matthew Lease. 2018. Your Behavior Signals Your Reliability: Modeling Crowd Behavioral Traces to Ensure Quality Relevance Annotations. 6 (2018), 41–49. doi:10.1609/hcomp.v6i1.13331

  37. [45]

    Hui Guo, Boyu Wang, and Grace Yi. 2023. Label Correction of Crowdsourced Noisy Annotations with an Instance- Dependent Noise Transition Model. 36 (2023), 347–386. https://proceedings.neurips.cc/paper_files/paper/2023/hash/ 015a8c69bedcb0a7b2ed2e1678f34399-Abstract-Conference.html

  38. [46]

    Margeret Hall, Mohammad Farhad Afzali, Markus Krause, and Simon Caton. 2022. What Quality Control Mecha- nisms Do We Need for High-Quality Crowd Work? 10 (2022), 99709–99723. doi:10.1109/ACCESS.2022.3207292

  39. [47]

    Jawad Haqbeen, Takayuki Ito, Rafik Hadfi, Tomohiro Nishida, Zoia Sahab, Sofia Sahab, Shafiq Roghmal, and Moham- mad Amiryar. 2020. Promoting Discussion with AI-based Facilitation: Urban Dialogue with Kabul City

  40. [48]

    Sandra G Hart. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. Human mental workload/Elsevier (1988)

  41. [49]

    Yueh-Ren Ho, Bao-Yu Chen, and Chien-Ming Li. 2023. Thinking More Wisely: Using the Socratic Method to Develop Critical Thinking Skills amongst Healthcare Students. 23, 1 (2023), 173. doi:10.1186/s12909-023-04134-2

  42. [50]

    James Hollan, Edwin Hutchins, and David Kirsh. 2000. Distributed cognition: toward a new foundation for human- computer interaction research. ACM Transactions on Computer-Human Interaction (TOCHI) 7, 2 (2000), 174–196

  43. [51]

    All of the White People Went First

    Mo Houtti, Moyan Zhou, Loren Terveen, and Stevie Chancellor. 2023. " All of the White People Went First": How Video Conferencing Consolidates Control and Exacerbates Workplace Bias. Proceedings of the ACM on Human- Computer Interaction 7, CSCW1 (2023), 1–25

  44. [52]

    Popescu, Saurabh Chatterjee, and Thad Starner

    Jui-Tse Hung, Christopher Cui, Diana M. Popescu, Saurabh Chatterjee, and Thad Starner. 2024. Socratic Mind: Scal- able Oral Assessment Powered By AI. In Proceedings of the Eleventh ACM Conference on Learning @ Scale (L@S ’24) . Association for Computing Machinery, New York, NY...

  45. [53]

    Edwin Hutchins. 1995. How a cockpit remembers its speeds. Cognitive science 19, 3 (1995), 265–288

  46. [54]

    Traganitis, Xiao Fu, and Georgios B

    Shahana Ibrahim, Panagiotis A. Traganitis, Xiao Fu, and Georgios B. Giannakis. 2024. Learning From Crowdsourced Noisy Labels: A Signal Processing Perspective . arXiv:2407.06902 doi:10.48550/arXiv.2407.06902

  47. [55]

    Muhammad Okky Ibrohim and Indra Budi. 2019. Multi-label hate speech and abusive language detection in Indone- sian Twitter. In Proceedings of the third workshop on abusive language online . 46–57

  48. [56]

    Oana Inel, Khalid Khamkham, Tatiana Cristea, Anca Dumitrache, Arne Rutjes, Jelle van der Ploeg, Lukasz Romaszko, Lora Aroyo, and Robert-Jan Sips. 2014. Crowdtruth: Machine-human computation framework for harnessing dis- agreement in gathering annotated data. InThe Semantic Web...

  49. [57]

    Panagiotis G Ipeirotis and Evgeniy Gabrilovich. 2014. Quizz: targeted crowdsourcing with a billion (potential) users. In Proceedings of the 23rd international conference on World wide web . 143–154

  50. [58]

    Shamsi T Iqbal and Brian P Bailey. 2005. Investigating the effectiveness of mental workload as a predictor of oppor- tune moments for interruption. In CHI’05 extended abstracts on Human factors in computing systems . 1489–1492

  51. [59]

    Deniz Iren and Semih Bilgen. 2014. Cost of Quality in Crowdsourcing. 1, 2 (2014). Issue 2. doi:10.15346/hc.v1i2.14

  52. [60]

    V. K. Chaithanya Manam, Dwarakanath Jampani, Mariam Zaim, Meng-Han Wu, and Alexander J. Quinn. 2019. Task- Mate: A Mechanism to Improve the Quality of Instructions in Crowdsourcing. InCompanion Proceedings of The 2019 World Wide Web Conference (WWW ’19) . Association for Compu...

  53. [61]

    Shivani Kapania, Alex S Taylor, and Ding Wang. 2023. A hunt for the snark: Annotator diversity in data practices. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–15

  54. [62]

    Martin F Kaplan and Ana M Martin. 1999. Effects of differential status of group members on process and outcome of deliberation. Group Processes & Intergroup Relations 2, 4 (1999), 347–364

  55. [63]

    Georgi Karadzhov, Tom Stafford, and Andreas Vlachos. 2023. DeliData: A Dataset for Deliberation in Multi-party Problem Solving. Proc. ACM Hum.-Comput. Interact. 7, CSCW2 (Oct. 2023), 265:1–265:25. doi:10.1145/3610056

  56. [64]

    Priyanka Kargupta, Ishika Agarwal, Dilek Hakkani-Tur, and Jiawei Han. 2024. Instruct, Not Assist: LLM-based Multi- Turn Planning and Hierarchical Questioning for Socratic Code Debugging . arXiv:2406.11709 doi:10.48550/arXiv.2406. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7...

  57. [65]

    Harmanpreet Kaur, Alex C Williams, Anne Loomis Thompson, Walter S Lasecki, Shamsi T Iqbal, and Jaime Teevan

  58. [66]

    Ashish Khetan, Zachary C Lipton, and Anima Anandkumar. 2018. Learning from noisy singly-labeled data. In Pro- ceedings of ICLR 2018

  59. [67]

    Nickerson, Michael Bernstein, Elizabeth Gerber, Aaron Shaw, John Zimmerman, Matt Lease, and John Horton

    Aniket Kittur, Jeffrey V. Nickerson, Michael Bernstein, Elizabeth Gerber, Aaron Shaw, John Zimmerman, Matt Lease, and John Horton. 2013. The Future of Crowd Work. In Proceedings of the 2013 Conference on Computer Supported Cooperative Work. ACM, San Antonio Texas USA, 1301–131...

  60. [68]

    Charles Koutcheme, Nicola Dainese, Arto Hellas, Sami Sarsa, Juho Leinonen, Syed Ashraf, and Paul Denny. 2024. Evaluating Language Models for Generating and Judging Programming Feedback. arXiv:2407.04873 [cs] doi:10. 48550/arXiv.2407.04873

  61. [69]

    Travis Kriplean, Michael Toomim, Jonathan Morgan, Alan Borning, and Amy J. Ko. 2012. Is This What You Meant?: Promoting Listening on the Web with Reflect. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM, Austin Texas USA, 1559–1568. doi:10.114...

  62. [70]

    Tzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein, and Haiyi Zhu. 2024. Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia. InProceedings of the CHI Conference on Human Factors in Computing Systems (C...

  63. [71]

    Hélène Landemore and Scott E Page. 2015. Deliberation and disagreement: Problem solving, prediction, and positive dissensus. Politics, philosophy & economics 14, 3 (2015), 229–254

  64. [72]

    Susan Leigh Star. 2010. This is not a boundary object: Reflections on the origin of a concept. Science, technology, & human values 35, 5 (2010), 601–617

  65. [73]

    Elisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini, and Sara Tonelli. 2021. Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ Disagreement. In Proceedings of the 2021 Con- ference on Empirical Methods in Natural Language Proce...

  66. [74]

    Franklin

    Guoliang Li, Jiannan Wang, Yudian Zheng, and Michael J. Franklin. 2016. Crowdsourced Data Management: A Survey. 28, 9 (2016), 2296–2319. doi:10.1109/TKDE.2016.2535242

  67. [75]

    Zihao Li. 2023. The dark side of chatgpt: Legal and ethical challenges from stochastic parrots and hallucination. arXiv preprint arXiv:2304.14347 (2023)

  68. [76]

    Cindy Kaiying Lin and Steven J. Jackson. 2023. From Bias to Repair: Error as a Site of Collaboration and Negotiation in Applied Data Science Work. Proceedings of the ACM on Human-Computer Interaction 7, CSCW1 (April 2023), 131:1–131:32. doi:10.1145/3579607

  69. [77]

    Jieli Liu and Pengyi Zhang. 2020. How to Initiate a Discussion Thread?: Exploring Factors Influencing Engagement Level of Online Deliberation. In Sustainable Digital Communities , Anneli Sundqvist, Gerd Berget, Jan Nolin, and Kjell Ivar Skjerdingstad (Eds.). Vol. 12051. Spring...

  70. [78]

    Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, and Sanjiv Kumar. 2020. Does Label Smoothing Mitigate Label Noise?. In Proceedings of the 37th International Conference on Machine Learning (ICML’20, Vol. 119) . JMLR.org, 6448–6458

  71. [79]

    Fenglong Ma, Yaliang Li, Qi Li, Minghui Qiu, Jing Gao, Shi Zhi, Lu Su, Bo Zhao, Heng Ji, and Jiawei Han. 2015. FaitCrowd: Fine Grained Truth Discovery for Crowdsourced Data Aggregation. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and D...

  72. [80]

    Ru Ma, Jiachen Zhao, Chenghang Huo, Xiaodong Zhan, and Fuzhi Zhang. 2023. Spammer Groups Detection Based on Hypergraph Embedding And Autoencoder Classifier Model. InProceedings of the 2023 7th International Conference on Electronic Information Technology and Computer Engineeri...

  73. [81]

    Shuai Ma, Qiaoyi Chen, Xinru Wang, Chengbo Zheng, Zhenhui Peng, Ming Yin, and Xiaojuan Ma. 2024. Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making . arXiv:2403.16812 doi:10.48550/arXiv.2403.16812

  74. [82]

    Walid Magdy, Kareem Darwish, and Norah Abokhodair. 2015. Quantifying public response towards Islam on Twitter after Paris attacks. arXiv preprint arXiv:1512.04570 (2015)

  75. [83]

    Bernard Manin. 2005. Democratic deliberation: Why we should promote debate rather than discussion. In Paper delivered at the program in ethics and public affairs seminar, Princeton University , Vol. 13. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW526. Publicat...

  76. [84]

    Samuel Mayworm, Kendra Albert, and Oliver L. Haimson. 2024. Misgendered During Moderation: How Transgender Bodies Make Visible Cisnormative Content Moderation Policies and Enforcement in a Meta Oversight Board Case. In Proceedings of the 2024 ACM Conference on Fairness, Accoun...

  77. [85]

    David Alvarez Melis, Harmanpreet Kaur, Hal Daumé III, Hanna Wallach, and Jennifer Wortman Vaughan. 2021. From human explanation to model interpretability: A framework based on weight of evidence. InProceedings of the AAAI Conference on Human Computation and Crowdsourcing , Vol...

  78. [86]

    Tali Mendelberg, Christopher F Karpowitz, and J Baxter Oliphant. [n. d.]. Gender inequality in deliberation: Unpack- ing the black box of interaction. Perspectives on Politics 12, 1 ([n. d.]), 18–44

  79. [87]

    David Miller. 2018. Is deliberative democracy unfair to disadvantaged groups? In Democracy as Public Deliberation . Routledge, 201–226

  80. [88]

    Chantal Mouffe. 1999. Deliberative Democracy or Agonistic Pluralism? 66, 3 (1999), 745–758. jstor:40971349 https: //www.jstor.org/stable/40971349

  81. [89]

    Sania Nayab, Giulio Rossolini, Marco Simoni, Andrea Saracino, Giorgio Buttazzo, Nicolamaria Manes, and Fab- rizio Giacomelli. 2024. Concise thoughts: Impact of output length on llm reasoning and cost. arXiv preprint arXiv:2407.19825 (2024)

  82. [90]

    Stefanie Nowak and Stefan Rüger. 2010. How Reliable Are Annotations via Crowdsourcing: A Study about Inter- Annotator Agreement for Multi-Label Image Annotation. InProceedings of the International Conference on Multimedia Information Retrieval. ACM, Philadelphia Pennsylvania U...

  83. [91]

    Jeongeon Park, Eun-Young Ko, Yeon Su Park, Jinyeong Yim, and Juho Kim. 2024. DynamicLabels: Supporting In- formed Construction of Machine Learning Label Sets with Crowd Feedback. In Proceedings of the 29th International Conference on Intelligent User Interfaces (IUI ’24). Asso...

  84. [92]

    Weiping Pei, Arthur Mayer, Kaylynn Tu, and Chuan Yue. 2020. Attention please: Your attention check questions in survey studies can be automatically answered. In Proceedings of The Web Conference 2020 . 1182–1193

  85. [93]

    Mark Perry. 2003. Distributed cognition. HCI models, theories, and frameworks: Toward a multidisciplinary science (2003), 193–223

  86. [94]

    Barbara Plank. 2022. The ’Problem’ of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation. arXiv. doi:10.48550/ARXIV.2211.02570

  87. [95]

    2013.An Evaluation of Aggregation Techniques in Crowdsourcing

    Nguyen Quoc Viet Hung, Nguyen Thanh Tam, Lam Ngoc Tran, and Karl Aberer. 2013.An Evaluation of Aggregation Techniques in Crowdsourcing. Springer Berlin Heidelberg, 1–15. doi:10.1007/978-3-642-41154-0_1

  88. [96]

    Cyrus Rashtchian, Peter Young, Micah Hodosh, and Julia Hockenmaier. 2010. Collecting Image Annotations Using Amazon’s Mechanical Turk. In Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk (CSLDAMT ’10) . Association f...

  89. [97]

    Alexander J Ratner, Christopher M De Sa, Sen Wu, Daniel Selsam, and Christopher Ré. 2016. Data programming: Creating large training sets, quickly. Advances in neural information processing systems 29 (2016)

  90. [98]

    Weingarten, Lilith Fury, Constanza Eliana Chinea, Tuck J

    Yim Register, Izzi Grasso, Lauren N. Weingarten, Lilith Fury, Constanza Eliana Chinea, Tuck J. Malloy, and Emma S. Spiro. 2024. Beyond Initial Removal: Lasting Impacts of Discriminatory Content Moderation to Marginalized Cre- ators on Instagram. 8 (2024), 23:1–23:28. Issue CSC...

  91. [99]

    Everyone Wants to Do the Model Work, Not the Data Work

    Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021. “Everyone Wants to Do the Model Work, Not the Data Work”: Data Cascades in High-Stakes AI. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Syste...

  92. [100]

    Marta Sandri, Elisa Leonardelli, Sara Tonelli, and Elisabetta Jezek. 2023. Why Don’t You Do It Right? Analysing Annotators’ Disagreement in Subjective Tasks. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (May 2023),...

  93. [101]

    Yisi Sang and Jeffrey Stanton. 2022. The Origin and Value of Disagreement Among Data Labelers: A Case Study of Individual Differences in Hate Speech Annotation. In Information for a Better World: Shaping the Global Future (Lecture Notes in Computer Science) , Malte Smits (Ed.)...

  94. [102]

    Mike Schaekermann, Joslin Goh, Kate Larson, and Edith Law. 2018. Resolvable vs. Irresolvable Disagreement: A Study on Worker Deliberation in Crowd Work. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (Nov. 2018), 1–19. doi:10.1145/3274423

  95. [103]

    Aashish Sheshadri and Matthew Lease. 2013. Square: A benchmark for research on computing crowd consensus. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , Vol. 1. 156–164

  96. [104]

    Hedderich, AndréS Lucero, and Antti Oulasvirta

    Joongi Shin, Michael A. Hedderich, AndréS Lucero, and Antti Oulasvirta. 2022. Chatbots Facilitating Consensus- Building in Asynchronous Co-Design. In Proceedings of the 35th Annual ACM Symposium on User Interface Software Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Articl...

  97. [105]

    Susan Leigh Star and James R Griesemer. 1989. Institutional ecology,translations’ and boundary objects: Amateurs and professionals in Berkeley’s Museum of Vertebrate Zoology, 1907-39.Social studies of science 19, 3 (1989), 387–420

  98. [106]

    Thitaree Tanprasert, Sidney S Fels, Luanne Sinnamon, and Dongwook Yoon. 2024. Debate Chatbots to Facilitate Critical Thinking on YouTube: Social Identity and Conversational Style Make A Difference. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems ...

  99. [107]

    Dapeng Tao, Jun Cheng, Zhengtao Yu, Kun Yue, and Lizhen Wang. 2018. Domain-weighted majority voting for crowdsourcing. IEEE transactions on neural networks and learning systems 30, 1 (2018), 163–174

  100. [108]

    Stephen E Toulmin. 2003. The uses of argument . Cambridge university press

  101. [109]

    Bernstein, and Ranjay Krishna

    Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna. 2023. Explanations Can Reduce Overreliance on AI Systems During Decision-Making. 7 (2023), 129:1–129:38. Issue CSCW1. doi:10.1145/3579605

  102. [110]

    Melanie A Wakefield, Barbara Loken, and Robert C Hornik. 2010. Use of mass media campaigns to change health behaviour. The lancet 376, 9748 (2010), 1261–1271

  103. [111]

    Stacy E. Walker. 2003. Active Learning Strategies to Promote Critical Thinking. Journal of Athletic Training 38, 3 (2003), 263–267

  104. [112]

    Shaun Wallace, Tianyuan Cai, Brendan Le, and Luis A. Leiva. 2022. Debiased Label Aggregation for Subjective Crowdsourcing Tasks. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (CHI EA ’22). Association for Computing Machinery, New York, ...

  105. [113]

    Douglas N Walton. 1998. The new dialectic: Conversational contexts of argument . University of Toronto Press

  106. [114]

    Qi Wang, Mulin Chen, Feiping Nie, and Xuelong Li. 2018. Detecting coherent groups in crowd scenes by multiview clustering. IEEE transactions on pattern analysis and machine intelligence 42, 1 (2018), 46–58

  107. [115]

    Tharindu Cyril Weerasooriya, Alexander Ororbia, Raj Bhensadadia, Ashiqur KhudaBukhsh, and Christopher Homan

  108. [116]

    Guy Williams et al. 2014. Harkness learning: principles of a radical American pedagogy. (2014)

  109. [117]

    Xuansheng Wu, Haiyan Zhao, Yaochen Zhu, Yucheng Shi, Fan Yang, Tianming Liu, Xiaoming Zhai, Wenlin Yao, Jundong Li, Mengnan Du, et al. 2024. Usable XAI: 10 strategies towards exploiting explainability in the LLM era. arXiv preprint arXiv:2403.08946 (2024)

  110. [118]

    Jingru Yang, Ju Fan, Zhewei Wei, Guoliang Li, Tongyu Liu, and Xiaoyong Du. 2018. Cost-effective data annotation using game-based crowdsourcing. Proceedings of the VLDB Endowment 12, 1 (2018), 57–70

  111. [119]

    Ming Yin and Yiling Chen. 2016. Predicting Crowd Work Quality under Monetary Interventions. 4 (2016), 259–268. doi:10.1609/hcomp.v4i1.13282

  112. [120]

    annotator rationales

    Omar Zaidan, Jason Eisner, and Christine Piatko. 2007. Using “annotator rationales” to improve machine learning for text categorization. In Human language technologies 2007: The conference of the North American chapter of the association for computational linguistics; proceedi...

  113. [121]

    Angie Zhang, Olympia Walker, Kaci Nguyen, Jiajun Dai, Anqing Chen, and Min Kyung Lee. 2023. Deliberating with AI: Improving Decision-Making for the Future through Participatory AI Design and Stakeholder Deliberation. Proc. ACM Hum.-Comput. Interact. 7, CSCW1 (April 2023), 125:...

  114. [122]

    Scheng, Tao Li, and Xindong Wu

    Jing Zhang, Victor S. Scheng, Tao Li, and Xindong Wu. 2017. Improving Crowdsourced Label Quality Using Noise Correction. 29, 5 (2017), 1675–1688. doi:10.1109/TNNLS.2017.2677468

  115. [123]

    Honglei Zhuang and Joel Young. 2015. Leveraging in-batch annotation bias for crowdsourced active learning. In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining . 243–252

  116. [124]

    I can’t provide any additional information outside what was given for the task. You should use your own knowledge and experience to help inform your choice

    Öznur Göçmen and Hamit Coşkun. 2019. The effects of the six thinking hats and speed on creativity in brainstorming. Thinking Skills and Creativity 31 (2019), 284–295. doi:10.1016/j.tsc.2019.02.006 Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW526. Publication da...

  117. [2018]

    Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 1–22

    Creating better action plans for writing tasks via vocabulary-based planning. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 1–22

  118. [2021]

    40, 2 (2021), 72–80

    Bias and Discrimination in AI: A Cross-Disciplinary Perspective. 40, 2 (2021), 72–80. doi:10.1109/MTS.2021. 3056293

  119. [2023]

    In Findings of the Association for Computational Linguistics: ACL 2023

    Disagreement matters: Preserving label diversity by jointly modeling item and annotator label distributions with DisCo. In Findings of the Association for Computational Linguistics: ACL 2023 . 4679–4695

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.