REVIEW 4 major objections 6 minor 127 references
Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist Data
T0 review · 4 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This paper claims that a large language model instructed to run a Socratic dialogue can stand in for synchronous human deliberation in crowd annotation, improving accuracy and confidence while avoiding the coordination costs of real-time di
desk verdict Useful new system, but the accuracy win over synchronous deliberation is not yet proven—the paper's own failure analysis may explain the effect. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the Socratic dialogue loop, specifically the elenchus stage: after the annotator asserts a claim, the LLM asks one question at a time, probes for counter-evidence and hypothetical boundary cases, and only approves the claim if the reasoning holds. The system prompt encodes five steps of the Socratic method, a temperament (humility, respect, joy, mutual understanding), and guardrails that forbid outside knowledge, limit responses to three sentences, and require a minimum of two rounds. The wording of the prompt matters: it is the entire mechanism that turns a general-purpose chat model (Claude 3 Haiku) into a deliberation partner, and the paper reports that six prompt iterations
What would settle it
Take the same 40 datapoints, recruit one pool of annotators, and randomly assign each datapoint to either Socratic-LLM deliberation or synchronous human deliberation, with identical instructions, the same number of annotators per datapoint, and identical pre/post confidence scales. If the Relation post-deliberation accuracy gap (64.79% vs 48.86%) does not reproduce in that matched design, the paper's central comparison is confounded; additionally, auditing conversation logs for any use of outside factual knowledge would test whether the accuracy gain came from the LLM leaking answers rather th
Extended reading notes
Core claim
The paper's central discovery is that a constrained Socratic LLM outperforms a synchronous human-deliberation benchmark on key annotation metrics. In the Relation task, participants changed labels more often (23.85% annotation-level flips versus 7.63%), and those changes moved overall accuracy from 52.52% to 64.79%, whereas the benchmark moved only from 47.16% to 48.86%; the improvement came mostly from correcting over-eager 'Expressed' relations. Confidence also increased: post-deliberation, 85.34% of annotations were marked high confidence versus 66.39% in the benchmark, with 28% of annotations moving from medium to high confidence. Qualitatively, the LLM took on four roles—argument evalua
Load-bearing premise
The headline results stand only if the asynchronous Socratic study and the synchronous human benchmark are genuinely comparable in participant pools, task instructions, and annotation counts; if they are not, the accuracy and confidence differences are confounded.
Editorial extensions
If this is right
- Asynchronous Socratic deliberation can replace synchronous human deliberation for at least some annotation tasks, cutting coordination costs while improving or matching outcomes.
- On the Relation task, deliberation with the Socratic LLM raised post-deliberation accuracy to 64.79% from 52.52%, beating the benchmark's 48.86% and mostly by correcting false positives.
- Annotator confidence rose substantially: 28% of annotations moved from medium to high confidence, versus no net change in the benchmark.
- Annotators engaged more (7.6 vs 5.4 messages; 104.7 vs 75.3 characters per message), suggesting the format sustains deliberation depth.
- The system's Socratic roles—argument evaluator, boundary negotiator, cognitive support tool, and validator—give dataset curators a way to inspect reasoning and class boundaries in the resulting conversation logs.
Reading between the lines
- An untested extension of the role analysis: deliberately switching the LLM's persona (evaluator, boundary negotiator, cognitive support, validator) across dialogue states could improve deliberation, but would place new demands on guardrails.
- The confidence metric could be repurposed: on subjective tasks without ground truth, a 28% medium-to-high confidence shift may serve as a quality proxy where accuracy is unavailable.
- The asymmetry in Relation flips (18.46% Expressed to Not Expressed vs 4.41% in the benchmark) raises a question the paper does not settle: whether Socratic questioning biases annotators toward a more conservative reading of relations.
- The paper's comparison suggests a direct test: a matched between-subjects study with identical participant pools and annotation counts would reveal whether the accuracy gain is attributable to the Socratic mechanism itself rather than to differences between the two studies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a Socratic LLM dialogue system that replaces a human deliberation partner during crowd annotation, allowing asynchronous reflection on binary labels. The authors benchmark against Schaekermann et al. (2018), using the same Sarcasm and Relation datasets, and report that their intervention produces more annotation flips on the Relation task, higher post-deliberation accuracy on 21 ground-truth Relation datapoints (64.79% vs. 48.86%), increased annotator confidence, and longer/more numerous discussion messages. They also present a qualitative analysis of LLM roles (argument evaluator, boundary negotiator, cognitive support tool, validator) and a candid discussion of failures, including LLM misrepresentations of the Relation task.
Significance. If its central claims hold, the paper makes a useful contribution to perspectivist data annotation by showing that a constrained Socratic LLM can partially substitute for costly synchronous deliberation, with the release of the full system prompt and the use of the benchmark's public dataset enabling replication. The qualitative analysis of how the LLM supports argumentation and the explicit treatment of failures are strengths, as is the authors' willingness to publish a complete, human-readable prompt (Appendix A). However, the main quantitative evidence is currently not strong enough to support the headline accuracy and confidence claims: the comparison is to a prior study rather than a matched control, the decisive accuracy effect is small (roughly 9 annotations), and a known LLM error is aligned with the observed flip direction. The paper's value is more persuasive as a design exploration and qualitative study than as a demonstration of improved annotation accuracy.
major comments (4)
- [§5.2, Table 2] The headline accuracy improvement is not established statistically. The increase from 52.52% to 64.79% (a 12.27 pp gain on n=71 annotations, i.e., about 8.7 flips) is reported without confidence intervals, a significance test, or an effect size. The comparison to the benchmark's n=1003 annotations is a cross-study comparison with different participant pools, task presentation, and annotation counts per datapoint, so the apparent advantage could reflect population or design differences. The text also contains an internal inconsistency: pre-deliberation accuracy is given as 53.52% in one sentence and 52.52% in the next. Please report a binomial confidence interval for the post-deliberation accuracy, test the pre/post change, and clearly state the limitations of the cross-study comparison.
- [§5.7.1, Table 2, Fig. 3] The accuracy gain may be attributable to the LLM's task misrepresentation rather than to Socratic deliberation. Section 5.7.1 states that the LLM 'occasionally' misrepresented the Relation task, enforcing the incorrect rule that relations must be explicitly stated rather than implied, and that this 'led to some annotation flips.' The dominant flip direction in Fig. 3 (Expressed→Not Expressed: 18.46% vs. 4.41% in the benchmark) and the drop in false positives in Table 2 (35.21%→22.54%) account for essentially the entire accuracy improvement. Given that only about 9 decisive flips are needed for the headline gain, even a handful of misrepresentation-driven flips could explain the result. The paper reports neither the number of such flips nor a sensitivity analysis excluding or reclassifying them. This is a load-bearing gap.
- [§4.2.3, §6.2, footnote 12] The re-annotation phase introduces a 'Not Sure' option that was absent in the first annotation phase, and the analysis excludes flips to 'Not Sure' in both conditions. This is a deliberate design choice with a perspectivist rationale, but it changes the response surface between pre- and post-deliberation and may inflate flip rates or confidence changes. The benchmark's 'Irresolvable' category is not identical to 'Not Sure,' so the exclusion is not symmetric. Please report how many annotations were affected, and show that the main flip-rate and accuracy conclusions are robust to including these responses or treating them as missing rather than excluded.
- [§5.5.4, Appendix A] The 'Validator' role is described as emergent, but it is directly instructed in the system prompt. Appendix A, step 5, tells the LLM: 'If the reasoning is sound based on the discussion, you should encourage them to continue on to re-annotate the item below their chat.' Similarly, the 'Argument Evaluator' and 'Classification Boundary Negotiator' behaviors closely follow prompt steps 2–4. The qualitative finding is better framed as 'prompt-induced behaviors' rather than emergent roles. This does not invalidate the qualitative analysis, but the current framing overstates the serendipity of the design.
minor comments (6)
- [Throughout] Statistical notation: p-values are printed as 'p « 0.001' (e.g., §5.1.1, §5.1.2, §5.3). This should be 'p < 0.001.'
- [§5.2] Inconsistent accuracy values: the text says 'Pre-deliberation, our participants had slightly higher accuracy in labeling than the benchmark (53.52% vs. 47.16%),' then later says 'Our overall accuracy increased substantially to 64.79% (from 52.52%).' Please correct the pre-deliberation value.
- [§5.3] The sentence 'The change in confidence (pre-intervention confidence minus post-intervention confidence) was significant...' appears to have the sign reversed for the reported increase. Clarify whether the t-test was computed on post−pre.
- [§5.1.1, Fig. 3] The Sankey diagrams are informative, but the figure does not show marginal totals. Adding per-arrow counts or percentages would make the comparison of flip asymmetry easier to verify.
- [§4.3] The authors report excluding participants who used outside LLMs or completed tasks too quickly, but do not state how many participants were excluded for each reason. Reporting these numbers would help assess data quality.
- [§5.2] The comparison group (n=1003 annotation-level observations from the benchmark) is much larger than the treatment group (n=71). While this is a consequence of the benchmark design, the authors should note that the large n gives the benchmark narrow confidence intervals and that the comparison is underpowered on the treatment side.
Circularity Check
Quantitative accuracy/confidence comparisons are measured and self-contained; the qualitative 'emergent roles' are partly prompt instructions relabeled as discoveries.
-
self definitional
[Section 5.5 ('LLM Roles for Supporting Data Annotation') and Appendix A prompt steps 2/5; contradicted by Section 7 limitation sentence]
"Through our qualitative analysis of the conversation logs ... we identified that the Socratic LLM’s adherence to the prompt instructions (provided in Appendix A) cause distinct patterns of LLM behavior to emerge. We present these patterns as emergent roles ... Validator, approving the annotator’s label once sufficient deliberation has occurred. ... [Appendix A step 5:] If the reasoning is sound based on the discussion, you should encourage them to continue on to re-annotate the item below their chat. ... [Section 7:] These roles were largely effective and appreciated by participants, despite n"
The claimed 'emergent roles' are textually equivalent to the prompt steps that prescribe them. The Validator role (approving a label once reasoning is sound) is exactly Appendix A step 5; the Classification Boundary Negotiator (using counter-examples to find boundaries) is exactly Appendix A step 2 ('ask about how they might categorize counter-examples to help identify boundaries in their logic'). Section 7 asserts the roles were 'not explicitly designed in the prompt instructions,' but they appear verbatim there. Thus the qualitative role taxonomy reduces to a summary of the system prompt by construction, rather than an independent empirical discovery. This is presentational circularity only; the accuracy and confidence outcomes are measured separately and do not depend on this taxonomy.
full rationale
The paper's headline numerical claims are not circular: Relation-task accuracy (52.52% pre vs. 64.79% post, vs. benchmark 47.16% to 48.86%) and confidence changes (28% medium-to-high) come from collected annotations compared against Schaekermann et al.'s public benchmark. No parameter is fitted and renamed a prediction, no prior result by these authors is invoked as a load-bearing uniqueness argument, and the derivation chain for the quantitative findings is external and measured. The only circularity found is presentational: the 'emergent roles' of Section 5.5 are relabeled versions of the system prompt's own Socratic steps (Validator = step 5; Boundary Negotiator = step 2), and the paper's own limitation statement disclaims that they were 'not explicitly designed' when the appendix shows they were. This affects a qualitative framing contribution, not the central accuracy/confidence comparisons. The skeptic concern about LLM task misrepresentation driving flips (§5.7.1) is a serious internal-validity threat, but it is not circularity: the accuracy outcome is still measured, not derived from the prompt by construction. Overall score reflects one minor, non-load-bearing circular framing.
Assumptions & free parameters
free parameters (1)
- Minimum required discussion messages per datapoint =
2
assumptions (5)
- domain assumption The prior synchronous-deliberation study by Schaekermann et al. [102] is a valid baseline for comparison despite differences in participant pool, task conditions, and data size.
- domain assumption Self-reported confidence measures genuine certainty, not social desirability or the effect of the LLM's supportive temperament.
- domain assumption The 21 Relation datapoints with ground truth are representative and their ground-truth labels are correct.
- standard math Statistical tests treat individual annotations as independent observations, although two annotations come from the same participant.
- ad hoc to paper The 'Not Sure' re-annotation option added only in the post-deliberation phase is a valid design, and flips to this option can be excluded from analysis.
Cite this review
Pith. "Pith review of Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist Data." pith.science (2026). https://pith.science/paper/7DVLTFCS
@misc{pith2026250809911,
author = {Pith},
title = {Pith review of: Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist Data},
year = {2026},
howpublished = {\url{https://pith.science/paper/7DVLTFCS}},
note = {Machine review of arXiv:2508.09911}
}
read the original abstract
Data annotation underpins the success of modern AI, but the aggregation of crowd-collected datasets can harm the preservation of diverse perspectives in data. Difficult and ambiguous tasks cannot easily be collapsed into unitary labels. Prior work has shown that deliberation and discussion improve data quality and preserve diverse perspectives -- however, synchronous deliberation through crowdsourcing platforms is time-intensive and costly. In this work, we create a Socratic dialog system using Large Language Models (LLMs) to act as a deliberation partner in place of other crowdworkers. Against a benchmark of synchronous deliberation on two tasks (Sarcasm and Relation detection), our Socratic LLM encouraged participants to consider alternate annotation perspectives, update their labels as needed (with higher confidence), and resulted in higher annotation accuracy (for the Relation task where ground truth is available). Qualitative findings show that our agent's Socratic approach was effective at encouraging reasoned arguments from our participants, and that the intervention was well-received. Our methodology lays the groundwork for building scalable systems that preserve individual perspectives in generating more representative datasets.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
[n. d.]. Socratic Methods - Wikiversity — en.wikiversity.org. https://en.wikiversity.org/wiki/Socratic_Methods. [Ac- cessed 2024-10-25]
2024
- [2]
-
[3]
Erfan Al-Hossami, Razvan Bunescu, Justin Smith, and Ryan Teehan. 2024. Can Language Models Employ the Socratic Method? Experiments with Code Debugging. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE 2024) . Association for Computing Machinery, New York, NY, USA, 53–59. doi:10. 1145/3626252.3630799
-
[4]
Reham Al Tamime, Joni Salminen, Soon-Gyo Jung, and Bernard Jansen. 2024. Evaluating LLM-Generated Topics from Survey Responses: Identifying Challenges in Recruiting Participants through Crowdsourcing. In 2024 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) . IEEE, 412–416
2024
-
[5]
Maya Aloni and Christine Harrington. 2018. Research Based Practices for Improving the Effectiveness of Asyn- chronous Online Discussion Boards. Scholarship of Teaching and Learning in Psychology 4, 4 (Dec. 2018), 271–289. doi:10.1037/stl0000121
-
[6]
Omar Alonso and Stefano Mizzaro. 2012. Using crowdsourcing for TREC relevance assessment. Information process- ing & management 48, 6 (2012), 1053–1066
2012
-
[7]
Paul André, Aniket Kittur, and Steven P Dow. 2014. Crowd synthesis: Extracting categories and clusters from complex data. In Proceedings of the 17th ACM conference on Computer supported cooperative work & social computing. 989–998
2014
-
[8]
Lora Aroyo and Chris Welty. 2015. Truth Is a Lie: Crowd Truth and the Seven Myths of Human Annotation. AI Magazine 36, 1 (March 2015), 15–24. doi:10.1609/aimag.v36i1.2564
Show all 127 references
-
[9]
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of sto- chastic parrots: Can language models be too big?. InProceedings of the 2021 ACM conference on fairness, accountability, and transparency. 610–623
2021
-
[10]
Karim Benharrak, Tim Zindulka, Florian Lehmann, Hendrik Heuer, and Daniel Buschek. 2024. Writer-Defined AI Personas for On-Demand Feedback Generation. In Proceedings of the 2024 CHI Conference on Human Factors in Com- puting Systems (Honolulu, HI, USA) (CHI ’24). Association f...
2024
-
[11]
Reuben Binns, Michael Veale, Max Van Kleek, and Nigel Shadbolt. 2017. Like Trainer, Like Bot? Inheritance of Bias in Algorithmic Content Moderation. In Social Informatics (Cham, 2017), Giovanni Luca Ciampaglia, Afra Mashhadi, and Taha Yasseri (Eds.). Springer International Pub...
2017 doi
-
[12]
Eugenia Arazo Boa, Amornrat Wattanatorn, and Kanchit Tagong. 2018. The Development and Validation of the Blended Socratic Method of Teaching (BSMT): An Instructional Model to Enhance Critical Thinking Skills of Under- graduate Business Students. 39, 1 (2018), 81–89. doi:10.101...
2018 doi
-
[13]
Jonathan Bragg, Mausam, and Daniel S. Weld. 2018. Sprout: Crowd-Powered Task Design for Crowdsourcing. In Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology (UIST ’18) . Association for Computing Machinery, New York, NY, USA, 165–176. doi:10...
2018
-
[14]
Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101
2006
-
[15]
Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. 5 (2021), 188:1–188:21. Issue CSCW1. doi:10.1145/ 3449287
2021
-
[16]
Federico Cabitza, Andrea Campagner, and Valerio Basile. 2023. Toward a Perspectivist Turn in Ground Truthing for Predictive Computing. Proceedings of the AAAI Conference on Artificial Intelligence 37, 6 (June 2023), 6860–6868. doi:10.1609/aaai.v37i6.25840
2023 doi
-
[17]
Carrie J Cai, Shamsi T Iqbal, and Jaime Teevan. 2016. Chain reactions: The impact of order on microtask chains. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems . 3143–3154
2016
-
[18]
Scott Allen Cambo and Darren Gergle. 2022. Model Positionality and Computational Reflexivity: Promoting Reflex- ivity in Data Science. In CHI Conference on Human Factors in Computing Systems . ACM, New Orleans LA USA, 1–19. doi:10.1145/3491102.3501998 Proc. ACM Hum.-Comput. In...
2022
-
[19]
Joseph Chee Chang, Saleema Amershi, and Ece Kamar. 2017. Revolt: Collaborative Crowdsourcing for Labeling Machine Learning Datasets. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems (CHI ’17). Association for Computing Machinery, New York, NY, US...
2017
-
[20]
Adriane Chapman, Philip Grylls, Pamela Ugwudike, David Gammack, and Jacqui Ayling. 2022. A Data-driven Anal- ysis of the Interplay between Criminological Theory and Predictive Policing Algorithms. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Trans...
2022
-
[21]
Ana Paula Chaves and Marco Aurelio Gerosa. 2021. How should my chatbot interact? A survey on social charac- teristics in human–chatbot interaction design. International Journal of Human–Computer Interaction 37, 8 (2021), 729–758
2021
-
[22]
Quanze Chen, Jonathan Bragg, Lydia B Chilton, and Dan S Weld. 2019. Cicero: Multi-turn, contextual argumentation for accurate crowdsourcing. In Proceedings of the 2019 chi conference on human factors in computing systems . 1–14
2019
-
[23]
Weld, and Amy X
Quan Ze Chen, Daniel S. Weld, and Amy X. Zhang. 2021. Goldilocks: Consistent Crowdsourced Scalar Annotations with Relative Uncertainty. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (Oct. 2021), 335:1– 335:25. doi:10.1145/3476076
2021 doi
-
[24]
Quan Ze Chen and Amy X. Zhang. 2023. Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus Disagreement. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2 (Oct. 2023), 283:1–283:26. doi:10.1145/3610074
2023 doi
-
[25]
Julian Chingoma and Adrian Haret. 2023. Deliberation as Evidence Disclosure: A Tale of Two Protocol Types. In Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems (AAMAS ’23) . Inter- national Foundation for Autonomous Agents and Multiag...
2023
-
[26]
Florian Daniel, Pavel Kucherbaev, Cinzia Cappiello, Boualem Benatallah, and Mohammad Allahbakhsh. 2018. Quality Control in Crowdsourcing: A Survey of Quality Attributes, Assessment Techniques, and Assurance Actions. 51, 1 (2018), 7:1–7:40. doi:10.1145/3148148
2018 doi
-
[27]
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran. 2022. Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations. Transactions of the Association for Computational Linguistics 10 (Jan. 2022), 92–110. doi:10.1162/tacl_a_00449
2022 doi
-
[28]
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated hate speech detection and the problem of offensive language. InProceedings of the international AAAI conference on web and social media, Vol. 11. 512–515
2017
-
[29]
Valerio De Stefano. 2015. The Rise of the ’Just-in-Time Workforce’: On-Demand Work, Crowd Work and Labour Protec- tion in the ’Gig-Economy’ . Social Science Research Network:2682602 doi:10.2139/ssrn.2682602
2015 doi
-
[30]
Amanda Delaney, Bella Lough, Michelle Whelan, Max Cameron, et al. 2004. A review of mass media campaigns in road safety. Monash University Accident Research Centre Reports 220 (2004), 85
2004
-
[31]
Haris Delić and Senad Bećirović. 2016. Socratic Method as an Approach to Teaching. European Researcher 111, 10 (Oct. 2016). doi:10.13187/er.2016.111.511
2016 doi
-
[32]
Yuyang Ding, Hanglei Hu, Jie Zhou, Qin Chen, Bo Jiang, and Liang He. 2024. Boosting Large Language Models with Socratic Method for Conversational Mathematics Teaching. arXiv:2407.17349 [cs]
2024 arXiv
-
[33]
Carl DiSalvo. 2012. Adversarial Design. The MIT Press. doi:10.7551/mitpress/8732.001.0001
2012 doi
-
[34]
Ryan Drapeau, Lydia Chilton, Jonathan Bragg, and Daniel Weld. 2016. Microtalk: Using argumentation to improve crowdsourcing accuracy. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , Vol. 4. 32–41
2016
-
[35]
Such, Mark Coté, and Natalia Criado
Xavier Ferrer, prefix=van useprefix=false family=Nuenen, given=Tom, Jose M. Such, Mark Coté, and Natalia Criado
-
[36]
Elena Filatova. 2012. Irony and Sarcasm: Corpus Generation and Analysis Using Crowdsourcing.. In Lrec. Citeseer, 392–398
2012
- [37]
-
[38]
Soon Yen Foo and Choon Lang Quek. 2019. Developing Students’ Critical Thinking through Asynchronous Online Discussions: A Literature Review. Malaysian Online Journal of Educational Technology 7, 2 (2019), 37–58. doi:10. 17220/mojet.2019.02.003
2019
-
[39]
Giles Foody, Linda See, Steffen Fritz, Inian Moorthy, Christoph Perger, Christian Schill, and Doreen Boyd. 2018. Increasing the accuracy of crowdsourced information on land cover via a voting procedure weighted by information inferred from the contributed data. ISPRS Internati...
2018
-
[40]
Benoît Frénay and Michel Verleysen. 2014. Classification in the Presence of Label Noise: A Survey.IEEE Transactions on Neural Networks and Learning Systems 25, 5 (May 2014), 845–869. doi:10.1109/TNNLS.2013.2292894 Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW52...
2014
-
[41]
Simona Frenda, Gavin Abercrombie, Valerio Basile, Alessandro Pedrani, Raffaella Panizzon, Alessandra Teresa Cignarella, Cristina Marco, and Davide Bernardi. 2024. Perspectivist approaches to natural language processing: a survey. Language Resources and Evaluation (2024), 1–28
2024
-
[42]
Gordon, Michelle S
Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel, Jeff Hancock, Tatsunori Hashimoto, and Michael S. Bernstein. 2022. Jury Learning: Integrating Dissenting Voices into Machine Learning Models. In CHI Conference on Human Factors in Computing Systems . ACM, New Or...
2022
-
[43]
Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S
Mitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S. Bernstein. 2021. The Disagree- ment Deconvolution: Bringing Machine Learning Performance Metrics In Line With Reality. InProceedings of the 2021 CHI Conference on Human Factors in Computing Syst...
2021
-
[44]
Tanya Goyal, Tyler McDonnell, Mucahid Kutlu, Tamer Elsayed, and Matthew Lease. 2018. Your Behavior Signals Your Reliability: Modeling Crowd Behavioral Traces to Ensure Quality Relevance Annotations. 6 (2018), 41–49. doi:10.1609/hcomp.v6i1.13331
2018 doi
-
[45]
Hui Guo, Boyu Wang, and Grace Yi. 2023. Label Correction of Crowdsourced Noisy Annotations with an Instance- Dependent Noise Transition Model. 36 (2023), 347–386. https://proceedings.neurips.cc/paper_files/paper/2023/hash/ 015a8c69bedcb0a7b2ed2e1678f34399-Abstract-Conference.html
2023
-
[46]
Margeret Hall, Mohammad Farhad Afzali, Markus Krause, and Simon Caton. 2022. What Quality Control Mecha- nisms Do We Need for High-Quality Crowd Work? 10 (2022), 99709–99723. doi:10.1109/ACCESS.2022.3207292
2022
-
[47]
Jawad Haqbeen, Takayuki Ito, Rafik Hadfi, Tomohiro Nishida, Zoia Sahab, Sofia Sahab, Shafiq Roghmal, and Moham- mad Amiryar. 2020. Promoting Discussion with AI-based Facilitation: Urban Dialogue with Kabul City
2020
-
[48]
Sandra G Hart. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. Human mental workload/Elsevier (1988)
1988
-
[49]
Yueh-Ren Ho, Bao-Yu Chen, and Chien-Ming Li. 2023. Thinking More Wisely: Using the Socratic Method to Develop Critical Thinking Skills amongst Healthcare Students. 23, 1 (2023), 173. doi:10.1186/s12909-023-04134-2
2023 doi
-
[50]
James Hollan, Edwin Hutchins, and David Kirsh. 2000. Distributed cognition: toward a new foundation for human- computer interaction research. ACM Transactions on Computer-Human Interaction (TOCHI) 7, 2 (2000), 174–196
2000
-
[51]
All of the White People Went First
Mo Houtti, Moyan Zhou, Loren Terveen, and Stevie Chancellor. 2023. " All of the White People Went First": How Video Conferencing Consolidates Control and Exacerbates Workplace Bias. Proceedings of the ACM on Human- Computer Interaction 7, CSCW1 (2023), 1–25
2023
-
[52]
Popescu, Saurabh Chatterjee, and Thad Starner
Jui-Tse Hung, Christopher Cui, Diana M. Popescu, Saurabh Chatterjee, and Thad Starner. 2024. Socratic Mind: Scal- able Oral Assessment Powered By AI. In Proceedings of the Eleventh ACM Conference on Learning @ Scale (L@S ’24) . Association for Computing Machinery, New York, NY...
2024
-
[53]
Edwin Hutchins. 1995. How a cockpit remembers its speeds. Cognitive science 19, 3 (1995), 265–288
1995
- [54]
-
[55]
Muhammad Okky Ibrohim and Indra Budi. 2019. Multi-label hate speech and abusive language detection in Indone- sian Twitter. In Proceedings of the third workshop on abusive language online . 46–57
2019
-
[56]
Oana Inel, Khalid Khamkham, Tatiana Cristea, Anca Dumitrache, Arne Rutjes, Jelle van der Ploeg, Lukasz Romaszko, Lora Aroyo, and Robert-Jan Sips. 2014. Crowdtruth: Machine-human computation framework for harnessing dis- agreement in gathering annotated data. InThe Semantic Web...
2014
-
[57]
Panagiotis G Ipeirotis and Evgeniy Gabrilovich. 2014. Quizz: targeted crowdsourcing with a billion (potential) users. In Proceedings of the 23rd international conference on World wide web . 143–154
2014
-
[58]
Shamsi T Iqbal and Brian P Bailey. 2005. Investigating the effectiveness of mental workload as a predictor of oppor- tune moments for interruption. In CHI’05 extended abstracts on Human factors in computing systems . 1489–1492
2005
-
[59]
Deniz Iren and Semih Bilgen. 2014. Cost of Quality in Crowdsourcing. 1, 2 (2014). Issue 2. doi:10.15346/hc.v1i2.14
2014 doi
-
[60]
V. K. Chaithanya Manam, Dwarakanath Jampani, Mariam Zaim, Meng-Han Wu, and Alexander J. Quinn. 2019. Task- Mate: A Mechanism to Improve the Quality of Instructions in Crowdsourcing. InCompanion Proceedings of The 2019 World Wide Web Conference (WWW ’19) . Association for Compu...
2019
-
[61]
Shivani Kapania, Alex S Taylor, and Ding Wang. 2023. A hunt for the snark: Annotator diversity in data practices. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–15
2023
-
[62]
Martin F Kaplan and Ana M Martin. 1999. Effects of differential status of group members on process and outcome of deliberation. Group Processes & Intergroup Relations 2, 4 (1999), 347–364
1999
-
[63]
Georgi Karadzhov, Tom Stafford, and Andreas Vlachos. 2023. DeliData: A Dataset for Deliberation in Multi-party Problem Solving. Proc. ACM Hum.-Comput. Interact. 7, CSCW2 (Oct. 2023), 265:1–265:25. doi:10.1145/3610056
2023 doi
- [64]
-
[65]
Harmanpreet Kaur, Alex C Williams, Anne Loomis Thompson, Walter S Lasecki, Shamsi T Iqbal, and Jaime Teevan
-
[66]
Ashish Khetan, Zachary C Lipton, and Anima Anandkumar. 2018. Learning from noisy singly-labeled data. In Pro- ceedings of ICLR 2018
2018
-
[67]
Nickerson, Michael Bernstein, Elizabeth Gerber, Aaron Shaw, John Zimmerman, Matt Lease, and John Horton
Aniket Kittur, Jeffrey V. Nickerson, Michael Bernstein, Elizabeth Gerber, Aaron Shaw, John Zimmerman, Matt Lease, and John Horton. 2013. The Future of Crowd Work. In Proceedings of the 2013 Conference on Computer Supported Cooperative Work. ACM, San Antonio Texas USA, 1301–131...
2013
- [68]
-
[69]
Travis Kriplean, Michael Toomim, Jonathan Morgan, Alan Borning, and Amy J. Ko. 2012. Is This What You Meant?: Promoting Listening on the Web with Reflect. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM, Austin Texas USA, 1559–1568. doi:10.114...
2012
-
[70]
Tzu-Sheng Kuo, Aaron Lee Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein, and Haiyi Zhu. 2024. Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia. InProceedings of the CHI Conference on Human Factors in Computing Systems (C...
2024
-
[71]
Hélène Landemore and Scott E Page. 2015. Deliberation and disagreement: Problem solving, prediction, and positive dissensus. Politics, philosophy & economics 14, 3 (2015), 229–254
2015
-
[72]
Susan Leigh Star. 2010. This is not a boundary object: Reflections on the origin of a concept. Science, technology, & human values 35, 5 (2010), 601–617
2010
-
[73]
Elisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini, and Sara Tonelli. 2021. Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators’ Disagreement. In Proceedings of the 2021 Con- ference on Empirical Methods in Natural Language Proce...
2021 doi
-
[74]
Franklin
Guoliang Li, Jiannan Wang, Yudian Zheng, and Michael J. Franklin. 2016. Crowdsourced Data Management: A Survey. 28, 9 (2016), 2296–2319. doi:10.1109/TKDE.2016.2535242
2016
-
[75]
Zihao Li. 2023. The dark side of chatgpt: Legal and ethical challenges from stochastic parrots and hallucination. arXiv preprint arXiv:2304.14347 (2023)
2023
-
[76]
Cindy Kaiying Lin and Steven J. Jackson. 2023. From Bias to Repair: Error as a Site of Collaboration and Negotiation in Applied Data Science Work. Proceedings of the ACM on Human-Computer Interaction 7, CSCW1 (April 2023), 131:1–131:32. doi:10.1145/3579607
2023 doi
-
[77]
Jieli Liu and Pengyi Zhang. 2020. How to Initiate a Discussion Thread?: Exploring Factors Influencing Engagement Level of Online Deliberation. In Sustainable Digital Communities , Anneli Sundqvist, Gerd Berget, Jan Nolin, and Kjell Ivar Skjerdingstad (Eds.). Vol. 12051. Spring...
2020 doi
-
[78]
Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, and Sanjiv Kumar. 2020. Does Label Smoothing Mitigate Label Noise?. In Proceedings of the 37th International Conference on Machine Learning (ICML’20, Vol. 119) . JMLR.org, 6448–6458
2020
-
[79]
Fenglong Ma, Yaliang Li, Qi Li, Minghui Qiu, Jing Gao, Shi Zhi, Lu Su, Bo Zhao, Heng Ji, and Jiawei Han. 2015. FaitCrowd: Fine Grained Truth Discovery for Crowdsourced Data Aggregation. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and D...
2015 doi
-
[80]
Ru Ma, Jiachen Zhao, Chenghang Huo, Xiaodong Zhan, and Fuzhi Zhang. 2023. Spammer Groups Detection Based on Hypergraph Embedding And Autoencoder Classifier Model. InProceedings of the 2023 7th International Conference on Electronic Information Technology and Computer Engineeri...
2023 doi
- [81]
-
[82]
Walid Magdy, Kareem Darwish, and Norah Abokhodair. 2015. Quantifying public response towards Islam on Twitter after Paris attacks. arXiv preprint arXiv:1512.04570 (2015)
2015 arXiv
-
[83]
Bernard Manin. 2005. Democratic deliberation: Why we should promote debate rather than discussion. In Paper delivered at the program in ethics and public affairs seminar, Princeton University , Vol. 13. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW526. Publicat...
2005
-
[84]
Samuel Mayworm, Kendra Albert, and Oliver L. Haimson. 2024. Misgendered During Moderation: How Transgender Bodies Make Visible Cisnormative Content Moderation Policies and Enforcement in a Meta Oversight Board Case. In Proceedings of the 2024 ACM Conference on Fairness, Accoun...
2024
-
[85]
David Alvarez Melis, Harmanpreet Kaur, Hal Daumé III, Hanna Wallach, and Jennifer Wortman Vaughan. 2021. From human explanation to model interpretability: A framework based on weight of evidence. InProceedings of the AAAI Conference on Human Computation and Crowdsourcing , Vol...
2021
-
[86]
Tali Mendelberg, Christopher F Karpowitz, and J Baxter Oliphant. [n. d.]. Gender inequality in deliberation: Unpack- ing the black box of interaction. Perspectives on Politics 12, 1 ([n. d.]), 18–44
-
[87]
David Miller. 2018. Is deliberative democracy unfair to disadvantaged groups? In Democracy as Public Deliberation . Routledge, 201–226
2018
-
[88]
Chantal Mouffe. 1999. Deliberative Democracy or Agonistic Pluralism? 66, 3 (1999), 745–758. jstor:40971349 https: //www.jstor.org/stable/40971349
1999
-
[89]
Sania Nayab, Giulio Rossolini, Marco Simoni, Andrea Saracino, Giorgio Buttazzo, Nicolamaria Manes, and Fab- rizio Giacomelli. 2024. Concise thoughts: Impact of output length on llm reasoning and cost. arXiv preprint arXiv:2407.19825 (2024)
2024 arXiv
-
[90]
Stefanie Nowak and Stefan Rüger. 2010. How Reliable Are Annotations via Crowdsourcing: A Study about Inter- Annotator Agreement for Multi-Label Image Annotation. InProceedings of the International Conference on Multimedia Information Retrieval. ACM, Philadelphia Pennsylvania U...
2010
-
[91]
Jeongeon Park, Eun-Young Ko, Yeon Su Park, Jinyeong Yim, and Juho Kim. 2024. DynamicLabels: Supporting In- formed Construction of Machine Learning Label Sets with Crowd Feedback. In Proceedings of the 29th International Conference on Intelligent User Interfaces (IUI ’24). Asso...
2024
-
[92]
Weiping Pei, Arthur Mayer, Kaylynn Tu, and Chuan Yue. 2020. Attention please: Your attention check questions in survey studies can be automatically answered. In Proceedings of The Web Conference 2020 . 1182–1193
2020
-
[93]
Mark Perry. 2003. Distributed cognition. HCI models, theories, and frameworks: Toward a multidisciplinary science (2003), 193–223
2003
- [94]
-
[95]
2013.An Evaluation of Aggregation Techniques in Crowdsourcing
Nguyen Quoc Viet Hung, Nguyen Thanh Tam, Lam Ngoc Tran, and Karl Aberer. 2013.An Evaluation of Aggregation Techniques in Crowdsourcing. Springer Berlin Heidelberg, 1–15. doi:10.1007/978-3-642-41154-0_1
2013 doi
-
[96]
Cyrus Rashtchian, Peter Young, Micah Hodosh, and Julia Hockenmaier. 2010. Collecting Image Annotations Using Amazon’s Mechanical Turk. In Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk (CSLDAMT ’10) . Association f...
2010
-
[97]
Alexander J Ratner, Christopher M De Sa, Sen Wu, Daniel Selsam, and Christopher Ré. 2016. Data programming: Creating large training sets, quickly. Advances in neural information processing systems 29 (2016)
2016
-
[98]
Weingarten, Lilith Fury, Constanza Eliana Chinea, Tuck J
Yim Register, Izzi Grasso, Lauren N. Weingarten, Lilith Fury, Constanza Eliana Chinea, Tuck J. Malloy, and Emma S. Spiro. 2024. Beyond Initial Removal: Lasting Impacts of Discriminatory Content Moderation to Marginalized Cre- ators on Instagram. 8 (2024), 23:1–23:28. Issue CSC...
2024 doi
-
[99]
Everyone Wants to Do the Model Work, Not the Data Work
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021. “Everyone Wants to Do the Model Work, Not the Data Work”: Data Cascades in High-Stakes AI. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Syste...
2021 doi
-
[100]
Marta Sandri, Elisa Leonardelli, Sara Tonelli, and Elisabetta Jezek. 2023. Why Don’t You Do It Right? Analysing Annotators’ Disagreement in Subjective Tasks. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (May 2023),...
2023 doi
-
[101]
Yisi Sang and Jeffrey Stanton. 2022. The Origin and Value of Disagreement Among Data Labelers: A Case Study of Individual Differences in Hate Speech Annotation. In Information for a Better World: Shaping the Global Future (Lecture Notes in Computer Science) , Malte Smits (Ed.)...
2022
-
[102]
Mike Schaekermann, Joslin Goh, Kate Larson, and Edith Law. 2018. Resolvable vs. Irresolvable Disagreement: A Study on Worker Deliberation in Crowd Work. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (Nov. 2018), 1–19. doi:10.1145/3274423
2018 doi
-
[103]
Aashish Sheshadri and Matthew Lease. 2013. Square: A benchmark for research on computing crowd consensus. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , Vol. 1. 156–164
2013
-
[104]
Hedderich, AndréS Lucero, and Antti Oulasvirta
Joongi Shin, Michael A. Hedderich, AndréS Lucero, and Antti Oulasvirta. 2022. Chatbots Facilitating Consensus- Building in Asynchronous Co-Design. In Proceedings of the 35th Annual ACM Symposium on User Interface Software Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Articl...
2022
-
[105]
Susan Leigh Star and James R Griesemer. 1989. Institutional ecology,translations’ and boundary objects: Amateurs and professionals in Berkeley’s Museum of Vertebrate Zoology, 1907-39.Social studies of science 19, 3 (1989), 387–420
1989
-
[106]
Thitaree Tanprasert, Sidney S Fels, Luanne Sinnamon, and Dongwook Yoon. 2024. Debate Chatbots to Facilitate Critical Thinking on YouTube: Social Identity and Conversational Style Make A Difference. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems ...
2024
-
[107]
Dapeng Tao, Jun Cheng, Zhengtao Yu, Kun Yue, and Lizhen Wang. 2018. Domain-weighted majority voting for crowdsourcing. IEEE transactions on neural networks and learning systems 30, 1 (2018), 163–174
2018
-
[108]
Stephen E Toulmin. 2003. The uses of argument . Cambridge university press
2003
-
[109]
Bernstein, and Ranjay Krishna
Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna. 2023. Explanations Can Reduce Overreliance on AI Systems During Decision-Making. 7 (2023), 129:1–129:38. Issue CSCW1. doi:10.1145/3579605
2023 doi
-
[110]
Melanie A Wakefield, Barbara Loken, and Robert C Hornik. 2010. Use of mass media campaigns to change health behaviour. The lancet 376, 9748 (2010), 1261–1271
2010
-
[111]
Stacy E. Walker. 2003. Active Learning Strategies to Promote Critical Thinking. Journal of Athletic Training 38, 3 (2003), 263–267
2003
-
[112]
Shaun Wallace, Tianyuan Cai, Brendan Le, and Luis A. Leiva. 2022. Debiased Label Aggregation for Subjective Crowdsourcing Tasks. In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems (CHI EA ’22). Association for Computing Machinery, New York, ...
2022
-
[113]
Douglas N Walton. 1998. The new dialectic: Conversational contexts of argument . University of Toronto Press
1998
-
[114]
Qi Wang, Mulin Chen, Feiping Nie, and Xuelong Li. 2018. Detecting coherent groups in crowd scenes by multiview clustering. IEEE transactions on pattern analysis and machine intelligence 42, 1 (2018), 46–58
2018
-
[115]
Tharindu Cyril Weerasooriya, Alexander Ororbia, Raj Bhensadadia, Ashiqur KhudaBukhsh, and Christopher Homan
-
[116]
Guy Williams et al. 2014. Harkness learning: principles of a radical American pedagogy. (2014)
2014
-
[117]
Xuansheng Wu, Haiyan Zhao, Yaochen Zhu, Yucheng Shi, Fan Yang, Tianming Liu, Xiaoming Zhai, Wenlin Yao, Jundong Li, Mengnan Du, et al. 2024. Usable XAI: 10 strategies towards exploiting explainability in the LLM era. arXiv preprint arXiv:2403.08946 (2024)
2024 arXiv
-
[118]
Jingru Yang, Ju Fan, Zhewei Wei, Guoliang Li, Tongyu Liu, and Xiaoyong Du. 2018. Cost-effective data annotation using game-based crowdsourcing. Proceedings of the VLDB Endowment 12, 1 (2018), 57–70
2018
-
[119]
Ming Yin and Yiling Chen. 2016. Predicting Crowd Work Quality under Monetary Interventions. 4 (2016), 259–268. doi:10.1609/hcomp.v4i1.13282
2016 doi
-
[120]
annotator rationales
Omar Zaidan, Jason Eisner, and Christine Piatko. 2007. Using “annotator rationales” to improve machine learning for text categorization. In Human language technologies 2007: The conference of the North American chapter of the association for computational linguistics; proceedi...
2007
-
[121]
Angie Zhang, Olympia Walker, Kaci Nguyen, Jiajun Dai, Anqing Chen, and Min Kyung Lee. 2023. Deliberating with AI: Improving Decision-Making for the Future through Participatory AI Design and Stakeholder Deliberation. Proc. ACM Hum.-Comput. Interact. 7, CSCW1 (April 2023), 125:...
2023 doi
-
[122]
Scheng, Tao Li, and Xindong Wu
Jing Zhang, Victor S. Scheng, Tao Li, and Xindong Wu. 2017. Improving Crowdsourced Label Quality Using Noise Correction. 29, 5 (2017), 1675–1688. doi:10.1109/TNNLS.2017.2677468
2017
-
[123]
Honglei Zhuang and Joel Young. 2015. Leveraging in-batch annotation bias for crowdsourced active learning. In Proceedings of the Eighth ACM International Conference on Web Search and Data Mining . 243–252
2015
-
[124]
I can’t provide any additional information outside what was given for the task. You should use your own knowledge and experience to help inform your choice
Öznur Göçmen and Hamit Coşkun. 2019. The effects of the six thinking hats and speed on creativity in brainstorming. Thinking Skills and Creativity 31 (2019), 284–295. doi:10.1016/j.tsc.2019.02.006 Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW526. Publication da...
2019 doi
-
[2018]
Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 1–22
Creating better action plans for writing tasks via vocabulary-based planning. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 1–22
2018
-
[2021]
40, 2 (2021), 72–80
Bias and Discrimination in AI: A Cross-Disciplinary Perspective. 40, 2 (2021), 72–80. doi:10.1109/MTS.2021. 3056293
2021 doi
-
[2023]
In Findings of the Association for Computational Linguistics: ACL 2023
Disagreement matters: Preserving label diversity by jointly modeling item and annotator label distributions with DisCo. In Findings of the Association for Computational Linguistics: ACL 2023 . 4679–4695
2023
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.