REVIEW 3 major objections 4 minor 54 references
Data Annotations as Pedagogical Hints: From Subjective Labels to Critical Thinking
T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Manual data annotation teaches students that human judgment shapes AI labels and that disagreement reflects domain complexity, not just noise.
desk verdict An honest exploratory study whose own qualitative data undercut the abstract's central claim — worth engaging for the consensus-illusion result, but the paper needs a major reframing before publication. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the annotation exercise itself, designed around a deliberately ambiguous 3-point scale (none/some/a lot of hair) applied to the same images by all group members. Comparing labels within groups and computing a standard agreement statistic (how often annotators give the same score) surfaces disagreement that a single-ground-truth mindset cannot explain. Conceptually, the exercise operationalizes the shift from an 'agreement paradigm'—where disagreement is noise—to a 'disagreement paradigm' in which multiple valid labels coexist; the paper calls the resulting labels 'hints' rather than ground truth. The learning-persona analytical framework translates this shift into me
What would settle it
A controlled replication in which the same images are annotated under deliberately vague versus detailed, illustrated instructions would settle it: if disagreement largely vanishes under detailed instructions, the phenomenon is task ambiguity, not subjectivity; if substantial disagreement persists even among expert annotators given identical guidelines, the subjectivity reading is supported.
Extended reading notes
Core claim
The paper's discovery is that assigning students to manually annotate an inherently subjective visual feature—hair coverage in skin-lesion images—makes the constructed nature of training data tangible. Across two university courses, 43 students who annotated in groups reported sharp increases in self-reported familiarity with annotation subjectivity, inter-annotator agreement measures, bias, and data quality, and most said the activity beat lectures for understanding bias. The paper's sharper finding, however, is that this awareness does not automatically turn into acceptance: when asked how to improve the task, students most often requested clearer guidelines and calibration to remove disag
Load-bearing premise
The central finding rests on the assumption that the moderate inter-annotator agreement (around 0.62) reflects genuine interpretive subjectivity rather than under-specified instructions—if students disagree mainly because the guidelines are unclear, then requesting clearer guidelines is a sensible request, not a sign of an unlearned lesson.
Editorial extensions
If this is right
- Data-annotation exercises can be embedded in machine-learning courses as a low-cost, first-hand route to teaching where bias enters AI: at the data-creation stage, not only in the model.
- Educators should explicitly frame disagreement as a learning feature; the study's own data show that experiencing ambiguity without that framing leaves many students seeking to eliminate it.
- Computing agreement statistics alone does not shift students' mental model: despite moderate measured agreement, groups reported consensus, indicating a 'consensus illusion' that needs direct pedagogical attention.
- Course designs should reduce repetitive labeling (a common complaint) while preserving group discussion and analysis, which students rated as the most enjoyable and instructive part.
- When using sensitive material such as medical images, instructors must plan for emotional discomfort, the most frequent negative theme, so it does not overshadow the lesson.
Reading between the lines
- The 'consensus illusion' could be investigated as a metacognitive construct in its own right—recording group discussions would reveal whether perceived agreement stems from metric confusion, social conformity, or a lingering single-truth belief.
- The mechanistic claim can be transferred to non-medical domains with meaningful disagreement (e.g., sentiment, content moderation, essay scoring), where the emotional-discomfort confound is absent and learning gains could be isolated more cleanly.
- If LLMs take over annotation, the demonstrated pedagogical value may be bypassed; a concrete alternative is having students critique or adjudicate LLM-generated labels, which may preserve the subjectivity lesson while removing drudgery.
- Because familiarity was measured by retrospective self-report, the durability of the shift is untested; a delayed follow-up in which students design their own labeling guidelines would show whether the conceptual change persists.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes that having students manually annotate medical images can teach them about the subjectivity of human labels, bias in datasets, and the idea that annotator disagreement can reflect genuine domain complexity rather than mere noise. The study implements a skin-lesion hair-coverage annotation task in two university courses (Fontys and ITU Copenhagen), collects surveys from 43 students, computes Fleiss' kappa for 23 student groups, and analyzes open-ended responses with thematic coding. The authors report self-perceived gains in familiarity with subjectivity, data quality, and bias, and interpret moderate inter-annotator agreement as evidence that the task exhibited the intended interpretive ambiguity. They also report a central tension: despite recognizing subjectivity, many students requested clearer guidelines or calibration to reduce disagreement, suggesting that the intended conceptual shift was only partially achieved. The paper closes with design recommendations and proposes a 'learning persona' framework for analyzing student responses.
Significance. If the findings are taken as an exploratory case study, the work addresses a timely and underexplored pedagogical question: whether annotation tasks can bridge the gap between modern disagreement-aware annotation research and AI education. The cross-institutional implementation, the public release of data/code, and the transparent discussion of a negative or partially negative result are strengths. The 'consensus illusion' observation—students reporting consensus despite only moderate measured agreement—is a useful, falsifiable phenomenon that could inspire follow-up work. The main contribution is a design insight rather than a proven method: annotation activities can raise awareness of subjectivity, but explicit scaffolding is needed if the goal is to teach students to value disagreement as information.
major comments (3)
- [Abstract and Section 5.1.5] The abstract claims that manual annotation 'effectively teach[es] students that ... disagreement reflects domain complexity, not just noise.' This is contradicted by the paper's own data. In Section 4.4.2, the most frequent improvement suggestion was to make annotations more consistent (n=20), while Section 4.5 reports that only n=2 explicitly described subjectivity as beneficial. Section 5.1.5 concedes that 'participants still search for the one correct answer... treat disagreement as a problem to be solved.' The reframing as a 'productive pedagogical tension' (Section 5) is reasonable, but the abstract and the final sentence of the introduction overstate the effect. The central claim should be rephrased to distinguish 'raising awareness of subjectivity' (supported) from 'teaching students to reframe disagreement as valuable' (not supported by the majority of responses).
- [Section 4.3] The statement that 'the moderate level of agreement indicates that the intended degree of subjectivity is indeed present in the dataset' is not justified. The task uses a 3-point ordinal scale for hair coverage, and disagreement among student annotators could equally result from under-specified instructions, ambiguous operationalization of 'some' versus 'a lot,' or insufficient training. The paper itself acknowledges in Section 5.2.1 that students must not mistake misunderstandings for subjectivity, but no evidence is presented to distinguish these sources. This is load-bearing because the pedagogical value of the disagreement depends on it being genuine interpretive divergence rather than task ambiguity.
- [Section 3.3 and Table 3] The central evidence for learning is a retrospective pre/post self-report without a control group, baseline measurement, or inferential statistics. The authors acknowledge these limitations in Section 5.3, but the abstract and Section 5.1.1 use causal language ('effectively teach,' 'facilitates a deeper understanding'). For an exploratory case study, this is acceptable if the claims are explicitly framed as perceived, self-reported gains. The abstract should be tempered accordingly, or the study should include at least a non-annotation comparison condition to support the causal wording.
minor comments (4)
- [Throughout] The text contains several typos and inconsistent typography: 'Pedagocical effectiviness' (Table 2), 'sufficient' and 'efficiency' (Sections 5.2.1 and elsewhere), 'F airness' (Section 5.1.1 heading), and 'T ask' in Figure 1 caption. A careful proofread is needed.
- [Figure 5] The heatmap is difficult to interpret as printed; row labels are cramped and several numeric entries appear garbled or overlapping. The figure would benefit from a larger layout or a table of frequencies.
- [Section 2.3] The claim that 'no prior work ... systematically investigated how students conceptualise annotation subjectivity' is strong; consider softening it to 'to our knowledge, no prior work in this specific setting' and citing any adjacent work in critical data studies or human-centered AI education.
- [Section 3.4.2] The relaxed F1 of 0.71 for the qualitative coding is reported, but the coding scheme is not included. Since the central qualitative claims (e.g., n=20 for consistent-annotation suggestions) depend on this coding, including the code definitions or a sample of coded responses would improve transparency.
Circularity Check
No significant circularity: the paper's claims rest on survey and annotation data, not on a derivation that reduces to its inputs.
full rationale
This is an empirical education study with no mathematical derivation chain, so the main circularity patterns (fitted parameters renamed as predictions, uniqueness theorems, ansatz smuggling) do not apply. The central claim that manual annotation teaches students about subjectivity is supported by self-reported survey data, qualitative coding, and Fleiss' kappa; it is not defined in terms of the conclusion. The 'learning persona' is an explicitly stated analytical ideal, not a quantity fitted from the data, so judging student responses against it is a substantive comparison rather than a tautology. The closest candidate for circularity is Section 4.3's statement that 'The moderate level of agreement indicates that the intended degree of subjectivity is indeed present in the dataset,' but this is an interpretive claim about measured inter-annotator agreement, not a definitional equivalence; moreover, the paper itself flags the competing explanation in Section 5.2.1, where task designers must ensure 'students do not mistake misunderstandings resulting from missing annotation guidelines or task descriptions for subjectivity.' That is a validity threat, not a circularity. The self-citations ([8], [42], [43]) provide background on disagreement being informative and a recommendation to close the annotation-model loop; none is load-bearing for the measured outcomes. The abstract's 'effectively teach' may overstate the qualitative finding that most students sought to eliminate disagreement, but that is an evidentiary/interpretation concern, not a circular derivation. Overall, the paper is self-contained against its own empirical benchmarks and exhibits no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Self-reported familiarity and perceived learning are valid proxies for actual learning of subjectivity concepts.
- domain assumption Moderate inter-annotator disagreement (mean Fleiss' kappa=0.62) arises from genuine interpretive subjectivity, not task or instruction ambiguity.
- ad hoc to paper The 'learning persona' trajectory is an appropriate normative framework for judging student responses.
- domain assumption Thematic analysis with two coders (relaxed F1=0.71) adequately captures the content of open-ended responses.
invented entities (2)
-
Learning persona
-
Consensus illusion
Cite this review
Pith. "Pith review of Data Annotations as Pedagogical Hints: From Subjective Labels to Critical Thinking." pith.science (2026). https://pith.science/paper/EXQRDBGT
@misc{pith2026260720149,
author = {Pith},
title = {Pith review of: Data Annotations as Pedagogical Hints: From Subjective Labels to Critical Thinking},
year = {2026},
howpublished = {\url{https://pith.science/paper/EXQRDBGT}},
note = {Machine review of arXiv:2607.20149}
}
read the original abstract
Machine learning courses often use pre-labeled datasets, hiding the subjectivity of human annotation. This creates students with an overly trusting view of AI data and models, undervaluing interpretive diversity. We investigated whether manual data annotation tasks teach students about subjective labeling. Study Design: An annotation activity was implemented at two universities: Fontys (Netherlands) and IT University Copenhagen (Denmark). Students annotated skin lesion images for hair coverage on a 3-point scale. Surveys were collected from 43 participants measuring their understanding of annotation ambiguity, data quality, bias, fairness, implementation barriers, and pedagogical effectiveness. Key Findings: Self-reported familiarity with course content increased substantially across all concepts. Most students recognised that personal interpretation affects annotations. Students rated the activity as more effective than traditional lectures for understanding bias. Participants were motivated to learn more. Main Drawbacks: Emotional unease from viewing medical images was the primary issue. Many students still requested clearer guidelines to reduce disagreement, suggesting they hadn't internalised that disagreement from different perspectives is a learning feature, not a bug. Recommendations for Future Iterations: Ensure sufficient interpretive ambiguity in materials. Reduce repetitive annotation workload. Mitigate emotional unease from sensitive content. Explicitly frame disagreement as a learning opportunity rather than a problem to solve. Manual data annotations effectively teach students that human judgment shapes model behavior and that disagreement reflects domain complexity, not just noise.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Exploring lan- guage representation through a resource in- ventory project
Carolyn Jane Anderson. Exploring lan- guage representation through a resource in- ventory project. In Sana Al-azzawi, Laura Biester, György Kovács, Ana Marasović, Leena Mathur, Margot Mieskes, and Leonie Weissweiler, editors, Proceedings of the Sixth Workshop on Teaching NLP , pages 91–93, Bangkok, Thailand, August 2024. Associa- tion for Computational Li...
2024
-
[2]
Truth is a lie: Crowd truth and the seven myths of human annotation
Lora Aroyo and Chris Welty. Truth is a lie: Crowd truth and the seven myths of human annotation. AI Mag. , 36(1):15–24, March 2015
2015
-
[3]
Collaborative development of modular open source educational resources for natural language processing
Matthias Aßenmacher, Andreas Stephan, Leonie Weissweiler, Erion Çano, Ingo Ziegler, Marwin Härttrich, Bernd Bischl, Benjamin Roth, Christian Heumann, and Hinrich Schütze. Collaborative development of modular open source educational resources for natural language processing. In Sana Al- azzawi, Laura Biester, György Kovács, Ana Marasović, Leena Mathur, Mar...
-
[4]
LEA – Linguistic Exercises with Annotation Tools
Fabian Barteld and Johanna Flick. LEA – Linguistic Exercises with Annotation Tools. In Peggy Bockwinkel, Thierry Declerck, San- dra Kübler, and Heike Zinsmeister, edi- tors, Proceedings of the Workshop on Teach- ing NLP for Digital Humanities (Teach4DH 2017), volume 1918 of CEUR Workshop Pro- ceedings, pages 11–16. CEUR-WS.org, 2017
2017
-
[5]
Tightly coupled worksheets and homework assign- ments for NLP
Laura Biester and Winston Wu. Tightly coupled worksheets and homework assign- ments for NLP. In Sana Al-azzawi, Laura Bi- ester, György Kovács, Ana Marasović, Leena Mathur, Margot Mieskes, and Leonie Weis- sweiler, editors, Proceedings of the Sixth Workshop on Teaching NLP , pages 66–68, Bangkok, Thailand, August 2024. Associa- tion for Computational Linguistics
2024
-
[6]
Using thematic analysis in psychology
Virginia Braun and Victoria Clarke. Using thematic analysis in psychology. Qual. Res. Psychol., 3(2):77–101, January 2006
2006
-
[7]
Martin, Christine Pre- ston, Peter Rutledge, and Alice Motion
Larissa Braz Sousa, Ciara Kenneally, Yaela Golumbic, John M. Martin, Christine Pre- ston, Peter Rutledge, and Alice Motion. Teacher experiences and understanding of citizen science in australian classrooms. PLOS ONE , 19(11):1–22, 11 2024. 18 Raumanns et al. Data Annotations as Pedagogical Hints
2024
-
[8]
Veronika Cheplygina and Josien P. W. Pluim. Crowd disagreement about medi- cal images is informative. In Danail Stoy- anov, Zeike Taylor, Simone Balocco, Raphael Sznitman, Anne Martel, Lena Maier-Hein, Luc Duong, Guillaume Zahnd, Stefanie Demirci, Shadi Albarqouni, Su-Lin Lee, Ste- fano Moriconi, Veronika Cheplygina, Diana Mateus, Emanuele Trucco, Eric Gr...
2018
Show all 54 references
-
[9]
From hate speech to societal empowerment: A pedagogical journey through computational thinking and NLP for high school students
Alessandra Teresa Cignarella, Elisa Chier- chiello, Chiara Ferrando, Simona Frenda, Soda Marem Lo, and Andrea Marra. From hate speech to societal empowerment: A pedagogical journey through computational thinking and NLP for high school students. In Sana Al-azzawi, Laura Bieste...
2024
-
[10]
Emre Celebi, Stephen Dusza, David Gutman, Brian Helba, Aadi Kalloo, Konstantinos Liopyris, Michael Marchetti, Harald Kittler, and Allan Halpern
Noel Codella, Veronica Rotemberg, Philipp Tschandl, M. Emre Celebi, Stephen Dusza, David Gutman, Brian Helba, Aadi Kalloo, Konstantinos Liopyris, Michael Marchetti, Harald Kittler, and Allan Halpern. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by th...
2018
-
[11]
Noel C. F. Codella, David Gutman, M. Emre Celebi, Brian Helba, Michael A. Marchetti, Stephen W. Dusza, Aadi Kalloo, Konstanti- nos Liopyris, Nabin Mishra, Harald Kit- tler, and Allan Halpern. Skin lesion analy- sis toward melanoma detection: A challenge at the 2017 internation...
2017
-
[12]
A coefficient of agreement for nominal scales
Jacob Cohen. A coefficient of agreement for nominal scales. Educational and Psychologi- cal Measurement, 20:37 – 46, 1960
1960
-
[13]
Marc Combalia, Noel C. F. Codella, Veron- ica Rotemberg, Brian Helba, Veronica Vi- laplana, Ofer Reiter, Cristina Carrera, Ali- cia Barreiro, Allan C. Halpern, Susana Puig, and Josep Malvehy. Bcn20000: Dermoscopic lesions in the wild, 2019
2019
-
[14]
Crowdtruth 2.0: Quality metrics for crowd- sourcing with disagreement
Anca Dumitrache, Oana Inel, Lora Aroyo, Benjamin Timmermans, and Chris Welty. Crowdtruth 2.0: Quality metrics for crowd- sourcing with disagreement. In Proceedings of the 1st Workshop on Subjectivity, Ambigu- ity and Disagreement in Crowdsourcing and the 1st Workshop on Disent...
2018
-
[15]
Classification of shared tasks used in teaching
Theresa Elstner, Bärbel Hanle, Frank Loebe, Maik Fröbe, Nikolay Kolyada, Janis Mohr, Jörg Frochte, Sven Hofmann, Benno Stein, and Martin Potthast. Classification of shared tasks used in teaching. In Proceedings of the 2024 on Innovation and Technology in Com- puter Science Edu...
2024
-
[16]
Active learning: An introduction
Richard Felder and Rebecca Brent. Active learning: An introduction. ASQ Higher Ed- ucation Brief , 2, 01 2009
2009
-
[17]
Measuring nominal scale agreement among many raters
Joseph Fleiss. Measuring nominal scale agreement among many raters. Psycholog- ical Bulletin , 76:378–382, 11 1971
1971
-
[18]
Resources for Combining Teaching and Re- search in Information Retrieval Coursework
Maik Fröbe, Harrisen Scells, Theresa Elst- ner, Christopher Akiki, Lukas Gienapp, Jan Heinrich Reimer, Sean MacA vaney, Benno Stein, Matthias Hagen, and Martin Potthast. Resources for Combining Teaching and Re- search in Information Retrieval Coursework. In Grace Hui Yang, Hon...
2024
-
[19]
From Annotation to Reflection: How Participatory AI Training Enhances Critical Thinking
Sassan Gholiagha, Jürgen Neyer, Mitja Sienknecht, Magdalena Anna Wolska, Dora Kiesel, Patrick Riehmann, Julius Voigt, Matti Wiegmann, Irene López García, Ka- trin Girgensohn, Benno Stein, and Bernd Fröhlich. From Annotation to Reflection: How Participatory AI Training Enhances...
2025
-
[20]
BELT: Building endangered language technology
Michael Ginn, David Saavedra-Beltrán, Camilo Robayo, and Alexis Palmer. BELT: Building endangered language technology. In Sana Al-azzawi, Laura Biester, György Kovács, Ana Marasović, Leena Mathur, Mar- got Mieskes, and Leonie Weissweiler, editors, Proceedings of the Sixth Work...
2024
-
[21]
David Gutman, Noel C F Codella, Emre Celebi, Brian Helba, Michael Marchetti, Nabin Mishra, and Allan Halpern. Skin le- sion analysis toward melanoma detection: A challenge at the international symposium on biomedical imaging (ISBI) 2016, hosted by the international skin imagin...
2016
-
[22]
Hands-on machine learn- ing with scikit-learn, keras, and tensorflow
Aurélien Géron. Hands-on machine learn- ing with scikit-learn, keras, and tensorflow . O’Reilly Media, 3 edition, October 2022
2022
-
[23]
Teaching LLMs at Charles University: Assignments and activities
Jindřich Helcl, Zdeněk Kasner, Ondřej Dušek, Tomasz Limisiewicz, Dominik Macháček, Tomáš Musil, and Jindřich Libovický. Teaching LLMs at Charles University: Assignments and activities. In Sana Al-azzawi, Laura Biester, György Kovács, Ana Marasović, Leena Mathur, Margot Mieskes...
2024
-
[24]
A course shared task on evaluating LLM output for clinical questions
Yufang Hou, Thy Thy Tran, Doan Nam Long Vu, Yiwen Cao, Kai Li, Lukas Rohde, and Iryna Gurevych. A course shared task on evaluating LLM output for clinical questions. In Sana Al-azzawi, Laura Biester, György Kovács, Ana Marasović, Leena Mathur, Mar- got Mieskes, and Leonie Weis...
2024
-
[25]
Open source data labeling and AI evaluation
HumanSignal. Open source data labeling and AI evaluation. https://labelstud.io. Accessed: 2026-6-4
2026
-
[26]
Striking a balance between classical and deep learning approaches in natural language pro- cessing pedagogy
Aditya Joshi, Jake Renzella, Pushpak Bhat- tacharyya, Saurav Jha, and Xiangyu Zhang. Striking a balance between classical and deep learning approaches in natural language pro- cessing pedagogy. In Sana Al-azzawi, Laura Biester, György Kovács, Ana Marasović, Leena Mathur, Margo...
2024
-
[27]
Testing stylistic interventions to re- duce emotional impact of content modera- tion workers
Sowmya Karunakaran and Rashmi Ramakr- ishan. Testing stylistic interventions to re- duce emotional impact of content modera- tion workers. Proceedings of the AAAI Con- ference on Human Computation and Crowd- sourcing, 7:50–58, October 2019
2019
-
[28]
Maurice G. Kendall. A new measure of rank correlation. Biometrika, 30(1–2):81–93, June 1938
1938
-
[29]
Estimating the reliabil- ity, systematic error and random error of in- terval data
Klaus Krippendorff. Estimating the reliabil- ity, systematic error and random error of in- terval data. Educational and Psychological Measurement, 30(1):61–70, 1970
1970
-
[30]
Empowering the future with multilinguality and language diversity
En-Shiun Annie Lee, Kosei Uemura, Syed Mekael Wasti, and Mason Shipton. Empowering the future with multilinguality and language diversity. In Sana Al-azzawi, Laura Biester, György Kovács, Ana Maraso- vić, Leena Mathur, Margot Mieskes, and Leonie Weissweiler, editors, Proceedin...
2024
-
[31]
D. J. Leiner. SoSci survey: professional online surveys, european product, GDPR compliant. https://www.soscisurvey.de. Ac- cessed: 2026-5-24
2026
-
[32]
Modeling annotator preference and stochastic annotation error for medical image segmentation
Zehui Liao, Shishuai Hu, Yutong Xie, and Yong Xia. Modeling annotator preference and stochastic annotation error for medical image segmentation. Medical Image Analy- sis, 92:103028, 2024
2024
-
[33]
Novel or drivel? variants of invariants for teaching NLP in the LLM era
Marius Micluța-Câmpeanu. Novel or drivel? variants of invariants for teaching NLP in the LLM era. In Matthias Aßenmacher, Laura Biester, Claudia Borg, György Kovács, Margot Mieskes, and Sofia Serrano, edi- tors, Proceedings of the Seventh Workshop on Teaching Natural Language ...
2026
-
[34]
doccano: Text annotation tool for hu- man, 2018
Hiroki Nakayama, Takahiro Kubo, Junya Kamura, Yasufumi Taniguchi, and Xu Liang. doccano: Text annotation tool for hu- man, 2018. Software available from https://github.com/doccano/doccano
2018
-
[35]
Industry vs academia: Running a course on trans- formers in two setups
Irina Nikishina, Maria Tikhonova, Viktoriia Chekalina, Alexey Zaytsev, Artem Vazhent- sev, and Alexander Panchenko. Industry vs academia: Running a course on trans- formers in two setups. In Sana Al-azzawi, Laura Biester, György Kovács, Ana Maraso- vić, Leena Mathur, Margot Mi...
2024
-
[36]
Example-driven course slides on natural language processing concepts
Natalie Parde. Example-driven course slides on natural language processing concepts. In Sana Al-azzawi, Laura Biester, György Kovács, Ana Marasović, Leena Mathur, Mar- got Mieskes, and Leonie Weissweiler, edi- tors, Proceedings of the Sixth Workshop on Teaching NLP , pages 4–6...
2024
-
[37]
Ex- perimental analysis of the effective compo- nents of problem‐based learning
María A Pease and Deanna Kuhn. Ex- perimental analysis of the effective compo- nents of problem‐based learning. Sci. Educ. , 95(1):57–86, January 2011
2011
-
[38]
The “problem” of human la- bel variation: On ground truth in data, mod- eling and evaluation
Barbara Plank. The “problem” of human la- bel variation: On ground truth in data, mod- eling and evaluation. In Yoav Goldberg, Zor- nitsa Kozareva, and Yue Zhang, editors, Pro- ceedings of the 2022 Conference on Empir- ical Methods in Natural Language Process- ing, pages 10671...
2022
-
[39]
Training an NLP scholar at a small liberal arts col- lege: A backwards designed course proposal
Grusha Prasad and Forrest Davis. Training an NLP scholar at a small liberal arts col- lege: A backwards designed course proposal. In Sana Al-azzawi, Laura Biester, György Kovács, Ana Marasović, Leena Mathur, Mar- got Mieskes, and Leonie Weissweiler, editors, Proceedings of the...
2024
-
[40]
The Persona Lifecycle
John Pruitt. The Persona Lifecycle . Elsevier S & T, August 2010
2010
-
[41]
Deeper learning by do- ing: Integrating hands-on research projects into a machine learning course
Sebastian Raschka. Deeper learning by do- ing: Integrating hands-on research projects into a machine learning course. arXiv [cs.CY], July 2021
2021
-
[42]
Multi- task learning with crowdsourced features im- proves skin lesion diagnosis
Ralf Raumanns, Elif K Contar, Gerard Schouten, and Veronika Cheplygina. Multi- task learning with crowdsourced features im- proves skin lesion diagnosis. arXiv preprint arXiv:2004.14745, 2020
2004 arXiv
-
[43]
Ralf Raumanns, Gerard Schouten, Max Joosten, Josien P. W. Pluim, and Veronika Cheplygina. Enhance (enriching health data by annotations of crowd and experts): A case study for skin lesion classification. Ma- chine Learning for Biomedical Imaging , 1:1– 26, 2021
2021
-
[44]
Two contrasting data annotation paradigms for subjective NLP tasks
Paul Röttger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. Two contrasting data annotation paradigms for subjective NLP tasks. In Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz, editors, Proceedings of the 2022 Conference of the North American C...
2022
-
[45]
Why don’t you do it right? analysing annotators’ disagree- ment in subjective tasks
Marta Sandri, Elisa Leonardelli, Sara Tonelli, and Elisabetta Jezek. Why don’t you do it right? analysing annotators’ disagree- ment in subjective tasks. In Andreas Vlachos and Isabelle Augenstein, editors, Proceedings of the 17th Conference of the European Chap- ter of the As...
2023
-
[46]
Turning software engineers into machine learning engineers
Alexander Schiendorfer, Carola Gajek, and W Reif. Turning software engineers into machine learning engineers. In Bernd Bis- chl, Oliver Guhr, Heidi Seibold, and Peter Steinbach, editors, Proceedings of Machine Learning Research, volume 141 of Proceed- ings of Machine Learning ...
2020
-
[47]
Projectindata- science2026_examtemplate (data folder),
Project In Data Science. Projectindata- science2026_examtemplate (data folder),
-
[48]
ABCD rule of dermatoscopy : a new practical method for early recogni- tion of malignant melanoma
STOLZ and W. ABCD rule of dermatoscopy : a new practical method for early recogni- tion of malignant melanoma. Eur. J. Der- matol., 4:521–527, 1994
1994
-
[49]
The HAM10000 dataset, a large collection of multi-source dermatoscopic im- ages of common pigmented skin lesions, 2018
Philipp Tschandl, Cliff Rosendahl, and Har- ald Kittler. The HAM10000 dataset, a large collection of multi-source dermatoscopic im- ages of common pigmented skin lesions, 2018
2018
-
[50]
Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio
Alexandra N. Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. Learning from disagree- ment: A survey. J. Artif. Int. Res. , 72:1385– 1470, January 2022
2022
-
[51]
Rotemberg Veronica, Kurtansky Nicholas, Betz-Stablein Brigid, Caffery Liam, Chousakos Emmanouil, Codella Noel, Combalia Marc, Stephen Dusza, Guitera Pascale, David Gutman, Allan Halpern, Helba Brian, Kittler Harald, Kose Ki- vanc, Steve Langer, Lioprys Konstantinos, Malvehy Jo...
-
[54]
(ISNI:0000 0001 2155 0800), SUNY Downstate Medical School, New York, USA (GRID:grid. 262863. b) (ISNI:0000 0001 0693 2202), Stony Brook Medical School, Stony Brook, USA (GRID:grid. 51462. 34), and Rabin Medical Center, Tel A viv, Israel (GRID:grid. 413156. 4) (ISNI:0000 0004 0...
2021
-
[2024]
Association for Computational Lin- guistics
-
[2026]
Accessed: 2026-06-28
2026
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.