REVIEW 3 major objections 6 minor 46 references
Donate or Create? Comparing Data Collection Strategies for Emotion-labeled Multimodal Social Media Posts
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that study-created posts train emotion models as well as donated real posts, but test scores on them are inflated.
desk verdict First direct comparison of created vs donated author-labeled multimodal posts; the practical recommendation holds up, but 'created vs genuine' is a quasi-experimental claim, not a settled causal one. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a paired collection design: the same annotation survey, covering emotion labels, event appraisals, text-image relations, and participant demographics, is attached to three collection strategies, and the same classifiers are trained and tested across strategies. The decisive instrument is the two-by-two training/test matrix, where models trained on one strategy are evaluated on both, isolating whether any performance gap is a training-data effect or a test-set effect. The design also includes a RECENT condition that collects the five most recent posts without an emotion prompt, which reveals how common multi-emotion posts are and how prompt-driven label selection may bias donated annotations.
What would settle it
Randomly assign a single participant pool to CREATION and DONATION in the same calendar period; if models trained on created posts no longer match donation-trained models on donated test data, or if the gap between created and donated test scores disappears, then the paper's central result is an artifact of who enrolled and when.
Extended reading notes
Core claim
The central discovery is that the data collection strategy changes what the corpus looks like without changing which training data transfers. Trained on 800 posts each, a text-only, image-only, and fused text-image classifier perform on par whether trained on study-created or donated posts when evaluated on donated posts; the same models score higher on study-created test posts, and zero-shot multimodal models confirm that this is a property of the test set. Study-created posts are about 26 percent longer than donated posts after controlling for emotion, use images that are less necessary for understanding the text, and contain more intense and more prototypical emotion-triggering events. Donated and recent-post samples skew younger, include more students and more non-European participants, and declined participation far more often, mostly citing privacy. The paper's answer to its title is therefore both: create for training, donate for evaluation.
Load-bearing premise
The comparison treats the collection task as the cause of the observed differences, but the three arms ran at different times with different samples who declined at different rates, so recruitment and timing could produce the same results.
Editorial extensions
If this is right
- Researchers can build larger training corpora cheaply and with fewer privacy risks by asking participants to create posts, without sacrificing performance on genuine posts.
- Any reported accuracy measured on study-created test data should be treated as an upper bound; published effectiveness numbers should come from genuine donated posts.
- Corpus collection strategy is a confound in emotion research: it changes post length, image use, event prototypicality, and who participates, so mixing strategies inside one corpus can distort analyses.
- Multi-emotion labeling, as done in the RECENT condition, should become standard practice, because single-emotion prompts may bias which label a participant assigns to a post.
Reading between the lines
- The paper's sequential design means the creation-donation contrast is entangled with recruitment period and participant pool; a within-subject replication, where the same participants create and donate posts, is the natural decisive test of the causal claim.
- If the training-transfer result holds under such a test, a division of labor could become standard: creation data for model development and genuine data as a small fixed evaluation benchmark, lowering cost while keeping estimates honest.
- Because the creation condition forced participants to choose images from a stock library, the observed multimodal differences may partly reflect image availability rather than the act of creating; allowing participants to supply their own images in a creation condition would separate these factors.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper compares three strategies for collecting author-labeled emotion data from social media posts: CREATION (participants write a post about a remembered emotional event and select a stock image), DONATION (participants donate a genuine past post matching a prompted emotion), and RECENT (participants donate their five most recent posts and annotate all emotions). The authors collect 2,507 posts from 522 participants and compare the resulting corpora along content features (length, image style, text-image relation), event appraisals and intensity, participant demographics, label structure, and downstream emotion-classification performance. The central claims are that study-created posts are longer, more text-dependent, more prototypical in their emotion triggers, and come from a demographically different participant pool; that models trained on CREATION data generalize to genuine (DONATION) test data as well as models trained on DONATION data; and that CREATION test data yields unrealistically optimistic performance estimates, so realistic evaluation requires genuine data.
Significance. If the central empirical comparison is valid, the paper offers actionable guidance for building author-labeled emotion corpora: CREATION-style data can be used for model training, while DONATION-style data should be reserved for evaluation. The paper also contributes a novel multimodal, author-labeled emotion corpus with rich appraisal annotations, and it is unusually transparent: data, code, and surveys are released, the descriptive analyses use Bonferroni-corrected ANOVAs and mixed-effects models, and the limitations section explicitly discusses temporal and design confounds. These strengths make the paper a useful reference for future data-collection decisions in affective NLP. However, the headline recommendation is only as strong as the causal attribution of observed differences to the 'create versus donate' manipulation, and the model-comparison evidence currently lacks inferential support and may be affected by leakage across participants.
major comments (3)
- [Section 4.5, Table 4] The central comparative claims—that CREATION-trained and DONATION-trained models perform equally on DONATION test data and that CREATION test data yields higher, less realistic scores—are based on macro-F1 point estimates averaged over five runs, with no reported variance, confidence intervals, or significance tests. Several of the differences most relevant to the first claim are small (e.g., on DONATION test, T+V: .38 vs .40; T: .41 vs .42; V: .18 vs .19), and with 300 test posts these gaps could easily be within sampling noise. Conversely, the test-set gap (CREATION test T+V .60–.62 vs DONATION test .38–.43) also needs uncertainty quantification. Please report per-run scores with standard errors or confidence intervals and run paired statistical tests (e.g., paired bootstrap on posts or a per-seed paired comparison) for both the training-strategy comparison and the test-set comparison. Without this, the abstract's recommendation is not supported by the reported numbers.
- [Section 4.5, 'Setup' paragraph] The train/dev/test split is by post, not by participant. Because each participant contributes up to five posts and may complete the study up to once per emotion, posts from the same participant can appear in both training and test sets. This can leak author-specific lexical style, content choices, and sociodemographic signals, inflating absolute performance and potentially biasing the CREATION-versus-DONATION comparison if the number of repeated participants differs by strategy. Please split by participant (e.g., assign all posts of a participant to one split) and re-run the modeling, or explicitly justify why such leakage is negligible. Relatedly, the logistic regressions in Table 11 cluster standard errors by post, not by participant; with multiple posts per participant the standard errors may be anticonservative. Please cluster by participant or use a mixed-effects model with a participant random intercept.
- [Section 4.3 and Limitations] The comparison between CREATION and DONATION is quasi-experimental rather than a clean manipulation of 'created versus genuine': the participant pools differ significantly in age, student status, and ethnicity (Appendix Table 6), decline rates differ sharply by strategy (Figure 7), and the studies were run sequentially in multiple phases, so posts come from different time windows (Limitations). The authors acknowledge these confounds, but the central recommendation—use CREATION for training and DONATION for testing—presupposes that residual differences are caused by the collection strategy rather than by selection or temporal content. The logistic regression analysis in Section 4.5 (Table 11) controls for measured demographics, yet a significant negative coefficient for DONATION (versus CREATION) remains for both CLIP models, which the text itself reads as 'further study differences that are not accounted for here.' To support the causal attribution, please add a robustness check restricted to an overlapping collection period and/or a demographic-matched subsample; at minimum, the conclusions should be softened to describe associations rather than strategy-caused effects.
minor comments (6)
- [Table 1] The line 'The event was familar' contains a typo; it should be 'familiar'. The same table also has 'occuring' instead of 'occurring'.
- [Limitations] The Limitations section contains the typo 'guarentee'; it should be 'guarantee'.
- [Appendix A.3.1] The phrase 'Images are preproccessed by resizing them to 224x224' contains a typo; it should be 'preprocessed'.
- [Figure 3] The legend entry 'Pro_Photo' uses an underscore, while Table 11 and the text use 'Pro Photo'; please unify the naming.
- [Section 3.3] The sentence 'CREATION and DONATION take 30 minutes and we pay participants £4.50' is grammatically awkward; consider 'CREATION and DONATION each take about 30 minutes, and participants are paid £4.50.'
- [Table 2] The note says that in RECENT, posts with tied emotions are counted in both categories; this should be stated more prominently because it makes the RECENT column totals exceed the number of posts and complicates any direct comparison of emotion distributions.
Circularity Check
No significant circularity: the central comparison is empirical, uses held-out test data, and does not reduce to a fitted parameter or a self-citation chain.
full rationale
The paper's central claims are empirical comparisons between independently collected corpora, and no parameter is fitted to the conclusion and then renamed as a prediction. Supervised models are fine-tuned on fixed training splits and evaluated on held-out test splits from both CREATION and DONATION; zero-shot models are prompted without training on either corpus, so the observed score differences reflect test-set composition rather than any fitted quantity. The statistical analyses (ANOVA, chi-square tests, mixed-effects models, logistic regressions) estimate differences and associations from the data rather than assuming the findings. The self-citations to Troiano et al. and related work by Klinger are background references about corpus-construction methodology and are not load-bearing; no uniqueness theorem or ansatz is imported to force the conclusion. The documented confounds (demographic differences in who completed each study and sequential data-collection time windows) are limitations on causal attribution and external validity, but they are transparently reported and do not make the derivation circular. Therefore no circular step is present, and the paper deserves a score of 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Author self-reports are ground-truth emotion labels for posts.
- ad hoc to paper The collection conditions differ only in strategy, not in unmeasured participant or temporal factors.
- domain assumption Appraisal questionnaire items measure event characteristics reliably.
Cite this review
Pith. "Pith review of Donate or Create? Comparing Data Collection Strategies for Emotion-labeled Multimodal Social Media Posts." pith.science (2026). https://pith.science/paper/TO2JVDDB
@misc{pith2026250524427,
author = {Pith},
title = {Pith review of: Donate or Create? Comparing Data Collection Strategies for Emotion-labeled Multimodal Social Media Posts},
year = {2026},
howpublished = {\url{https://pith.science/paper/TO2JVDDB}},
note = {Machine review of arXiv:2505.24427}
}
read the original abstract
Accurate modeling of subjective phenomena such as emotion expression requires data annotated with authors' intentions. Commonly such data is collected by asking study participants to donate and label genuine content produced in the real world, or create content fitting particular labels during the study. Asking participants to create content is often simpler to implement and presents fewer risks to participant privacy than data donation. However, it is unclear if and how study-created content may differ from genuine content, and how differences may impact models. We collect study-created and genuine multimodal social media posts labeled for emotion and compare them on several dimensions, including model performance. We find that compared to genuine posts, study-created posts are longer, rely more on their text and less on their images for emotion expression, and focus more on emotion-prototypical events. The samples of participants willing to donate versus create posts are demographically different. Study-created data is valuable to train models that generalize well to genuine data, but realistic effectiveness estimates require genuine data.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Ines Abbes, Wajdi Zaghouani, Omaima El-Hardlo, and Faten Ashour. 2020. https://aclanthology.org/2020.lrec-1.768/ DAICT : A dialectal A rabic irony corpus extracted from T witter . In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 6265--6271, Marseille, France. European Language Resources Association
work page 2020
-
[2]
Ibrahim Abu Farha, Silviu Vlad Oprea, Steven Wilson, and Walid Magdy. 2022. https://doi.org/10.18653/v1/2022.semeval-1.111 SemEval-2022 Task 6: iSarcasmEval , Intended Sarcasm Detection in English and Arabic . In Proceedings of the 16th International Workshop on Semantic Evaluation ( SemEval-2022 ) , pages 802--814. Association for Computational Linguistics
-
[3]
Möller, Theo Araujo, and Daniel L
Laura Boeschoten, Jef Ausloos, Judith E. Möller, Theo Araujo, and Daniel L. Oberski. 2022. https://doi.org/10.5117/CCR2022.2.002.BOES A framework for privacy preserving digital trace data collection through data donation . Computational Communication Research, 4(2):388--423
-
[4]
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008. https://doi.org/10.1007/s10579-008-9076-6 IEMOCAP: Interactive emotional dyadic motion capture database . Language resources and evaluation, 42:335--359
-
[5]
Carlos Busso, Srinivas Parthasarathy, Alec Burmania, Mohammed AbdelWahab, Najmeh Sadoughi, and Emily Mower Provost. 2016. https://ieeexplore.ieee.org/document/7374697 MSP-IMPROV: An acted corpus of dyadic interactions to study emotion perception . IEEE Transactions on Affective Computing, 8(1):67--80
-
[6]
Pasquale Capuozzo, Ivano Lauriola, Carlo Strapparava, Fabio Aiolli, and Giuseppe Sartori. 2020. https://aclanthology.org/2020.lrec-1.178/ D ec O p: A multilingual and multi-domain corpus for detecting deception in typed text . In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 1423--1430, Marseille, France. European Language...
work page 2020
-
[7]
Carrière, Laura Boeschoten, Bella Struminskaya, Heleen L
Thijs C. Carrière, Laura Boeschoten, Bella Struminskaya, Heleen L. Janssen, Niek C. De Schipper, and Theo Araujo. 2024. https://doi.org/10.1007/s11135-024-01983-x Best practices for studies using digital data donation . Quality & Quantity
-
[8]
Junhan Chen, Yumin Yan, and John Leach. 2022. https://doi.org/10.12840/ISSN.2255-4165.034 Are emotion-expressing messages more shared on social media? a meta-analytic review . Review of Communication Research, 10:--
Show all 46 references
-
[9]
Wingyan Chung and Daniel Zeng. 2020. https://doi.org/10.1016/j.im.2018.09.008 Dissecting emotion and user influence in social media communities: An interaction modeling approach . Information & Management, 57(1):103108. Big data and business analytics: A research agenda for re...
2020 doi
-
[10]
Irina Degtiar and Sherri Rose. 2023. https://doi.org/10.1146/annurev-statistics-042522-103837 A review of generalizability and transportability . Annual Review of Statistics and Its Application, 10(Volume 10, 2023):501--524
2023 doi
-
[11]
Fischer, and Arjan E.R
Daantje Derks, Agneta H. Fischer, and Arjan E.R. Bos. 2008. https://doi.org/10.1016/j.chb.2007.04.004 The role of emotion in computer-mediated communication: A review . Computers in Human Behavior, 24(3):766--785. Instructional Support for Enhancing Students' Information Probl...
2008 doi
-
[12]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. https://openreview.net/forum?id=YicbFdNTTy An image is worth 1...
2021
-
[13]
Aparna Elangovan, Jiayuan He, Yuan Li, and Karin Verspoor. 2024. https://doi.org/10.18653/v1/2024.naacl-long.127 Principles from clinical research for NLP model generalization . In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computat...
2024 doi
-
[14]
Alejandra Gomez Ortega, Jacky Bourgeois, Wiebke Toussaint Hutiri, and Gerd Kortuem. 2023. https://doi.org/10.1007/s00146-023-01755-5 Beyond data transactions: A framework for meaningfully informed data donation . AI & Society
2023 doi
-
[15]
Anurag Illendula and Amit Sheth. 2019. https://doi.org/10.1145/3308560.3316549 Multimodal emotion classification . In Companion Proceedings of The 2019 World Wide Web Conference, WWW '19, page 439–449, New York, NY, USA. Association for Computing Machinery
2019
-
[16]
Sasha Shen Johfre and Jeremy Freese. 2021. Reconsidering the reference category. Sociological Methodology, 51(2):253--269
2021
-
[17]
Tomoyuki Kajiwara, Chenhui Chu, Noriko Takemura, Yuta Nakashima, and Hajime Nagahara. 2021. https://doi.org/10.18653/v1/2021.naacl-main.169 WRIME : A New Dataset for Emotional Intensity Estimation with Subjective and Objective Annotations . In Proceedings of the 2021 Conferenc...
2021 doi
-
[18]
Kensinger
Elizabeth A. Kensinger. 2009. https://doi.org/10.1177/1754073908100432 Remembering the Details : Effects of Emotion . Emotion Review, 1(2):99--113
2009 doi
-
[19]
Pankowska, Alexandru Cernat, and Ruben L
Florian Keusch, Paulina K. Pankowska, Alexandru Cernat, and Ruben L. Bach. 2024. https://doi.org/10.1177/1525822X231225907 Do You Have Two Minutes to Talk about Your Data ? Willingness to Participate and Nonparticipation Bias in Facebook Data Donation . Field Methods, 36(4):279--293
2024 doi
-
[20]
Yiyi Li and Ying Xie. 2020. https://doi.org/10.1177/0022243719881113 Is a picture worth a thousand words? an empirical study of image content and social media engagement . Journal of Marketing Research, 57(1):1--19
2020 doi
-
[21]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. http://arxiv.org/abs/1907.11692 Roberta: A robustly optimized bert pretraining approach . Cite arxiv:1907.11692
2019 arXiv
-
[22]
Paige Lloyd, Jason C
E. Paige Lloyd, Jason C. Deska, Kurt Hugenberg, Allen R. McConnell, Brandon T. Humphrey, and Jonathan W. Kunstman. 2019. https://doi.org/10.3758/s13428-018-1061-4 Miami University deception detection database . Behavior Research Methods, 51(1):429--439
2019 doi
-
[23]
Anthony Lyons and Yoshihisa Kashima. 2003. https://doi.org/10.1037/0022-3514.85.6.989 How Are Stereotypes Maintained Through Communication ? The Influence of Stereotype Sharedness . Journal of Personality and Social Psychology, 85(6):989--1005
2003 doi
-
[24]
Anthony Lyons and Yoshihisa Kashima. 2006. https://doi.org/10.1111/j.1467-839X.2006.00184.x Maintaining stereotypes in communication: Investigating memory biases and coherence-seeking in storytelling . Asian Journal of Social Psychology, 9(1):59--71
2006
-
[25]
Saif Mohammad. 2012. https://aclanthology.org/S12-1033/ \# emotional tweets . In * SEM 2012: The First Joint Conference on Lexical and Computational Semantics -- Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth Internatio...
2012
-
[26]
Jonathan Mummolo and Erik Peterson. 2019. https://doi.org/10.1017/S0003055418000837 Demand effects in survey experiments: An empirical assessment . American Political Science Review, 113(2):517–529
2019 doi
-
[27]
Tsubasa Nakagawa, Shunsuke Kitada, and Hitoshi Iyatomi. 2022. https://doi.org/10.1145/3511808.3557599 Expressions Causing Differences in Emotion Recognition in Social Networking Service Documents . In Proceedings of the 31st ACM International Conference on Information & Knowle...
2022
-
[28]
Austin Lee Nichols and Jon K. Maner. 2008. https://doi.org/10.3200/GENP.135.2.151-166 The good-subject effect: Investigating participant demand characteristics . The Journal of General Psychology, 135(2):151--166. PMID: 18507315
2008 doi
-
[29]
Silviu Oprea and Walid Magdy. 2020. https://doi.org/10.18653/v1/2020.acl-main.118 iSarcasm : A Dataset of Intended Sarcasm . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 1279--1289. Association for Computational Linguistics
2020 doi
-
[30]
Myle Ott, Claire Cardie, and Jeffrey T. Hancock. 2013. https://aclanthology.org/N13-1053/ Negative deceptive opinion spam . In Proceedings of the 2013 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies , page...
2013
-
[31]
Nico Pfiffner, Pim Witlox, and Thomas N. Friemel . 2024. https://doi.org/10.5117/CCR2024.2.4.PFIF Data Donation Module : A Web Application for Collecting and Enriching Data Donations . Computational Communication Research, 6(2):1
2024 doi
-
[32]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. https://proceedings.mlr.press/v139/radford21a.html Learning transferable visual model...
2021
-
[33]
Wisniewski
Afsaneh Razi, Ashwaq Alsoubai, Seunghyun Kim, Nurun Naher, Shiza Ali, Gianluca Stringhini, Munmun De Choudhury, and Pamela J. Wisniewski. 2022. https://doi.org/10.1145/3491101.3503569 Instagram data donation: A case study on collecting ecologically valid social media data for ...
2022
-
[34]
Matthew Rhodes-Purdy, Rachel Navarre, and Stephen M Utych. 2021. https://doi.org/doi:10.1017/XPS.2019.35 Measuring simultaneous emotions: Existing problems and a new way forward . Journal of Experimental Political Science, 8(1):1--14
2021 doi
-
[35]
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. https://doi.org/10.18653/v1/2020.acl-main.442 Beyond accuracy: Behavioral testing of NLP models with C heck L ist . In Proceedings of the 58th Annual Meeting of the Association for Computational Lingu...
2020 doi
-
[36]
Andrea Scarantino. 2016. The philosophy of emotions and its impact on affective science. Handbook of emotions, 4:3--48
2016
-
[37]
Klaus Scherer and Harald Wallbott. 1997. The ISEAR questionnaire and codebook. Geneva Emotion Research Group
1997
-
[38]
Klaus R. Scherer. 2005. https://doi.org/10.1177/0539018405058216 What are emotions? and how can they be measured? Social Science Information, 44(4):695--729
2005 doi
-
[39]
Marco Antonio Stranisci, Simona Frenda, Eleonora Ceccaldi, Valerio Basile, Rossana Damiano, and Viviana Patti. 2022. https://aclanthology.org/2022.lrec-1.406 APPR eddit: a corpus of R eddit posts annotated for appraisal . In Proceedings of the Thirteenth Language Resources and...
2022
-
[40]
Enrica Troiano, Sofie Labat, Marco Antonio Stranisci, Rossana Damiano, Viviana Patti, and Roman Klinger. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.89 Dealing with controversy: An emotion and coping strategy corpus based on role playing . In Findings of the Associat...
2024 doi
-
[41]
Enrica Troiano, Laura Oberl\"ander, and Roman Klinger. 2023. https://doi.org/10.1162/coli_a_00461 Dimensional modeling of emotions in text with appraisal theories: Corpus creation, annotation reliability, and prediction . Computational Linguistics, 49(1)
2023 doi
-
[42]
Enrica Troiano, Sebastian Pad \'o , and Roman Klinger. 2019. https://doi.org/10.18653/v1/P19-1391 Crowdsourcing and validating event-focused emotion corpora for G erman and E nglish . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, p...
2019 doi
-
[43]
van Driel, Anastasia Giachanou, J
Irene I. van Driel, Anastasia Giachanou, J. Loes Pouwels, Laura Boeschoten, Ine Beyens, and Patti M. Valkenburg. 2022. https://doi.org/10.1080/19312458.2022.2109608 Promises and pitfalls of social media data donations . Communication Methods and Measures, 16(4):266--282
2022
-
[44]
Clara Vania, Ruijie Chen, and Samuel R. Bowman. 2020. https://doi.org/10.18653/v1/2020.aacl-main.68 Asking Crowdworkers to Write Entailment Examples : The Best of Bad Options . In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computationa...
2020 doi
-
[45]
Aswathy Velutharambath, Amelie W \"u hrl, and Roman Klinger. 2024. https://aclanthology.org/2024.lrec-main.243/ Can factual statements be deceptive? the D e F a B el corpus of belief-based deception . In Proceedings of the 2024 Joint International Conference on Computational L...
2024
-
[46]
Linyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu, Yidong Wang, Jingming Zhuo, Lingqiao Liu, Jindong Wang, Jennifer Foster, and Yue Zhang. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.276 Out-of-distribution generalization in natural language processing: Past, present, and...
2023 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.