REVIEW 3 major objections 4 minor 117 references
Sentiment Analysis Tools in Software Engineering: A Systematic Mapping Study
T0 review · 3 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read A 106-paper map of sentiment analysis in software engineering ranks neural networks and BERT at the top, while open-source data and unsolved sarcasm problems shape the field.
desk verdict Useful updated map of sentiment analysis in SE, but the 'BERT is best' ranking is an unweighted average of incomparable evaluations and needs reframing before it should guide tool choice. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the mapping-study corpus plus its aggregation tables. A systematic mapping study is a structured review that classifies a literature corpus into categories to answer research questions; here, the corpus of 106 papers was assembled from five databases with forward and backward snowballing. The central objects carrying the argument are Table 7, which pools reported accuracy and F1 scores by machine-learning approach, and Table 10, which pools them by named tool, so that heterogeneous primary studies are compressed into a single average ranking. The data-source tallies in Table 5 and the difficulty tallies in Table 11 do the same for the landscape claims about where sentiment analysis is applied and what problems remain.
What would settle it
Run a controlled benchmark that applies BERT, SentiStrength-SE, Senti4SD, and SentiCR to the same held-out comment sets from Stack Overflow, JIRA, and GitHub with balanced labels; if a non-neural tool matches or beats BERT on any platform, or the ordering changes by platform, the paper's cross-study average ranking would not survive.
Extended reading notes
Core claim
The central discovery, on the authors' terms, is that the research literature supports a clear performance hierarchy. Among approaches used to develop sentiment analysis tools, neural networks average the highest accuracy (0.87 over 15 data points) and the highest F1 score (0.80 over 15 data points), with SVM second in accuracy (0.82) but lower in F1 (0.64) and Bayes at 0.71 accuracy and 0.67 F1. Among the 34 existing tools identified in the corpus, BERT stands out with an average accuracy of 0.94 and F1 of 0.83, followed by RoBERTa and XLNet; the SE-specific tools Senti4SD, SentiCR, and SentiStrength-SE cluster around 0.73 to 0.77 accuracy, while generic tools NLTK and CoreNLP are at the bottom. The authors therefore conclude that a neural-network-based tool such as BERT is the best current choice and that future work should target objectively labeled data, industry data, and unresolved difficulties such as irony and sarcasm.
Load-bearing premise
The load-bearing premise is that accuracy and F1 values reported across different studies, datasets, and label distributions can be meaningfully averaged, even though the authors state they did not distinguish which tool was evaluated on which data set or whether a tool was pre-trained on the test data.
Editorial extensions
If this is right
- Tool selection guidance: teams should prefer neural-network-based tools, with BERT as the top performer, over lexicon-based tools like SentiStrength and generic NLP tools like NLTK for software-engineering text.
- Research priority guidance: the pooled numbers point future work toward objectively labeled datasets, industry-domain data, and the unsolved problems of irony and sarcasm and cross-platform drift.
- The map's skew is a finding: because 71 of 106 papers apply existing tools, the field's evidence base is application-heavy and thin on controlled comparison and development.
- Domain-adapted tools are common but not dominant in performance, and four of the top five tools are neural-network-based, including three BERT-family models.
Reading between the lines
- Editorial extension: the pooled ranking likely overstates BERT's edge for SE-specific text because BERT entries in primary studies are in most cases fine-tuned on in-domain data, while lexicon tools like SentiStrength are often applied off-the-shelf; a like-for-like comparison would probably shrink the gap.
- Editorial extension: unweighted averaging across studies treats a 90:10 positive-negative dataset and a balanced dataset as equal, so the ordering in Tables 7 and 10 could change if results were weighted by dataset difficulty or label balance.
- Testable extension: a standardized benchmark with fixed training and test splits on Stack Overflow, JIRA, and GitHub, reporting accuracy and F1 per platform, would turn the paper's average ranking into a decision table that practitioners could use directly.
- Implicit consequence: the scarcity of industry data is probably not only a technical gap but a data-governance one, since workplace communication analysis raises legal and privacy barriers; tool guidance for industrial teams should be paired with guidance on lawful data access.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript presents a systematic mapping study (SMS) of sentiment analysis tools and approaches in software engineering, extending the authors' earlier SLR with a new search covering January 2021 to October 2021. From 106 included papers, the study maps application domains (open-source, industry, academia), purposes (development, comparison, application), data sources (most frequently Stack Overflow, JIRA, GitHub), development approaches (Bayes, neural networks, SVM), existing tools (most frequently SentiStrength, Senti4SD, SentiStrength-SE), and reported difficulties (domain adaptation, sarcasm, subjective labeling, cross-platform performance). The paper's headline performance claim is that neural networks are the best-performing approach (average accuracy 0.87, F1 0.80) and that BERT is the best-performing tool (average accuracy 0.94, F1 0.83).
Significance. If the descriptive census is reliable, the study provides a useful map of sentiment analysis in SE, with a clearly documented search process and a published dataset. The demographic and taxonomical findings—such as the dominance of open-source data, the prevalence of SentiStrength as a comparison baseline, and the recurring difficulties of sarcasm, subjectivity, and cross-platform transfer—are valuable for researchers and practitioners. However, the headline performance ranking is not supported by the evidence presented: it averages accuracy and F1 values from heterogeneous primary studies without matched comparisons, variance estimates, or statistical tests, and the authors explicitly concede in Section 5.4 that they do not know which tools were evaluated on which datasets or whether tools were pre-trained on test data. The descriptive mapping is the core contribution; the performance ranking, as currently justified, cannot be accepted as a reliable basis for tool selection.
major comments (3)
- [Section 4.6, Table 10; abstract] The claim that BERT performs 'significantly better' than Senti4SD, SentiStrength, and SentiStrength-SE is not supported by any significance test or matched comparison. The reported means are unweighted averages of accuracy and F1 values collected from primary studies that differ in datasets, label distributions, training configurations, and evaluation protocols. Section 5.4 explicitly concedes that the authors do not distinguish which tools were evaluated on which datasets and cannot rule out pre-training on test data. Given the small and unequal data points (BERT: 10 accuracy/12 F1; Senti4SD: 30/44; SentiStrength-SE: 32/44), the observed differences could be artifacts of evaluation-set difficulty or class distribution rather than genuine performance advantages. The abstract's conclusion that 'the best tool is BERT' therefore does not follow from the presented evidence.
- [Section 4.5, Table 7] The same aggregation problem applies to the approach ranking. Neural networks are reported as best with average accuracy 0.87 and F1 0.80 over 15 data points, while SVM has 0.82 accuracy over 8 points and 0.64 F1 over 9 points. No statistical test is provided, and the text states that decision tree and Bayes 'are significantly worse' than neural networks without any variance or uncertainty measure. Since the data points come from non-comparable studies with different class distributions and evaluation settings, the ranking may reflect dataset selection rather than true approach quality. This weakens the claim that neural networks are the best-performing approach.
- [Table 4 and Section 4.3] The application subcategory counts in Table 4 are inconsistent with the text. The text states that 71 papers are of the application type, but the Application row in Table 4 sums to 95 (45 correlations + 25 social aspects + 17 values measurements + 8 values predictions). This is a substantial discrepancy that casts doubt on the reliability of the classification data. The authors should reconcile the counts or explain the apparent arithmetic error in the table.
minor comments (4)
- [Section 4.6] The sentence 'From the 106 papers, 21 overall provided information about accuracy or F1 score of their used machine learning approach' appears to be a copy-paste error from Section 4.5; in Section 4.6 it should refer to tools, not approaches.
- [Throughout] There are several typos and inconsistencies, including 'SentiStrenght-SE' for 'SentiStrength-SE' (Section 4.6), 'a objective framework' (Section 3.2.6), 'the later represents' for 'the latter' (Section 3.2.2), and 'the publication has not been peer-reviewed' for 'has not' (Table 1).
- [Table 2] The formatting of Table 2 is unclear: the header lists 2012 through 2021 and a Total column, but the body appears to show only ten numbers. This should be reformatted so that each year has its own column and the total is clearly separated.
- [Section 5.4] The limitations section is candid about the comparability problem, but the abstract and conclusions do not reflect those caveats. If the performance ranking is retained, the abstract and Section 6 should be reworded to state that the results are aggregated averages from non-comparable studies and do not establish a definitive ranking.
Circularity Check
No circularity: this is a descriptive systematic mapping study whose performance conclusions are direct aggregations of reported primary-study values, and its self-citation is a transparent starting point rather than a load-bearing derivation.
full rationale
This paper is a systematic mapping study with no fitted parameters and no equations. Its central results are descriptive aggregations of values reported in 106 primary studies: the counts in Tables 4 and 5, the average accuracy and F1 values in Tables 7 and 10, and the narrative finding that neural networks and BERT have the highest average scores. Those averages are computed directly from the extracted per-paper values, so the conclusion 'the best performing approach in our analysis is neural networks and the best tool is BERT' is a summary of the data, not a consequence of a definition that presupposes it. The only self-citation of note is the authors' prior SLR [20], which supplies the initial paper set for the updated search (Section 3). That is transparently stated ('The reference to the initial data set can be found in our previous paper [20]'), and the present claims do not depend on an unverified premise from [20] to force the conclusions; the prior review is a published, externally checkable evidence base. Section 5.4 contains a candid limitation: the authors do not distinguish which tools were evaluated on which datasets or whether a tool was pre-trained on the test data, and they warn that a 90/10 label distribution could change results. That is a threat to the comparability and soundness of the performance ranking, not a circularity: the averaged values come from the primary studies rather than being re-derived from the paper's own inputs. The apparent arithmetic inconsistency in Table 4 (subcategory counts summing to 95 rather than the stated 71) is an extraction and reporting concern, not a circularity concern. No step in the paper reduces an alleged prediction to its own input by construction, so the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Accuracy and F1 values reported in different primary studies can be meaningfully averaged across datasets and evaluation protocols to rank approaches and tools.
- domain assumption The five selected databases and the search string capture the relevant sentiment-analysis-in-SE literature.
- domain assumption Screening and data extraction by one researcher, reviewed by another, is sufficiently reliable.
Cite this review
Pith. "Pith review of Sentiment Analysis Tools in Software Engineering: A Systematic Mapping Study." pith.science (2026). https://pith.science/paper/EHL2HCOC
@misc{pith2026250207893,
author = {Pith},
title = {Pith review of: Sentiment Analysis Tools in Software Engineering: A Systematic Mapping Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/EHL2HCOC}},
note = {Machine review of arXiv:2502.07893}
}
read the original abstract
Software development is a collaborative task. Previous research has shown social aspects within development teams to be highly relevant for the success of software projects. A team's mood has been proven to be particularly important. It is paramount for project managers to be aware of negative moods within their teams, as such awareness enables them to intervene. Sentiment analysis tools offer a way to determine the mood of a team based on textual communication. We aim to help developers or stakeholders in their choice of sentiment analysis tools for their specific purpose. Therefore, we conducted a systematic mapping study (SMS). We present the results of our SMS of sentiment analysis tools developed for or applied in the context of software engineering (SE). Our results summarize insights from 106 papers with respect to (1) the application domain, (2) the purpose, (3) the used data sets, (4) the approaches for developing sentiment analysis tools, (5) the usage of already existing tools, and (6) the difficulties researchers face. We analyzed in more detail which tools and approaches perform how in terms of their performance. According to our results, sentiment analysis is frequently applied to open-source software projects, and most approaches are neural networks or support-vector machines. The best performing approach in our analysis is neural networks and the best tool is BERT. Despite the frequent use of sentiment analysis in SE, there are open issues, e.g. regarding the identification of irony or sarcasm, pointing to future research directions. We conducted an SMS to gain an overview of the current state of sentiment analysis in order to help developers or stakeholders in this matter. Our results include interesting findings e.g. on the used tools and their difficulties. We present several suggestions on how to solve these identified problems.
Figures
Reference graph
Works this paper leans on
-
[1]
R. E. Kraut, L. A. Streeter, Coordination in software development, Com- mun. ACM 38 (3) (1995) 69–81. doi:10.1145/203330.203345
arXiv 1995
-
[2]
D. E. Perry, N. A. Staudenmayer, L. G. Votta, People, organizations, and process improvement, IEEE Software 11 (4) (1994) 36–45. doi:10.1109/ 52.300082
1994
-
[3]
M. Kuhrmann, P. Tell, J. Kl¨ under, R. Hebig, Sherlock A. Licorish, S. G. MacDonell, Helena stage 2 results (2018). doi:10.13140/RG.2.2.14807. 52649
-
[4]
T. Niinimaki, A. Piri, C. Lassenius, Factors affecting audio and text-based communication media choice in global software development projects, in: 2009 Fourth IEEE International Conference on Global Software Engineer- ing, 2009, pp. 153–162. doi:10.1109/ICGSE.2009.23
-
[5]
M.-A. Storey, C. Treude, A. van Deursen, L.-T. Cheng, The impact of social media on software engineering practices and tools, in: Proceedings of the FSE/SDP Workshop on Future of Software Engineering Research, FoSER ’10, Association for Computing Machinery, New York, NY, USA, 2010, p. 359–364. doi:10.1145/1882362.1882435
arXiv 2010
-
[6]
D. Graziotin, X. Wang, P. Abrahamsson, Happy software developers solve problems better: psychological measurements in empirical software engi- neering, PeerJ 2 (2014) e289. doi:10.7717/peerj.289. 37
-
[7]
D. Graziotin, X. Wang, P. Abrahamsson, Do feelings matter? on the cor- relation of affects and the self-assessed productivity in software engineer- ing, Journal of Software: Evolution and Process 27 (7) (2015) 467–487. doi:10.1002/smr.1673
doi:10.1002/smr.1673 2015
-
[8]
De Choudhury, S
M. De Choudhury, S. Counts, Understanding Affect in the Workplace via Social Media, Association for Computing Machinery, New York, NY, USA, 2013, p. 303–316
2013
Show all 117 references
-
[9]
Guzman, D
E. Guzman, D. Az´ ocar, Y. Li, Sentiment analysis of commit comments in github: an empirical study, in: S. Kim, M. Pinzger, P. Devanbu (Eds.), 11th Working Conference on Mining Software Repositories : proceedings : May 31 - June 1, 2014, Hyderabad, India, ACM, [Place of public...
2014
-
[10]
Novielli, F
N. Novielli, F. Calefato, D. Dongiovanni, D. Girardi, F. Lanubile, Can we use se-specific sentiment analysis tools in a cross-platform setting?, Pro- ceedings of the 17th International Conference on Mining Software Repos- itoriesdoi:10.1145/3379597.3387446
-
[11]
Calefato, F
F. Calefato, F. Lanubile, N. Novielli, Emotxt: A toolkit for emotion recog- nition from text, in: 2017 Seventh International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW), IEEE, Piscataway, NJ, 2017, pp. 79–80. doi:10.1109/ACIIW.2017...
2017 doi
-
[12]
Ahmed, A
T. Ahmed, A. Bosu, A. Iqbal, S. Rahimi, Senticr: A customized sentiment analysis tool for code review interactions, in: 2017 32nd IEEE/ACM In- ternational Conference on Automated Software Engineering (ASE), IEEE, 2017, pp. 106–111. doi:10.1109/ASE.2017.8115623
2017
-
[13]
Z. Chen, Y. Cao, X. Lu, Q. Mei, X. Liu, Sentimoji: An emoji-powered learning approach for sentiment analysis in software engineering, in: Pro- ceedings of the 2019 27th ACM Joint Meeting on European Software En- 38 gineering Conference and Symposium on the Foundations of Softw...
2019
-
[14]
M. R. Islam, M. F. Zibran, Leveraging automated sentiment analysis in software engineering, in: 2017 IEEE/ACM 14th International Conference on Mining Software Repositories, IEEE, Piscataway, NJ, 2017. doi:10. 1109/msr.2017.9
2017
-
[15]
M. R. Islam, M. F. Zibran, Deva: sensing emotions in the valence arousal space in software engineering text, in: H. M. Haddad, R. L. Wainwright, R. Chbeir (Eds.), Applied computing 2018, Association for Computing Machinery Inc. (ACM), New York, NY, 2018, pp. 1536–1543. doi:10....
2018
-
[16]
Murgia, P
A. Murgia, P. Tourani, B. Adams, M. Ortu, Do developers feel emo- tions? an exploratory analysis of emotions in software artifacts, in: S. Kim, M. Pinzger, P. Devanbu (Eds.), 11th Working Conference on Mining Software Repositories : proceedings : May 31 - June 1, 2014, Hyderab...
2014
-
[17]
Novielli, F
N. Novielli, F. Calefato, F. Lanubile, Towards discovering the role of emo- tions in stack overflow, in: Proceedings of the 6th International Workshop on Social Software Engineering, SSE 2014, Association for Computing Ma- chinery, New York, NY, USA, 2014, p. 33–36
2014
-
[18]
Calefato, F
F. Calefato, F. Lanubile, F. Maiorano, N. Novielli, Sentiment polarity detection for software development, Empirical Software Engineering 23 (3) (2018) 1352–1382. doi:10.1007/s10664-017-9546-9
2018 doi
-
[19]
M. R. Islam, M. F. Zibran, Sentistrength-se: Exploiting domain specificity for improved sentiment analysis in software engineering text, Journal of Systems and Software 145 (2018) 125–146. doi:10.1016/j.jss.2018. 08.030. 39
2018 doi
-
[20]
Obaidi, J
M. Obaidi, J. Kl¨ under, Development and application of sentiment analy- sis tools in software engineering: A systematic literature review, in: Eval- uation and Assessment in Software Engineering, EASE 2021, Associa- tion for Computing Machinery, New York, NY, USA, 2021, p. 80...
2021
-
[21]
Kumar, A
A. Kumar, A. Jaiswal, Systematic literature review of sentiment analysis on twitter using soft computing techniques, Concurrency and Computa- tion: Practice and Experience 32 (1) (2020) e5107, e5107 CPE-18-1167.R1. doi:10.1002/cpe.5107
2020 doi
-
[22]
M. E. M. Abo, R. G. Raj, A. Qazi, A. Zakari, Sentiment analysis for arabic in social media network: A systematic mapping study (2019)
2019
-
[23]
Devika, C
M. Devika, C. Sunitha, A. Ganesh, Sentiment analysis: A comparative study on different approaches, Procedia Computer Science 87 (2016) 44 – 49, fourth International Conference on Recent Trends in Computer Science & Engineering (ICRTCSE 2016). doi:10.1016/j.procs.2016.05.124
2016 doi
-
[24]
J. Z. Maitama, N. Idris, A. Zakari, A systematic mapping study of the empirical explicit aspect extractions in sentiment analysis, IEEE Access 8 (2020) 113878–113899. doi:10.1109/ACCESS.2020.3003625
2020
-
[25]
Kastrati, F
Z. Kastrati, F. Dalipi, A. S. Imran, K. Pireva Nuci, M. A. Wani, Sentiment analysis of students’ feedback with nlp and deep learning: A systematic mapping study, Applied Sciences 11 (9). doi:10.3390/app11093986
-
[26]
Baragash, H
R. Baragash, H. Aldowah, Sentiment analysis in higher education: a sys- tematic mapping review, Journal of Physics: Conference Series 1860 (1) (2021) 012002. doi:10.1088/1742-6596/1860/1/012002
2021 doi
-
[28]
Imtiaz, J
N. Imtiaz, J. Middleton, P. Girouard, E. Murphy-Hill, Sentiment and politeness analysis tools on developer discussions are unreliable, but so are people, in: Proceedings of the 3rd International Workshop on Emotion Awareness in Software Engineering, SEmotion ’18, Associa- tion...
2018
-
[29]
A. S. M. Venigalla, S. Chimalakonda, Stackemo: Towards enhancing user experience by augmenting stack overflow with emojis, in: Proceed- ings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, ESEC...
2021
-
[30]
Calefato, F
F. Calefato, F. Lanubile, M. C. Marasciulo, N. Novielli, Mining successful answers in stack overflow, in: 2015 IEEE/ACM 12th Working Conference on Mining Software Repositories, 2015, pp. 430–433. doi:10.1109/MSR. 2015.56
2015 doi
-
[31]
Q. Umer, H. Liu, I. Illahi, Cnn-based automatic prioritization of bug reports, IEEE Transactions on Reliability (2020) 1–14 doi:10.1109/TR. 2019.2959624
2020
-
[32]
Werner, G
C. Werner, G. Tapuc, L. Montgomery, D. Sharma, S. Dodos, D. Damian, How angry are your customers? sentiment analysis of support tickets that escalate, in: D. Fucci, N. Novielli, E. Guzm´ an (Eds.), 2018 1st Interna- tional Workshop on Affective Computing for Requirements Engin...
2018
-
[33]
Zhang, B
T. Zhang, B. Xu, F. Thung, S. A. Haryono, D. Lo, L. Jiang, Senti- ment analysis for software engineering: How far can pre-trained trans- former models go?, in: 2020 IEEE International Conference on Software Maintenance and Evolution (ICSME), 2020, pp. 70–80. doi:10.1109/ ICSME...
2020
-
[34]
Devlin, M.-W
J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional transformers for language understanding (2019)
2019
-
[35]
Novielli, F
N. Novielli, F. Calefato, F. Lanubile, A. Serebrenik, Assessment of off- the-shelf se-specific sentiment analysis tools: An extended replication study, Empirical Software Engineering 26 (4) (2021) 77. doi:10.1007/ s10664-021-09960-w
2021
-
[36]
Nicole Novielli, Daniela Girardi, Filippo Lanubile, A benchmark study on sentiment analysis for software engineering research, in: 2018 IEEE/ACM 15th International Conference on Mining Software Repositories (MSR), 2018, pp. 364–375
2018
-
[37]
Biswas, K
E. Biswas, K. Vijay-Shanker, L. Pollock, Exploring word embedding techniques to improve sentiment analysis of software engineering texts, in: 2019 IEEE/ACM 16th International Conference on Mining Soft- ware Repositories, IEEE, Piscataway, NJ, 2019. doi:10.1109/msr.2019. 00020
2019 doi
-
[38]
M. R. Islam, M. F. Zibran, A comparison of software engineering domain specific sentiment analysis tools, in: 25th IEEE International Conference on Software Analysis, Evolution and Reengineering, IEEE, Piscataway, NJ, 2018. doi:10.1109/saner.2018.8330245
2018
-
[39]
Bin Lin, Nathan W. Cassee, Alexander Serebrenik, Gabriele Bavota, Nicole Novielli, Michele Lanza, Opinion mining for software development: A systematic literature review, ACM Transactions on Software Engineer- ing and Methodology XX (X)
-
[40]
S´ anchez-Gord´ on, R
M. S´ anchez-Gord´ on, R. Colomo-Palacios, Taking the emotional pulse of software engineering — a systematic literature review of empirical studies, Information and Software Technology 115 (2019) 23–43. doi:10.1016/j. infsof.2019.08.002. 42
2019 doi
-
[41]
Wohlin, P
C. Wohlin, P. Runeson, M. H¨ ost, M. C. Ohlsson, B. Regnell, A. Wessl´ en, Experimentation in software engineering, Springer, Berlin, 2012. doi: 10.1007/978-3-642-29044-2
2012 doi
-
[42]
Kitchenham, O
B. Kitchenham, O. Pearl Brereton, D. Budgen, M. Turner, J. Bailey, S. Linkman, Systematic literature reviews in software engineering – a sys- tematic literature review, Information and Software Technology 51 (1) (2009) 7 – 15, special Section - Most Cited Articles in 2002 and ...
2009 doi
-
[43]
Charters, Guidelines for performing Sys- tematic Literature Reviews in Software Engineering, Vol
Barbara Kitchenham, Stuart M. Charters, Guidelines for performing Sys- tematic Literature Reviews in Software Engineering, Vol. 2, 2007
2007
-
[44]
Petersen, R
K. Petersen, R. Feldt, S. Mujtaba, M. Mattsson, Systematic mapping studies in software engineering, BCS Learning & Development, 2008.doi: 10.14236/ewic/ease2008.8
2008 doi
-
[45]
Unger-Windeler, J
C. Unger-Windeler, J. Kl¨ under, K. Schneider, A mapping study on prod- uct owners in industry: Identifying future research directions, in: 2019 IEEE/ACM International Conference on Software and System Processes (ICSSP), 2019, pp. 135–144. doi:10.1109/ICSSP.2019.00026
2019
-
[46]
Yasin, R
A. Yasin, R. Fatima, L. Wen, W. Afzal, M. Azhar, R. Torkar, On us- ing grey literature and google scholar in systematic literature reviews in software engineering, IEEE Access 8 (2020) 36226–36243. doi:10.1109/ ACCESS.2020.2971712
2020
-
[47]
Shevtsov, M
S. Shevtsov, M. Berekmeri, D. Weyns, M. Maggio, Control-theoretical software adaptation: A systematic literature review, IEEE Transactions on Software Engineering 44 (8) (2018) 784–810. doi:10.1109/TSE.2017. 2704579
2018 doi
-
[48]
M. Kosa, M. Yilmaz, R. O’Connor, P. Clarke, Software engineering ed- ucation and games: A systematic literature review, Journal of Universal Computer Science 22 (2016) 1558–1574. 43
2016
-
[49]
Garousi, K
V. Garousi, K. Petersen, B. Ozkan, Challenges and best practices in industry-academia collaborations in software engineering: A systematic literature review, Information and Software Technology 79 (2016) 106–
2016
-
[50]
J. A.-C. Kl¨ under, P. Hohl, N. Prenner, K. Schneider, Transformation to- wards agile software product line engineering in large companies: A lit- erature review, Journal of Software: Evolution and Process 31 (5) (2019) e2168. doi:10.1002/smr.2168
2019 doi
-
[51]
Zahra, F
K. Zahra, F. Azam, F. Ilyas, H. Faisal, N. Ambreen, N. Gondal, Suc- cess factors of organizational change in software process improvement: A systematic literature review, in: Proceedings of the 5th International Conference on Information and Education Technology, ICIET ’17, As...
2017
-
[52]
Prenner, C
N. Prenner, C. Unger-Windeler, K. Schneider, How are hybrid develop- ment approaches organized? a systematic literature review, in: Proceed- ings of the International Conference on Software and System Processes, ICSSP ’20, Association for Computing Machinery, New York, NY, USA...
2020
-
[53]
Deshpande, V
M. Deshpande, V. Rao, Depression detection using emotion artificial in- telligence, in: 2017 International Conference on Intelligent Sustainable Systems (ICISS), 2017, pp. 858–862. doi:10.1109/ISS1.2017.8389299
2017
-
[54]
Liu, Sentiment analysis and opinion mining, Synthesis Lectures on Human Language Technologies 5 (1) (2012) 1–167
B. Liu, Sentiment analysis and opinion mining, Synthesis Lectures on Human Language Technologies 5 (1) (2012) 1–167. doi:10.2200/ S00416ED1V01Y201204HLT016
2012
-
[55]
B. Liu, L. Zhang, A Survey of Opinion Mining and Sentiment Anal- ysis, Springer US, Boston, MA, 2012, pp. 415–463. doi:10.1007/ 978-1-4614-3223-4_13 . 44
2012
-
[56]
X. Fang, J. Zhan, Sentiment analysis using product review data, Journal of Big Data 2 (1) (2015) 1–14. doi:10.1186/s40537-015-0015-2
2015 doi
-
[57]
M. Ortu, A. Murgia, G. Destefanis, P. Tourani, R. Tonelli, M. Marchesi, B. Adams, The emotional side of software developers in jira, in: Pro- ceedings of the 13th International Conference on Mining Software Repos- itories, MSR ’16, Association for Computing Machinery, New York...
2016
-
[58]
W. G. Parrott, Emotions in social psychology: Essential readings, psy- chology press, 2001
2001
-
[59]
C. Wohlin, Guidelines for snowballing in systematic literature studies and a replication in software engineering, in: Proceedings of the 18th Interna- tional Conference on Evaluation and Assessment in Software Engineering, EASE ’14, Association for Computing Machinery, New Yor...
-
[60]
Obaidi, L
M. Obaidi, L. Nagel, A. Specht, J. Kl¨ under, Dataset: Systematic Mapping Study on the Development and Application of Sentiment Analysis Tools in Software Engineering (Mar. 2022). doi:10.5281/zenodo.4726650
2022 doi
-
[61]
M. R. Islam, M. F. Zibran, Towards understanding and exploiting devel- opers’ emotional variations in software engineering, in: Y.-T. Song (Ed.), 2016 IEEE/ACIS 14th International Conference on Software Engineering Research, Management and Applications (SERA), IEEE, Piscataway...
2016
-
[62]
Z. Chen, Y. Cao, H. Yao, X. Lu, X. Peng, H. Mei, X. Liu, Emoji- powered sentiment and emotion detection from software developers’ com- munication data, ACM Trans. Softw. Eng. Methodol. 30 (2). doi: 10.1145/3424308
-
[63]
J. Wu, C. Ye, H. Zhou, Bert for sentiment classification in software en- 45 gineering, in: 2021 International Conference on Service Science (ICSS), 2021, pp. 115–121. doi:10.1109/ICSS53362.2021.00026
2021
-
[64]
Calefato, F
F. Calefato, F. Lanubile, N. Novielli, L. Quaranta, Emtk - the emotion mining toolkit, in: 2019 IEEE/ACM 4th International Workshop on Emo- tion Awareness in Software Engineering (SEmotion 2019), IEEE, Piscat- away, NJ, 2019, pp. 34–37. doi:10.1109/SEmotion.2019.00014
2019
-
[65]
M. E. Whiting, I. Gao, M. Xing, N. J. Diarrassouba, T. Nguyen, M. S. Bernstein, Parallel worlds: Repeated initializations of the same team to improve team viability 4 (CSCW1). doi:10.1145/3392877
-
[66]
A. F. Gkontzis, C. V. Karachristos, C. T. Panagiotakopoulos, E. C. Stavropoulos, V. S. Verykios, Sentiment analysis to track emotion and polarity in student fora, in: V. Vlachos (Ed.), Proceedings of the 21st Pan-Hellenic Conference on Informatics, ACM, New York, NY, 2017. doi...
2017
-
[67]
Guzman, Visualizing emotions in software development projects, in: A
E. Guzman, Visualizing emotions in software development projects, in: A. Telea (Ed.), 2013 First IEEE Working Conference on Software Visual- ization (VISSOFT), IEEE, Piscataway, NJ, 2013.doi:10.1109/vissoft. 2013.6650529
2013
-
[68]
Guzman, B
E. Guzman, B. Bruegge, Towards emotional awareness in software de- velopment teams, in: B. Meyer, M. Mezini, L. Baresi (Eds.), 2013 9th Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on the Foundations of Software Engineering (ESEC/...
2013 doi
-
[69]
A. M. El-Halees, Software usability evaluation using opinion mining, Jour- nal of Software 9 (2). doi:10.4304/jsw.9.2.343-349. 46
-
[70]
Nayebi, H
M. Nayebi, H. Farahi, G. Ruhe, Which version should be released to app store?, in: 11th ACM/IEEE International Symposium on Empirical Soft- ware Engineering and Measurement, IEEE, Piscataway, NJ, 2017, pp. 324–333. doi:10.1109/ESEM.2017.46
2017 doi
-
[71]
Kaewyong, A
P. Kaewyong, A. Sukprasert, N. Salim, F. Phang, The possibility of stu- dents’ comments automatic interpret using lexicon based sentiment anal- ysis to teacher evaluation, 2015
2015
-
[72]
K. Z. Aung, N. N. Myo, Sentiment analysis of students’ comment using lexicon based approach, in: G. Zhu (Ed.), 16th IEEE/ACIS International Conference on Computer and Information Science (ICIS 2017), IEEE, Piscataway, NJ, 2017, pp. 149–154. doi:10.1109/icis.2017.7959985
2017
-
[73]
Thelwall, K
M. Thelwall, K. Buckley, G. Paltoglou, Di Cai, A. Kappas, Sentiment strength detection in short informal text, Journal of the American Society for Information Science and Technology 61 (12) (2010) 2544–2558. doi: 10.1002/asi.21416
2010 doi
-
[74]
Thelwall, K
M. Thelwall, K. Buckley, G. Paltoglou, Sentiment strength detection for the social web, Journal of the American Society for Information Science and Technology 63 (1) (2012) 163–173. doi:10.1002/asi.21662
2012 doi
-
[75]
Cagnoni, L
S. Cagnoni, L. Cozzini, G. Lombardo, M. Mordonini, A. Poggi, M. Tomaiuolo, Emotion-based analysis of programming languages on stack overflow, ICT Express 6 (3) (2020) 238–242. doi:10.1016/j.icte. 2020.07.002
2020 doi
-
[76]
B. Lin, F. Zampetti, G. Bavota, M. Di Penta, M. Lanza, Pattern-based mining of opinions in q&a websites, in: 2019 IEEE/ACM 41st Interna- tional Conference on Software Engineering, IEEE, Piscataway, NJ, 2019. doi:10.1109/icse.2019.00066
2019
-
[77]
Kl¨ under, J
J. Kl¨ under, J. Horstmann, O. Karras, Identifying the mood of a software development team by analyzing text-based communication in chats with 47 machine learning, in: R. Bernhaupt, C. Ardito, S. Sauer (Eds.), Human- Centered Software Engineering, Springer International Publis...
2020
-
[78]
T. Yang, C. Gao, J. Zang, D. Lo, M. Lyu, Tour: Dynamic topic and sentiment analysis of user reviews for assisting app release, in: Com- panion Proceedings of the Web Conference 2021, WWW ’21, Associa- tion for Computing Machinery, New York, NY, USA, 2021, p. 708–712. doi:10.11...
2021
-
[79]
Kumar, A
A. Kumar, A. Abraham, Opinion mining to assist user acceptance test- ing for open-beta versions., Journal of Information Assurance & Security 12 (4)
-
[80]
M. R. Islam, M. K. Ahmmed, M. F. Zibran, Marvalous - machine learn- ing based detection of emotions in the valence-arousal space in software engineering text, in: C.-C. Hung (Ed.), Proceedings of the 34th ACMSI- GAPP Symposium on Applied Computing, ACM Digital Library, ACM, Ne...
2019
-
[81]
J. Shen, O. Baysal, M. O. Shafiq, Evaluating the performance of machine learning sentiment analysis algorithms in software engineering, in: 2019 IEEE Intl Conf on Dependable, Autonomic and Secure Computing, Intl Conf on Pervasive Intelligence and Computing, Intl Conf on Cloud ...
2019
-
[82]
Mostafa, M
L. Mostafa, M. Abd Elghany, Investigating game developers’ guilt emo- tions using sentiment analysis, Int. J. Softw. Eng. Appl.(IJSEA) 9 (2018)
2018
-
[83]
Murgia, M
A. Murgia, M. Ortu, P. Tourani, B. Adams, S. Demeyer, An exploratory qualitative and quantitative analysis of emotions in issue report comments 48 of open source systems, Empirical Software Engineering 23 (1) (2018) 521–
2018
-
[84]
A. Patwardhan, Sentiment identification for collaborative, geographically dispersed, cross-functional software development teams, in: 2017 IEEE 3rd International Conference on Collaboration and Internet Computing, IEEE, Piscataway, NJ, 2017. doi:10.1109/cic.2017.00014
2017
-
[85]
doi:10.5121/ijsea.2018.9604
2018
-
[86]
Cheriyan, B
J. Cheriyan, B. T. R. Savarimuthu, S. Cranefield, Towards offensive lan- guage detection and reduction in four software engineering communities, EASE 2021, Association for Computing Machinery, New York, NY, USA, 2021, p. 254–259. doi:10.1145/3463274.3463805
2021
-
[87]
Loper, S
E. Loper, S. Bird, Nltk: The natural language toolkit (2002)
2002
-
[88]
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, V. Stoyanov, Roberta: A robustly optimized bert pre- training approach (2019)
2019
-
[89]
Shaver, J
P. Shaver, J. Schwartz, D. Kirson, C. O’Connor, Emotion knowledge: Further exploration of a prototype approach, Journal of Personality and Social Psychology 52 (6) (1987) 1061–1086. doi:10.1037/0022-3514. 52.6.1061
1987 doi
-
[90]
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, R. Soricut, Albert: A lite bert for self-supervised learning of language representations (2020). arXiv:1909.11942
2020 arXiv
-
[91]
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, Q. V. Le, Xlnet: Generalized autoregressive pretraining for language understanding, 49 in: H. Wallach, H. Larochelle, A. Beygelzimer, F. d 'Alch´ e-Buc, E. Fox, R. Garnett (Eds.), Advances in Neural Information Proce...
2019
-
[92]
Manning, M
C. Manning, M. Surdeanu, J. Bauer, J. Finkel, S. Bethard, D. McClosky, The stanford corenlp natural language processing toolkit, in: Proceedings of 52nd Annual Meeting of the Association for Computational Linguis- tics: System Demonstrations, Association for Computational Ling...
2014 doi
-
[93]
Holmes, A
G. Holmes, A. Donkin, I. Witten, Weka: a machine learning workbench, in: Proceedings of ANZIIS ’94 - Australian New Zealnd Intelligent Infor- mation Systems Conference, 1994, pp. 357–361. doi:10.1109/ANZIIS. 1994.396988
1994
-
[94]
A. Kaur, A. P. Singh, G. S. Dhillon, D. Bisht, Emotion mining and sentiment analysis in software engineering domain, in: Proceedings of the Second International Conference on Electronics, Communication and Aerospace Technology (ICECA 2018), IEEE, Piscataway, NJ, 2018, pp. 1170...
2018
-
[95]
Ferreira, J
I. Ferreira, J. Cheng, B. Adams, The ”shut the f**k up” phe- nomenon: Characterizing incivility in open source code review discussions 5 (CSCW2). doi:10.1145/3479497
-
[96]
Uddin, F
G. Uddin, F. Khomh, Automatic mining of opinions expressed about apis in stack overflow, IEEE Transactions on Software Engineering 47 (3) (2021) 522–559. doi:10.1109/TSE.2019.2900245
2021
-
[97]
Jongeling, P
R. Jongeling, P. Sarkar, S. Datta, A. Serebrenik, On negative results when using sentiment analysis tools for software engineering research, Empirical Software Engineering 22 (5) (2017) 2543–2584. doi:10.1007/ s10664-016-9493-x
2017
-
[98]
K.-i. Park, B. Sharif, Assessing perceived sentiment in pull requests with emoji: Evidence from tools and developer eye movements, in: 50 2021 IEEE/ACM Sixth International Workshop on Emotion Awareness in Software Engineering (SEmotion), 2021, pp. 1–6. doi:10.1109/ SEmotion525...
2021
-
[99]
Novielli, F
N. Novielli, F. Calefato, D. Dongiovanni, D. Girardi, F. Lanubile, A gold standard for polarity of emotions of software developers in github (2020). doi:10.6084/m9.figshare.11604597
2020 doi
-
[100]
M. R. Islam, M. F. Zibran, Sentiment Analysis of Software Bug Related Commit Messages, ISCA, 2018
2018
-
[101]
L. A. Cabrera-Diego, N. Bessis, I. Korkontzelos, Classifying emotions in stack overflow and jira using a multi-label approach, Knowledge-Based Systems 195 (2020) 105633. doi:10.1016/j.knosys.2020.105633
2020
-
[102]
K. Sun, H. Gao, H. Kuang, X. Ma, G. Rong, D. Shao, H. Zhang, Ex- ploiting the unique expression for improved sentiment analysis in soft- ware engineering text, in: 2021 IEEE/ACM 29th International Con- ference on Program Comprehension (ICPC), 2021, pp. 149–159. doi: 10.1109/IC...
2021
-
[103]
R. E. Fabry, Getting it: A predictive processing approach to irony comprehension, Synthese 198 (7) (2021) 6455–6489. doi:10.1007/ s11229-019-02470-9
2021
-
[104]
Akula, I
R. Akula, I. Garibay, Interpretable multi-head self-attention architecture for sarcasm detection in social media, Entropy 23 (4). doi:10.3390/ e23040394
-
[105]
B. Lin, F. Zampetti, G. Bavota, M. Di Penta, M. Lanza, R. Oliveto, Sentiment analysis for software engineering: How far can we go?, in: Pro- ceedings of the 40th International Conference on Software Engineering, ICSE ’18, Association for Computing Machinery, New York, NY, USA,...
2018
-
[106]
Ahasanuzzaman, M
M. Ahasanuzzaman, M. Asaduzzaman, C. K. Roy, K. A. Schneider, Caps: a supervised technique for classifying stack overflow posts concerning api issues, Empirical Software Engineering 25 (2) (2020) 1493–1532. doi: 10.1007/s10664-019-09743-4
2020 doi
-
[107]
Guzman, R
E. Guzman, R. Alkadhi, N. Seyff, An exploratory study of twitter messages about software applications, Requirements Engineering 22 (3) (2017) 387–
2017
-
[108]
Novielli, F
N. Novielli, F. Calefato, F. Lanubile, The challenges of sentiment detec- tion in the social programmer ecosystem, in: I. Hammouda, A. Sillitti (Eds.), Proceedings of the 7th International Workshop on Social Software Engineering - SSE 2015, SSE 2015, ACM Press, New York, New Y...
2015
-
[109]
Tourani, Y
P. Tourani, Y. Jiang, B. Adams, Monitoring sentiment in open source mailing lists: Exploratory study on the apache ecosystem, in: Proceed- ings of 24th Annual International Conference on Computer Science and Software Engineering, CASCON ’14, IBM Corp., USA, 2014, p. 34–44. 51
2014
-
[110]
Mansoor, C
N. Mansoor, C. S. Peterson, B. Sharif, How developers and tools categorize sentiment in stack overflow questions - a pilot study, in: 2021 IEEE/ACM Sixth International Workshop on Emotion Awareness in Software En- gineering (SEmotion), 2021, pp. 19–22. doi:10.1109/SEmotion5256...
2021
-
[111]
Plutchik, Chapter 1 - a general psychoevolutionary theory of emotion, in: R
R. Plutchik, Chapter 1 - a general psychoevolutionary theory of emotion, in: R. Plutchik, H. Kellerman (Eds.), Theories of Emotion, Academic Press, 1980, pp. 3–33. doi:10.1016/B978-0-12-558701-3.50007-7
1980 doi
-
[112]
Watson, L
D. Watson, L. Clark, A. Tellegen, Development and validation of brief measures of positive and negative affect: the panas scales, Journal of personality and social psychology 54 (6) (1988) 1063–1070. doi:10.1037/ /0022-3514.54.6.1063. 52
1988
-
[113]
M. H. Asyrofi, Z. Yang, I. N. B. Yusuf, H. J. Kang, F. Thung, D. Lo, Biasfinder: Metamorphic test generation to uncover bias for sentiment analysis systems (2021). arXiv:2102.01859. 53
2021 arXiv
-
[114]
Panichella, A
S. Panichella, A. Di Sorbo, E. Guzman, C. A. Visaggio, G. Canfora, H. C. Gall, How can i improve my app? classifying user reviews for software maintenance and evolution, in: 2015 IEEE International Conference on Software Maintenance and Evolution (ICSME), 2015, pp. 281–290. do...
2015
-
[127]
doi:10.1016/j.infsof.2016.07.006
2016 doi
-
[412]
doi:10.1007/s00766-017-0274-x
-
[564]
doi:10.1007/s10664-017-9526-0
-
[2014]
doi:10.1145/2601248.2601268
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.