REVIEW 4 major objections 5 minor 37 references
Large-Scale Analysis of Discussions by CS Educators Across the Stack Exchange Network
T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Across 748,084 posts by CS educators on Stack Exchange, 55 discussion topics emerge, with IT at 64.88% and non-IT at 35.12%, and non-technical topics growing steadily from 2013 to 2018.
desk verdict Plausible first map of CS-educator talk across Stack Exchange, but the corpus definition is sloppy enough that the headline split isn't trustworthy yet. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine of the analysis is Latent Dirichlet Allocation (LDA), a probabilistic topic model run through the MALLET library. The authors preprocess all posts by stripping code and HTML, removing stop words, and lemmatizing with spaCy, then choose the number of topics K by measuring c_v coherence across K values from 5 to 70; K=55 scores highest (0.6318) and is adopted. Each of the 55 topics is manually labeled using open card sorting, and the labels are organized hierarchically into two categories and nine subcategories. To measure change over time, the authors assign each post to its dominant topic and compute monthly absolute and relative impact metrics for the categories.
What would settle it
Take a random sample of posts excluded by the accepted-answer filter, apply the same preprocessing and LDA model, and compare the resulting category shares against the reported 64.88% IT / 35.12% Non-IT split; a material shift would show the filter distorted the claimed distribution.
Extended reading notes
Core claim
The paper's central claim is that CS educators are not confined to technical discussion: across the Stack Exchange network they participate in a broad and evolving topic space, with IT topics (64.88%) dominating but non-IT topics (35.12%) forming a substantial and growing share. Using LDA topic modeling on a filtered corpus of 397,061 questions and 351,023 accepted answers from 12,949 CS Educators users, the authors identify 55 topics, manually label them, and organize them into two categories and nine subcategories. The largest IT subcategory, Programming and Software Development, accounts for about 23% of posts; the largest non-IT subcategory, Mathematics, accounts for about 8.22%. The temporal analysis shows IT leading absolute volume throughout 2008-2024, while non-IT topics gained relative ground after 2014, with steady growth from 2013 to 2018.
Load-bearing premise
The findings rest on the assumption that posts by users of the CS Educators site, filtered to questions that received an accepted answer, fairly represent what CS educators talk about across the network; if unanswered questions or unaccepted answers cover different topics, the 64.88/35.12 split and the 2013-2018 growth trend would be biased.
Editorial extensions
If this is right
- If the taxonomy holds, the nine-subcategory hierarchy gives educators, trainers, and platform designers a ready-made map of where CS educators actually spend their attention.
- If the relative-impact trend continues, non-technical topics, especially mathematics, humanities, and lifestyle, will claim a growing share of CS educators' discussion, so professional development that ignores them will miss a real part of the job.
- The 55-topic hierarchy can be reused as a baseline for comparing other expert communities, such as math educators or engineering educators, on the same network.
- Because the absolute impact of IT topics remained highest throughout 2008-2024, any support system for CS educators must keep technical topics as the core while expanding to interdisciplinary ones.
Reading between the lines
- Beyond the paper: the accepted-answer filter probably excludes questions that never got resolved, often the hardest or most niche ones, so the reported 64.88/35.12 split may understate how much of CS educators' discussion is non-IT or emerging.
- Beyond the paper: the 2013-2018 non-IT growth could partly be a platform-wide trend; comparing CS educators against a matched sample of other Stack Exchange users would isolate what is specific to educators.
- Beyond the paper: the substantial mathematics, physics, and humanities activity suggests CS educators use the network for their own subject-matter learning, not only for teaching advice; a follow-up study could test whether such cross-disciplinary posting correlates with curriculum or course-design choices.
- Beyond the paper: a direct extension would be an inter-rater reliability study of the manual labels; quantifying agreement on a sample would show how stable the 55-topic taxonomy is.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper analyzes posts by users of the CS Educators Stack Exchange across 169 English-language Stack Exchange sites, using LDA topic modeling (K=55) on a corpus described inconsistently as 748,084 or 784,048 posts, manually labeling topics and grouping them into IT and Non-IT categories. For RQ1, the paper reports that 64.88% of posts are IT-related and 35.12% are Non-IT, with programming and software development dominant in IT and mathematics dominant in Non-IT. For RQ2, it reports absolute and relative impact trends, including a claim of steady Non-IT growth from 2013 to 2018. The paper positions this as the first cross-network analysis of CS educators' discussions on Stack Exchange.
Significance. If the empirical findings withstand scrutiny, the study provides useful descriptive evidence about the breadth of CS educators' participation across Stack Exchange, going beyond prior single-site studies. The cross-network user tracing via AccountId, the large dataset, and the coherence-based selection of the LDA topic count K are notable strengths. However, the central percentages and trend claims are currently undermined by unresolved corpus count inconsistencies, an unvalidated accepted-answer filter, absent inter-rater reliability, and flaws in the relative-impact definition. These issues are fixable, but they are load-bearing for the paper's main claims, so the manuscript requires major revision.
major comments (4)
- [Section 3.1, Table 1] The corpus size and selection rule are inconsistent. The text says the final dataset comprised 748,084 entries (397,061 questions + 351,023 accepted answers), while Table 1 lists 784,048 posts used for topic modeling; since every percentage in RQ1 is computed on this corpus, the discrepancy must be resolved. Moreover, if only questions with accepted answers are retained, the number of questions should equal the number of accepted answers; the 46,038 difference between 397,061 and 351,023 suggests either questions without accepted answers were kept, accepted answers include answers by non-CS-educators, or the selection rule is misstated. Please clarify the exact filtering and report a single corpus count.
- [Section 3.1, RQ1] The decision to analyze only accepted answers (and possibly only questions with accepted answers) is not validated. Excluding unaccepted answers removes a large fraction of CS educators' posts, and if accepted answers are more likely to be attached to well-specified technical questions, the reported 64.88% IT share would be biased upward. The paper should compare topic distributions on the full corpus of CS-educator posts, or at least on questions plus all answers, and report whether the IT/Non-IT split is stable under this filter.
- [Section 3.2, manual labeling] The manual labeling and hierarchical grouping drive all quantitative claims in RQ1 and RQ2, but no inter-rater reliability is reported. The text says 'two raters' and 'over 20 Zoom-based iterations,' yet no agreement metric (e.g., Cohen's kappa) or confusion analysis is given. Please report quantified agreement for topic labels and category assignments, or justify why label noise cannot materially affect the 64.88/35.12 split and the trend results.
- [Section 5, Equations (4) and (5)] Equations (4) and (5) are internally inconsistent. Equation (4) defines P(m) as the total number of posts in month m that contain topic x_n; with that definition the ratio is identically 1. The indicator in Equation (4) refers to 'topic x_n is present' whereas Equation (2) defines dominance, so the two metrics are not aligned. Equation (5) sums relative impacts over topics, which may exceed 1 if a post can contain multiple topics. Please correct the definitions and clarify whether relative impact is a proportion of posts whose dominant topic is in the category. In addition, the claim of 'steady growth from 2013 to 2018' in Section 5.1.1 is stated without trend fitting or confidence intervals, so the reader cannot assess whether the growth is meaningful.
minor comments (5)
- [Section 2] There is a typo in 'CS Eucators'; it should be 'CS Educators'.
- [Section 4.1 and Summary of RQ1] The subcategory percentage for Programming and Software Development is given as 23.00% in Section 4.1 and 23.04% in the Summary of RQ1; also, the listed subcategory percentages for IT and Non-IT do not sum to exactly 100%. Please reconcile the numbers.
- [References] References [33] and [34] are the same paper and should be deduplicated.
- [Abstract and Introduction] The figure 79,854,463 refers to the full Stack Exchange corpus, not the analyzed subset; please state this explicitly to avoid confusion with the 748,084/784,048 entries actually used.
- [Figures 5 and 6] The figures for absolute and relative impact are not included in the provided manuscript; when finalized, ensure the axes are labeled, the legend is legible, and the lines are clearly identified as raw counts or smoothed trend lines.
Circularity Check
No significant circularity: the study makes no predictive claim that reduces to its LDA parameters, corpus filter, or author-defined topic labels.
full rationale
The paper reports a descriptive topic-model analysis. LDA is fitted to the CS-educator post corpus and the resulting topic proportions, category shares (IT 64.88%, Non-IT 35.12%), and time trends are summaries of that fitted model on the same data; because no out-of-sample or causal prediction is claimed, there is no fitted input being renamed as a prediction. The manual grouping of labeled topics into IT and Non-IT categories is a taxonomy choice, not a derivation that presupposes the reported percentages. The only self-citation involving an author (ref. [34], used to justify the absolute/relative impact metrics, with overlapping author Omar Alam) is methodological and standard, and the central quantitative claims do not rest on it. The accepted-answer-only corpus filter and the discrepancy between 748,084 entries and the 784,048 posts used for topic modeling are potential threats to external validity and reporting consistency, but they are not circular reasoning: the study's conclusions describe the filtered corpus it analyzed, and no claim is made that this filter was independently validated against the unfiltered corpus. No step in the derivation chain is equivalent by construction to its input.
Assumptions & free parameters
free parameters (1)
- Number of LDA topics K =
55
assumptions (3)
- domain assumption Users of the CS Educators Stack Exchange site are treated as a proxy for CS educators generally.
- domain assumption Posts with accepted answers are representative of all CS educator discussions.
- domain assumption Manual labels assigned by two raters using card sorting are valid topic interpretations.
Cite this review
Pith. "Pith review of Large-Scale Analysis of Discussions by CS Educators Across the Stack Exchange Network." pith.science (2026). https://pith.science/paper/X7Z6EUVE
@misc{pith2026260804352,
author = {Pith},
title = {Pith review of: Large-Scale Analysis of Discussions by CS Educators Across the Stack Exchange Network},
year = {2026},
howpublished = {\url{https://pith.science/paper/X7Z6EUVE}},
note = {Machine review of arXiv:2608.04352}
}
read the original abstract
Stack Exchange is a widely used question-and-answer network that facilitates knowledge exchange across diverse domains. Within this network, the Computer Science (CS) Educators Stack Exchange provides a dedicated platform where CS educators exchange ideas, seek advice, and discuss teaching practices. In this study, we analyzed 79,854,463 Stack Exchange posts, comprising 32,187,805 questions and 47,666,658 answers, with a particular focus on English-language posts contributed by CS Educators participants. Using topic modeling, we identified, manually labeled, and hierarchically organized the underlying discussion topics, then examined their distribution and complexity. Our findings reveal evolving discussion patterns spanning both technical (IT) and non-technical (Non-IT) domains. Within the IT category, programming and software development were the most prominent topics, whereas mathematics, education, and the humanities received substantial attention within the Non-IT category. These results highlight the broad range of interests and expertise shared by CS educators and provide insights into their evolving priorities. We hope this work contributes to a better understanding of CS educators' knowledge-sharing practices and informs future efforts to better support their professional and educational needs.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Ahmad Abdellatif, Diego Costa, Khaled Badran, Rabe Abdalkareem, and Emad Shihab. 2020. Challenges in Chatbot Development: A Study of Stack Overflow Posts. InProceedings of the 17th International Conference on Mining Software Repositories(Seoul, Republic of Korea)(MSR ’20). Association for Computing Machinery, New York, NY, USA, 174–185. doi:10.1145/337959...
-
[2]
A. Agrawal, W. Fu, and T. Menzies. 2018. What is wrong with topic modeling? and how to fix it using search-based software engineering.Information and Software Technology98 (2018), 74–88
work page 2018
- [3]
-
[4]
R. Arun, V. Suresh, C. E. V. Madhavan, and M. N. N. Murthy. 2010. On Finding the Natural Number of Topics with Latent Dirichlet Allocation: Some Observations. InProceedings of the 14th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining. 391–402
work page 2010
-
[5]
Ask Ubuntu. 2024. Ask Ubuntu. https://askubuntu.com/
work page 2024
-
[6]
M. Bagherzadeh and R. Khatchadourian. 2019. Going big: A large-scale study on what big data developers ask. InProceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE). 432–442
work page 2019
-
[7]
Alan Bandeira, Carlos Alberto Medeiros, Matheus Paixao, and Paulo Henrique Maia. 2019. We Need to Talk About Microservices: an Analysis from the Dis- cussions on StackOverflow. In2019 IEEE/ACM 16th International Conference on Mining Software Repositories (MSR). 255–259. doi:10.1109/MSR.2019.00051
- [8]
Show all 37 references
-
[9]
Steven Bird, Ewan Klein, and Edward Loper. 2016. NLTK: Natural Language Toolkit. http://www.nltk.org/howto/sentiment.html
2016
-
[10]
Guillermo Blanco, Roi Pérez-López, Florentino Fdez-Riverola, and Anália Maria Garcia Lourenço. 2020. Understanding the social evolution of the Java com- munity in Stack Overflow: A 10-year study of developer interactions.Future Gen- eration Computer Systems105 (2020), 446–454....
2020 doi
-
[11]
Blei, Andrew Y
David M. Blei, Andrew Y. Ng, and Michael I. Jordan. 2003. Latent Dirichlet allocation.Journal of Machine Learning Research3 (2003), 993–1022. http: //www.jmlr.org/papers/v3/blei03a.html
2003
-
[12]
Partha Chakraborty, Rifat Shahriyar, Anindya Iqbal, and Gias Uddin. 2021. How do developers discuss and support new programming languages in technical Q&A site? An empirical study of Go, Swift, and Rust in Stack Overflow.Information and Software Technology137 (2021), 106603. d...
2021
-
[13]
T.-H. P. Chen, S. W. Thomas, and A. E. Hassan. 2016. A survey on the use of topic models when mining software repositories. InProceedings of the 2016 ACM SIGSOFT International Symposium on Empirical Software Engineering and Measurement (ESEM). 1843–1919
2016
-
[14]
CS Educators Stack Exchange. 2024. CS Educators Stack Exchange. https:// cseducators.stackexchange.com
2024
-
[15]
Explosion AI. 2025. Lemmatizer — spaCy API Documentation. https://spacy.io/ api/lemmatizer. Accessed: 2025-02-01
2025
-
[16]
Nikolas Gordon and Omar Alam. 2020. The Role of Race and Gender in Teach- ing Evaluation of Computer Science Professors: A Large Scale Analysis on RateMyProfessor Data. InProceedings of the 51st ACM Technical Symposium on Computer Science Education. Association for Computing M...
2020
-
[17]
Moshe Hazoom, Vibhor Malik, and Ben Bogin. 2021. Text-to-SQL in the Wild: A Naturally-Occurring Dataset Based on Stack Exchange Data. arXiv:2106.05006 [cs.CL] https://arxiv.org/abs/2106.05006
2021 arXiv
-
[18]
W. Hudson. 2013.Card Sorting(2nd ed.). The Interaction Design Foundation
2013
-
[19]
S Lal and R Mourya. 2022. For CS Educators, by CS Educators: An Exploratory Analysis of Issues and Recommendations for Online Teaching in Computer Science. Societies 2022, 12, 116
2022
-
[20]
Andrew McCallum. 2002. MALLET: A Machine Learning for Language Toolkit. http://mallet.cs.umass.edu/
2002
-
[21]
Sukanya Kannan Moudgalya, Kathryn M Rich, Aman Yadav, and Matthew J Koehler. 2019. Computer science educators stack exchange: Perceptions of equity and gender diversity in computer science. InProceedings of the 50th ACM Technical Symposium on Computer Science Education. 1197–1203
2019
-
[23]
Röder, A
M. Röder, A. Both, and A. Hinneburg. 2015. Exploring the Space of Topic Coher- ence Measures. InProceedings of the Eighth ACM International Conference on Web Search and Data Mining. 399–408
2015
-
[24]
Eder Santos, Felipe Gomes, Sávio Freire, Manoel Mendonça, Thiago Mendes, and Rodrigo Spínola. 2023. Technical Debt on Agile Projects: Managers’ point of view at Stack Exchange. InProceedings of the 2023 ACM/IEEE International Conference on Technical Debt (TechDebt). 1–9. doi:1...
2023
-
[25]
Stack Exchange. 2024. Stack Exchange API. https://api.stackexchange.com/. https://api.stackexchange.com/
2024
-
[26]
Stack Exchange. 2024. Stack Exchange Sites. https://stackexchange.com/sites# oldest. Online; accessed 2023-07
2024
-
[27]
Stack Exchange, Inc. 2024. Stack Exchange Data Dump. https://archive.org/ details/stackexchange. https://archive.org/details/stackexchange
2024
-
[28]
Stack Overflow. 2024. Stack Overflow. https://stackoverflow.com
2024
-
[29]
X. Sun, B. Li, Y. Li, and Y. Chen. 2015. What information in software historical repositories do we need to support software maintenance tasks? An approach based on topic model.Computer and Information Science(2015), 22–37
2015
-
[30]
Super User. 2024. Super User. https://superuser.com
2024
-
[31]
Mohammad Tahaei, Kami Vaniea, and Naomi Saphra. 2020. Understanding Privacy-Related Questions on Stack Overflow. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA)(CHI ’20). Association for Computing Machinery, New York, NY, USA, ...
2020
-
[32]
Amjed Tahir, Jens Dietrich, Steve Counsell, Sherlock Licorish, and Aiko Yamashita
-
[34]
Gias Uddin, Fatima Sabir, Yann-Gaël Guéhéneuc, Omar Alam, and Foutse Khomh
-
[35]
Z. Xie, J. Zhao, and H. Liu. 2020. Understanding content dynamics in online Q&A communities.Computer Networks168 (2020), 107104
2020
-
[36]
X.-L. Yang, D. Lo, X. Xia, Z.-Y. Wan, and J.-L. Sun. 2016. What security questions do developers ask? A large-scale study of Stack Overflow posts.Journal of Computer Science and Technology31, 5 (2016), 910–924
2016
-
[37]
doi:10.1007/s10664-021- 10021-5
An Empirical Study of IoT Topics in IoT Developer Discussions on Stack Overflow.Empirical Software Engineering26, 6 (2021). doi:10.1007/s10664-021- 10021-5
2021 doi
-
[40]
Mansooreh Zahedi, Roshan Namal Rajapakse, and Muhammad Ali Babar. 2020. Mining Questions Asked about Continuous Software Engineering: A Case Study of Stack Overflow. InProceedings of the 24th International Conference on Evaluation and Assessment in Software Engineering(Trondhe...
2020
-
[2020]
doi:10.1016/j.infsof.2020.106333
A large scale study on how developers discuss code smells and anti-pattern in Stack Exchange sites.Information and Software Technology125 (2020), 106333. doi:10.1016/j.infsof.2020.106333
2020
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.