REVIEW 4 major objections 6 minor 30 references
Characterizing Data Scientists in the Real World
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A 116-respondent survey finds data scientists are highly qualified and that their academic background has little impact on how they work.
desk verdict Useful descriptive snapshot of 116 data science workers, but the central claim that academic background has little impact is not supported by the reported analysis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The analytical engine is the 'CS background?' attribute, a derived binary variable that records whether a respondent had any formal training in computer science. The authors cross this variable with every work-related response—self-rated task confidence, reported difficulties, time spent coding, and technology choices—using contingency tables and Fisher's exact tests to decide which observed differences are statistically meaningful. The same machinery is applied to years of experience and gender, making it the device that carries the paper's main argument that background has little impact.
What would settle it
Run the same survey on a stratified sample drawn from company employment records, professional certifications, or national labor statistics instead of open online forums, and compare the proportions for CS background, task confidence, and reported difficulties. If the background effects or difficulty rankings differ materially from the paper's numbers, the claim that academic background has little impact on work would fail as a general statement.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that the stereotype of under-trained data science workers does not hold: the respondents are mostly men aged 26 to 45, overwhelmingly degree-holders, and about 62 percent have formal computer science training. Their academic past, however, does not decide how they work: confidence in most tasks, the difficulties they report, and their analytical goals are broadly shared, with only a few statistically significant differences tied to a CS background, such as deep-learning confidence and some language choices. The professionals' main shared problems are poor-quality data, difficult access to relevant data, and unclear questions to answer, and they report spending up to half their working time coding. The paper also reports a gender gap: many fewer women than men responded, and women report lower job satisfaction.
Load-bearing premise
The load-bearing assumption is that the 116 people who filled out the survey, recruited through online forums and personal contacts, represent data science professionals worldwide; if the sample is skewed, the profile and the conclusion that academic background has little impact describe only this respondent pool.
Editorial extensions
If this is right
- Tool and training efforts should target the shared bottlenecks—access to quality data and applying deep learning—rather than remedial computer science education.
- Because Python, SQL, and spreadsheet editors dominate across both groups, improving interoperability around these tools would serve most practitioners.
- Self-reported confidence in deep learning is the clearest skill gap, and it is larger among professionals without a CS background, so focused support there would reach a real need.
- The gender imbalance and lower satisfaction reported by women point to an equity issue that additional technical training alone would not address.
Reading between the lines
- Beyond the paper, clustering respondents on the tasks and technologies they report would test whether academic background truly leaves no signature; if it does not, the clusters should not separate by degree field.
- Beyond the paper, the recruitment through programming-centric forums and personal contacts may over-represent code-comfortable practitioners, so the high qualification level could be an upper-bound estimate; a registry-based or employer-based sample would check this.
- Beyond the paper, the deep-learning gap is read as a skills gap, but it could equally be an exposure gap; comparing confidence among practitioners who have completed deep-learning projects, regardless of degree, would separate the two.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents results of a survey of 116 self-selected data science practitioners, conducted from April to December 2020 via online forums, social media, and personal contacts. The survey collected academic background, professional experience, self-assessed task proficiency, work difficulties, and technology use. The authors use descriptive statistics and Fisher's exact tests to answer three research questions about the profile of data scientists, the impact of profile on work, and technology choices. The main conclusion is that data scientists are highly qualified and that academic background (CS vs. non-CS) has little impact on how they work; they also report a gender gap and lower satisfaction among women.
Significance. If the descriptive patterns are representative, the paper provides a useful snapshot for tool designers and educators: data quality and data access are the most common difficulties, deep learning is the least self-confident task, and Python/SQL dominate. The authors are transparent about their data preparation and they emphasize that the survey was public. The paper also gains credibility from using a reproducible R script for Fisher's tests (Section 3.3) and from reporting frequencies and modes transparently. However, the paper's central null claim about academic background is not statistically supported, and the sampling strategy limits population-level generalizations. The manuscript would be a worthwhile descriptive report after major revisions to align conclusions with the evidence.
major comments (4)
- [Section 7 and Sections 4.3/4.4] The conclusion that 'academic past has little impact on the way they work' is an inference from null results. In Section 4.3, the authors state that for satisfaction crossed with experience and background 'we could not find significant differences' but report no p-values, effect sizes, or confidence intervals, and no equivalence tests. With n=116 and subgroups as small as 29 women and 44 non-CS respondents, non-significant Fisher tests do not demonstrate that population differences are small or absent. The authors should either run equivalence tests with pre-specified bounds, report effect sizes for all comparisons, or restrict the conclusion to the sample.
- [Section 4.4 and Section 3.3] The only reported inferential statistic is a Fisher's exact test for deep learning self-efficacy with p=0.049. This is borderline, and no correction for multiple comparisons is described even though many contingency tables were tested across Sections 4.3-4.6. The conclusion that 'those with a CS background feel more apt to apply deep learning techniques' should be supported by adjusted p-values, effect sizes (e.g., odds ratio), and confidence intervals.
- [Section 3.1 and Section 6] The representativeness of the sample is not established. The survey was distributed through convenience channels, with no sampling frame, response rate, or comparison to known population characteristics of data scientists. The responses are heavily concentrated in Portugal (44/116) and male (87/116). The Threats to validity section acknowledges sampling errors but only asserts that multi-platform distribution helped; it does not quantify non-response or self-selection. Consequently, RQ2 and RQ3's population-level claims are not supported. The authors should reframe findings as descriptive of the sample or provide external benchmarks.
- [Sections 4.5-4.6 and Section 7] The descriptive results contain numerous CS/non-CS differences that contradict the 'little impact' summary: people without a CS background spend more time coding (Figure 9), the two groups differ in IDE, language, and statistics tool choices (Figures 11-15), and self-reported lack of data science skills is distributed differently (Figure 7). The conclusion that academic background has little impact is not reconciled with these descriptive differences. The authors need to either qualify the conclusion (e.g., 'little impact on job satisfaction' or 'limited to specific tasks/tools') or conduct a multivariate analysis that accounts for background along with other variables.
minor comments (6)
- [Section 4.4] The text uses 'Fischer's test'; this should be 'Fisher's exact test'.
- [Table 1] The column headers 'F emale' and 'T otal' contain typos, and the table layout would benefit from clearer alignment.
- [Section 4.2] The phrase 'background inNatural Sciences' is missing a space, and the italicized conclusion 'data science professionals reported being highly qualified' is presented without a supporting statistical test or benchmark.
- [Section 6] The sentence 'Participants may fill influenced to answer questions' is ungrammatical; it should likely read 'feel influenced'.
- [Section 3.3] The R script used for Fisher's tests is mentioned but not included; providing the script and the anonymized dataset would strengthen reproducibility.
- [Abstract] The opening statistic (64.2 zettabytes) is not cited in the abstract; a reference to the IDC report should be added.
Circularity Check
No circularity: the survey's conclusions are descriptive and inferential summaries of collected data, not quantities fitted to or derived from themselves.
full rationale
This paper is an empirical survey study; it contains no derivation, no fitted parameters, no formulas, and no reliance on prior results by the same authors. The conclusions in Sections 5 and 7 (e.g., 'people in data science are generally highly qualified... their academic past has little impact on the way they work') are summaries of the 116 survey responses and of Fisher's exact tests applied to contingency tables. Those tests compare observed response distributions across groups; they are not constructed so that the tested quantity equals the group definition. The only identified fragility is statistical and sampling validity: the convenience sample is not shown to be representative, and non-significant Fisher tests are treated as evidence of absence in RQ2. These are threats to external validity and to the support for a null conclusion, not circularity under the definitions used here. There are no self-citations, no imported uniqueness theorems, and no ansatz smuggled in via citation.
Assumptions & free parameters
assumptions (3)
- domain assumption Respondents self-identifying as data science workers are indeed data science professionals
- domain assumption Self-reported skills, satisfaction, and difficulties accurately reflect actual behavior
- domain assumption The 116 responses are representative enough for Fisher's exact test to support generalizations
Cite this review
Pith. "Pith review of Characterizing Data Scientists in the Real World." pith.science (2026). https://pith.science/paper/RKNM4RFZ
@misc{pith2026241112225,
author = {Pith},
title = {Pith review of: Characterizing Data Scientists in the Real World},
year = {2026},
howpublished = {\url{https://pith.science/paper/RKNM4RFZ}},
note = {Machine review of arXiv:2411.12225}
}
read the original abstract
Data collection is pervasively bound to our digital lifestyle. A recent study by the IDC reports that the growth of the data created and replicated in 2020 was even higher than in the previous years due to pandemic-related confinements to an astonishing global amount of 64.2 zettabytes of data. While not all the produced data is meant to be analyzed, there are numerous companies whose services/products rely heavily on data analysis. That is to say that mining the produced data has already revealed great value for businesses in different sectors. But to be able to fully realize this value, companies need to be able to hire professionals that are capable of gleaning insights and extracting value from the available data. We hypothesize that people nowadays conducting data-science-related tasks in practice may not have adequate training or formation. So in order to be able to fully support them in being productive in their duties, e.g. by building appropriate tools that increase their productivity, we first need to characterize the current generation of data scientists. To contribute towards this characterization, we conducted a public survey to fully understand who is doing data science, how they work, what are the skills they hold and lack, and which tools they use and need.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
Shane Allua and Cheryl Bagley Thompson. Inferential Statistics. Air Medical Journal, 28(4):168–171, 7 2009
work page 2009
-
[2]
Leonard Bickman and Debra J. Rog. The SAGE Handbook of Applied Social Research Methods. SAGE Publications, Inc, 2 edition, 2008
work page 2008
-
[3]
Data science: A comprehensive overview
Longbing Cao. Data science: A comprehensive overview. ACM Comput. Surv., 50(3), jun 2017
work page 2017
-
[4]
Ilyas, Sanjay Krishnan, and Jiannan Wang
Xu Chu, Ihab F. Ilyas, Sanjay Krishnan, and Jiannan Wang. Data cleaning: Overview and emerging challenges. In Proceedings of the ACM SIGMOD 24 International Conference on Management of Data , pages 2201–2206. Asso- ciation for Computing Machinery, 6 2016
work page 2016
-
[5]
Thomas H. Davenport and D. J. Patil. Data scientist: The sexiest job of the 21st century. Harvard Business Review , 90(10):5, 2012
work page 2012
- [6]
-
[7]
Murray J. Fisher and Andrea P. Marshall. Understanding descriptive statis- tics. Australian Critical Care, 22(2):93–97, 5 2009
work page 2009
-
[8]
Fritz H. Grupe and M. Mehdi Owrang. Data base mining: Discovering new knowledge and competitive advantage. Information Systems Management , 12(4):26–31, 1995
work page 1995
Show all 30 references
-
[9]
Analyzing the Analyz- ers: An Introspective Survey of Data Scientists and Their Work
Harlan Harris, Sean Murphy, and Marck Vaisman. Analyzing the Analyz- ers: An Introspective Survey of Data Scientists and Their Work . O’Reilly Media, Inc., 1 edition, 2013
2013
-
[10]
Data created worldwide 2010-2025 — Statista, 2019
Arne Holst. Data created worldwide 2010-2025 — Statista, 2019
2010
-
[11]
Research directions in data wrangling: Visual- izations and transformations for usable and credible data
Sean Kandel, Jeffrey Heer, Catherine Plaisant, Jessie Kennedy, Frank van Ham, Nathalie Henry Riche, Chris Weaver, Bongshin Lee, Dominique Brod- beck, and Paolo Buono. Research directions in data wrangling: Visual- izations and transformations for usable and credible data. Info...
2011
-
[12]
Statistical notes for clinical researchers: Chi-squared test and Fisher’s exact test
Hae-Young Kim. Statistical notes for clinical researchers: Chi-squared test and Fisher’s exact test. Restorative Dentistry & Endodontics , 42(2):152, 3 2017
2017
-
[13]
The emerging role of data scientists on software development teams
Miryung Kim, Thomas Zimmermann, Robert DeLine, and Andrew Begel. The emerging role of data scientists on software development teams. Pro- ceedings - International Conference on Software Engineering , 14-22-May- :96–107, 2016
2016
-
[14]
Principles of survey research part 6
Barbara Kitchenham and Shari Lawrence Pfleeger. Principles of survey research part 6. ACM SIGSOFT Software Engineering Notes , 28(2):24, 2003
2003
-
[15]
Use of Big Data for Competitive Advantage of Company
Milan Kubina, Michal Varmus, and Irena Kubinova. Use of Big Data for Competitive Advantage of Company. Procedia Economics and Finance , 26:561–565, 2015
2015
-
[16]
D1.4 Study Evaluation Report 2
Leonard Mack and David Tarrant. D1.4 Study Evaluation Report 2. Tech- nical report, European Commission, 2015
2015
-
[17]
Professoren des Inst
Heiko M¨ uller and Johann Christoph Freytag.Problems, methods, and chal- lenges in comprehensive data cleansing . Professoren des Inst. F¨ ur Infor- matik, 2005. 25
2005
-
[18]
Vera Liao, Casey Dugan, and Thomas Erickson
Michael Muller, Ingrid Lange, Dakuo Wang, David Piorkowski, Jason Tsay, Q. Vera Liao, Casey Dugan, and Thomas Erickson. How data science workers work with data: Discovery, capture, curation, design, creation. In Proceedings of the 2019 CHI Conference on Human Factors in Comput...
2019
-
[19]
Regulating the internet giants: The world’s most valuable resource is no longer oil, but data
David Parkins. Regulating the internet giants: The world’s most valuable resource is no longer oil, but data. The Economist (United Kingdom), 2017
2017
-
[20]
Tekla S. Perry. Demand and Salaries for Data Scientists Continue to Climb, 2019
2019
-
[21]
Data Science and its Relationship to Big Data and Data-Driven Decision Making
Foster Provost and Tom Fawcett. Data Science and its Relationship to Big Data and Data-Driven Decision Making. Big Data , 1(1):51–59, 3 2013
2013
-
[22]
Data Cleaning: Problems and Current Approaches Erhard
Erhard Rahm and Hong Hai Do. Data Cleaning: Problems and Current Approaches Erhard. IEEE Transactions on Cloud Computing , 2(1):1–1, 2014
2014
-
[23]
2015 Data Science Survey
Karl Rexer, Paul Gearan, and Heather Allen. 2015 Data Science Survey. Technical report, Rexer Analytics, 6 2015
2015
-
[24]
Analytics: The real-world use of big data
Michael Schroeck, Rebecca Shockley, Janet Smart, Dolores Romero- Morales, and Peter Tufano. Analytics: The real-world use of big data. Technical report, IBM Institute for Business Value, 2012
2012
-
[25]
Data cleaning: Detecting, diagnosing, and editing data abnormalities, 2005
Jan Van Den Broeck, Solveig Argeseanu Cunningham, Roger Eeckels, and Kobus Herbst. Data cleaning: Detecting, diagnosing, and editing data abnormalities, 2005
2005
-
[26]
How Big Data is Creating Competitive Advantage — K2 Partnering Solutions, 2018
Ayshea Williams. How Big Data is Creating Competitive Advantage — K2 Partnering Solutions, 2018
2018
-
[27]
Ohlsson, Bj¨ orn Regnell, and Anders Wessl´ en
Claes Wohlin, Per Runeson, Martin H¨ ost, Magnus C. Ohlsson, Bj¨ orn Regnell, and Anders Wessl´ en. Experimentation in Software Engineering . Springer Berlin Heidelberg, Berlin, Heidelberg, 1 edition, 2012
2012
-
[28]
Goals, process, and challenges of exploratory data analysis: An interview study
Kanit Wongsuphasawat, Yang Liu, and Jeffrey Heer. Goals, process, and challenges of exploratory data analysis: An interview study. CoRR, abs/1911.00568, 2019
1911 arXiv
-
[29]
The Data Science Revolution That’s Transforming Avia- tion, 2017
Oliver Wyman. The Data Science Revolution That’s Transforming Avia- tion, 2017
2017
-
[30]
Zhang, Michael J
Amy X. Zhang, Michael J. Muller, and Dakuo Wang. How do data science workers collaborate? roles, workflows, and tools. CoRR, abs/2001.06684, 2020. 26
2001 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.