Pith. sign in

REVIEW 4 major objections 6 minor 30 references

Characterizing Data Scientists in the Real World

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A 116-respondent survey finds data scientists are highly qualified and that their academic background has little impact on how they work.

desk verdict Useful descriptive snapshot of 116 data science workers, but the central claim that academic background has little impact is not supported by the reported analysis. read the letter →

arxiv 2411.12225 v1 pith:RKNM4RFZ submitted 2024-11-19 cs.HC

classification cs.HC
keywords datascienceprofessionalsempiricalsurveyworkflowscomputerbackgroundFisher'sexacttesttechnologyusegendergap
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper starts from the hypothesis that people doing data science in practice may lack adequate training, then tests it with a public survey of 116 practitioners. The central finding is the opposite of that hypothesis: the workforce is highly qualified, most respondents have a computer science background, and academic background has little impact on how they actually work. The same difficulties dominate across backgrounds—access to quality data and applying deep learning techniques—while Python, SQL, and spreadsheet editors form the common technological core. The authors argue this characterization should guide tools and methodologies toward the workflow challenges everyone shares.

What carries the argument

The analytical engine is the 'CS background?' attribute, a derived binary variable that records whether a respondent had any formal training in computer science. The authors cross this variable with every work-related response—self-rated task confidence, reported difficulties, time spent coding, and technology choices—using contingency tables and Fisher's exact tests to decide which observed differences are statistically meaningful. The same machinery is applied to years of experience and gender, making it the device that carries the paper's main argument that background has little impact.

What would settle it

Run the same survey on a stratified sample drawn from company employment records, professional certifications, or national labor statistics instead of open online forums, and compare the proportions for CS background, task confidence, and reported difficulties. If the background effects or difficulty rankings differ materially from the paper's numbers, the claim that academic background has little impact on work would fail as a general statement.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that the stereotype of under-trained data science workers does not hold: the respondents are mostly men aged 26 to 45, overwhelmingly degree-holders, and about 62 percent have formal computer science training. Their academic past, however, does not decide how they work: confidence in most tasks, the difficulties they report, and their analytical goals are broadly shared, with only a few statistically significant differences tied to a CS background, such as deep-learning confidence and some language choices. The professionals' main shared problems are poor-quality data, difficult access to relevant data, and unclear questions to answer, and they report spending up to half their working time coding. The paper also reports a gender gap: many fewer women than men responded, and women report lower job satisfaction.

Load-bearing premise

The load-bearing assumption is that the 116 people who filled out the survey, recruited through online forums and personal contacts, represent data science professionals worldwide; if the sample is skewed, the profile and the conclusion that academic background has little impact describe only this respondent pool.

Editorial extensions

If this is right

  • Tool and training efforts should target the shared bottlenecks—access to quality data and applying deep learning—rather than remedial computer science education.
  • Because Python, SQL, and spreadsheet editors dominate across both groups, improving interoperability around these tools would serve most practitioners.
  • Self-reported confidence in deep learning is the clearest skill gap, and it is larger among professionals without a CS background, so focused support there would reach a real need.
  • The gender imbalance and lower satisfaction reported by women point to an equity issue that additional technical training alone would not address.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, clustering respondents on the tasks and technologies they report would test whether academic background truly leaves no signature; if it does not, the clusters should not separate by degree field.
  • Beyond the paper, the recruitment through programming-centric forums and personal contacts may over-represent code-comfortable practitioners, so the high qualification level could be an upper-bound estimate; a registry-based or employer-based sample would check this.
  • Beyond the paper, the deep-learning gap is read as a skills gap, but it could equally be an exposure gap; comparing confidence among practitioners who have completed deep-learning projects, regardless of degree, would separate the two.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents results of a survey of 116 self-selected data science practitioners, conducted from April to December 2020 via online forums, social media, and personal contacts. The survey collected academic background, professional experience, self-assessed task proficiency, work difficulties, and technology use. The authors use descriptive statistics and Fisher's exact tests to answer three research questions about the profile of data scientists, the impact of profile on work, and technology choices. The main conclusion is that data scientists are highly qualified and that academic background (CS vs. non-CS) has little impact on how they work; they also report a gender gap and lower satisfaction among women.

Significance. If the descriptive patterns are representative, the paper provides a useful snapshot for tool designers and educators: data quality and data access are the most common difficulties, deep learning is the least self-confident task, and Python/SQL dominate. The authors are transparent about their data preparation and they emphasize that the survey was public. The paper also gains credibility from using a reproducible R script for Fisher's tests (Section 3.3) and from reporting frequencies and modes transparently. However, the paper's central null claim about academic background is not statistically supported, and the sampling strategy limits population-level generalizations. The manuscript would be a worthwhile descriptive report after major revisions to align conclusions with the evidence.

major comments (4)
  1. [Section 7 and Sections 4.3/4.4] The conclusion that 'academic past has little impact on the way they work' is an inference from null results. In Section 4.3, the authors state that for satisfaction crossed with experience and background 'we could not find significant differences' but report no p-values, effect sizes, or confidence intervals, and no equivalence tests. With n=116 and subgroups as small as 29 women and 44 non-CS respondents, non-significant Fisher tests do not demonstrate that population differences are small or absent. The authors should either run equivalence tests with pre-specified bounds, report effect sizes for all comparisons, or restrict the conclusion to the sample.
  2. [Section 4.4 and Section 3.3] The only reported inferential statistic is a Fisher's exact test for deep learning self-efficacy with p=0.049. This is borderline, and no correction for multiple comparisons is described even though many contingency tables were tested across Sections 4.3-4.6. The conclusion that 'those with a CS background feel more apt to apply deep learning techniques' should be supported by adjusted p-values, effect sizes (e.g., odds ratio), and confidence intervals.
  3. [Section 3.1 and Section 6] The representativeness of the sample is not established. The survey was distributed through convenience channels, with no sampling frame, response rate, or comparison to known population characteristics of data scientists. The responses are heavily concentrated in Portugal (44/116) and male (87/116). The Threats to validity section acknowledges sampling errors but only asserts that multi-platform distribution helped; it does not quantify non-response or self-selection. Consequently, RQ2 and RQ3's population-level claims are not supported. The authors should reframe findings as descriptive of the sample or provide external benchmarks.
  4. [Sections 4.5-4.6 and Section 7] The descriptive results contain numerous CS/non-CS differences that contradict the 'little impact' summary: people without a CS background spend more time coding (Figure 9), the two groups differ in IDE, language, and statistics tool choices (Figures 11-15), and self-reported lack of data science skills is distributed differently (Figure 7). The conclusion that academic background has little impact is not reconciled with these descriptive differences. The authors need to either qualify the conclusion (e.g., 'little impact on job satisfaction' or 'limited to specific tasks/tools') or conduct a multivariate analysis that accounts for background along with other variables.
minor comments (6)
  1. [Section 4.4] The text uses 'Fischer's test'; this should be 'Fisher's exact test'.
  2. [Table 1] The column headers 'F emale' and 'T otal' contain typos, and the table layout would benefit from clearer alignment.
  3. [Section 4.2] The phrase 'background inNatural Sciences' is missing a space, and the italicized conclusion 'data science professionals reported being highly qualified' is presented without a supporting statistical test or benchmark.
  4. [Section 6] The sentence 'Participants may fill influenced to answer questions' is ungrammatical; it should likely read 'feel influenced'.
  5. [Section 3.3] The R script used for Fisher's tests is mentioned but not included; providing the script and the anonymized dataset would strengthen reproducibility.
  6. [Abstract] The opening statistic (64.2 zettabytes) is not cited in the abstract; a reference to the IDC report should be added.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey's conclusions are descriptive and inferential summaries of collected data, not quantities fitted to or derived from themselves.

full rationale

This paper is an empirical survey study; it contains no derivation, no fitted parameters, no formulas, and no reliance on prior results by the same authors. The conclusions in Sections 5 and 7 (e.g., 'people in data science are generally highly qualified... their academic past has little impact on the way they work') are summaries of the 116 survey responses and of Fisher's exact tests applied to contingency tables. Those tests compare observed response distributions across groups; they are not constructed so that the tested quantity equals the group definition. The only identified fragility is statistical and sampling validity: the convenience sample is not shown to be representative, and non-significant Fisher tests are treated as evidence of absence in RQ2. These are threats to external validity and to the support for a null conclusion, not circularity under the definitions used here. There are no self-citations, no imported uniqueness theorems, and no ansatz smuggled in via citation.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper has no mathematical derivation and no free parameters. Its central claims rest on three domain assumptions about respondent identity, self-report accuracy, and sample representativeness, none of which are independently verified.

assumptions (3)
  • domain assumption Respondents self-identifying as data science workers are indeed data science professionals
    The survey was distributed through data-science-related forums and personal contacts, but there was no verification of the respondents' job roles. This underlies all profile claims in Section 4.
  • domain assumption Self-reported skills, satisfaction, and difficulties accurately reflect actual behavior
    The paper relies on self-evaluation in Section 4.4 and job satisfaction in Section 4.3 without external validation, which is common in surveys but an unverified assumption.
  • domain assumption The 116 responses are representative enough for Fisher's exact test to support generalizations
    Fisher's exact test assumes random sampling from a defined population; the convenience sample from Section 3.1 does not meet this, yet the test is used in Section 3.3 to generalize findings.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Characterizing Data Scientists in the Real World." pith.science (2026). https://pith.science/paper/RKNM4RFZ

@misc{pith2026241112225,
  author       = {Pith},
  title        = {Pith review of: Characterizing Data Scientists in the Real World},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RKNM4RFZ}},
  note         = {Machine review of arXiv:2411.12225}
}
read the original abstract

Data collection is pervasively bound to our digital lifestyle. A recent study by the IDC reports that the growth of the data created and replicated in 2020 was even higher than in the previous years due to pandemic-related confinements to an astonishing global amount of 64.2 zettabytes of data. While not all the produced data is meant to be analyzed, there are numerous companies whose services/products rely heavily on data analysis. That is to say that mining the produced data has already revealed great value for businesses in different sectors. But to be able to fully realize this value, companies need to be able to hire professionals that are capable of gleaning insights and extracting value from the available data. We hypothesize that people nowadays conducting data-science-related tasks in practice may not have adequate training or formation. So in order to be able to fully support them in being productive in their duties, e.g. by building appropriate tools that increase their productivity, we first need to characterize the current generation of data scientists. To contribute towards this characterization, we conducted a public survey to fully understand who is doing data science, how they work, what are the skills they hold and lack, and which tools they use and need.

Figures

Figures reproduced from arXiv: 2411.12225 by the authors.

Figure 1
Figure 1. , the platform Coursera3 was mentioned 49 times by the participants. In fact, almost every answer to this question mentioned some type of online resource, from online courses to blogs about data science and online lectures from ivy league universities. Also, as demonstrated in [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Satisfaction by background [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Satisfaction by gender. 4.4 Self-evaluation on data science tasks To understand how participants assess their competence in tasks related to data manipulation, data processing, and data analysis, they rated their experience in a set of tasks as: Very Poor - Little or no knowledge/expertise; Poor - Exper￾imental/vague knowledge; Ok - Familiar and competent user; Good - Regular and confident user; Very Good - Leading … view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: Applying deep learning techniques by Background. [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: Access to relevant data by background. 13 [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: Lack of clear questions to answer by background. [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]
Figure 7
Figure 7. Figure 7: Lack of data science skills by background. [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: Time spent actively coding [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]
Figure 9
Figure 9. Figure 9: Time spent coding by background [PITH_FULL_IMAGE:figures/full_fig_p017_9.png]
Figure 10
Figure 10. Figure 10: Participants analytical goals. 17 [PITH_FULL_IMAGE:figures/full_fig_p017_10.png]
Figure 11
Figure 11. Figure 11: IDE or Editor by background. Programming, Scripting or Markup Language [PITH_FULL_IMAGE:figures/full_fig_p018_11.png]
Figure 12
Figure 12. Figure 12: Programming, Scripting or Markup Language by background. [PITH_FULL_IMAGE:figures/full_fig_p019_12.png]
Figure 13
Figure 13. Figure 13: Machine Learning Frameworks/Libraries/Tools by background. [PITH_FULL_IMAGE:figures/full_fig_p020_13.png]
Figure 14
Figure 14. Figure 14: Statistics packages/tools by background. [PITH_FULL_IMAGE:figures/full_fig_p021_14.png]
Figure 15
Figure 15. Figure 15: Data visualization libraries/tools by background. [PITH_FULL_IMAGE:figures/full_fig_p021_15.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

30 extracted references · 29 canonical work pages

  1. [1]

    Inferential Statistics

    Shane Allua and Cheryl Bagley Thompson. Inferential Statistics. Air Medical Journal, 28(4):168–171, 7 2009

  2. [2]

    Leonard Bickman and Debra J. Rog. The SAGE Handbook of Applied Social Research Methods. SAGE Publications, Inc, 2 edition, 2008

  3. [3]

    Data science: A comprehensive overview

    Longbing Cao. Data science: A comprehensive overview. ACM Comput. Surv., 50(3), jun 2017

  4. [4]

    Ilyas, Sanjay Krishnan, and Jiannan Wang

    Xu Chu, Ihab F. Ilyas, Sanjay Krishnan, and Jiannan Wang. Data cleaning: Overview and emerging challenges. In Proceedings of the ACM SIGMOD 24 International Conference on Management of Data , pages 2201–2206. Asso- ciation for Computing Machinery, 6 2016

  5. [5]

    Davenport and D

    Thomas H. Davenport and D. J. Patil. Data scientist: The sexiest job of the 21st century. Harvard Business Review , 90(10):5, 2012

  6. [6]

    Big Data

    Vasant Dhar, Matthias Jarke, and J¨ urgen Laartz. Big Data. Business & Information Systems Engineering , 6(5):257–259, 10 2014

  7. [7]

    Fisher and Andrea P

    Murray J. Fisher and Andrea P. Marshall. Understanding descriptive statis- tics. Australian Critical Care, 22(2):93–97, 5 2009

  8. [8]

    Grupe and M

    Fritz H. Grupe and M. Mehdi Owrang. Data base mining: Discovering new knowledge and competitive advantage. Information Systems Management , 12(4):26–31, 1995

Show all 30 references
  1. [9]

    Analyzing the Analyz- ers: An Introspective Survey of Data Scientists and Their Work

    Harlan Harris, Sean Murphy, and Marck Vaisman. Analyzing the Analyz- ers: An Introspective Survey of Data Scientists and Their Work . O’Reilly Media, Inc., 1 edition, 2013

  2. [10]

    Data created worldwide 2010-2025 — Statista, 2019

    Arne Holst. Data created worldwide 2010-2025 — Statista, 2019

  3. [11]

    Research directions in data wrangling: Visual- izations and transformations for usable and credible data

    Sean Kandel, Jeffrey Heer, Catherine Plaisant, Jessie Kennedy, Frank van Ham, Nathalie Henry Riche, Chris Weaver, Bongshin Lee, Dominique Brod- beck, and Paolo Buono. Research directions in data wrangling: Visual- izations and transformations for usable and credible data. Info...

  4. [12]

    Statistical notes for clinical researchers: Chi-squared test and Fisher’s exact test

    Hae-Young Kim. Statistical notes for clinical researchers: Chi-squared test and Fisher’s exact test. Restorative Dentistry & Endodontics , 42(2):152, 3 2017

  5. [13]

    The emerging role of data scientists on software development teams

    Miryung Kim, Thomas Zimmermann, Robert DeLine, and Andrew Begel. The emerging role of data scientists on software development teams. Pro- ceedings - International Conference on Software Engineering , 14-22-May- :96–107, 2016

  6. [14]

    Principles of survey research part 6

    Barbara Kitchenham and Shari Lawrence Pfleeger. Principles of survey research part 6. ACM SIGSOFT Software Engineering Notes , 28(2):24, 2003

  7. [15]

    Use of Big Data for Competitive Advantage of Company

    Milan Kubina, Michal Varmus, and Irena Kubinova. Use of Big Data for Competitive Advantage of Company. Procedia Economics and Finance , 26:561–565, 2015

  8. [16]

    D1.4 Study Evaluation Report 2

    Leonard Mack and David Tarrant. D1.4 Study Evaluation Report 2. Tech- nical report, European Commission, 2015

  9. [17]

    Professoren des Inst

    Heiko M¨ uller and Johann Christoph Freytag.Problems, methods, and chal- lenges in comprehensive data cleansing . Professoren des Inst. F¨ ur Infor- matik, 2005. 25

  10. [18]

    Vera Liao, Casey Dugan, and Thomas Erickson

    Michael Muller, Ingrid Lange, Dakuo Wang, David Piorkowski, Jason Tsay, Q. Vera Liao, Casey Dugan, and Thomas Erickson. How data science workers work with data: Discovery, capture, curation, design, creation. In Proceedings of the 2019 CHI Conference on Human Factors in Comput...

  11. [19]

    Regulating the internet giants: The world’s most valuable resource is no longer oil, but data

    David Parkins. Regulating the internet giants: The world’s most valuable resource is no longer oil, but data. The Economist (United Kingdom), 2017

  12. [20]

    Tekla S. Perry. Demand and Salaries for Data Scientists Continue to Climb, 2019

  13. [21]

    Data Science and its Relationship to Big Data and Data-Driven Decision Making

    Foster Provost and Tom Fawcett. Data Science and its Relationship to Big Data and Data-Driven Decision Making. Big Data , 1(1):51–59, 3 2013

  14. [22]

    Data Cleaning: Problems and Current Approaches Erhard

    Erhard Rahm and Hong Hai Do. Data Cleaning: Problems and Current Approaches Erhard. IEEE Transactions on Cloud Computing , 2(1):1–1, 2014

  15. [23]

    2015 Data Science Survey

    Karl Rexer, Paul Gearan, and Heather Allen. 2015 Data Science Survey. Technical report, Rexer Analytics, 6 2015

  16. [24]

    Analytics: The real-world use of big data

    Michael Schroeck, Rebecca Shockley, Janet Smart, Dolores Romero- Morales, and Peter Tufano. Analytics: The real-world use of big data. Technical report, IBM Institute for Business Value, 2012

  17. [25]

    Data cleaning: Detecting, diagnosing, and editing data abnormalities, 2005

    Jan Van Den Broeck, Solveig Argeseanu Cunningham, Roger Eeckels, and Kobus Herbst. Data cleaning: Detecting, diagnosing, and editing data abnormalities, 2005

  18. [26]

    How Big Data is Creating Competitive Advantage — K2 Partnering Solutions, 2018

    Ayshea Williams. How Big Data is Creating Competitive Advantage — K2 Partnering Solutions, 2018

  19. [27]

    Ohlsson, Bj¨ orn Regnell, and Anders Wessl´ en

    Claes Wohlin, Per Runeson, Martin H¨ ost, Magnus C. Ohlsson, Bj¨ orn Regnell, and Anders Wessl´ en. Experimentation in Software Engineering . Springer Berlin Heidelberg, Berlin, Heidelberg, 1 edition, 2012

  20. [28]

    Goals, process, and challenges of exploratory data analysis: An interview study

    Kanit Wongsuphasawat, Yang Liu, and Jeffrey Heer. Goals, process, and challenges of exploratory data analysis: An interview study. CoRR, abs/1911.00568, 2019

  21. [29]

    The Data Science Revolution That’s Transforming Avia- tion, 2017

    Oliver Wyman. The Data Science Revolution That’s Transforming Avia- tion, 2017

  22. [30]

    Zhang, Michael J

    Amy X. Zhang, Michael J. Muller, and Dakuo Wang. How do data science workers collaborate? roles, workflows, and tools. CoRR, abs/2001.06684, 2020. 26

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.