{"id":"3b4949af-5788-4e8b-b171-0e60b5632d01","arxiv_id":"2411.12225","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of 116 data science workers characterizes their education, skills, difficulties, and tool use, showing a highly qualified but gender-skewed group with heavy reliance on Python and online learning.","lead":"The authors surveyed 116 self-selected data science workers about their backgrounds, skills, tools, and job satisfaction. The results describe a mostly male, highly educated group that relies heavily on Python and online courses, but the small convenience sample limits generalization.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'academic past has little impact' claim is not established: the paper treats non-significant Fisher tests as evidence of absence, reports only one uncorrected p-value (0.049), and lacks equivalence tests or effect sizes.","rationale":"The reader's weakest-assumption concern about representativeness is valid: the convenience sample with no sampling frame or response rate limits any population generalization. However, the single most load-bearing problem for the central claim is internal: even if the sample were representative, the conclusion that academic background has little impact is not supported by the inferential analysis as reported. Non-significant Fisher tests are treated as evidence of absence, only one borderline p-value is reported, and no effect sizes or equivalence bounds are provided. This is a correctness risk, not just a generalizability risk. The descriptive differences that do appear in the paper make the 'little impact' wording especially hard to defend. Because the paper could be revised to condition its conclusions on the sample and to add proper equivalence/effect-size analysis, the reader's CONDITIONAL verdict remains appropriate; no verdict change is needed.","tokens_in":11587,"tokens_out":6199,"duration_ms":66005,"concrete_test":"Ask the authors to release the anonymized survey data and R script as a condition of acceptance. Re-run Fisher's exact tests on every CS-background by outcome contingency table, apply a multiple-comparison correction, and compute standardized effect sizes (e.g., odds ratios or Cramér's V) with 95% confidence intervals. Then run a two one-sided (TOST) equivalence test with a pre-committed smallest effect size of interest on the key outcomes. For example, the reported deep-learning self-efficacy percentages (77.78% vs 56.82%, Figure 4) correspond to an odds ratio of roughly 2.7; treating that as 'little impact' requires justification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Section 7) has two parts: data scientists are highly qualified, and academic background has little impact on how they work. The first part is a descriptive claim about a convenience sample, so generalizability is already uncertain. The more load-bearing problem is the second part, which is a null claim about a population effect. The analysis in Sections 3.3 and 4 applies Fisher's exact tests across many CS-background by outcome tables, yet the only p-value actually reported is 0.049 for deep-learning self-efficacy (Section 4.4). In Section 4.3, for satisfaction by years of experience and background, the text says 'we could not find significant differences' without reporting p-values, effect sizes, or confidence intervals. This non-significance arises in samples of 116 with subgroups as small as 29 women and 44 non-CS respondents; it does not imply that population differences are small. No equivalence test, pre-specified effect size, or multiple-comparison correction is mentioned. The conclusion 'academic past has little impact on the way they work' is therefore not supported by the analysis actually reported, even setting aside the sampling frame. The paper's own descriptive results also show several CS/non-CS differences (time spent coding, language and tool choices, self-reported skill gaps), which at minimum require careful reconciliation with the claim of little impact.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents results of a survey of 116 self-selected data science practitioners, conducted from April to December 2020 via online forums, social media, and personal contacts. The survey collected academic background, professional experience, self-assessed task proficiency, work difficulties, and technology use. The authors use descriptive statistics and Fisher's exact tests to answer three research questions about the profile of data scientists, the impact of profile on work, and technology choices. The main conclusion is that data scientists are highly qualified and that academic background (CS vs. non-CS) has little impact on how they work; they also report a gender gap and lower satisfaction among women.","tokens_in":11812,"tokens_out":4324,"duration_ms":40834,"significance":"If the descriptive patterns are representative, the paper provides a useful snapshot for tool designers and educators: data quality and data access are the most common difficulties, deep learning is the least self-confident task, and Python/SQL dominate. The authors are transparent about their data preparation and they emphasize that the survey was public. The paper also gains credibility from using a reproducible R script for Fisher's tests (Section 3.3) and from reporting frequencies and modes transparently. However, the paper's central null claim about academic background is not statistically supported, and the sampling strategy limits population-level generalizations. The manuscript would be a worthwhile descriptive report after major revisions to align conclusions with the evidence.","major_comments":[{"comment":"The conclusion that 'academic past has little impact on the way they work' is an inference from null results. In Section 4.3, the authors state that for satisfaction crossed with experience and background 'we could not find significant differences' but report no p-values, effect sizes, or confidence intervals, and no equivalence tests. With n=116 and subgroups as small as 29 women and 44 non-CS respondents, non-significant Fisher tests do not demonstrate that population differences are small or absent. The authors should either run equivalence tests with pre-specified bounds, report effect sizes for all comparisons, or restrict the conclusion to the sample.","section":"Section 7 and Sections 4.3/4.4"},{"comment":"The only reported inferential statistic is a Fisher's exact test for deep learning self-efficacy with p=0.049. This is borderline, and no correction for multiple comparisons is described even though many contingency tables were tested across Sections 4.3-4.6. The conclusion that 'those with a CS background feel more apt to apply deep learning techniques' should be supported by adjusted p-values, effect sizes (e.g., odds ratio), and confidence intervals.","section":"Section 4.4 and Section 3.3"},{"comment":"The representativeness of the sample is not established. The survey was distributed through convenience channels, with no sampling frame, response rate, or comparison to known population characteristics of data scientists. The responses are heavily concentrated in Portugal (44/116) and male (87/116). The Threats to validity section acknowledges sampling errors but only asserts that multi-platform distribution helped; it does not quantify non-response or self-selection. Consequently, RQ2 and RQ3's population-level claims are not supported. The authors should reframe findings as descriptive of the sample or provide external benchmarks.","section":"Section 3.1 and Section 6"},{"comment":"The descriptive results contain numerous CS/non-CS differences that contradict the 'little impact' summary: people without a CS background spend more time coding (Figure 9), the two groups differ in IDE, language, and statistics tool choices (Figures 11-15), and self-reported lack of data science skills is distributed differently (Figure 7). The conclusion that academic background has little impact is not reconciled with these descriptive differences. The authors need to either qualify the conclusion (e.g., 'little impact on job satisfaction' or 'limited to specific tasks/tools') or conduct a multivariate analysis that accounts for background along with other variables.","section":"Sections 4.5-4.6 and Section 7"}],"minor_comments":[{"comment":"The text uses 'Fischer's test'; this should be 'Fisher's exact test'.","section":"Section 4.4"},{"comment":"The column headers 'F emale' and 'T otal' contain typos, and the table layout would benefit from clearer alignment.","section":"Table 1"},{"comment":"The phrase 'background inNatural Sciences' is missing a space, and the italicized conclusion 'data science professionals reported being highly qualified' is presented without a supporting statistical test or benchmark.","section":"Section 4.2"},{"comment":"The sentence 'Participants may fill influenced to answer questions' is ungrammatical; it should likely read 'feel influenced'.","section":"Section 6"},{"comment":"The R script used for Fisher's tests is mentioned but not included; providing the script and the anonymized dataset would strengthen reproducibility.","section":"Section 3.3"},{"comment":"The opening statistic (64.2 zettabytes) is not cited in the abstract; a reference to the IDC report should be added.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper's descriptive material is a potentially useful contribution to the HCI/CSCW literature on data science work. The main concern is that the authors over-interpret null results as evidence of no effect. If the revision reframes the conclusions as descriptive of the sample and adds appropriate statistical reporting, the paper could be acceptable. I also note the paper does not cite several recent large-scale surveys of data scientists (e.g., Kaggle's State of Data Science), which would help contextualize the sample and the generalizability of the findings."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Pereira et al. report a 2020 survey of 116 data science workers, focusing on how CS background relates to tools, self-rated skills, and difficulties. The descriptive core is honest and usable: the cross-tabulations by CS background on language choice, IDE use, and task confidence are a genuinely specific angle that earlier surveys (Rexer, Harris, Zhang) largely didn't cut this way. I also give them credit for a fair related-work review and for describing their data-cleaning decisions. The finding that online learning resources dominated books is a small but concrete data point.\n\nThe problem is the load-bearing conclusion. Section 7 states that academic background 'has little impact on the way they work.' The analysis does not support this. Fisher's exact tests are used across many comparisons, but only one p-value is actually reported (0.049 for deep-learning self-efficacy, Section 4.4), and it is not corrected for multiple comparisons. Section 4.3 says 'we could not find significant differences' for satisfaction by experience and background, with no p-values, effect sizes, or confidence intervals. In subgroups of 29 women and 44 non-CS respondents, non-significance means the test is underpowered; it does not mean the population difference is small. No equivalence tests, no pre-specified effect sizes, nothing. On top of that, the sample is a convenience sample—forums and personal contacts—with no response rate or representativeness check, so even the descriptive generalizing statements in Sections 4 and 5 are shaky.\n\nThe paper's own descriptive results also undercut the 'little impact' narrative: people without a CS background spend more time coding, use R more, and report different self-efficacy on deep learning. That is impact on how they work, even if it doesn't reach statistical significance. So the conclusion should be softened to 'we did not find strong evidence of impact' or, better, re-analyzed with effect sizes and a pre-registered plan.\n\nMinor stuff: the gender satisfaction difference is stated without a test, and the data and R script are not shared, which makes replication impossible. The paper doesn't claim to be a formal derivation, so circularity isn't an issue.\n\nWho is this for? Tool builders and educators who want a quick descriptive read on self-reported tool use and skill gaps will find the tables useful. It deserves a serious referee because it has original data and a clear question, but the referee should treat the central claim as unsupported and require a major revision—either restrict conclusions to the sample, or provide proper inferential support.","headline":"Useful descriptive snapshot of 116 data science workers, but the central claim that academic background has little impact is not supported by the reported analysis.","tokens_in":12360,"tokens_out":2516,"would_cite":false,"duration_ms":26319,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 116-respondent survey finds data scientists are highly qualified and that their academic background has little impact on how they work.","keywords":["data science professionals","empirical survey","data science workflows","computer science background","Fisher's exact test","technology use","gender gap"],"falsifier":"Run the same survey on a stratified sample drawn from company employment records, professional certifications, or national labor statistics instead of open online forums, and compare the proportions for CS background, task confidence, and reported difficulties. If the background effects or difficulty rankings differ materially from the paper's numbers, the claim that academic background has little impact on work would fail as a general statement.","tokens_in":11369,"feed_emoji":"📊","tokens_out":9140,"duration_ms":89928,"temperature":0.7,"pith_summary":"The paper starts from the hypothesis that people doing data science in practice may lack adequate training, then tests it with a public survey of 116 practitioners. The central finding is the opposite of that hypothesis: the workforce is highly qualified, most respondents have a computer science background, and academic background has little impact on how they actually work. The same difficulties dominate across backgrounds—access to quality data and applying deep learning techniques—while Python, SQL, and spreadsheet editors form the common technological core. The authors argue this characterization should guide tools and methodologies toward the workflow challenges everyone shares.","feed_headline":"Data science work: same with or without a CS degree","feed_subtitle":"The real bottlenecks are data access and deep learning, so tools should target shared pain points, not education gaps.","key_machinery":"The analytical engine is the 'CS background?' attribute, a derived binary variable that records whether a respondent had any formal training in computer science. The authors cross this variable with every work-related response—self-rated task confidence, reported difficulties, time spent coding, and technology choices—using contingency tables and Fisher's exact tests to decide which observed differences are statistically meaningful. The same machinery is applied to years of experience and gender, making it the device that carries the paper's main argument that background has little impact.","core_discovery":"On the paper's own terms, the discovery is that the stereotype of under-trained data science workers does not hold: the respondents are mostly men aged 26 to 45, overwhelmingly degree-holders, and about 62 percent have formal computer science training. Their academic past, however, does not decide how they work: confidence in most tasks, the difficulties they report, and their analytical goals are broadly shared, with only a few statistically significant differences tied to a CS background, such as deep-learning confidence and some language choices. The professionals' main shared problems are poor-quality data, difficult access to relevant data, and unclear questions to answer, and they report spending up to half their working time coding. The paper also reports a gender gap: many fewer women than men responded, and women report lower job satisfaction.","pith_inferences":["Beyond the paper, clustering respondents on the tasks and technologies they report would test whether academic background truly leaves no signature; if it does not, the clusters should not separate by degree field.","Beyond the paper, the recruitment through programming-centric forums and personal contacts may over-represent code-comfortable practitioners, so the high qualification level could be an upper-bound estimate; a registry-based or employer-based sample would check this.","Beyond the paper, the deep-learning gap is read as a skills gap, but it could equally be an exposure gap; comparing confidence among practitioners who have completed deep-learning projects, regardless of degree, would separate the two."],"forward_implications":["Tool and training efforts should target the shared bottlenecks—access to quality data and applying deep learning—rather than remedial computer science education.","Because Python, SQL, and spreadsheet editors dominate across both groups, improving interoperability around these tools would serve most practitioners.","Self-reported confidence in deep learning is the clearest skill gap, and it is larger among professionals without a CS background, so focused support there would reach a real need.","The gender imbalance and lower satisfaction reported by women point to an equity issue that additional technical training alone would not address."],"supporting_citations":[{"why":"Supplies the definition and rationale of inferential statistics that the paper uses to generalize from the sample to the population.","marker":"[1]"},{"why":"Provides the Fisher's exact test method used to decide which differences between groups are statistically significant.","marker":"[12]"},{"why":"Underpins the survey-question design choices on which the quality of all collected data depends.","marker":"[2]"},{"why":"Establishes the prior introspective survey of data scientists that the paper extends by adding academic background and tool preferences.","marker":"[9]"},{"why":"Contributes the earlier role characterization of data scientists that motivates the paper's background-aware analysis of work styles.","marker":"[13]"},{"why":"Provides the prior finding that certain data science skills are most desired by employers, against which the paper's self-rated competence results are read.","marker":"[16]"},{"why":"Supplies interview-based evidence on the concrete steps of data work that the paper's task self-evaluation questions operationalize.","marker":"[18]"},{"why":"Provides a large-scale prior survey of data analytics professionals that anchors the survey-based characterization tradition the paper continues.","marker":"[23]"}],"fun_headline_variants":["Data scientists aren't all self-taught; CS degree doesn't matter much","CS degree doesn't predict data science struggles; data access does","Real data science pain: poor data access, not lack of CS training","CS background doesn't shape data scientists' daily work; data quality does"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the 116 people who filled out the survey, recruited through online forums and personal contacts, represent data science professionals worldwide; if the sample is skewed, the profile and the conclusion that academic background has little impact describe only this respondent pool.","fun_headline_variants_meta":{"raw":{"variants":["Data scientists aren't all self-taught; CS degree doesn't matter much","CS degree doesn't predict data science struggles; data access does","Real data science pain: poor data access, not lack of CS training","CS background doesn't shape data scientists' daily work; data quality does"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00052,"raw_usage":{"total_tokens":2503,"prompt_tokens":913,"completion_tokens":1590,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":1512}},"tokens_in":529,"tokens_out":1590,"duration_ms":11523,"temperature":1.0,"reasoning_tokens":1512,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:45:39.596364+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same survey on a stratified sample drawn from company employment records, professional certifications, or national labor statistics instead of open online forums, and compare the proportions for CS background, task confidence, and reported difficulties. If the background effects or difficulty rankings differ materially from the paper's numbers, the claim that academic background has little impact on work would fail as a general statement.","supporting_citations":[{"cited_title":"Inferential Statistics","cited_arxiv_id":null,"evidence_quote":"Supplies the definition and rationale of inferential statistics that the paper uses to generalize from the sample to the population."},{"cited_title":"Statistical notes for clinical researchers: Chi-squared test and Fisher’s exact test","cited_arxiv_id":null,"evidence_quote":"Provides the Fisher's exact test method used to decide which differences between groups are statistically significant."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Underpins the survey-question design choices on which the quality of all collected data depends."},{"cited_title":"Analyzing the Analyz- ers: An Introspective Survey of Data Scientists and Their Work","cited_arxiv_id":null,"evidence_quote":"Establishes the prior introspective survey of data scientists that the paper extends by adding academic background and tool preferences."},{"cited_title":"The emerging role of data scientists on software development teams","cited_arxiv_id":null,"evidence_quote":"Contributes the earlier role characterization of data scientists that motivates the paper's background-aware analysis of work styles."},{"cited_title":"D1.4 Study Evaluation Report 2","cited_arxiv_id":null,"evidence_quote":"Provides the prior finding that certain data science skills are most desired by employers, against which the paper's self-rated competence results are read."},{"cited_title":"Vera Liao, Casey Dugan, and Thomas Erickson","cited_arxiv_id":null,"evidence_quote":"Supplies interview-based evidence on the concrete steps of data work that the paper's task self-evaluation questions operationalize."},{"cited_title":"2015 Data Science Survey","cited_arxiv_id":null,"evidence_quote":"Provides a large-scale prior survey of data analytics professionals that anchors the survey-based characterization tradition the paper continues."}],"review_version":1}