Pith. sign in

REVIEW 3 major objections 4 minor 59 references

Detecting Soft Skills in ML Engineering Roles CVs

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Soft skills in technical CVs appear about three times as often in narrative prose as in keyword lists, so keyword-based screening systematically misses them.

desk verdict First real candidate-side evidence on soft-skill disclosure in ML-adjacent CVs, with careful inferential statistics; the role and seniority findings are likely solid, but the headline 3:1 narrative-over-keyword ratio rests on an unvalidated explicit/implicit measurement split that needs a major revision. read the letter →

arxiv 2608.10046 v1 pith:ZP2DIKW3 submitted 2026-08-10 cs.LG cs.CY

classification cs.LGcs.CY
keywords softskillsCVanalysisMLengineersdatascientistssoftwareLLMextractionnarrativedisclosurekeywordscreening
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper studies how ML engineers, data scientists, and software engineers articulate soft skills in their CVs, using a balanced corpus of 300 documents and an LLM-based extractor that distinguishes explicitly listed from implicitly narrated mentions. It converts the demand-side literature's claims about these roles into 13 falsifiable hypotheses and tests them with effect sizes and family-wise error control. The central finding is that candidates disclose soft skills through narrative rather than keyword lists by roughly three to one, and most so for the competencies employers value most: leadership, coordination, and mentoring are 88–96% narrative. Seniority nearly triples the odds of articulating leadership, and leadership is not role-invariant, as software engineers mention it at roughly half the rate of their peers. The practical conclusion is that a keyword-based screening pipeline systematically under-detects the soft-skill evidence that differentiates candidates.

What carries the argument

The central object is a two-stage LLM-based extraction pipeline that separates explicit from implicit soft-skill mentions. Stage one extracts skills only from dedicated skills sections (set $E$); stage two extracts all skills from the full text (set $T$); implicit narrative mentions are defined deductively as $I = T - E$. Skills are mapped to the TaxoSoft taxonomy, and each extraction is tied to a supporting phrase from the CV, allowing verification. This design makes explicit and implicit disclosure comparable within the same CV as a paired observation, which is what the disclosure-style hypotheses require.

What would settle it

Manually annotate a fresh held-out set of CVs for narrative versus keyword disclosure, or measure the pipeline's sensitivity separately on skills sections and on work-experience sections; if human raters find narrative and keyword mentions near equal, or if the pipeline's sensitivity differs sharply between the two sections, the 1,007-to-347 dominance and the claim that keyword screening misses most soft skills would be overturned.

Watch

Extended reading notes

Core claim

The paper claims that the candidate side of the soft-skills picture is measurable and that it partly corroborates, partly corrects, the demand side. Candidates articulate soft skills predominantly through narrative: 1,007 of 1,354 unique skill–CV pairs are implicit, versus 347 explicit, and 94.7% of CVs whose two channels disagree disclose only through narrative. Narrative dominance is concentrated in exactly the competencies that discriminate between candidates—coaching, coordination, mentoring, presentation, and leadership are 88–100% narrative—while generic labels like communication and analytical ability are as often keyword-listed. Seniority roughly triples the odds of articulating leadership (adjusted OR 3.06), and the effect is homogeneous across roles; collaboration accumulates alongside leadership instead of being displaced. The prediction that leadership would be role-invariant is refuted: software engineers articulate it at 26%, versus 47% for data scientists and 42% for ML engineers. Eleven of 13 hypotheses are supported, one partially, and one refuted.

Load-bearing premise

The whole narrative-dominance result depends on the assumption that subtracting skills-section mentions from total mentions ($I = T - E$) correctly measures narrative disclosure, and that the LLM does not systematically over-detect narrative language while under-detecting skills sections; if that asymmetry is a pipeline artifact rather than a property of the CVs, the 3-to-1 ratio and the per-competency narrative shares could be wrong.

Editorial extensions

If this is right

  • Keyword-based applicant tracking systems will under-detect soft skills, missing roughly 74% of the unique skill–CV pairs in this corpus.
  • A keyword-only extractor cannot detect the seniority effect on leadership, the strongest signal in the study, because only 14 of 115 leadership disclosures appear as explicit labels.
  • Screening results will be systematically biased toward generic competencies (communication, analytical ability) and away from evidence-bearing ones (leadership, coordination, mentoring), distorting candidate ranking.
  • Role-specific guidance follows: data scientists' communication emphasis and ML engineers' mentoring emphasis are candidate-side signatures that role-tailored screening or coaching could exploit.
  • The refutation of leadership universality means software engineering candidates who lead without saying so are less visible than equally leading peers in other roles.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the narrative-dominance pattern replicates in other corpora, the soft-skills gap is partly an articulation gap: competencies are present in candidates' own words but absent from the vocabularies of machine screening, which reframes the problem as teachable CV-writing skill rather than only competency development.
  • The leadership gap for software engineers could be partly a vocabulary effect if those candidates use titles like 'tech lead' or describe mentoring instead of 'leadership'; a follow-up expanding the taxonomy's variants could test whether the 26% figure is a true role difference or a labeling artifact.
  • The two-stage subtraction design could be validated directly by collecting human annotations of implicit mentions on held-out CVs; if it holds, the same machinery transfers to LinkedIn profiles or cover letters to test whether narrative dominance is a property of the document type.
  • A corpus stratified by collection period within each role, as the paper itself suggests, could separate period effects in CV-writing conventions from genuine role and seniority signals.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies how soft skills are articulated in 300 curated CVs from ML engineers, data scientists, and software engineers. The authors build an LLM-based extraction pipeline that distinguishes explicitly listed soft skills (in dedicated skills sections, E) from implicitly narrated ones (I = T − E), validate the pipeline against a human-annotated ground truth (F1 = 0.72), and test 13 hypotheses derived from the demand-side literature with effect sizes, confidence intervals, Holm–Bonferroni correction, and TOST equivalence testing. Eleven hypotheses are supported, one partially, one refuted. Key findings are that narrative disclosure outweighs keyword disclosure by about 1,007 to 347 unique skill–CV pairs; leadership, coordination, mentoring, and coaching are 88–100% narrative; seniority nearly triples the odds of articulating leadership (adjusted OR 3.06); and leadership articulation is not role-invariant, with software engineers at roughly half the rate of the other two roles.

Significance. If the measurement assumptions hold, this is a valuable and rare candidate-side complement to the job-posting and interview literature on soft skills. The study is unusually careful on the statistical side: it reports effect sizes with confidence intervals, controls family-wise error, uses TOST for the invariance hypothesis, and provides sensitivity analyses for seniority dichotomization, document length, and non-differential misclassification. The hypotheses are derived from external literature and several could have failed, which strengthens the confirmatory interpretation. The paper also ships a replication package with prompts, extraction outputs, and analysis scripts, which supports reproducibility. However, the headline practical claim — that keyword-based screening systematically misses soft skills because candidates write them narratively — rests on the unvalidated asymmetry between the constrained skills-section extraction (E) and full-text extraction (T). Because the paper does not separately measure the recall of these two stages, the central 3:1 narrative-to-keyword ratio could be an artifact of the measurement design.

major comments (3)
  1. [§3.5, §4, §5.5] The implicit set I = T − E is the load-bearing construction for H3a, the 1,007 vs. 347 narrative-to-keyword ratio, and the narrative shares in Table 6, but the paper validates only the aggregate pipeline F1 of 0.72 on 100 CVs (Section 4) and never reports recall for the constrained skills-section stage (E) and the full-text stage (T) separately. If the E stage under-detects explicitly listed skills — because the skills-section locator misses nonstandard headings or because the constrained prompt is more conservative than the full-text prompt — those explicitly listed skills remain in T and are automatically counted as implicit, mechanically inflating the narrative ratio. The Section 5.5 assertion that systematic extraction bias “largely cancels” because both branches use the same model and taxonomy is not supported: the two stages use different prompts and different input spans, so their recall for the same underlying explicit skills can differ systematically. The quantitative bias analysis in Section 5.5 corrects prevalence estimates, not the E/I partition, so H3a and Table 6 are not protected by that sensitivity analysis. I recommend that the revision either (a) validate stage-specific recall and precision on the human-annotated ground truth with explicit/implicit labels, or (b) re-estimate H3a under plausible E-recall scenarios and report how the narrative ratio changes. If the ratio is sensitive, the screening-pipeline claim in Sections 6.5 and 8 should be softened accordingly.
  2. [§3.2, §5.1] The purposeful selection of 62 data scientist and 53 software engineer CVs from the GitHub corpus based on “text completeness” is a selection-on-outcome risk for the role comparisons, particularly H1c (no-skill rate) and the skill-density results in Section 5.1. Data scientists selected for CVs that contain the full set of expected sections (summary, experience, skills, education) could mechanically have lower no-skill rates than ML engineers, whose corpus subset was not selected on that criterion. The paper should provide a sensitivity analysis using a random or source-representative sample from the corpus, or otherwise quantify how much of the H1c contrast (4.0% vs. 18.0% no-skill rate) could be attributable to this selection rule.
  3. [§4] The extraction validation rests on a ground truth whose reliability is not quantified: the paper states that no formal inter-coder reliability statistic was computed (Section 4), and the 100-CV evaluation set is drawn from the same corpus on which the pipeline was tuned through 22 trials. This is especially limiting for the explicit-versus-implicit distinction, because the E stage and T stage are never separately validated and the two annotators resolved disagreements through joint calibration. Reporting Cohen’s kappa or a comparable agreement measure, ideally stratified by explicit vs. implicit mentions, would materially strengthen the claim that the gold standard supports the fine-grained disclosure-style analysis. At minimum, the absence of this statistic should be acknowledged as a limitation that affects H3a directly.
minor comments (4)
  1. [Abstract and throughout] Several passages are missing spaces between words, e.g., “Weclosebothgaps” in the abstract, “disclosesoftskillsthroughnarrativeratherthankeywordlists”, and “Figure 1:An overview of the study process”. A full proofreading pass is needed.
  2. [§5.3] The sentence “only 3.3% of candidates (10/300) relied exclusively on explicit keyword lists” is clear, but the later phrase “roughly three-quarters (74.4%) of the soft skill evidence” could be more carefully tied to the assumption that E and T have comparable recall; otherwise it risks overstating what the extraction pipeline established.
  3. [§3.6] The description of the sensitivity of the design is helpful, but the phrase “the smallest differences it can reliably detect” would be clearer with an explicit statement that these are minimum detectable differences under the stated baseline prevalence and power, not a guarantee about all competencies.
  4. [Table 6] The table caption says “Keyword” and “Narrative” partition the CVs in which the competency was detected, but the column sums can exceed the total prevalence if a CV contains both channels; consider adding a “Both” column or clarifying that the two columns are not mutually exclusive at the CV level.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: hypotheses are external, outcomes are falsifiable, and the narrative-versus-keyword measure is a data-dependent decomposition rather than a forced identity.

full rationale

The derivation chain is self-contained and falsifiable. The 13 hypotheses are imported from external demand-side studies (job-advertisement and interview literature: refs. 14, 30, 34, 37, 40, 6, 21), not from the paper's own fitted quantities, and the analysis can fail: H1e predicted no role difference in leadership and was refuted, and H2e explicitly adjudicated between competing substitution and accumulation accounts, with accumulation winning. The narrative-versus-keyword claim is built on the definitional decomposition I = T − E (Section 3.5), but that definition does not force the relative magnitude of the two disjoint channels; the empirical comparison was data-dependent (177 implicit-only vs. 10 explicit-only discordant CVs, McNemar p < 0.001). The same-corpus evaluation set and 22 tuning trials are a measurement-validation concern, not a circular reduction: the pipeline's F1 = 0.72 is benchmarked against human annotation and external systems (Table 4), and Section 5.5 propagates misclassification error through the estimates rather than assuming it away. The paper explicitly states that no formal inter-coder reliability statistic was computed (Section 4), which is an admitted limitation of the gold standard, but it does not make any result equivalent to its inputs by construction. Self-citations (refs. 3, 4, 21) support background motivation and one hypothesis grounding, but no load-bearing uniqueness claim or ansatz is imported from the authors' prior work; the hypotheses and outcome measures would stand independently of those citations. The unvalidated asymmetry between the constrained skills-section stage and the full-text stage is a construct-validity and differential-recall threat to the I = T − E estimator, but the paper does not define 'narrative' in terms of the conclusion, and no equation reduces the headline claim to its own inputs.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

No new physical or theoretical entities are introduced. The analysis depends on three hand-chosen thresholds, the validity of an external taxonomy, and assumptions about extraction error and sample representativeness, all of which the paper discusses in its threats to validity.

free parameters (3)
  • Seniority threshold = 5 years
    Binary junior/senior split used for all seniority hypotheses (H2a-H2f); chosen from industry leveling standards, not fitted.
  • Equivalence margin for H1e = ±20 percentage points
    Margin used in TOST for the leadership invariance hypothesis; chosen as the smallest practically meaningful role difference. Different margin could change the H1e outcome.
  • Exploratory inclusion threshold = 15 CVs
    Only competencies detected in at least 15 CVs entered the exploratory role and seniority contrasts, shaping the exploratory families.
assumptions (5)
  • domain assumption TaxoSoft taxonomy is a valid and sufficient inventory of soft skills for CV analysis
    All extraction and annotation rely on TaxoSoft [25] labels and variants (Section 3.3); skills outside its language are excluded, acknowledged in Section 7.1.
  • domain assumption The LLM extraction error is approximately non-differential across roles and seniority levels
    The quantitative bias analysis in Section 5.5 assumes group-invariant sensitivity and specificity; if errors were strongly differential, the corrected gaps could shrink.
  • domain assumption The 300-CV corpus is representative of the three target role populations
    The corpus is a convenience sample of public, English-language CVs with purposeful selection (Section 3.2); external validity is limited (Section 7.4).
  • domain assumption Demand-side literature used to derive hypotheses was interpreted correctly
    Hypotheses and predictions in Table 1 are grounded in [37,6,14,40,34]; an incorrect reading would change the falsifiable content.
  • standard math Standard frequentist tests are valid for the observed cell sizes
    Fisher's exact, chi-square, TOST, McNemar, logistic regression and rank tests are used as specified in Section 3.6, with expected cell counts checked.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Detecting Soft Skills in ML Engineering Roles CVs." pith.science (2026). https://pith.science/paper/ZP2DIKW3

@misc{pith2026260810046,
  author       = {Pith},
  title        = {Pith review of: Detecting Soft Skills in ML Engineering Roles CVs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZP2DIKW3}},
  note         = {Machine review of arXiv:2608.10046}
}
read the original abstract

Soft skills shape collaboration among ML engineers, data scientists, and software engineers building ML-enabled systems, yet what we know about them comes almost entirely from the demand side. Job advertisements, surveys, and hiring manager interviews capture what employers ask for. How candidates themselves articulate these competencies has not been studied, and existing CV-mining work is both keyword-based, so it cannot see skills conveyed through narrative, and descriptive, reporting frequency rankings without testing whether group differences exceed sampling variation. We close both gaps. Using a balanced corpus of 300 curated CVs spanning the three roles, we extract explicitly listed and implicitly narrated soft skills with an LLM-based pipeline validated against a human-annotated ground truth, a distinction that existing extractors were not designed to make. We then convert the demand-side literature's claims into 13 falsifiable hypotheses about role signatures, seniority progression, and disclosure style, and test them with effect sizes under family-wise error control, so that candidate-side data can corroborate or contradict the demand-side account rather than merely illustrate it. Eleven hypotheses are supported, one partially, and one refuted. Candidates disclose soft skills through narrative rather than keyword lists by roughly three to one, and most so for the competencies employers value most: leadership, coordination, and mentoring (88-96% narrative). Seniority nearly triples the odds of articulating leadership. That competency, assumed universal in prior work, is articulated by software engineers at half the rate of their peers. Technical candidates do articulate soft skills, but a keyword-based screening systematically misses them.

Figures

Figures reproduced from arXiv: 2608.10046 by the authors.

Figure 1
Figure 1. An overview of the study process This question investigates the rhetorical patterns candi￾dates use to present soft skills, distinguishing between ex￾plicit mentions, in which skills are listed as keywords in ded￾icated sections, and implicit mentions, in which behavioral competencies are conveyed through narrative descriptions of professional experience. The answer indicates whether candidates rely primarily on key… view at source ↗
Figure 2
Figure 2. An example of ML engineer/data scientist role ambiguity in a CV current labor-market practices, a supplementary dataset of contemporary CVs was manually curated. For that, publicly hosted CVs were collected through targeted Google queries using the filetype:pdf operator to restrict results to PDF documents. Then, they were manually checked to remove the sample ones. The specific queries used were: • ML CV filetype:p… view at source ↗
Figure 3
Figure 3. Top 10 soft skills mentioned across all roles Collaboration Leadership Communication Problem solving Presentation Coordination Mentoring Time management Creativity Self-organized Coaching Writing Passion Initiative Speaking Soft Skill 0 10 20 30 40 50 60 70 Frequency 68 47 54 34 40 30 18 10 21 21 11 13 16 9 3 50 42 18 15 30 22 28 6 10 6 16 14 9 7 14 66 26 31 35 10 25 13 29 9 12 6 2 4 11 1 Role Data Scientist ML Engi… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: The number of soft skills mentioned by each role 3.94]), ℎ = 0.31, 𝑝 = 0.013 (Holm-adjusted 𝑝 = 0.029). H1b is supported, though with a considerably smaller effect than H1a, and the confidence interval admits differences as small as three percentage points. H1c (no-ski…
Figure 5
Figure 5. Figure 5: Distribution of soft skill disclosure styles across the CVs. examining each disclosure style in more detail, explicit mentions were identified in 28.3% of CVs (85/300). The most frequently listed skills within these structured sections were communication (15.7%) and co…
Figure 6
Figure 6. Figure 6: Soft skill disclosure styles by role [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

59 extracted references · 25 canonical work pages

  1. [1]

    Can llms replace manual annotation of software engineering artifacts?, in: 2025IEEE/ACM22ndInternationalConferenceonMiningSoftware Repositories(MSR),IEEE.p.526–538

    Ahmed, T., Devanbu, P., Treude, C., Pradel, M., 2025. Can llms replace manual annotation of software engineering artifacts?, in: 2025IEEE/ACM22ndInternationalConferenceonMiningSoftware Repositories(MSR),IEEE.p.526–538. URL:http://dx.doi.org/10. 1109/msr66628.2025.00086, doi:10.1109/msr66628.2025.00086

  2. [2]

    Job description parsing with explainable trans- former based ensemble models to extract the technical and non- technical skills

    Akkasi, A., 2024. Job description parsing with explainable trans- former based ensemble models to extract the technical and non- technical skills. Natural Language Processing Journal 9, 100102. URL:http://dx.doi.org/10.1016/j.nlp.2024.100102, doi:10.1016/j. nlp.2024.100102

  3. [3]

    Azamnouri,A.,2025. Cocochallengesinmlengineeringteams:How tocollaborativelybuildml-enabledsystems,in:2025IEEE/ACM4th International Conference on AI Engineering – Software Engineering forAI(CAIN),IEEE.p.241–243. URL:http://dx.doi.org/10.1109/ cain66642.2025.00036, doi:10.1109/cain66642.2025.00036

  4. [4]

    Teaching ai competencies: Experiences from coaching interdisciplinary teams to develop ai prototypes

    Azamnouri, A., Hörauf, D., Schönberger, B., Bogner, J., Wagner, S., 2026. Teaching ai competencies: Experiences from coaching interdisciplinary teams to develop ai prototypes. Computers and Education: Artificial Intelligence 10, 100580. URL:http://dx.doi. org/10.1016/j.caeai.2026.100580, doi:10.1016/j.caeai.2026.100580

  5. [5]

    Guidelines for empirical studies in software engineering involving large language models

    Baltes, S., Angermeir, F., Arora, C., Barón, M.M., Chen, C., Böhme, L., Calefato, F., Ernst, N., Falessi, D., Fitzgerald, B., et al., 2025. Guidelines for empirical studies in software engineering involving large language models. arXiv preprint arXiv:2508.15503

  6. [6]

    Busquim, G., Araújo, A.A., Lima, M.J., Kalinowski, M., 2024a. Towards effective collaboration between software engineers and data scientists developing machine learning-enabled systems, in: Anais do XXXVIII Simpósio Brasileiro de Engenharia de Software (SBES 2024), Sociedade Brasileira de Computação. p. 24–34. URL:http: //dx.doi.org/10.5753/sbes.2024.3027...

  7. [7]

    On the Interaction Between Software Engineers and Data Scien- tists When Building Machine Learning-Enabled Systems

    Busquim, G., Villamizar, H., Lima, M.J., Kalinowski, M., 2024b. On the Interaction Between Software Engineers and Data Scien- tists When Building Machine Learning-Enabled Systems. Springer Nature Switzerland. p. 55–75. URL:http://dx.doi.org/10.1007/ 978-3-031-56281-5_4, doi:10.1007/978-3-031-56281-5_4

  8. [8]

    Calanca, F., Sayfullina, L., Minkus, L., Wagner, C., Malmi, E.,

Show all 59 references
  1. [9]

    Evaluating clinical ai summaries with large language models as judges

    Croxford,E.,Gao,Y.,First,E.,Pellegrino,N.,Schnier,M.,Caskey,J., Oguss,M.,Wills,G.,Chen,G.,Dligach,D.,Churpek,M.M.,Mayam- purath, A., Liao, F., Goswami, C., Wong, K.K., Patterson, B.W., Af- shar, M., 2025. Evaluating clinical ai summaries with large language models as judges. n...

  2. [10]

    De Morais Leça, M., De Souza Santos, R., 2025. Curious, critical thinker, empathetic, and ethically responsible: Essential soft skills for data scientists in software engineering, in: 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in ...

  3. [11]

    Esco: Towards a semantic web for the european labor market

    De Smedt, J., le Vrang, M., Papantoniou, A., 2015. Esco: Towards a semantic web for the european labor market. Ldow@ www 144

  4. [12]

    Extreme multi-label skill extraction training using large language models

    Decorte, J.J., Verlinden, S., Hautte, J.V., Deleu, J., Develder, C., Demeester, T., 2023. Extreme multi-label skill extraction training using large language models. URL:https://arxiv.org/abs/2307. 10778,arXiv:2307.10778

  5. [13]

    The growing importance of social skills in the labor market*

    Deming, D.J., 2017. The growing importance of social skills in the labor market*. The Quarterly Journal of Economics 132, 1593–1640. URL:http://dx.doi.org/10.1093/qje/qjx022, doi:10. 1093/qje/qjx022

  6. [14]

    Galster, M., Mitrovic, A., Malinen, S., Holland, J., 2022. What soft skills does the software industry *really* want? an exploratory study ofsoftwarepositionsinnewzealand,in:Proceedingsofthe16thACM / IEEE International Symposium on Empirical Software Engineering and Measuremen...

  7. [15]

    Senior data scientist jobs in berlin.https: //www.glassdoor.com/Job/berlin-senior-data-scientist-jobs-SRCH_ IL.0,6_IM1020_KO7,28.htm

    Glassdoor, 2026. Senior data scientist jobs in berlin.https: //www.glassdoor.com/Job/berlin-senior-data-scientist-jobs-SRCH_ IL.0,6_IM1020_KO7,28.htm. Accessed: 2026-01-24

  8. [16]

    A Probabilistic Interpretation of Pre- cision, Recall and F-Score, with Implication for Evaluation

    Goutte, C., Gaussier, E., 2005. A Probabilistic Interpretation of Pre- cision, Recall and F-Score, with Implication for Evaluation. Springer Berlin Heidelberg. p. 345–359. URL:http://dx.doi.org/10.1007/ 978-3-540-31865-1_25, doi:10.1007/978-3-540-31865-1_25

  9. [17]

    Applied Thematic Analysis

    Guest, G., MacQueen, K., Namey, E., 2012. Applied Thematic Analysis. SAGEPublications,Inc. URL:http://dx.doi.org/10.4135/ 9781483384436, doi:10.4135/9781483384436

  10. [18]

    Implicit skills extraction using document embedding and its use in job recommendation

    Gugnani, A., Misra, H., 2020. Implicit skills extraction using document embedding and its use in job recommendation. Pro- ceedings of the AAAI Conference on Artificial Intelligence 34, 13286–13293. URL:http://dx.doi.org/10.1609/aaai.v34i08.7038, doi:10.1609/aaai.v34i08.7038

  11. [19]

    Gener- ating unified candidate skill graph for career path recommendation, in: 2018 IEEE International Conference on Data Mining Workshops (ICDMW), IEEE

    Gugnani, A., Reddy Kasireddy, V.K., Ponnalagu, K., 2018. Gener- ating unified candidate skill graph for career path recommendation, in: 2018 IEEE International Conference on Data Mining Workshops (ICDMW), IEEE. p. 328–333. URL:http://dx.doi.org/10.1109/ icdmw.2018.00054, doi:1...

  12. [20]

    Language writ large: Llms, chatgpt, grounding, meaning and understanding

    Harnad, S., 2025. Language writ large: Llms, chatgpt, grounding, meaning and understanding. URL:https://arxiv.org/abs/2402. 02243,arXiv:2402.02243

  13. [21]

    pp.16–36

    Haug, M., Azamnouri, A., Fritz, M., Woltmann, L., Schriever, C., Wagner,S.,2025.MLOpsAdoptionintheManufacturingIndustry:A CaseStudywithZeissSMT.SpringerNatureSwitzerland. pp.16–36. URL:http://dx.doi.org/10.1007/978-3-032-07313-6_2, doi:10.1007/ 978-3-032-07313-6_2

  14. [22]

    Fromcodetocourtroom:Llmsasthenewsoftwarejudges

    He, J., Shi, J., Zhuo, T.Y., Treude, C., Sun, J., Xing, Z., Du, X., Lo, D.,2025. Fromcodetocourtroom:Llmsasthenewsoftwarejudges. URL:https://arxiv.org/abs/2503.02246,arXiv:2503.02246

  15. [23]

    Seniordatascientistjobsinberlin.https://de.indeed

    Indeed,2026. Seniordatascientistjobsinberlin.https://de.indeed. com/q-senior-data-scientist-l-berlin-jobs.html. Accessed: 2026- 01-24

  16. [24]

    Skills prediction based on multi-label resume classification using cnn with model predictions explanation

    Jiechieu, K.F.F., Tsopze, N., 2020. Skills prediction based on multi-label resume classification using cnn with model predictions explanation. Neural Computing and Applications 33, 5069–5087. URL:http://dx.doi.org/10.1007/s00521-020-05302-x, doi:10.1007/ s00521-020-05302-x. A....

  17. [25]

    Building a soft skill taxonomy from job openings

    Khaouja, I., Mezzour, G., Carley, K.M., Kassou, I., 2019. Building a soft skill taxonomy from job openings. Social Network Analysis and Mining 9. URL:http://dx.doi.org/10.1007/s13278-019-0583-9, doi:10.1007/s13278-019-0583-9

  18. [26]

    Kolluru, K., Adlakha, V., Aggarwal, S., Mausam, Chakrabarti, S.,

  19. [27]

    Digital marketing employ- ability skills in job advertisements – must-have soft skills for entry level workers: A content analysis

    Kovacs, I., Vamosi Zarandne, K., 2022. Digital marketing employ- ability skills in job advertisements – must-have soft skills for entry level workers: A content analysis. Economics &amp; Sociology 15, 178–192. URL:http://dx.doi.org/10.14254/2071-789x.2022/15-1/ 11, doi:10.1425...

  20. [28]

    The levels.fyi standard: Career level framework

    Levels.fyi, 2026. The levels.fyi standard: Career level framework. https://www.levels.fyi/2020/. Accessed: 2026-01-24

  21. [29]

    URL: https://arxiv.org/abs/2304.11060,arXiv:2304.11060

    Li,N.,Kang,B.,Bie,T.D.,2023.Skillgpt:arestfulapiserviceforskill extraction and standardization using a large language model. URL: https://arxiv.org/abs/2304.11060,arXiv:2304.11060

  22. [30]

    Loufek, S.B., Santos, F., Trinkenreich, B., 2025. Beyond the job posting:Whathiringmanagersseekinentry-levelsoftwareengineer- ing candidates, in: 2025 ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM), IEEE. p. 01–11. URL:http://dx.doi.o...

  23. [31]

    Softskills,hardskills:Whatmattersmost?evi- dencefromjobpostings

    Lyu,W.,Liu,J.,2021. Softskills,hardskills:Whatmattersmost?evi- dencefromjobpostings. AppliedEnergy300,117307. URL:http:// dx.doi.org/10.1016/j.apenergy.2021.117307,doi:10.1016/j.apenergy. 2021.117307

  24. [32]

    Mailach, A., Siegmund, N., 2023. Socio-technical anti-patterns in building ml-enabled software: Insights from leaders on the forefront, in: 2023 IEEE/ACM 45th International Conference on Software En- gineering (ICSE), IEEE. p. 690–702. URL:http://dx.doi.org/10. 1109/icse48619....

  25. [33]

    Malherbe, E., Aufaure, M.A., 2016. Bridge the terminology gap between recruiters and candidates: A multilingual skills base built fromsocialmediaandlinkeddata,in:2016IEEE/ACMInternational Conference on Advances in Social Networks Analysis and Mining (ASONAM), IEEE. p. 583–590....

  26. [34]

    Malinen,S.,Galster,M.,Mitrovic,A.,Iyer,S.S.,Peirisand,P.,Clarke, A., 2025. Soft skills in software engineering: Insights from the trenches,in:2025IEEE/ACM47thInternationalConferenceonSoft- ware Engineering: Software Engineering in Practice (ICSE-SEIP), IEEE.p.296–306. URL:http...

  27. [35]

    In-context learning inlargelanguagemodels(llms):Mechanisms,capabilities,andimpli- cations for advanced knowledge representation and reasoning

    Mohamed, A., Rashid, M.E., Shaalan, K., 2025. In-context learning inlargelanguagemodels(llms):Mechanisms,capabilities,andimpli- cations for advanced knowledge representation and reasoning. IEEE Access 13, 95574–95593. URL:http://dx.doi.org/10.1109/access. 2025.3575303, doi:10....

  28. [36]

    Behavioral Sciences 14, 894

    Mohammed,F.S.,Ozdamli,F.,2024.Asystematicliteraturereviewof soft skills in information technology education. Behavioral Sciences 14, 894. URL:http://dx.doi.org/10.3390/bs14100894, doi:10.3390/ bs14100894

  29. [37]

    Nahar, N., Zhou, S., Lewis, G., Kästner, C., 2022. Collabora- tion challenges in building ml-enabled systems: communication, documentation, engineering, and process, in: Proceedings of the 44th International Conference on Software Engineering, ACM. p. 413–425. URL:http://dx.do...

  30. [38]

    In The AI Era, Soft Skills Are The New Hard Skills

    Niva, D., Yariv, I., 2020. In The AI Era, Soft Skills Are The New Hard Skills. Emerald Publishing Limited. p. 55–78. URL:http: //dx.doi.org/10.1108/book-978-1-64802-075-920251007, doi:10.1108/ book-978-1-64802-075-920251007

  31. [39]

    Exploringtheimportanceofsoft and hard skills as perceived by it internship students and industry: A gap analysis

    Patacsil,F.,S.Tablatin,C.L.,2017. Exploringtheimportanceofsoft and hard skills as perceived by it internship students and industry: A gap analysis. Journal of Technology and Science Education 7, 347. URL:http://dx.doi.org/10.3926/jotse.271, doi:10.3926/jotse.271

  32. [40]

    From junior to senior: Skill require- ments for ai professionals across career stages

    Peretz, O., Nakash, M., 2025. From junior to senior: Skill require- ments for ai professionals across career stages. Proceedings of the International Conference on Research in Business, Management and Finance 2, 9–20. URL:http://dx.doi.org/10.33422/icrbmf.v2i1. 1178, doi:10.33...

  33. [41]

    Ontology-Based Resume Searching System for Job Applicants in Information Technology

    Phan, T.T., Pham, V.Q., Nguyen, H.D., Huynh, A.T., Tran, D.A., Pham, V.T., 2021. Ontology-Based Resume Searching System for Job Applicants in Information Technology. Springer Interna- tional Publishing. p. 261–273. URL:http://dx.doi.org/10.1007/ 978-3-030-79457-6_23, doi:10.10...

  34. [42]

    How ai developers overcome communication challenges in a multidisciplinary team: A case study

    Piorkowski,D.,Park,S.,Wang,A.Y.,Wang,D.,Muller,M.,Portnoy, F., 2021. How ai developers overcome communication challenges in a multidisciplinary team: A case study. Proceedings of the ACM on Human-Computer Interaction 5, 1–25. URL:http://dx.doi.org/10. 1145/3449205, doi:10.1145/3449205

  35. [43]

    Senior Software Engineer Salary (Up- dated for 2026).https://www.roberthalf.com/us/en/job-details/ senior-software-engineer

    Robert Half, 2026. Senior Software Engineer Salary (Up- dated for 2026).https://www.roberthalf.com/us/en/job-details/ senior-software-engineer. 2026SalaryGuide;accessed2026-01-24

  36. [44]

    Soft skills: students and employers crave

    Romanenko, Y.N., Stepanova, M., Maksimenko, N., 2024. Soft skills: students and employers crave. Humanities and Social Sciences Communications 11. URL:http://dx.doi.org/10.1057/ s41599-024-03250-8, doi:10.1057/s41599-024-03250-8

  37. [45]

    Resume parsing framework for e-recruitment, in: 2022 16th International Conference on Ubiqui- tousInformationManagementandCommunication(IMCOM),IEEE

    Sajid, H., Kanwal, J., Bhatti, S.U.R., Qureshi, S.A., Basharat, A., Hussain, S., Khan, K.U., 2022. Resume parsing framework for e-recruitment, in: 2022 16th International Conference on Ubiqui- tousInformationManagementandCommunication(IMCOM),IEEE. p. 1–8. URL:http://dx.doi.org...

  38. [46]

    The Coding Manual for Qualitative Re- searchers

    Saldana, J., 2021. The Coding Manual for Qualitative Re- searchers. SAGEPublicationsLtd. URL:http://dx.doi.org/10.4135/ 9781036235611, doi:10.4135/9781036235611

  39. [47]

    Re- sume Screening Using Natural Language Processing and Ma- chine Learning: A Systematic Review

    Sinha, A.K., Amir Khusru Akhtar, M., Kumar, A., 2021. Re- sume Screening Using Natural Language Processing and Ma- chine Learning: A Systematic Review. Springer Singapore. p. 207–214. URL:http://dx.doi.org/10.1007/978-981-33-4859-2_21, doi:10.1007/978-981-33-4859-2_21

  40. [48]

    Syntax- based skill extractor for job advertisements, in: 2019 6th Swiss Conference on Data Science (SDS), IEEE

    Smith, E., Braschler, M., Weiler, A., Haberthuer, T., 2019. Syntax- based skill extractor for job advertisements, in: 2019 6th Swiss Conference on Data Science (SDS), IEEE. p. 80–81. URL:http: //dx.doi.org/10.1109/sds.2019.000-3, doi:10.1109/sds.2019.000-3

  41. [49]

    Are you ready to find a job ranking of a list of soft skills to enhance graduates’ employability

    Succi, C., 2019. Are you ready to find a job ranking of a list of soft skills to enhance graduates’ employability. International Journal of Human Resources Development and Management 19,

  42. [50]

    Promptengineer:Analyzinghardand softskillrequirementsintheaijobmarket

    Vu,A.,Oppenlaender,J.,2026. Promptengineer:Analyzinghardand softskillrequirementsintheaijobmarket. URL:https://arxiv.org/ abs/2506.00058,arXiv:2506.00058

  43. [51]

    Can llms replace human evaluators? an empirical study of llm-as-a-judge in software engineering

    Wang,R.,Guo,J.,Gao,C.,Fan,G.,Chong,C.Y.,Xia,X.,2025. Can llms replace human evaluators? an empirical study of llm-as-a-judge in software engineering. Proceedings of the ACM on Software En- gineering 2, 1955–1977. URL:http://dx.doi.org/10.1145/3728963, doi:10.1145/3728963

  44. [52]

    Wang, Y., Allouache, Y., Joubert, C., 2021. Analysing cv corpus for finding suitable candidates using knowledge graph and bert, in: DBKDA2021TheThirteenthInternationalConferenceonAdvances in Databases, Knowledge,and Data Applications, Valencia, Spain. URL:https://hal.science/h...

  45. [53]

    Experimentation in Software Engineering

    Wohlin, C., Runeson, P., Höst, M., Ohlsson, M.C., Regnell, B., Wesslén, A., 2024. Experimentation in Software Engineering. Springer Berlin Heidelberg. URL:http://dx.doi.org/10.1007/ 978-3-662-69306-3, doi:10.1007/978-3-662-69306-3

  46. [54]

    Automated extraction of information from polish resume documents in the it recruitment process

    Wosiak, A., 2021. Automated extraction of information from polish resume documents in the it recruitment process. Procedia Computer A. Azamnouri et al.:Preprint submitted to ElsevierPage 27 of 28 Detecting Soft Skills in ML Engineering Roles CVs Science 192, 2432–2439. URL:htt...

  47. [55]

    Large language models for generative information extraction: a survey

    Xu, D., Chen, W., Peng, W., Zhang, C., Xu, T., Zhao, X., Wu, X., Zheng, Y., Wang, Y., Chen, E., 2024. Large language models for generative information extraction: a survey. Frontiers of Computer Science 18. URL:http://dx.doi.org/10.1007/s11704-024-40555-y, doi:10.1007/s11704-0...

  48. [56]

    Zhang, M., Jensen, K., Sonniks, S., Plank, B., 2022. Skillspan: Hard and soft skill extraction from english job postings, in: Pro- ceedings of the 2022 Conference of the North American Chap- ter of the Association for Computational Linguistics: Human Lan- guage Technologies, A...

  49. [281]

    1504/ijhrdm.2019.100638

    URL:http://dx.doi.org/10.1504/ijhrdm.2019.100638, doi:10. 1504/ijhrdm.2019.100638

  50. [2019]

    EPJ Data Science 8

    Responsible team players wanted: an analysis of soft skill requirements in job advertisements. EPJ Data Science 8. URL:http://dx.doi.org/10.1140/epjds/s13688-019-0190-z, doi:10. 1140/epjds/s13688-019-0190-z

  51. [2020]

    URL:http://dx.doi.org/10.18653/v1/2020.emnlp-main.306, doi:10

    Openie6: Iterative grid labeling and coordination analy- sis for open information extraction, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP),AssociationforComputationalLinguistics.p.3748–3761. URL:http://dx.doi.org/10.18653/v...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.