{"id":"f67a374a-7267-4866-a950-546fb8966a0f","arxiv_id":"2608.10046","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In 300 public CVs, soft skills appear in narrative prose about three times as often as in keyword lists, leadership mentions nearly triple with seniority, and software engineers articulate leadership at half the rate of the other roles.","lead":"An analysis of 300 public CVs shows that ML engineers, data scientists, and software engineers mostly communicate soft skills through career stories rather than keyword lists, and that leadership language roughly triples with seniority. It is the first candidate-side test of how these roles present collaboration and leadership, and it suggests keyword-based resume screening misses most of the evidence.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 3:1 narrative-vs-keyword claim rests on an unvalidated asymmetry between the two extraction stages: the constrained skills-section stage (E) may have lower recall than the full-text stage (T), inflating I = T − E.","rationale":"The strongest practical claim — the roughly three-to-one narrative dominance of soft-skill disclosure — is the load-bearing conclusion because it licenses the screening implication that keyword-based systems systematically miss technical candidates' soft skills. The seniority (H2a: OR 2.88, adjusted OR 3.06) and role-signature (H1a, H1d, H1e) findings are computed from the total extraction T and would survive even if the explicit/implicit split were wrong; they are additionally supported by the continuous-experience robustness check and the quantitative bias analysis. The narrative dominance, by contrast, is computed by subtracting a separately prompted, separately scoped extraction output from the full-text output (Section 3.5), and the paper provides no validation of the asymmetry between those two stages. The only validation is an aggregate F1 on an evaluation set drawn from the same corpus and used for 22 prompt-tuning trials (Section 4); no statistic reports the constrained stage's recall on true explicit mentions, and the paper explicitly acknowledges the absence of a formal inter-coder reliability coefficient. The failure mode is concrete: if the skills-section locator misses nonstandard section headings such as 'Areas of Expertise' or 'Core Competencies', E shrinks, T retains those mentions, and they are reclassified as implicit, mechanically producing exactly the 3:1 ratio and the high narrative shares reported in Table 6. The Section 5.5 argument that shared-model bias 'largely cancels' is an assertion, since the two prompts and input spans differ and bias need not cancel. This is the same weakest assumption the reader identified, and the CONDITIONAL verdict remains appropriate: the seniority and role findings are likely robust, but the headline narrative-disclosure claim should not be treated as established until the two-stage asymmetry is validated on a held-out set with human-annotated explicit mentions.","tokens_in":38330,"tokens_out":6141,"duration_ms":57118,"concrete_test":"Take a held-out sample of 50–100 CVs from the same corpus, preferably not used during the 22 prompt-tuning trials. Have two annotators independently mark the set of soft-skill mentions that appear in dedicated skills sections (E_true), resolving disagreements by the paper's existing consensus rules and reporting inter-coder agreement. Run both pipeline stages on the same sample: (a) the constrained skills-section extraction producing E, and (b) the full-text extraction producing T. Compute recall of E on E_true and recall of T on E_true. If recall(E) is substantially below recall(T) on E_true, the implicit set I = T − E is biased upward; recompute the implicit/explicit ratio and Table 6 narrative shares using a corrected explicit set (or directly annotate implicit narrative mentions).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central practical claim — that candidates disclose soft skills through narrative rather than keyword lists (1,007 implicit vs. 347 explicit skill–CV pairs) and that keyword-based screening therefore systematically misses them — depends on the I = T − E construction in Section 3.5. The paper validates only an aggregate pipeline F1 of 0.72 on 100 CVs drawn from the same corpus and tuned through 22 trials (Section 4); it never separately measures recall of the constrained skills-section stage (E) or the full-text stage (T). If E under-detects explicitly listed skills — because the skills-section locator misses nonstandard headings, or because the constrained prompt is more conservative than the full-text prompt — those explicitly listed skills remain in T and are automatically counted as implicit, mechanically inflating the narrative ratio. The Section 5.5 assertion that systematic extraction bias 'largely cancels' because both branches use the same model and taxonomy is not supported: the two stages use different prompts and different input spans, so their recall for the same underlying explicit skills can differ systematically. Moreover, the quantitative bias analysis in Section 5.5 corrects prevalence estimates, not the E/I partition, so H3a and Table 6's narrative shares (e.g., leadership 87.8%, coaching 100%) are not protected by that sensitivity analysis. The paper itself states that no formal inter-coder reliability statistic was computed (Section 4), further weakening confidence in the gold standard used to validate the stages. If E's recall is materially below T's recall on true explicit mentions, the 3:1 ratio and the 88–96% narrative shares for leadership, coordination, and mentoring could be artifacts of the extraction design rather than properties of the CVs.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how soft skills are articulated in 300 curated CVs from ML engineers, data scientists, and software engineers. The authors build an LLM-based extraction pipeline that distinguishes explicitly listed soft skills (in dedicated skills sections, E) from implicitly narrated ones (I = T − E), validate the pipeline against a human-annotated ground truth (F1 = 0.72), and test 13 hypotheses derived from the demand-side literature with effect sizes, confidence intervals, Holm–Bonferroni correction, and TOST equivalence testing. Eleven hypotheses are supported, one partially, one refuted. Key findings are that narrative disclosure outweighs keyword disclosure by about 1,007 to 347 unique skill–CV pairs; leadership, coordination, mentoring, and coaching are 88–100% narrative; seniority nearly triples the odds of articulating leadership (adjusted OR 3.06); and leadership articulation is not role-invariant, with software engineers at roughly half the rate of the other two roles.","tokens_in":38600,"tokens_out":4562,"duration_ms":45036,"significance":"If the measurement assumptions hold, this is a valuable and rare candidate-side complement to the job-posting and interview literature on soft skills. The study is unusually careful on the statistical side: it reports effect sizes with confidence intervals, controls family-wise error, uses TOST for the invariance hypothesis, and provides sensitivity analyses for seniority dichotomization, document length, and non-differential misclassification. The hypotheses are derived from external literature and several could have failed, which strengthens the confirmatory interpretation. The paper also ships a replication package with prompts, extraction outputs, and analysis scripts, which supports reproducibility. However, the headline practical claim — that keyword-based screening systematically misses soft skills because candidates write them narratively — rests on the unvalidated asymmetry between the constrained skills-section extraction (E) and full-text extraction (T). Because the paper does not separately measure the recall of these two stages, the central 3:1 narrative-to-keyword ratio could be an artifact of the measurement design.","major_comments":[{"comment":"The implicit set I = T − E is the load-bearing construction for H3a, the 1,007 vs. 347 narrative-to-keyword ratio, and the narrative shares in Table 6, but the paper validates only the aggregate pipeline F1 of 0.72 on 100 CVs (Section 4) and never reports recall for the constrained skills-section stage (E) and the full-text stage (T) separately. If the E stage under-detects explicitly listed skills — because the skills-section locator misses nonstandard headings or because the constrained prompt is more conservative than the full-text prompt — those explicitly listed skills remain in T and are automatically counted as implicit, mechanically inflating the narrative ratio. The Section 5.5 assertion that systematic extraction bias “largely cancels” because both branches use the same model and taxonomy is not supported: the two stages use different prompts and different input spans, so their recall for the same underlying explicit skills can differ systematically. The quantitative bias analysis in Section 5.5 corrects prevalence estimates, not the E/I partition, so H3a and Table 6 are not protected by that sensitivity analysis. I recommend that the revision either (a) validate stage-specific recall and precision on the human-annotated ground truth with explicit/implicit labels, or (b) re-estimate H3a under plausible E-recall scenarios and report how the narrative ratio changes. If the ratio is sensitive, the screening-pipeline claim in Sections 6.5 and 8 should be softened accordingly.","section":"§3.5, §4, §5.5"},{"comment":"The purposeful selection of 62 data scientist and 53 software engineer CVs from the GitHub corpus based on “text completeness” is a selection-on-outcome risk for the role comparisons, particularly H1c (no-skill rate) and the skill-density results in Section 5.1. Data scientists selected for CVs that contain the full set of expected sections (summary, experience, skills, education) could mechanically have lower no-skill rates than ML engineers, whose corpus subset was not selected on that criterion. The paper should provide a sensitivity analysis using a random or source-representative sample from the corpus, or otherwise quantify how much of the H1c contrast (4.0% vs. 18.0% no-skill rate) could be attributable to this selection rule.","section":"§3.2, §5.1"},{"comment":"The extraction validation rests on a ground truth whose reliability is not quantified: the paper states that no formal inter-coder reliability statistic was computed (Section 4), and the 100-CV evaluation set is drawn from the same corpus on which the pipeline was tuned through 22 trials. This is especially limiting for the explicit-versus-implicit distinction, because the E stage and T stage are never separately validated and the two annotators resolved disagreements through joint calibration. Reporting Cohen’s kappa or a comparable agreement measure, ideally stratified by explicit vs. implicit mentions, would materially strengthen the claim that the gold standard supports the fine-grained disclosure-style analysis. At minimum, the absence of this statistic should be acknowledged as a limitation that affects H3a directly.","section":"§4"}],"minor_comments":[{"comment":"Several passages are missing spaces between words, e.g., “Weclosebothgaps” in the abstract, “disclosesoftskillsthroughnarrativeratherthankeywordlists”, and “Figure 1:An overview of the study process”. A full proofreading pass is needed.","section":"Abstract and throughout"},{"comment":"The sentence “only 3.3% of candidates (10/300) relied exclusively on explicit keyword lists” is clear, but the later phrase “roughly three-quarters (74.4%) of the soft skill evidence” could be more carefully tied to the assumption that E and T have comparable recall; otherwise it risks overstating what the extraction pipeline established.","section":"§5.3"},{"comment":"The description of the sensitivity of the design is helpful, but the phrase “the smallest differences it can reliably detect” would be clearer with an explicit statement that these are minimum detectable differences under the stated baseline prevalence and power, not a guarantee about all competencies.","section":"§3.6"},{"comment":"The table caption says “Keyword” and “Narrative” partition the CVs in which the competency was detected, but the column sums can exceed the total prevalence if a CV contains both channels; consider adding a “Both” column or clarifying that the two columns are not mutually exclusive at the CV level.","section":"Table 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong candidate-side empirical study with careful statistics, but the central practical claim about narrative-vs-keyword disclosure depends on an unvalidated asymmetry between the E and T extraction stages. The authors should be asked to validate the explicit/implicit partition directly or provide a sensitivity analysis; the corpus selection on text completeness is a second issue that needs addressing. With those fixed, the paper could be acceptable. No concerns about novelty or scope; the topic fits the journal well."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first real candidate-side look at soft-skill disclosure in ML-adjacent CVs, and the seniority and role-signature results are likely solid. But the headline 3:1 narrative-over-keyword ratio rests on an explicit/implicit partition that hasn't been validated, and that is the one thing standing between this paper and a confident accept.\n\nWhat the paper does well: it converts demand-side findings into 13 falsifiable hypotheses, sets a confirmatory family with Holm control, reports effect sizes and CIs, uses TOST for the predicted-null H1e, and runs three sensitivity analyses including a quantitative bias analysis that shows the main seniority effect would need implausible differential extraction error to reverse. The refutation of leadership universality (software engineers at half the rate) is a genuinely interesting result, and it survives adjustment for seniority. The replication package with per-CV outcome matrix lets anyone re-run the stats without re-running the LLM. Good transparency.\n\nThe soft spot is in the measurement of RQ3. Explicit mentions E are extracted only from section headings detected by the constrained stage, then implicit I = T - E. The paper reports an aggregate F1=0.72 but never measures the recall of the two stages separately. If the skills-section stage is more conservative than the full-text stage, true explicit skills get counted as implicit and the 3:1 ratio inflates. The Section 5.5 claim that bias 'largely cancels' because both branches share the same model and taxonomy doesn't cover this: the prompts and input spans differ. The bias analysis corrects prevalence estimates, not the E/I partition. Given no formal inter-coder reliability on the gold standard either, the narrative-dominance claim and Table 6's 88-100% narrative shares are not yet established.\n\nNone of this threatens the seniority and role findings, which are based on total detection and are much more robust. It also doesn't threaten the paper's useful practical point that technical candidates do surface soft skills. But the 'keyword screening misses three-quarters' claim needs a validation of the split on held-out data, or per-stage recall numbers.\n\nWho is this for: HR/ATS researchers, CS educators, and anyone studying skill signaling. It deserves peer review, but as a major revision, not as-is.","headline":"First real candidate-side evidence on soft-skill disclosure in ML-adjacent CVs, with careful inferential statistics; the role and seniority findings are likely solid, but the headline 3:1 narrative-over-keyword ratio rests on an unvalidated explicit/implicit measurement split that needs a major revision.","tokens_in":39189,"tokens_out":2554,"would_cite":true,"duration_ms":24103,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Soft skills in technical CVs appear about three times as often in narrative prose as in keyword lists, so keyword-based screening systematically misses them.","keywords":["soft skills","CV analysis","ML engineers","data scientists","software engineers","LLM extraction","narrative disclosure","keyword screening"],"falsifier":"Manually annotate a fresh held-out set of CVs for narrative versus keyword disclosure, or measure the pipeline's sensitivity separately on skills sections and on work-experience sections; if human raters find narrative and keyword mentions near equal, or if the pipeline's sensitivity differs sharply between the two sections, the 1,007-to-347 dominance and the claim that keyword screening misses most soft skills would be overturned.","tokens_in":38122,"feed_emoji":"📄","tokens_out":6800,"duration_ms":56658,"temperature":0.7,"pith_summary":"This paper studies how ML engineers, data scientists, and software engineers articulate soft skills in their CVs, using a balanced corpus of 300 documents and an LLM-based extractor that distinguishes explicitly listed from implicitly narrated mentions. It converts the demand-side literature's claims about these roles into 13 falsifiable hypotheses and tests them with effect sizes and family-wise error control. The central finding is that candidates disclose soft skills through narrative rather than keyword lists by roughly three to one, and most so for the competencies employers value most: leadership, coordination, and mentoring are 88–96% narrative. Seniority nearly triples the odds of articulating leadership, and leadership is not role-invariant, as software engineers mention it at roughly half the rate of their peers. The practical conclusion is that a keyword-based screening pipeline systematically under-detects the soft-skill evidence that differentiates candidates.","feed_headline":"Keyword screening misses 74% of soft-skill evidence","feed_subtitle":"A 300-CV study finds leadership, mentoring, and coordination are narrated, not listed, in tech CVs.","key_machinery":"The central object is a two-stage LLM-based extraction pipeline that separates explicit from implicit soft-skill mentions. Stage one extracts skills only from dedicated skills sections (set $E$); stage two extracts all skills from the full text (set $T$); implicit narrative mentions are defined deductively as $I = T - E$. Skills are mapped to the TaxoSoft taxonomy, and each extraction is tied to a supporting phrase from the CV, allowing verification. This design makes explicit and implicit disclosure comparable within the same CV as a paired observation, which is what the disclosure-style hypotheses require.","core_discovery":"The paper claims that the candidate side of the soft-skills picture is measurable and that it partly corroborates, partly corrects, the demand side. Candidates articulate soft skills predominantly through narrative: 1,007 of 1,354 unique skill–CV pairs are implicit, versus 347 explicit, and 94.7% of CVs whose two channels disagree disclose only through narrative. Narrative dominance is concentrated in exactly the competencies that discriminate between candidates—coaching, coordination, mentoring, presentation, and leadership are 88–100% narrative—while generic labels like communication and analytical ability are as often keyword-listed. Seniority roughly triples the odds of articulating leadership (adjusted OR 3.06), and the effect is homogeneous across roles; collaboration accumulates alongside leadership instead of being displaced. The prediction that leadership would be role-invariant is refuted: software engineers articulate it at 26%, versus 47% for data scientists and 42% for ML engineers. Eleven of 13 hypotheses are supported, one partially, and one refuted.","pith_inferences":["If the narrative-dominance pattern replicates in other corpora, the soft-skills gap is partly an articulation gap: competencies are present in candidates' own words but absent from the vocabularies of machine screening, which reframes the problem as teachable CV-writing skill rather than only competency development.","The leadership gap for software engineers could be partly a vocabulary effect if those candidates use titles like 'tech lead' or describe mentoring instead of 'leadership'; a follow-up expanding the taxonomy's variants could test whether the 26% figure is a true role difference or a labeling artifact.","The two-stage subtraction design could be validated directly by collecting human annotations of implicit mentions on held-out CVs; if it holds, the same machinery transfers to LinkedIn profiles or cover letters to test whether narrative dominance is a property of the document type.","A corpus stratified by collection period within each role, as the paper itself suggests, could separate period effects in CV-writing conventions from genuine role and seniority signals."],"forward_implications":["Keyword-based applicant tracking systems will under-detect soft skills, missing roughly 74% of the unique skill–CV pairs in this corpus.","A keyword-only extractor cannot detect the seniority effect on leadership, the strongest signal in the study, because only 14 of 115 leadership disclosures appear as explicit labels.","Screening results will be systematically biased toward generic competencies (communication, analytical ability) and away from evidence-bearing ones (leadership, coordination, mentoring), distorting candidate ranking.","Role-specific guidance follows: data scientists' communication emphasis and ML engineers' mentoring emphasis are candidate-side signatures that role-tailored screening or coaching could exploit.","The refutation of leadership universality means software engineering candidates who lead without saying so are less visible than equally leading peers in other roles."],"supporting_citations":[{"why":"Supplies the foundational multi-labeled CV corpus, including full texts and expert occupation tags that the balanced dataset builds on.","marker":"[24]"},{"why":"Provides the TaxoSoft taxonomy, the canonical soft-skill labels and alternative variants used for extraction, annotation, and standardization.","marker":"[25]"},{"why":"Reports junior-to-senior AI skill requirements from job advertisements and grounds the seniority hypotheses H2a–H2c.","marker":"[40]"},{"why":"Describes collaboration challenges in ML-enabled systems and grounds several role-signature hypotheses, including the translational data-scientist role.","marker":"[37]"},{"why":"Analyzes soft-skill requirements in software job advertisements and grounds the communication hypothesis H1a and the leadership-universality prediction H1e.","marker":"[14]"},{"why":"Interview study of software practitioners that grounds the leadership-universality prediction and the disclosure-style expectations behind H3a.","marker":"[34]"},{"why":"Hiring-manager study of entry-level software candidates that grounds the no-skill-rate hypothesis H2d and the disclosure-style role predictions of H3b.","marker":"[30]"},{"why":"Represents the implicit-skills-extraction and keyword-matching screening approach that H3a argues against.","marker":"[18]"},{"why":"Surveys LLM-based generative information extraction and motivates the choice of generative extraction over dictionary matching for narrative content.","marker":"[55]"}],"fun_headline_variants":["Narrative CVs: 74% of soft skills invisible to keyword screening","Leadership on CVs: told, not listed — and keyword screens miss it","Software engineers articulate leadership at half the data-scientist rate","Soft skills in tech CVs: 3-to-1 narrative over keyword, yet screens miss them"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole narrative-dominance result depends on the assumption that subtracting skills-section mentions from total mentions ($I = T - E$) correctly measures narrative disclosure, and that the LLM does not systematically over-detect narrative language while under-detecting skills sections; if that asymmetry is a pipeline artifact rather than a property of the CVs, the 3-to-1 ratio and the per-competency narrative shares could be wrong.","fun_headline_variants_meta":{"raw":{"variants":["Narrative CVs: 74% of soft skills invisible to keyword screening","Leadership on CVs: told, not listed — and keyword screens miss it","Software engineers articulate leadership at half the data-scientist rate","Soft skills in tech CVs: 3-to-1 narrative over keyword, yet screens miss them"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001088,"raw_usage":{"total_tokens":4590,"prompt_tokens":1034,"completion_tokens":3556,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":3471}},"tokens_in":650,"tokens_out":3556,"duration_ms":24489,"temperature":1.0,"reasoning_tokens":3471,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:14:16.585905+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Manually annotate a fresh held-out set of CVs for narrative versus keyword disclosure, or measure the pipeline's sensitivity separately on skills sections and on work-experience sections; if human raters find narrative and keyword mentions near equal, or if the pipeline's sensitivity differs sharply between the two sections, the 1,007-to-347 dominance and the claim that keyword screening misses most soft skills would be overturned.","supporting_citations":[{"cited_title":"Building a soft skill taxonomy from job openings","cited_arxiv_id":null,"evidence_quote":"Provides the TaxoSoft taxonomy, the canonical soft-skill labels and alternative variants used for extraction, annotation, and standardization."},{"cited_title":"From junior to senior: Skill require- ments for ai professionals across career stages","cited_arxiv_id":null,"evidence_quote":"Reports junior-to-senior AI skill requirements from job advertisements and grounds the seniority hypotheses H2a–H2c."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Interview study of software practitioners that grounds the leadership-universality prediction and the disclosure-style expectations behind H3a."}],"review_version":1}