{"id":"9827bb76-9d4e-4539-82e5-d2f2ed0d5276","arxiv_id":"1908.06731","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Model-assisted calibration of online job ads to official vacancy totals yields bias-corrected estimates of employer skill demand in Poland, showing large over-representation of interpersonal and managerial skills in online postings.","lead":"This paper combines a Polish official survey on job vacancies with online job ads to estimate which skills employers demand. It shows that online ads overstate social and organizational skills and understate technical and physical work, and that statistical calibration can partially correct this bias.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The bias-reduction claim rests on an untestable ignorability assumption; the paper states it in §1 but provides no sensitivity analysis showing that within-occupation differences in skill listing would not change the headline 53.8→35.1 adjustment.","rationale":"The paper is a careful application of established methods, and the reader's conditional verdict is appropriate. The most load-bearing condition for the main substantive conclusion is not the estimation machinery but the identifying assumption that online-ad selection is ignorable given occupation and NACE. The calibration estimators in §3.2–3.3 are internally coherent, but they are unbiased for the DL population only under missing-at-random given the auxiliary cells; no population benchmark for skills exists, so the \"bias reduction\" in Table 8 cannot be verified from the data alone. I agree with the reader's weakest_assumption that this is the central vulnerability. The paper is transparent about the assumption, which is a credit, but the abstract and conclusions state the result as an established reduction in representation bias rather than as a conditional estimate. A sensitivity analysis with an unobserved confounder is the standard way to check how much of the headline correction could be driven by violations of ignorability. I did not base the verdict on the secondary bootstrap-variance concern (fixed occupation-within-NACE shares), which mostly affects the standard-error comparison rather than the bias-reduction claim; that would be a partial-agreement point if it were the focus. Because the reader already conditions the verdict on addressing the ignorability issue, my stress-test does not change the verdict.","tokens_in":17366,"tokens_out":10125,"duration_ms":124307,"concrete_test":"Run a hidden-bias sensitivity analysis: introduce a binary unobserved confounder U associated with both online selection and skill presence, with odds ratios for selection and for the skill ranging over exp(0.5), exp(1.0), and exp(1.5). Re-estimate the ECLASSO1 share for interpersonal and physical skills under each combination, and compare the resulting range of estimates with the bootstrap confidence intervals around Table 8. If the plausible shifts are small relative to the observed raw-to-calibrated gap (e.g., 53.8→35.1 for interpersonal), the bias-reduction conclusion is robust; if they are comparable or larger, the claim that calibration reduces representation bias is not supported by the data alone.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim that the calibrated estimates \"reduce representation bias\" rests on an identifying assumption stated in Section 1: \"We assumed that the selection bias was ignorable given auxiliary variables.\" The calibration estimators in §3.2–3.3 reweight the Careerjet sample to match DL totals for occupation and NACE, so they can only remove bias explained by those cells. They cannot remove bias due to firms or channels whose skill-listing behaviour differs within the same occupation and NACE. Table 3 shows that job-ad source itself is strongly associated with skill listing (e.g., self-organization appears in 59.1% of Careerjet ads but only 7.6% of DEO ads), and Table 4 shows that the available auxiliaries are weak predictors for several skills (Cramer's V with occupation: mathematical 0.05, office 0.11, physical 0.17). Consequently, the large corrections in Table 8 (interpersonal from 53.8% raw to about 35% after calibration) are identified only by assumption. Without an external benchmark or a sensitivity analysis, the phrase \"reduces representation bias\" is not demonstrated; the paper's own conclusion acknowledges the assumption but does not probe its consequences.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a method to enhance Poland's Demand for Labour (DL) survey by adding skill information from online job advertisements (Careerjet.pl). Treating the online data as a non-probability sample, the authors apply traditional calibration (ECGREG) and model-assisted calibration with logistic regression, LASSO, and adaptive LASSO (ECMC, ECLASSO1, ECLASSO2, ECALASSO1), using only estimated population totals from the DL survey since unit-level data are unavailable. They introduce a bootstrap procedure that resamples the online sample and perturbs NACE totals to account for uncertainty in the reported estimates. Empirically, for 11 skill categories over 2011, 2013, and 2014, the calibrated estimates are substantially lower than raw online shares for interpersonal, managerial, and self-organization skills and higher for technical and physical skills. The paper claims that the LASSO-assisted calibration outperforms traditional calibration in terms of standard errors and reduces representation bias in online job-ad skills.","tokens_in":17638,"tokens_out":6388,"duration_ms":60038,"significance":"If the bias-reduction claim holds, this is a useful applied contribution showing how official vacancy surveys can be augmented with big data sources using recent model-assisted calibration methods. The manuscript is careful in documenting data processing, imputation, and coding quality, and it provides reproducible code and data. However, the headline result rests on an untestable ignorability assumption, and the variance estimation treats parts of the control totals as fixed; both issues limit the strength of the conclusions. As a methodological contribution the paper largely applies existing methods (Chen et al., 2018, 2019) to a new domain, so its value is primarily empirical and demonstrative.","major_comments":[{"comment":"Section 1 states 'We assumed that the selection bias was ignorable given auxiliary variables.' The calibration constraints used in Section 3.5 rely only on occupation and NACE (and, in preliminary analyses, province). Table 4 shows weak associations between these auxiliaries and several skills (e.g., Cramer's V with occupation: mathematical 0.05, office 0.11, physical 0.17), while Table 8 reports large corrections for interpersonal skills (53.8% raw vs. about 35% after calibration). Because the auxiliary variables are weak predictors for precisely some of the skills with large corrections, the claim that the calibrated estimates 'reduce representation bias' is not demonstrated without an external benchmark or a sensitivity analysis that assesses the impact of within-occupation selection. The paper's own conclusion (Section 5) acknowledges 'reduced bias in online data for several skills but not for all,' which is at odds with the abstract's unqualified claim.","section":"Section 1, Table 4, Table 8"},{"comment":"In the bootstrap procedure (Algorithm 1, step 2), a random NACE total is generated and then allocated to occupation-by-NACE cells using the fixed empirical shares \\hat T_NACE,OCCUP / \\hat T_NACE from the original point estimate. This leaves the conditional distribution of occupation within NACE fixed across bootstrap replicates, so the uncertainty in the cross-classified totals—and hence in the occupation totals used as calibration controls—is not propagated. The reported relative standard errors in Table 9 therefore likely understate the true uncertainty, and the very small RSEs for some skills (e.g., Availability 1.0% for ECMC) should be viewed cautiously. The authors should either sample the full joint distribution of occupation-NACE totals (e.g., via a multivariate normal or a survey bootstrap on the DL data) or explicitly state that the variance is conditional on the estimated cross-classification.","section":"Algorithm 1, step 2"},{"comment":"Variance estimation relies on the assumption that Q1 relative standard errors can be approximated by those published for Q4 of the same year, stated in Section 3.6: 'standard errors are similar in a given year and we can approximate standard errors from the 1st quarter based on information from the 4th quarter.' This assumption is not tested, and Table 6 provides RSEs only for NACE sections; the occupation RSEs in Table 7 are derived under the same assumption. Because the paper's headline claim that LASSO-assisted calibration outperforms traditional calibration in terms of standard errors is based on Table 9, the variance comparison inherits this untestable assumption. The authors should discuss the direction and potential magnitude of bias in the variance estimates if Q1 RSEs differ from Q4, or provide a sensitivity check using alternative assumed RSEs.","section":"Section 3.6 and Table 9"}],"minor_comments":[{"comment":"The column header 'MCGREG' is inconsistent with the estimator names (ECGREG, ECMC, ECLASSO1, ECLASSO2, ECALASSO1) defined in Section 3.5; use 'ECGREG' or clarify what is meant.","section":"Table 9"},{"comment":"The statement that 'ECMC is less efficient than estimators with LASSO' is not uniformly supported by Table 9: for Physical, ECMC has RSE 4.1 versus ECLASSO1 4.2, and for Technical, ECMC has 5.3 versus ECLASSO2 7.8. The claim should be qualified accordingly.","section":"Section 4, Table 9"},{"comment":"The sentence 'The AUC varies from 0.644 for cognitive skills to 0.829 for technical competences, which indicates that the standard LASSO model is better than the adaptive one' is confusing because Table 15 shows identical AUC values for ECLASSO1 and ECALASSO1 for every skill, and the higher AUCs belong to ECLASSO2. Rewrite to distinguish the comparison between ECLASSO1/ECALASSO1 and ECLASSO2.","section":"Section 4, Table 15"},{"comment":"The abstract claims the method 'reduces representation bias in skills observed in online job ads,' while Section 5 concludes 'reduced bias in online data for several skills but not for all.' These statements should be aligned.","section":"Abstract and Section 5"},{"comment":"The description of ECALASSO1 refers to 'adaptive LASSO regression with the seame settings as ECLASSO1' (typo: 'seame' should be 'same'), and the value used for the adaptive LASSO weight power gamma is not specified; state how gamma was chosen for reproducibility.","section":"Section 3.5"},{"comment":"There are several typographical errors: 'V oivodeship' in Table 13 header, 'deported' in Section 5, 'a attention' in Section 5, and 'masurement' in the McGuiness reference. These should be corrected.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a competent application of existing model-assisted calibration methods to a new empirical setting, with good data documentation and reproducible code. However, the central claim of bias reduction is not fully supported because the ignorability assumption is untestable with the available data and no sensitivity analysis is provided. The variance estimation has a structural issue (fixed occupation-NACE shares in the bootstrap) that affects the reported standard errors. These issues require substantive revision, not just copyediting. The paper may be suitable for an applied survey statistics or official statistics journal after the authors address these concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a competent, clearly written application of existing calibration machinery to a genuinely important problem, and it ships code and data. The advertised bias correction is the main product, but it is only as good as an ignorability assumption that the paper states and does not stress-test.\n\nWhat is actually new: the authors take the well-known idea of calibrating non-probability samples to external totals and push it into official statistics for skills demand. The empirical estimates for Poland are new, including the direction and size of the online-ad bias: interpersonal, managerial and self-organization skills are overrepresented, while technical and physical skills are underrepresented. That pattern is plausible and largely driven by occupation mix. The bootstrap treatment of estimated control totals, when only published aggregates are available, is a useful practical contribution. The authors also deserve credit for being honest in the conclusion: they say the correction reduced bias for several skills but not all, and they point to weak correlations with auxiliaries as the reason.\n\nThe soft spots are real but not fatal. The main one is the ignorability assumption stated in Section 1. Calibration to occupation, NACE and province can only remove selection bias that operates through those cells. If, within the same occupation and sector, firms that advertise online systematically describe skills differently from those that do not, the corrected estimates are still biased. The paper gives no sensitivity analysis for this. Given that Table 4 shows weak Cramer's V values for mathematical, office, and physical skills, and Table 3 shows huge differences between Careerjet and DEO ads, this is not a minor footnote. The large correction for interpersonal skills, from about 54% to 35%, needs more support than the current design provides.\n\nTwo smaller technical concerns. The bootstrap procedure draws NACE totals from a normal distribution but keeps the occupation-by-NACE shares fixed, so the reported standard errors likely understate the uncertainty in the cross-classified totals. And the paper applies Q4 standard errors to Q1 estimates because that is all Statistics Poland publishes; the authors flag this, but it is still a limitation. The heavy imputation of missing NACE, up to 57% in some waves, could also matter more than the paper suggests.\n\nOn circularity: the reader's low concern score is right. The calibrated estimates are not forced by construction; they use external vacancy totals, and the LASSO models are fitted to the online data. That is standard model-assisted estimation.\n\nWho should read this: survey statisticians and anyone at a national statistical institute thinking about using online job ads as a cheaper or richer supplement to an existing vacancy survey. It is a worthy peer-review candidate. I would send it out, not desk-reject it, and ask for a sensitivity analysis around the ignorability assumption, a less absolute claim about bias reduction, and more transparent treatment of the bootstrap's fixed shares.","headline":"Useful application of model-assisted calibration to a real official-statistics problem; the headline bias-correction claim is credible but rests on an untestable ignorability assumption that the paper does not stress-test.","tokens_in":18140,"tokens_out":2014,"would_cite":true,"duration_ms":23471,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62D05","62J07","62P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Calibration trims online job-ad skill bias from 54% to 35%","keywords":["online job advertisements","non-probability samples","model-assisted calibration","LASSO","adaptive LASSO","demand for labour survey","skills demand","representation bias"],"falsifier":"Take a probability sample of employers from the Demand for Labour survey, code the skill requirements of their actual vacancies exactly as the online ads were coded, and compare the resulting skill shares with the calibrated online estimates within occupation-by-NACE cells; a gap larger than the bootstrap standard errors would show the ignorability assumption fails.","tokens_in":17204,"feed_emoji":"📊","tokens_out":6815,"duration_ms":59822,"temperature":0.7,"pith_summary":"The paper's goal is to give the Polish Demand for Labour survey a way to report which skills employers demand, something the survey currently does not measure. The authors treat online job advertisements as a non-probability sample and correct its representation error by calibrating pseudo-weights to the survey's estimated vacancy totals, with a LASSO-assisted model for each skill. They show these calibrated estimates outperform traditional calibration in precision and produce skill shares that differ sharply from the raw online numbers: interpersonal, managerial and self-organization skills are overestimated online, while technical and physical skills are underestimated. The empirical conclusion is that raw online vacancy data cannot be taken at face value for measuring skill demand. If the approach holds, statistical agencies can enrich their labour-market statistics using big data without redesigning their probability surveys.","feed_headline":"Calibration trims online job-ad skill bias from 54% to 35%","feed_subtitle":"Raw online vacancy data overstate interpersonal skills; bias-corrected estimates put the share near 35%.","key_machinery":"The machinery is estimated-control calibration with a model-assisted twist. Pseudo-weights from the online sample are calibrated to external estimates of vacancy totals by occupation and NACE, not to known population totals, because only estimated totals are published by the statistical office; the calibration equation replaces the population total with the survey-based estimate. Because unit-level survey data are unavailable, a bootstrap procedure perturbs the estimated totals using their reported standard errors and resamples the online ads, producing variance estimates that account for both sources of uncertainty. For each skill separately, a logistic regression with a LASSO or adaptive LASSO penalty supplies model predictions that serve as calibration variables, and the LASSO-assisted weights yield smaller standard errors than traditional GREG calibration with the same auxiliary information.","core_discovery":"The central claim is that model-assisted calibration with LASSO removes a large and systematic representation bias in the skills mentioned in online job advertisements. Using two-digit occupation and NACE section as auxiliary variables, the authors build a separate logistic model for each of eleven skills, fit the model with LASSO or adaptive LASSO, and adjust the online sample's pseudo-weights to reproduce the estimated vacancy totals reported by Statistics Poland. The bias-corrected share for interpersonal skills falls from roughly 54% in raw online data to about 35%, managerial skills drop by about ten percentage points, and technical and physical skills rise from roughly 4–5% to 7–8%. The direction of the correction matches the under-representation of craft and plant-operator occupations in online vacancy data.","pith_inferences":["The calibration corrections are large enough that similar selection bias may affect commercial online vacancy databases used in other countries, so their skill-demand findings should be re-examined for the same occupation-mix distortion.","With automated occupation and sector coding, the same pipeline could be run on continuously scraped job ads to produce near-real-time indicators of skill demand, updating the official survey between waves.","A direct validation—code skills in a small probability sample of DL-survey vacancies and compare with the calibrated online estimates within the same occupation-by-NACE cells—would test the ignorability assumption the whole method rests on.","If additional auxiliary totals (such as firm size or ownership sector) become available, the residual bias for skills weakly correlated with occupation, like office and physical skills, could be reduced further."],"forward_implications":["National statistical institutes can add a skill-demand dimension to their vacancy surveys without adding questions, by combining existing totals with online ad text and the calibration machinery.","Analyses of skill demand based only on raw online job postings will overstate interpersonal, managerial and computer skills and understate physical and technical skills, at least in labour markets with the same occupation mix as Poland's.","LASSO-assisted calibration is the preferred estimator when many auxiliary categories are available, since it produced lower relative standard errors than traditional calibration in this application.","The bootstrap variance approach provides a template for propagating uncertainty about estimated control totals when micro-data from the reference survey cannot be released."],"supporting_citations":[{"why":"Develops calibration of non-probability surveys to estimated control totals using LASSO and supplies the estimation approach the paper adopts.","marker":"Chen et al. (2019)"},{"why":"Establishes model-assisted inference for non-probability samples, the basis for the LASSO-assisted estimator.","marker":"Chen et al. (2018)"},{"why":"Introduces calibration estimation in survey sampling, the foundation of the pseudo-weight adjustment.","marker":"Deville and Särndal (1992)"},{"why":"Introduces model-calibration using auxiliary information, the mechanism behind the model-assisted weights.","marker":"Wu and Sitter (2001)"},{"why":"Extends calibration to the case where control totals are estimated rather than known, which the paper needs because only estimated DL totals exist.","marker":"Dever and Valliant (2010)"},{"why":"Introduces LASSO, the penalty used to select and shrink auxiliary variables in the assisting models.","marker":"Tibshirani (1996)"},{"why":"Introduces adaptive LASSO, the variant compared against standard LASSO in the paper.","marker":"Zou (2006)"},{"why":"Source of the estimated vacancy totals and standard errors that serve as calibration controls.","marker":"Statistics Poland (2018)"}],"fun_headline_variants":["LASSO calibration corrects skill bias in online job ads","Calibration cuts job-ad interpersonal skills from 54% to 35%","Job-ad skill bias shrinks with LASSO-assisted calibration","Model-assisted calibration trims bias in online vacancy data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that, once occupation, sector and province are held constant, vacancies advertised online ask for the same skills as those posted offline; if the two channels attract employers with systematically different skill descriptions, every bias-corrected estimate inherits that difference.","fun_headline_variants_meta":{"raw":{"variants":["LASSO calibration corrects skill bias in online job ads","Calibration cuts job-ad interpersonal skills from 54% to 35%","Job-ad skill bias shrinks with LASSO-assisted calibration","Model-assisted calibration trims bias in online vacancy data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000737,"raw_usage":{"total_tokens":3264,"prompt_tokens":888,"completion_tokens":2376,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":504,"completion_tokens_details":{"reasoning_tokens":2306}},"tokens_in":504,"tokens_out":2376,"duration_ms":21106,"temperature":1.0,"reasoning_tokens":2306,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:26:49.491545+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a probability sample of employers from the Demand for Labour survey, code the skill requirements of their actual vacancies exactly as the online ads were coded, and compare the resulting skill shares with the calibrated online estimates within occupation-by-NACE cells; a gap larger than the bootstrap standard errors would show the ignorability assumption fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Develops calibration of non-probability surveys to estimated control totals using LASSO and supplies the estimation approach the paper adopts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes model-assisted inference for non-probability samples, the basis for the LASSO-assisted estimator."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends calibration to the case where control totals are estimated rather than known, which the paper needs because only estimated DL totals exist."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces LASSO, the penalty used to select and shrink auxiliary variables in the assisting models."},{"cited_title":"The demand for labour in 2017","cited_arxiv_id":null,"evidence_quote":"Source of the estimated vacancy totals and standard errors that serve as calibration controls."}],"review_version":1}