Pith. sign in

REVIEW 3 major objections 6 minor 54 references

Enhancing the Demand for Labour survey by including skills from online job advertisements using model-assisted calibration

T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Calibration trims online job-ad skill bias from 54% to 35%

desk verdict Useful application of model-assisted calibration to a real official-statistics problem; the headline bias-correction claim is credible but rests on an untestable ignorability assumption that the paper does not stress-test. read the letter →

arxiv 1908.06731 v1 pith:MKF3R5IW submitted 2019-08-08 econ.GN q-fin.ECstat.AP

classification econ.GNq-fin.ECstat.AP MSC 62D0562J0762P20
keywords onlinejobadvertisementsnon-probabilitysamplesmodel-assistedcalibrationLASSOadaptivedemandforlaboursurveyskillsrepresentationbias
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper's goal is to give the Polish Demand for Labour survey a way to report which skills employers demand, something the survey currently does not measure. The authors treat online job advertisements as a non-probability sample and correct its representation error by calibrating pseudo-weights to the survey's estimated vacancy totals, with a LASSO-assisted model for each skill. They show these calibrated estimates outperform traditional calibration in precision and produce skill shares that differ sharply from the raw online numbers: interpersonal, managerial and self-organization skills are overestimated online, while technical and physical skills are underestimated. The empirical conclusion is that raw online vacancy data cannot be taken at face value for measuring skill demand. If the approach holds, statistical agencies can enrich their labour-market statistics using big data without redesigning their probability surveys.

What carries the argument

The machinery is estimated-control calibration with a model-assisted twist. Pseudo-weights from the online sample are calibrated to external estimates of vacancy totals by occupation and NACE, not to known population totals, because only estimated totals are published by the statistical office; the calibration equation replaces the population total with the survey-based estimate. Because unit-level survey data are unavailable, a bootstrap procedure perturbs the estimated totals using their reported standard errors and resamples the online ads, producing variance estimates that account for both sources of uncertainty. For each skill separately, a logistic regression with a LASSO or adaptive LASSO penalty supplies model predictions that serve as calibration variables, and the LASSO-assisted weights yield smaller standard errors than traditional GREG calibration with the same auxiliary information.

What would settle it

Take a probability sample of employers from the Demand for Labour survey, code the skill requirements of their actual vacancies exactly as the online ads were coded, and compare the resulting skill shares with the calibrated online estimates within occupation-by-NACE cells; a gap larger than the bootstrap standard errors would show the ignorability assumption fails.

Watch

Extended reading notes

Core claim

The central claim is that model-assisted calibration with LASSO removes a large and systematic representation bias in the skills mentioned in online job advertisements. Using two-digit occupation and NACE section as auxiliary variables, the authors build a separate logistic model for each of eleven skills, fit the model with LASSO or adaptive LASSO, and adjust the online sample's pseudo-weights to reproduce the estimated vacancy totals reported by Statistics Poland. The bias-corrected share for interpersonal skills falls from roughly 54% in raw online data to about 35%, managerial skills drop by about ten percentage points, and technical and physical skills rise from roughly 4–5% to 7–8%. The direction of the correction matches the under-representation of craft and plant-operator occupations in online vacancy data.

Load-bearing premise

The load-bearing premise is that, once occupation, sector and province are held constant, vacancies advertised online ask for the same skills as those posted offline; if the two channels attract employers with systematically different skill descriptions, every bias-corrected estimate inherits that difference.

Editorial extensions

If this is right

  • National statistical institutes can add a skill-demand dimension to their vacancy surveys without adding questions, by combining existing totals with online ad text and the calibration machinery.
  • Analyses of skill demand based only on raw online job postings will overstate interpersonal, managerial and computer skills and understate physical and technical skills, at least in labour markets with the same occupation mix as Poland's.
  • LASSO-assisted calibration is the preferred estimator when many auxiliary categories are available, since it produced lower relative standard errors than traditional calibration in this application.
  • The bootstrap variance approach provides a template for propagating uncertainty about estimated control totals when micro-data from the reference survey cannot be released.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The calibration corrections are large enough that similar selection bias may affect commercial online vacancy databases used in other countries, so their skill-demand findings should be re-examined for the same occupation-mix distortion.
  • With automated occupation and sector coding, the same pipeline could be run on continuously scraped job ads to produce near-real-time indicators of skill demand, updating the official survey between waves.
  • A direct validation—code skills in a small probability sample of DL-survey vacancies and compare with the calibrated online estimates within the same occupation-by-NACE cells—would test the ignorability assumption the whole method rests on.
  • If additional auxiliary totals (such as firm size or ownership sector) become available, the residual bias for skills weakly correlated with occupation, like office and physical skills, could be reduced further.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes a method to enhance Poland's Demand for Labour (DL) survey by adding skill information from online job advertisements (Careerjet.pl). Treating the online data as a non-probability sample, the authors apply traditional calibration (ECGREG) and model-assisted calibration with logistic regression, LASSO, and adaptive LASSO (ECMC, ECLASSO1, ECLASSO2, ECALASSO1), using only estimated population totals from the DL survey since unit-level data are unavailable. They introduce a bootstrap procedure that resamples the online sample and perturbs NACE totals to account for uncertainty in the reported estimates. Empirically, for 11 skill categories over 2011, 2013, and 2014, the calibrated estimates are substantially lower than raw online shares for interpersonal, managerial, and self-organization skills and higher for technical and physical skills. The paper claims that the LASSO-assisted calibration outperforms traditional calibration in terms of standard errors and reduces representation bias in online job-ad skills.

Significance. If the bias-reduction claim holds, this is a useful applied contribution showing how official vacancy surveys can be augmented with big data sources using recent model-assisted calibration methods. The manuscript is careful in documenting data processing, imputation, and coding quality, and it provides reproducible code and data. However, the headline result rests on an untestable ignorability assumption, and the variance estimation treats parts of the control totals as fixed; both issues limit the strength of the conclusions. As a methodological contribution the paper largely applies existing methods (Chen et al., 2018, 2019) to a new domain, so its value is primarily empirical and demonstrative.

major comments (3)
  1. [Section 1, Table 4, Table 8] Section 1 states 'We assumed that the selection bias was ignorable given auxiliary variables.' The calibration constraints used in Section 3.5 rely only on occupation and NACE (and, in preliminary analyses, province). Table 4 shows weak associations between these auxiliaries and several skills (e.g., Cramer's V with occupation: mathematical 0.05, office 0.11, physical 0.17), while Table 8 reports large corrections for interpersonal skills (53.8% raw vs. about 35% after calibration). Because the auxiliary variables are weak predictors for precisely some of the skills with large corrections, the claim that the calibrated estimates 'reduce representation bias' is not demonstrated without an external benchmark or a sensitivity analysis that assesses the impact of within-occupation selection. The paper's own conclusion (Section 5) acknowledges 'reduced bias in online data for several skills but not for all,' which is at odds with the abstract's unqualified claim.
  2. [Algorithm 1, step 2] In the bootstrap procedure (Algorithm 1, step 2), a random NACE total is generated and then allocated to occupation-by-NACE cells using the fixed empirical shares \hat T_NACE,OCCUP / \hat T_NACE from the original point estimate. This leaves the conditional distribution of occupation within NACE fixed across bootstrap replicates, so the uncertainty in the cross-classified totals—and hence in the occupation totals used as calibration controls—is not propagated. The reported relative standard errors in Table 9 therefore likely understate the true uncertainty, and the very small RSEs for some skills (e.g., Availability 1.0% for ECMC) should be viewed cautiously. The authors should either sample the full joint distribution of occupation-NACE totals (e.g., via a multivariate normal or a survey bootstrap on the DL data) or explicitly state that the variance is conditional on the estimated cross-classification.
  3. [Section 3.6 and Table 9] Variance estimation relies on the assumption that Q1 relative standard errors can be approximated by those published for Q4 of the same year, stated in Section 3.6: 'standard errors are similar in a given year and we can approximate standard errors from the 1st quarter based on information from the 4th quarter.' This assumption is not tested, and Table 6 provides RSEs only for NACE sections; the occupation RSEs in Table 7 are derived under the same assumption. Because the paper's headline claim that LASSO-assisted calibration outperforms traditional calibration in terms of standard errors is based on Table 9, the variance comparison inherits this untestable assumption. The authors should discuss the direction and potential magnitude of bias in the variance estimates if Q1 RSEs differ from Q4, or provide a sensitivity check using alternative assumed RSEs.
minor comments (6)
  1. [Table 9] The column header 'MCGREG' is inconsistent with the estimator names (ECGREG, ECMC, ECLASSO1, ECLASSO2, ECALASSO1) defined in Section 3.5; use 'ECGREG' or clarify what is meant.
  2. [Section 4, Table 9] The statement that 'ECMC is less efficient than estimators with LASSO' is not uniformly supported by Table 9: for Physical, ECMC has RSE 4.1 versus ECLASSO1 4.2, and for Technical, ECMC has 5.3 versus ECLASSO2 7.8. The claim should be qualified accordingly.
  3. [Section 4, Table 15] The sentence 'The AUC varies from 0.644 for cognitive skills to 0.829 for technical competences, which indicates that the standard LASSO model is better than the adaptive one' is confusing because Table 15 shows identical AUC values for ECLASSO1 and ECALASSO1 for every skill, and the higher AUCs belong to ECLASSO2. Rewrite to distinguish the comparison between ECLASSO1/ECALASSO1 and ECLASSO2.
  4. [Abstract and Section 5] The abstract claims the method 'reduces representation bias in skills observed in online job ads,' while Section 5 concludes 'reduced bias in online data for several skills but not for all.' These statements should be aligned.
  5. [Section 3.5] The description of ECALASSO1 refers to 'adaptive LASSO regression with the seame settings as ECLASSO1' (typo: 'seame' should be 'same'), and the value used for the adaptive LASSO weight power gamma is not specified; state how gamma was chosen for reproducibility.
  6. [Throughout] There are several typographical errors: 'V oivodeship' in Table 13 header, 'deported' in Section 5, 'a attention' in Section 5, and 'masurement' in the McGuiness reference. These should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the calibrated skill estimates are anchored to independent Statistics Poland vacancy totals; the ignorability assumption is an identification condition, not a circular step.

full rationale

The paper's derivation chain is not circular. The target quantities—shares of online job advertisements containing each of 11 skills—are measured only in the Careerjet/DEO online sample (Section 2.2), while the calibration constraints in Eqs. (3)–(10) use estimated vacancy totals by occupation, NACE section, and province from the Demand for Labour survey (Section 2.1; Tables 1, 6), an independent probability sample conducted by Statistics Poland. The LASSO and adaptive LASSO models in Eq. (11) are fitted to the online data, but the resulting model-assisted weights are then calibrated to those external totals, so the headline estimates in Table 8 are not equal to fitted parameters by construction. The central identifying assumption—'We assumed that the selection bias was ignorable given auxiliary variables'—is an untestable identification condition, and the weak Cramer's V values in Table 4 and the strong source differences in Table 3 are legitimate threats to the bias-reduction interpretation. Those concerns belong to correctness or validity risk, not circularity, because an identifying assumption is not an input that is renamed as an output. The self-citations (Beresewicz 2017; Beresewicz et al. 2018; Pater et al. 2019) are used for data provenance and literature context, not as the load-bearing justification for the empirical estimates. The paper is self-contained against an external benchmark: the DL survey totals. No equation in the paper reduces the reported skill-demand estimates to the online sample's raw skill shares or to the fitted LASSO coefficients themselves.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new entities or physical quantities. Its central estimates depend on a chain of statistical assumptions: ignorability of selection, the job-ad-to-vacancy mapping, approximate standard errors for control totals, and a heuristic bootstrap. None of these are independently validated, so the empirical conclusions inherit uncertainty from each.

free parameters (1)
  • Adaptive LASSO weight power gamma = not reported
    The adaptive LASSO model in equation (11) includes a tuning parameter gamma that is set by the researcher. The paper does not report the chosen value, though results are nearly identical to standard LASSO, indicating low sensitivity.
assumptions (5)
  • domain assumption Selection into online job ads is independent of skill requirements given occupation, NACE and province.
    Stated in Section 1: 'We assumed that the selection bias was ignorable given auxiliary variables.' This is the key identifying assumption for bias correction.
  • domain assumption A job advertisement reflects a vacancy as defined by the DL survey.
    Section 2.2.1: 'we assumed that information included in the job ad can be taken to reflect the job vacancy.' The DL survey defines vacancies with specific conditions that online ads may not meet.
  • ad hoc to paper Relative standard errors for Q1 equal those published for Q4 of the same year.
    Section 3.6: 'standard errors are similar in a given year and we can approximate standard errors from the 1st quarter based on information from the 4th quarter.' This is acknowledged as unverifiable.
  • ad hoc to paper Occupation-by-NACE cross-tabulation totals can be generated by proportionally allocating randomized NACE totals.
    Algorithm 1 step 2: T_NACE,OCCUP* = T_NACE* times T_NACE,OCCUP / T_NACE. This fixes the joint distribution's shares and ignores their sampling uncertainty.
  • domain assumption Missing values in occupation, NACE and province are missing at random and the kNN imputation with random-forest-derived weights is correct.
    Section 2.2.2: missing NACE reached 56.86% in 2013, and imputation was performed with the VIM package. If missingness is informative, the imputed auxiliary variables bias the calibration.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Enhancing the Demand for Labour survey by including skills from online job advertisements using model-assisted calibration." pith.science (2026). https://pith.science/paper/MKF3R5IW

@misc{pith2026190806731,
  author       = {Pith},
  title        = {Pith review of: Enhancing the Demand for Labour survey by including skills from online job advertisements using model-assisted calibration},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MKF3R5IW}},
  note         = {Machine review of arXiv:1908.06731}
}
read the original abstract

In the article we describe an enhancement to the Demand for Labour (DL) survey conducted by Statistics Poland, which involves the inclusion of skills obtained from online job advertisements. The main goal is to provide estimates of the demand for skills (competences), which is missing in the DL survey. To achieve this, we apply a data integration approach combining traditional calibration with the LASSO-assisted approach to correct representation error in the online data. Faced with the lack of access to unit-level data from the DL survey, we use estimated population totals and propose a~bootstrap approach that accounts for the uncertainty of totals reported by Statistics Poland. We show that the calibration estimator assisted with LASSO outperforms traditional calibration in terms of standard errors and reduces representation bias in skills observed in online job ads. Our empirical results show that online data significantly overestimate interpersonal, managerial and self-organization skills while underestimating technical and physical skills. This is mainly due to the under-representation of occupations categorised as Craft and Related Trades Workers and Plant and Machine Operators and Assemblers.

Figures

Figures reproduced from arXiv: 1908.06731 by the authors.

Figure 1
Figure 1. Point estimates for the estimators and HTSRS estimator for each skill for 2011, 2013 and 2014 separately [PITH_FULL_IMAGE:figures/full_fig_p035_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

54 extracted references · 51 canonical work pages

  1. [1]

    Ber e sewicz, M. (2017). A two-step procedure to measure representativeness of internet data sources. International Statistical Review\/ 85\/ (3), 473--493

  2. [2]

    Lehtonen, F

    Beręsewicz, M., R. Lehtonen, F. Reis, L. Di Consiglio, and M. Karlberg (2018). An overview of methods for treating selectivity in big data sources. Statistical working papers, Eurostat

  3. [3]

    Boudarbat, B. and V. Chernoff (2012). Education–job match among recent canadian university graduates. Applied Economics Letters\/ 19\/ (18), 1923--1926

  4. [4]

    Burger, and J

    Buelens, B., J. Burger, and J. A. van den Brakel (2018). Comparing inference methods for non-probability samples. International Statistical Review\/ 86\/ (2), 322--343

  5. [5]

    Online job vacancies and skills analysis -- A Cedefop pan-European approach

    Cedefop (2019a). Online job vacancies and skills analysis -- A Cedefop pan-European approach

  6. [6]

    The online job vacancy market in the EU

    Cedefop (2019b). The online job vacancy market in the EU. Driving forces and emerging trends

  7. [7]

    Chen, J. K. T. (2016). Using LASSO to Calibrate Non-probability Samples using Probability Samples . Ph.\ D. thesis

  8. [8]

    Chen, J. K. T., M. R. Elliott, and R. Valliant (2018). Inference for nonprobability samples. Survey Methodology\/ 44\/ (1), 117--144

Show all 54 references
  1. [9]

    Chen, J. K. T., R. L. Valliant, and M. R. Elliott (2019). Calibrating non-probability surveys to estimated control totals using lasso, with an application to political polling. Journal of the Royal Statistical Society: Series C (Applied Statistics)\/ 68\/ (3), 657--681

  2. [10]

    Chevalier, A. (2011). Subject choice and earnings of uk graduates. Economics of Education Review\/ 30\/ (6), 1187--1201

  3. [11]

    Citro, C. F. (2014). From multiple modes for surveys to multiple data sources for estimates. Survey Methodology\/ 40\/ (2), 137--61

  4. [12]

    Mercorio, and M

    Colombo, E., F. Mercorio, and M. Mezzanzanica (2019). Ai meets labor market: Exploring the link between automation and skills. Information Economics and Policy\/ 47 , 27--37

  5. [13]

    Couper, M. P. (2013). Is the sky falling? new technology, changing media, and the future of surveys. Survey Research Methods\/ 7\/ (3), 145--156

  6. [14]

    Czarnik, S. (2011). Report concluding the 1st round of the Study conducted in 2010. Study of Human Capital in Poland

  7. [15]

    Daas, P. J., M. J. Puts, B. Buelens, and P. A. van den Hurk (2015). Big data as a source for official statistics. Journal of Official Statistics\/ 31\/ (2), 249--262

  8. [16]

    Deming, D. and L. B. Kahn (2018). Skill requirements across firms and labor markets: Evidence from job postings for professionals. Journal of Labor Economics\/ 36\/ (S1), S337--S369

  9. [17]

    Dever, J. A. and R. Valliant (2010). A comparison of variance estimators for poststratification to estimated control totals. Survey Methodology\/ 36\/ (1), 45--56

  10. [18]

    and C.-E

    Deville, J.-C. and C.-E. S \"a rndal (1992). Calibration estimators in survey sampling. Journal of the American Statistical Association\/ 87\/ (418), 376--382

  11. [19]

    Elliott, M. R. and R. Valliant (2017). Inference for nonprobability samples. Statistical Science\/ 32\/ (2), 249--264

  12. [20]

    ESSnet Big Data

    ESSnet on Big Data (2017). ESSnet Big Data. Final Technical Report, Work Package 1 - Web Scraping / Job Vacancies - Deliverable 1.3

  13. [21]

    ESSnet Big Data

    ESSnet on Big Data (2018). ESSnet Big Data. Final Technical Report, Work Package 1 - Web Scraping / Job Vacancies - Deliverable 1.3

  14. [22]

    Work Package B Overview

    ESSnet on Big Data (2019). Work Package B Overview

  15. [23]

    Hershbein, B. and L. B. Kahn (2018). Do recessions accelerate routine-biased technological change? evidence from vacancy postings. American Economic Review\/ 108\/ (7), 1737--1772

  16. [24]

    Skills mismatch in Europe

    International Labour Organization (2014). Skills mismatch in Europe

  17. [25]

    Kreuter, M

    Japec, L., F. Kreuter, M. Berg, P. Biemer, P. Decker, C. Lampe, J. Lane, C. O’Neil, and A. Usher (2015). Big data in survey researchaapor task force report. Public Opinion Quarterly\/ 79\/ (4), 839--880

  18. [26]

    Kim, J. K., S. Park, Y. Chen, and C. Wu (2018). Combining non-probability and probability survey samples through mass imputation. arXiv preprint arXiv:1812.10694\/

  19. [27]

    Kim, J. K. and Z. Wang (2018, aug). Sampling techniques for big data analysis. International Statistical Review\/ 87\/ (S1), S177--S191

  20. [28]

    Kowarik, A. and M. Templ (2016). Imputation with the R package VIM . Journal of Statistical Software\/ 74\/ (7), 1--16

  21. [29]

    Kuhn, P. and M. Skuterud (2004). Internet job search and unemployment durations. American Economic Review\/ 94\/ (1), 218--232

  22. [30]

    Marinescu, I. and R. Rathelot (2018). Mismatch unemployment and the geography of job search. American Economic Journal: Macroeconomics\/ 10\/ (3), 42--70

  23. [31]

    McConville, K. S., F. J. Breidt, T. C. Lee, and G. G. Moisen (2017). Model-assisted survey regression estimation with the lasso. Journal of Survey Statistics and Methodology\/ 5\/ (2), 131--158

  24. [32]

    McGowan, M. A. and D. Andrews (2015). Skill Mismatch and Public Policy in OECD Countries. OECD Economics Department Working Papers

  25. [33]

    McGuiness, S. and K. Pouliakas (2018). Skills mismatch: Concepts, masurement and policy approaches. Journal of Economic Surveys\/ 32\/ (4), 985--1015

  26. [34]

    Szko a, and M

    Pater, R., J. Szko a, and M. Kozak (2019). A method for measuring detailed demand for workers’ competences. Economics: The Open-Access, Open-Assessment E-Journal\/ 13\/ (2019-27), 1--29

  27. [35]

    Pfeffermann, D. (2015). Methodological Issues and Challenges in the Production of Official Statistics: 24th Annual Morris Hansen Lecture . Journal of Survey Statistics and Methodology\/ 3\/ (4), 425--483

  28. [36]

    Methodology of the Study of Human Capital in Poland (1st round)

    Polish Agency for Enterprise Development (2011). Methodology of the Study of Human Capital in Poland (1st round)

  29. [37]

    Methodology of the Study of Human Capital in Poland (2st round)

    Polish Agency for Enterprise Development (2012). Methodology of the Study of Human Capital in Poland (2st round)

  30. [38]

    Methodology of the Study of Human Capital in Poland (3st round)

    Polish Agency for Enterprise Development (2013). Methodology of the Study of Human Capital in Poland (3st round)

  31. [39]

    Methodology of the Study of Human Capital in Poland (4st round)

    Polish Agency for Enterprise Development (2014). Methodology of the Study of Human Capital in Poland (4st round)

  32. [40]

    Methodology of the Study of Human Capital in Poland (5st round)

    Polish Agency for Enterprise Development (2015). Methodology of the Study of Human Capital in Poland (5st round)

  33. [41]

    R: A Language and Environment for Statistical Computing

    R Core Team (2018). R: A Language and Environment for Statistical Computing . Vienna, Austria: R Foundation for Statistical Computing

  34. [42]

    Zabala, and A

    Reid, G., F. Zabala, and A. Holmberg (2017). Extending TSE to administrative data: a quality framework and case studies from stats nz. Journal of Official Statistics\/ 33\/ (2), 477--511

  35. [43]

    a rndal, C.-E. and S. Lundstr \

    S \"a rndal, C.-E. and S. Lundstr \"o m (2005). Estimation in surveys with nonresponse . John Wiley & Sons

  36. [44]

    Friedman, T

    Simon, N., J. Friedman, T. Hastie, and R. Tibshirani (2011). Regularization paths for cox's proportional hazards model via coordinate descent. Journal of Statistical Software\/ 39\/ (5), 1--13

  37. [45]

    Somers, M. A., S. J. Cabus, W. Groot, and H. M. van den Brink (2019). Horizontal mismatch between employment and field of education: Evidence from a systematic literature review. Journal of Economic Surveys\/ 33\/ (2), 567--603

  38. [46]

    The demand for labour in 2017

    Statistics Poland (2018). The demand for labour in 2017

  39. [47]

    Tibshirani, R. (1996). Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society. Series B (Methodological)\/ , 267--288

  40. [48]

    Valliant, R. (2019). Comparing Alternatives for Estimation from Nonprobability Samples . Journal of Survey Statistics and Methodology\/

  41. [49]

    Wu, C. and R. R. Sitter (2001). A model-calibration approach to using complete auxiliary information from survey data. Journal of the American Statistical Association\/ 96\/ (453), 185--193

  42. [50]

    Yang, S. and J. K. Kim (2018). Integration of survey data and big observational data for finite population inference using mass imputation. arXiv preprint arXiv:1807.02817\/

  43. [51]

    Yang, S. and J. K. Kim (2019). Nearest neighbor imputation for general parameter estimation in survey sampling. In The Econometrics of Complex Survey Data: Theory and Applications , pp.\ 209--234. Emerald Publishing Limited

  44. [52]

    Yang, S., J. K. Kim, and R. Song (2019). Doubly robust inference when combining probability and non-probability samples with high-dimensional data. arXiv preprint arXiv:1903.05212\/

  45. [53]

    Zhang, L.-C. (2012). Topics of statistical theory for register-based statistics and data integration. Statistica Neerlandica\/ 66\/ (1), 41--63

  46. [54]

    Zou, H. (2006). The adaptive lasso and its oracle properties. Journal of the American Statistical Association\/ 101\/ (476), 1418--1429

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.