Pith. sign in

REVIEW 4 major objections 6 minor 96 references

Trading off performance and human oversight in algorithmic policy: evidence from Danish college admissions

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Machine-learning models that rank Danish college applicants by predicted degree completion outperform the current GPA-based and human-assessment rankings, and even the simplest model yields a large estimated fiscal gain.

desk verdict The ranking result is genuine and worth refereeing; the 86M USD economic headline is a back-of-envelope extrapolation that needs heavy caveats or removal. read the letter →

arxiv 2411.15348 v2 pith:GX4T6RH6 submitted 2024-11-22 cs.CY

classification cs.CY
keywords collegeadmissionsdropoutpredictionalgorithmicpolicyriskscoresfairnessdegreecompletionhumanoversightmachinelearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Danish college admissions currently rank applicants by high-school GPA or, for those who opt in, by a human assessment. The paper claims that machine-learning models trained on pre-admission grade transcripts predict who will complete a degree better than either of those rankings, and that the biggest accuracy jump comes from replacing the GPA with any ML model rather than from choosing a deep architecture. On a nationwide dataset it reports 68.5% AUC for a simple logistic regression versus 64.6% for GPA-based ranking, with a transformer reaching 69.6%. Simulating the newly enacted 10% intake reduction, the authors estimate that the worst-performing model would cut dropout among the marginal rejected students by 9.4 to 12.7 percentage points and raise government revenue by about 86 million USD per cohort per year. They present the real price as a policy trade-off between performance, transparency, and human oversight, since deep learning buys only a small accuracy gain while complicating compliance with high-risk AI regulation.

What carries the argument

The load-bearing object is the risk score: a model-generated probability that an applicant completes the degree they applied for, used strictly as a ranking device. Paper-specific steps: each student's pre-admission record is converted into chronological event sequences (socio-demographic, grade, and enrollment events), then embedded and summed into fixed-dimensional inputs; the models are trained on 2006–2016 admissions and scored on 2017. Because Danish admissions uses a deferred-acceptance mechanism with two parallel rankings—GPA and human assessment—the paper can compare algorithmic rankings against both observed rankings on the same admitted population, sidestepping the selective-labels problem that usually blocks such evaluations.

What would settle it

Run a pilot admission lottery: for one cohort, fill some seats by random assignment among marginal applicants, and compare actual completion rates of the algorithm-selected and human/GPA-selected students. If the 9.4–12.7 percentage-point completion gap does not appear in the randomly assigned marginal group, the estimated 86M USD revenue gain collapses; a cheaper falsification is to test whether the model's ranking advantage persists on applicants who were not admitted under either ranking.

Watch

Extended reading notes

Core claim

The central claim is that degree-completion risk scores—computed only from data available before enrollment—rank applicants more accurately than the rankings Denmark actually uses. The paper tests this on the full population of Danish students admitted between 2006 and 2017, encoding each applicant's grade history as a chronological sequence of events and feeding it to transformer, LSTM, logistic-regression, and gradient-boosted-tree models. Every model outperforms the GPA-based and human-assessment rankings; the transformer achieves 69.6% out-of-sample AUC versus 64.6% for GPA, but logistic regression already reaches 68.5%, so the marginal value of the advanced architecture is about 1.1 percentage points. The paper further claims that under a policy contracting admissions by 10%, the models identify rejected subgroups whose completion rates are 9.4 to 12.7 percentage points below those of the students currently rejected, and that using the logistic regression would add 377 graduates per cohort and an estimated 86 million USD in yearly government revenue. It concludes that simple, transparent models capture most of the benefit, while fairness—measured by sufficiency and by the ABROCA ranking metric—is not systematically worse than current practice.

Load-bearing premise

The counterfactual evaluation assumes that the students an algorithm would reject would have had exactly the same completion outcomes—and no shift in program choices, applicant behavior, or pool composition—as they do under the current admission rules, even though the paper only observes outcomes for admitted students.

Editorial extensions

If this is right

  • A 10% intake contraction ranked by a logistic-regression risk score would reject students whose completion rates are 9.4 to 12.7 percentage points lower than those rejected under current GPA or human rankings.
  • Using any machine learning model on pre-admission grades captures most of the ranking improvement; a transformer adds roughly 1.1 percentage points of AUC over logistic regression.
  • The estimated 86 million USD annual revenue gain—from 377 additional graduates per cohort under the worst-performing model—exceeds the paper's estimated implementation and running costs of about 17.6 million USD.
  • Algorithmic rankings satisfy a calibration-based fairness criterion (sufficiency) across sex, nativity, and socioeconomic status, and the LSTM even shows lower ABROCA disparity than current GPA and human rankings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same grade-transcript features likely transfer to other centralized admissions systems, but the revenue figure is tied to Danish tax, subsidy, and funding rules; an equivalent estimate would need local parameters.
  • The paper's internal comparison implies a ready policy recommendation the authors only gesture at: a transparent logistic-regression score, not a deep model, is the cost-effective choice unless a context demands the extra 1.1 AUC points.
  • A real rollout should test for strategic response: once applicants know that grade patterns beyond the average matter, they may reshape their course choices, and the paper's counterfactual assumes no such behavioral change.
  • Because the evaluation only observes outcomes for admitted students, the 86M USD estimate is an extrapolation-like figure; a natural experiment or pilot admission lottery would be needed to verify it on the full applicant pool.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper uses Danish registry data covering 2006-2017 to predict degree completion conditional on admission, comparing logistic regression, gradient-boosted trees, LSTM, and transformer models against the current GPA-based and human-assessment ranking systems. Models are trained on 2006-2016 and evaluated on held-out 2017 data. The main descriptive findings are that all ML models using pre-admission academic grades rank applicants by predicted completion better than current GPA or human rankings (logistic regression AUC 68.5% vs. 64.6% for GPA; transformer 69.6%), and that a 10% intake contraction guided by ML rankings rejects students with 9.4-12.7 percentage-point lower completion rates than current rankings. The paper then estimates that adopting a logistic-regression-based admission rule would increase Danish government revenue by 86 million USD annually per cohort, and argues that the policy has an infinite Marginal Value of Public Funds. Fairness comparisons using ABROCA and related metrics are also reported.

Significance. The descriptive ranking comparison is a valuable contribution: it uses a large national dataset, evaluates on a genuinely held-out year, and exploits the centralized admission mechanism to observe full rankings of admitted students, thereby avoiding the selective-labels problem for the within-cohort comparison. The finding that the bulk of the gain comes from replacing current rankings with any ML model, rather than from model complexity, is policy-relevant and clearly presented. The economic estimate is less secure: it depends on externally borrowed parameters, an untested invariance assumption, and is reported without uncertainty quantification. If the authors add sensitivity analysis and reframe the economic claims proportionately, the paper could make a solid contribution to the algorithmic-policy literature. I also credit the authors for explicitly acknowledging the selective-labels limitation, the general-equilibrium assumption, and the uncertainty in the override-rate cost estimate.

major comments (4)
  1. [§2.2 and §A.9] The 86 million USD revenue estimate is a point estimate computed as 377 additional graduates times 230,000 USD per graduate, using the lowest borrowed return estimate (2.4 million DKK from Dalskov) and Danish tax parameters. The subtraction of 15.6 million USD in running costs is driven by an 18% human override rate imported from bail decisions (ref [70]) and applied to admissions, which the authors themselves call 'very uncertain.' No confidence intervals, plausible ranges, or alternative override rates are reported. Since the headline policy conclusion depends on net revenue remaining positive, please add a systematic sensitivity analysis over the override rate (e.g., 0-50%), per-graduate revenue, implementation costs, and development delay, and state whether the 86 million USD figure is robust over that range.
  2. [§1, §3, and §A.9] The counterfactual policy analysis assumes that outcomes observed for currently admitted students would be unchanged if admission decisions were made by the algorithm: no general-equilibrium effects on program switching, no applicant behavioral responses, and no change in program or peer composition. The Introduction states this as 'an assumption of no impact on switching between study programs,' and the Discussion acknowledges feedback loops and manipulation risks, but no evidence is provided that the ranking is robust to these responses. Because the revenue estimate is the product of ranking gains and per-graduate tax revenue, even modest degradation under deployment would directly shrink the headline figure. Please state the invariance assumption formally in A.9 and provide a sensitivity bound, for example by assuming marginal admitted students' completion rates differ by x percentage points and showing how the 86 million USD estimate changes, or by using a regression discontinuity around current GPA cutoffs to probe external validity.
  3. [Table 1 and Figure 2] The central descriptive claim that ML rankings outperform GPA and human rankings by 9.4-12.7 percentage points in the contracted decile is reported without uncertainty. The paper gives standard errors for the AUC comparisons (about 0.3 pp) but not for the contraction differences, which are load-bearing for the conclusion that any ML model improves on current policy. Please report standard errors or bootstrap confidence intervals for the contraction differences in Table 1 and Figure 2, clustered by study program if appropriate.
  4. [§2.2 and §A.9] The statement that algorithmic admissions yield an 'infinite Marginal Value of Public Funds' is an artifact of dividing a benefit estimate by a negative net government cost; with both the numerator and denominator estimated with substantial error, an infinite ratio is not informative for policy comparison. I recommend reporting net present value under the cost scenarios in Figures 4, 6, and 7, together with uncertainty ranges, and avoiding the infinite-MVPF formulation or clearly labeling it as a limiting statement under point estimates.
minor comments (6)
  1. [Appendix A header] The introduction to Appendix A says 'estimate the Marginal Value of Public Goods,' but the correct term used elsewhere is 'Marginal Value of Public Funds.'
  2. [§2.2] The phrase 'worst-best performing model (logistic regression)' is confusing; it should read 'worst-performing of the ML models' or 'best of the simple models,' as appropriate.
  3. [§A.9] In the infinite geometric series formula, the exponent should be k, not k-1, in the displayed expression S = sum ar^k.
  4. [Figure 4 and Figure 7 captions] The figure captions contain the typo 'V alue' in place of 'Value.'
  5. [§A.2] The sentence 'students admitted through the secondary quota being both older, having lower grade point averages and higher graduation rates' should read 'older, with lower grade point averages and higher graduation rates.'
  6. [Supplementary figures SI 4-SI 9] Several supplementary figure panels appear to have garbled or placeholder axis labels in the provided PDF; the authors should ensure the final version embeds readable vector text in all figure panels.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: out-of-sample evaluation and externally sourced revenue parameters keep the derivation self-contained; acknowledged general-equilibrium and manipulation limits are external-validity risks, not circular reductions.

full rationale

The paper's central ranking claims are self-contained against an external benchmark. Models are trained on 2006-2016 data and evaluated on held-out 2017 data ('we train our models on data from 2006 to 2016 and test the predictions on data from 2017', Section A.1), so the AUC comparisons (e.g., 68.5% logistic vs 64.6% GPA baseline, Section 2.1) are out-of-sample and not forced by fitting choices. The contraction analysis compares rejection sets within the observed 2017 admitted cohort; completion outcomes for every student in both the current-policy and algorithmic rejection sets are observed, so the 'reduction in dropout' of 377 graduates (Table 1) is a direct arithmetic comparison of two rankings of the same admitted population, not a fitted parameter relabeled as a prediction. The 86M USD figure multiplies this 377-graduate difference by externally sourced per-graduate government revenue estimates (2.4M DKK return from [90]; 37.7% income and 23% consumption tax rates from [92,93]), so the economic conclusion is parameterized by external evidence rather than by the model's own outputs. The acknowledged limitations—'an assumption of no impact on switching between study programs' (Introduction), 'we are unable to say anything about how our performance extends to the wider unadmitted population' (Discussion), and the feedback-loop and manipulation risks in the Discussion—are external-validity threats to the policy counterfactual, not circular reductions; they do not make any equation equal to its own input. Self-citations [13,38,63,66] appear only in institutional background and risk discussions; no load-bearing claim depends on an unverified self-citation, and no uniqueness theorem or ansatz is imported from the authors' prior work. The cross-program matching analysis (Section A.8) describes properties of the fitted model's predictions rather than newly identified causal effects, which is an interpretive caveat rather than a circular step. Accordingly, no step in the derivation reduces by construction.

Assumptions & free parameters 2 free parameters · 7 assumptions · 0 invented entities

The counterfactual and welfare estimates rest on two hand-chosen economic parameters (the 18% human-override rate and the 1M+1M USD implementation costs) and on the assumption that outcomes observed for admitted students under current rules would be unchanged under algorithmic rankings. The models themselves use only standard hyperparameters; the external return-to-education value (2.4M DKK per graduate) is an input from prior literature, not a fitted constant. No invented entities are introduced; 'risk scores' are model outputs.

free parameters (2)
  • Human override rate in cost model = 18%
    Borrowed from the judge override rate in bail decisions (Angelova et al. 2023) and applied to admissions overrides to estimate the 15.6M USD running cost; the paper calls this 'a very uncertain estimate' (Section A.9).
  • Implementation and operating costs = 1M USD fixed, 1M USD annual
    Projected costs for developing and running the centralized algorithm; used to compute net government revenue and the 'infinite MVPF' result (Section A.9).
assumptions (7)
  • domain assumption No general equilibrium effects: students completing programs under current admissions would complete at the same rates under algorithmic rankings, and applicants would not change their application behavior.
    Stated in the Introduction ('an assumption of no impact on switching between study programs') and in the Discussion's discussion of manipulation and feedback loops; load-bearing for the counterfactual contraction and revenue estimates.
  • domain assumption Selective labels are ignorable for the counterfactual: outcomes for the currently admitted population identify what would happen under the algorithm's ranking for the subpopulation that would be admitted under current rules.
    The paper argues selective labels are 'not an issue' because the counterfactual policy reduces intake from the admitted population (Discussion), but this holds only if the 10% rejected under the algorithm would otherwise have been admitted and their outcomes were not affected by the change.
  • domain assumption The external return-to-education estimate (2.4M DKK per graduate from Dalskov 2009) is valid and causal for the marginal students.
    Used in A.9 to convert 377 additional graduates into 86M USD of annual government revenue; the paper uses the lowest available estimate, but it is still an external parameter.
  • ad hoc to paper The 18% override rate from bail decisions applies to human oversight of algorithmic admissions.
    A.9 borrows the override rate from Angelova et al. (2023) on bail; the authors call this 'a very uncertain estimate' and use it to compute the 15.6M USD running cost.
  • domain assumption The 2017 cohort is representative of future cohorts for which the policy would be applied.
    Models are trained on 2006-2016 and tested on 2017 only (A.1); no forward validation on later cohorts is provided.
  • domain assumption The chosen fairness metrics (ABROCA, sufficiency, separation, independence) are appropriate for comparing rankings and predictions.
    A.7 operationalizes fairness; ABROCA requires ROC curves, and the sufficiency test 'cannot reject' is interpreted as evidence of sufficiency, which is a weak-form claim.
  • standard math Standard statistical machinery (AUC, ROC, z-tests, geometric series discounting, neural network training) is valid for the quantities computed.
    Used throughout Sections 2 and A.5-A.9.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Trading off performance and human oversight in algorithmic policy: evidence from Danish college admissions." pith.science (2026). https://pith.science/paper/GX4T6RH6

@misc{pith2026241115348,
  author       = {Pith},
  title        = {Pith review of: Trading off performance and human oversight in algorithmic policy: evidence from Danish college admissions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GX4T6RH6}},
  note         = {Machine review of arXiv:2411.15348}
}
read the original abstract

Student dropout is a significant concern for educational institutions due to its social and economic impact, driving the need for risk prediction systems to identify at-risk students before enrollment. We explore the accuracy of such systems in the context of higher education by predicting degree completion before admission, with potential applications for prioritizing admissions decisions. Using a large-scale dataset from Danish higher education admissions, we demonstrate that advanced sequential AI models offer more precise and fair predictions compared to current practices that rely on either high school grade point averages or human judgment. These models not only improve accuracy but also outperform simpler models, even when the simpler models use protected sociodemographic attributes. Importantly, our predictions reveal how certain student profiles are better matched with specific programs and fields, suggesting potential efficiency and welfare gains in public policy. We estimate that even the use of simple AI models to guide admissions decisions, particularly in response to a newly implemented nationwide policy reducing admissions by 10 percent, could yield significant economic benefits. However, this improvement would come at the cost of reduced human oversight and lower transparency. Our findings underscore both the potential and challenges of incorporating advanced AI into educational policymaking.

Figures

Figures reproduced from arXiv: 2411.15348 by the authors.

Figure 1
Figure 1. Admission to higher education and sequence representation. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Predictive performance of models and admission criteria. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Fairness of admission criteria and algorithms. [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Value of adopting prediction-based admission for different scenarios [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Sequential model architectures with aggregate embeddings [PITH_FULL_IMAGE:figures/full_fig_p025_5.png]
Figure 6
Figure 6. Figure 6: Net government revenue for different cost scenarios for GPA admission. [PITH_FULL_IMAGE:figures/full_fig_p032_6.png]
Figure 7
Figure 7. Figure 7: Net government revenue for different cost scenarios for human evaluation [PITH_FULL_IMAGE:figures/full_fig_p033_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

96 extracted references · 69 canonical work pages

  1. [70]

    Dobbie, and Crystal Yang

    Victoria Angelova, Will S. Dobbie, and Crystal Yang. Algorithmic Recommendations and Human Discretion. Working Paper. Sept. 2023

  2. [1]

    A Machine Learning Framework to Identify Students at Risk of Adverse Academic Outcomes

    Himabindu Lakkaraju et al. “A Machine Learning Framework to Identify Students at Risk of Adverse Academic Outcomes”. In: Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . KDD ’15. New York, NY, USA: Association for Computing Machinery, Aug. 2015, pp. 1909–1918

  3. [2]

    Predicting Student Dropout in Higher Education

    Lovenoor Aulck et al. Predicting Student Dropout in Higher Education . arXiv:1606.06364 [cs, stat]. Mar. 2017

  4. [3]

    Temporal and Between-Group Variability in College Dropout Prediction

    Dominik Glandorf et al. “Temporal and Between-Group Variability in College Dropout Prediction”. In: Proceedings of the 14th Learning Analytics and Knowledge Conference . 2024, pp. 486–497

  5. [4]

    A systematic review for MOOC dropout prediction from the perspective of machine learning

    Jing Chen et al. “A systematic review for MOOC dropout prediction from the perspective of machine learning”. In: Interactive Learning Environments 32.5 (May 2024). Publisher: Routledge eprint: https://doi.org/10.1080/10494820.2022.2124425, pp. 1642–1655

  6. [5]

    Dropout and completion in higher education in Europe : main report

    European Commission et al. Dropout and completion in higher education in Europe : main report. Publications Office, 2015. 14

  7. [6]

    Reimagining the machine learning life cycle to improve educational outcomes of students

    Lydia T. Liu et al. “Reimagining the machine learning life cycle to improve educational outcomes of students”. In: Proceedings of the National Academy of Sciences 120.9 (Feb. 2023). Publisher: Proceedings of the National Academy of Sciences, e2204781120

  8. [7]

    Dropping out of university: a literature review

    Andreas Behr et al. “Dropping out of university: a literature review”. en. In: Review of Education 8.2 (2020). eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/rev3.3202, pp. 614–652

Show all 96 references
  1. [8]

    A Major in Science? Initial Beliefs and Final Outcomes for College Major and Dropout

    Ralph Stinebrickner and Todd R. Stinebrickner. “A Major in Science? Initial Beliefs and Final Outcomes for College Major and Dropout”. In: The Review of Economic Studies 81.1 (Jan. 2014), pp. 426–472

  2. [9]

    Assessments Used in Higher Education Admissions

    Mar ´ ıa Elena Oliveri. “Assessments Used in Higher Education Admissions”. In: Higher Education Admissions Practices: An International Perspective. Ed. by Mar ´ ıa Elena Oliveri and CathyEditors Wendler. Educational and Psychological Testing in a Global Context. Cambridge Univ...

  3. [10]

    What grades and achievement tests measure

    Lex Borghans et al. “What grades and achievement tests measure”. eng. In: Proceedings of the National Academy of Sciences of the United States of America 113.47 (Nov. 2016), pp. 13354–13359

  4. [11]

    A Seven-College Experiment Using Algorithms to Track Students: Impacts and Implications for Equity and Fairness

    Peter Bergman, Elizabeth Kopko, and Julio Rodriguez. A Seven-College Experiment Using Algorithms to Track Students: Impacts and Implications for Equity and Fairness. en. SSRN Scholarly Paper. Rochester, NY, June 2021

  5. [12]

    Perdomo et al

    Juan C. Perdomo et al. Difficult Lessons on Social Prediction from Wisconsin Public Schools. arXiv:2304.06205 [cs, econ, q-fin, stat]. Sept. 2023

  6. [13]

    Task-specific information outperforms surveillance-style big data in predictive analytics

    Andreas Bjerre-Nielsen et al. “Task-specific information outperforms surveillance-style big data in predictive analytics”. In: Proceedings of the National Academy of Sciences 118.14 (Apr. 2021). Publisher: Proceedings of the National Academy of Sciences, e2020258118

  7. [14]

    Attention Is All You Need

    Ashish Vaswani et al. “Attention Is All You Need”. In: Advances in Neural Information Processing Systems 2017-December (June 2017). arXiv: 1706.03762 Publisher: Neural in- formation processing systems foundation, pp. 5999–6009

  8. [15]

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

    Jacob Devlin et al. “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”. In: NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Pro- ceedings of the Confe...

  9. [16]

    Introducing ChatGPT

    OpenAI. Introducing ChatGPT. en-US. Mar. 2023

  10. [17]

    Demystifying the Draft EU Artificial Intelligence Act — Analysing the good, the bad, and the unclear elements of the proposed approach

    Michael Veale and Frederik Zuiderveen Borgesius. “Demystifying the Draft EU Artificial Intelligence Act — Analysing the good, the bad, and the unclear elements of the proposed approach”. en. In: Computer Law Review International 22.4 (Aug. 2021). Publisher: Verlag Dr. Otto Sch...

  11. [18]

    The US Algorithmic Accountability Act of 2022 vs. The EU Arti- ficial Intelligence Act: what can they learn from each other?

    Jakob M¨ okander et al. “The US Algorithmic Accountability Act of 2022 vs. The EU Arti- ficial Intelligence Act: what can they learn from each other?” en. In: Minds and Machines 32.4 (Dec. 2022), pp. 751–758

  12. [19]

    Predicting academic success in higher education: literature review and best practices

    Eyman Alyahyan and Dilek D¨ u¸ steg¨ or. “Predicting academic success in higher education: literature review and best practices”. en. In: International Journal of Educational Tech- nology in Higher Education 17.1 (Feb. 2020), p. 3. 15

  13. [20]

    Belfield and Peter M

    Clive R. Belfield and Peter M. Crosta. Predicting Success in College: The Importance of Placement Tests and High School Transcripts. en. Tech. rep. Publication Title: Community College Research Center, Columbia University ERIC Number: ED529827. Community College Research Cente...

  14. [21]

    Systematic review: Predictors of students’ success in baccalaureate nursing programs

    Reem Al-Alawi, Gina Oliver, and Joe F. Donaldson. “Systematic review: Predictors of students’ success in baccalaureate nursing programs”. In: Nurse Education in Practice 48 (Oct. 2020), p. 102865

  15. [22]

    Predictive Analytics for University Student Admission: A Literature Review

    Kam Cheong Li, Billy Tak-Ming Wong, and Hon Tung Chan. “Predictive Analytics for University Student Admission: A Literature Review”. en. In: Blended Learning : Lessons Learned and Ways Forward . Ed. by Chen Li et al. Cham: Springer Nature Switzerland, 2023, pp. 250–259

  16. [23]

    Towards an Appropriate Query, Key, and Value Computation for Knowledge Tracing

    Youngduck Choi et al. “Towards an Appropriate Query, Key, and Value Computation for Knowledge Tracing”. In: Proceedings of the Seventh ACM Conference on Learning @ Scale. L@S ’20. New York, NY, USA: Association for Computing Machinery, Aug. 2020, pp. 341–344

  17. [24]

    BEHRT: Transformer for Electronic Health Records

    Yikuan Li et al. “BEHRT: Transformer for Electronic Health Records”. en. In: Scientific Reports 10.1 (Apr. 2020). Number: 1 Publisher: Nature Publishing Group, p. 7155

  18. [25]

    CAREER: Economic Prediction of Labor Sequence Data Under Dis- tribution Shift

    Keyon Vafa et al. “CAREER: Economic Prediction of Labor Sequence Data Under Dis- tribution Shift”. en. In: Oct. 2022

  19. [26]

    Using sequences of life-events to predict human lives

    Germans Savcisens et al. “Using sequences of life-events to predict human lives”. In: Nature Computational Science 4.1 (2024). Publisher: Nature Publishing Group US New York, pp. 43–56

  20. [27]

    The Unreasonable Effec- tiveness of Algorithms

    Jens Ludwig, Sendhil Mullainathan, and Ashesh Rambachan. “The Unreasonable Effec- tiveness of Algorithms”. en. In: AEA Papers and Proceedings 114 (May 2024), pp. 623– 627

  21. [28]

    Equality of opportunity in supervised learn- ing

    Moritz Hardt, Eric Price, and Nati Srebro. “Equality of opportunity in supervised learn- ing”. In: Advances in neural information processing systems 29 (2016)

  22. [29]

    Human Decisions and Machine Predictions

    Jon Kleinberg et al. “Human Decisions and Machine Predictions”. In: The Quarterly Jour- nal of Economics 133.1 (Feb. 2018), pp. 237–293

  23. [30]

    GRADE: Machine Learning Support for Graduate Admissions

    Austin Waters and Risto Miikkulainen. “GRADE: Machine Learning Support for Graduate Admissions”. en. In: AI Magazine 35.1 (Mar. 2014). Number: 1, pp. 64–64

  24. [31]

    The use of selective admissions tools to predict students’ success in an advanced standing baccalaureate nursing program

    Jennifer E. Timer and Marion I. Clauson. “The use of selective admissions tools to predict students’ success in an advanced standing baccalaureate nursing program”. In: Nurse Education Today 31.6 (Aug. 2011), pp. 601–606

  25. [32]

    Evaluating a Learned Admission- Prediction Model as a Replacement for Standardized Tests in College Admissions

    Hansol Lee, Ren´ e F. Kizilcec, and Thorsten Joachims. “Evaluating a Learned Admission- Prediction Model as a Replacement for Standardized Tests in College Admissions”. In: Proceedings of the Tenth ACM Conference on Learning @ Scale . L@S ’23. New York, NY, USA: Association fo...

  26. [33]

    Prediction Policy Problems

    Jon Kleinberg et al. “Prediction Policy Problems”. en. In: American Economic Review 105.5 (May 2015), pp. 491–495. 16

  27. [34]

    XGBoost: A Scalable Tree Boosting System

    Tianqi Chen and Carlos Guestrin. “XGBoost: A Scalable Tree Boosting System”. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . KDD ’16. New York, NY, USA: Association for Computing Machinery, Aug. 2016, pp. 785–794

  28. [35]

    Long short-term memory

    Sepp Hochreiter and J¨ urgen Schmidhuber. “Long short-term memory”. In: Neural com- putation 9.8 (1997). Publisher: MIT press, pp. 1735–1780

  29. [36]

    Algorithmic Fairness

    Jon Kleinberg et al. “Algorithmic Fairness”. en. In: AEA Papers and Proceedings 108 (May 2018), pp. 22–27

  30. [37]

    College Admission as a Screening and Sorting Device

    Mikkel Gandil and Edwin Leuven. College Admission as a Screening and Sorting Device . en. SSRN Scholarly Paper. Rochester, NY, Sept. 2022

  31. [38]

    Voluntary Information Disclosure in Cen- tralized Matching: Efficiency Gains and Strategic Properties

    Andreas Bjerre-Nielsen and Emil Chrisander. Voluntary Information Disclosure in Cen- tralized Matching: Efficiency Gains and Strategic Properties . arXiv:2206.15096 [econ, q- fin]. June 2022

  32. [39]

    Should college dropout prediction models include protected attributes?

    Renzhe Yu, Hansol Lee, and Ren´ e F Kizilcec. “Should college dropout prediction models include protected attributes?” In: Proceedings of the eighth ACM conference on learning@ scale. 2021, pp. 91–100

  33. [40]

    Bekendtgørelse om adgang til universitetsuddan- nelser tilrettelagt p ˚ a heltid

    Uddannelses- og Forskningsministeriet. Bekendtgørelse om adgang til universitetsuddan- nelser tilrettelagt p ˚ a heltid. Jan. 2022

  34. [41]

    Using AUC and accuracy in evaluating learning algorithms

    Jin Huang and C.X. Ling. “Using AUC and accuracy in evaluating learning algorithms”. In: IEEE Transactions on Knowledge and Data Engineering 17.3 (Mar. 2005). Conference Name: IEEE Transactions on Knowledge and Data Engineering, pp. 299–310

  35. [42]

    On the Stability of Fine- tuning BERT: Misconceptions, Explanations, and Strong Baselines

    Marius Mosbach, Maksym Andriushchenko, and Dietrich Klakow. On the Stability of Fine- tuning BERT: Misconceptions, Explanations, and Strong Baselines . en. arXiv:2006.04884 [cs, stat]. Mar. 2021

  36. [43]

    Udmøntning af sektordimensionering p ˚ a univer- siteterne

    Uddannelses og Forskningsministeriet. Udmøntning af sektordimensionering p ˚ a univer- siteterne. da. Apr. 2024

  37. [44]

    Evaluating the Fairness of Pre- dictive Student Models Through Slicing Analysis

    Josh Gardner, Christopher Brooks, and Ryan Baker. “Evaluating the Fairness of Pre- dictive Student Models Through Slicing Analysis”. en. In: Proceedings of the 9th Inter- national Conference on Learning Analytics & Knowledge . Tempe AZ USA: ACM, Mar. 2019, pp. 225–234

  38. [45]

    Algorithmic fairness in education

    Ren´ e F. Kizilcec and Hansol Lee. “Algorithmic fairness in education”. en. In:The Ethics of Artificial Intelligence in Education . 1st ed. New York: Routledge, Aug. 2022, pp. 174–202

  39. [46]

    Fairness and Machine Learning: Limitations and Opportunities

    Solon Barocas, Moritz Hardt, and Arvind Narayanan. Fairness and Machine Learning: Limitations and Opportunities . MIT Press, 2023

  40. [47]

    Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments

    Alexandra Chouldechova. “Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments”. In: Big Data 5.2 (June 2017). Publisher: Mary Ann Liebert, Inc., publishers, pp. 153–163

  41. [48]

    Inherent Trade-Offs in the Fair Determination of Risk Scores

    Jon Kleinberg, Sendhil Mullainathan, and Manish Raghavan. “Inherent Trade-Offs in the Fair Determination of Risk Scores”. In: 8th Innovations in Theoretical Computer Science Conference (ITCS 2017) . Ed. by Christos H. Papadimitriou. Vol. 67. Leibniz Interna- tional Proceedings...

  42. [49]

    A Unified Welfare Analysis of Government Policies

    Nathaniel Hendren and Ben Sprung-Keyser. “A Unified Welfare Analysis of Government Policies”. In: The Quarterly Journal of Economics 135.3 (Aug. 2020), pp. 1209–1318

  43. [50]

    Denmark: Country Profile

    International Monetary Fund. Denmark: Country Profile . 2024

  44. [51]

    It’s Just Not That Simple: An Empirical Study of the Accuracy- Explainability Trade-off in Machine Learning for Public Policy

    Andrew Bell et al. “It’s Just Not That Simple: An Empirical Study of the Accuracy- Explainability Trade-off in Machine Learning for Public Policy”. In: Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency . F AccT ’22. New York, NY, USA: Associa...

  45. [52]

    Fairness and Accountability Design Needs for Algorithmic Support in High-Stakes Public Sector Decision-Making

    Michael Veale, Max Van Kleek, and Reuben Binns. “Fairness and Accountability Design Needs for Algorithmic Support in High-Stakes Public Sector Decision-Making”. In: Pro- ceedings of the 2018 CHI Conference on Human Factors in Computing Systems . CHI ’18. New York, NY, USA: Ass...

  46. [53]

    How to build models for government: criteria driving model acceptance in policymaking

    Daniel Antony Kolkman et al. “How to build models for government: criteria driving model acceptance in policymaking”. en. In: Policy Sciences 49.4 (Dec. 2016), pp. 489–504

  47. [54]

    Stop explaining black box machine learning models for high stakes deci- sions and use interpretable models instead

    Cynthia Rudin. “Stop explaining black box machine learning models for high stakes deci- sions and use interpretable models instead”. en. In: Nature Machine Intelligence 1.5 (May 2019). Publisher: Nature Publishing Group, pp. 206–215

  48. [55]

    Artificial Intelligence Act

    European Parliament. Artificial Intelligence Act . en. Mar. 2024

  49. [56]

    Sanity Checks for Saliency Maps

    Julius Adebayo et al. “Sanity Checks for Saliency Maps”. In: Advances in Neural Infor- mation Processing Systems. Ed. by S. Bengio et al. Vol. 31. Curran Associates, Inc., 2018

  50. [57]

    Faithfulness Tests for Natural Language Explanations

    Pepa Atanasova et al. “Faithfulness Tests for Natural Language Explanations”. In: Pro- ceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Vol- ume 2: Short Papers) . Ed. by Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki. Toronto, Canada:...

  51. [58]

    Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness? en

    Alon Jacovi and Yoav Goldberg. Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness? en. arXiv:2004.03685 [cs]. Apr. 2020

  52. [59]

    The Disagreement Problem in Explainable Machine Learning: A Practitioner’s Perspective

    Satyapriya Krishna et al. The Disagreement Problem in Explainable Machine Learning: A Practitioner’s Perspective. arXiv:2202.01602 [cs]. July 2024

  53. [60]

    Explainability in AI Policies: A Critical Review of Communications, Reports, Regulations, and Standards in the EU, US, and UK

    Luca Nannini, Agathe Balayn, and Adam Leon Smith. “Explainability in AI Policies: A Critical Review of Communications, Reports, Regulations, and Standards in the EU, US, and UK”. In: Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. F AccT ’...

  54. [61]

    The EU AI Act: a summary of its significance and scope

    Lilian Edwards. “The EU AI Act: a summary of its significance and scope”. In: Artificial Intelligence (the EU AI Act) 1 (2021)

  55. [62]

    Fragile Algorithms and Fallible Decision-Makers: Lessons from the Justice System

    Jens Ludwig and Sendhil Mullainathan. “Fragile Algorithms and Fallible Decision-Makers: Lessons from the Justice System”. en. In: Journal of Economic Perspectives 35.4 (Nov. 2021), pp. 71–96

  56. [63]

    Why Do Students Lie and Should We Worry? An Analysis of Non-truthful Reporting

    Emil Chrisander and Andreas Bjerre-Nielsen. Why Do Students Lie and Should We Worry? An Analysis of Non-truthful Reporting . arXiv:2302.13718 [econ, q-fin]. Mar. 2023

  57. [64]

    Smart Matching Platforms and Heterogeneous Beliefs in Centralized School Choice*

    Felipe Arteaga et al. “Smart Matching Platforms and Heterogeneous Beliefs in Centralized School Choice*”. In: The Quarterly Journal of Economics 137.3 (Aug. 2022), pp. 1791– 1848. 18

  58. [65]

    The Economics of Matching: Stability and Incentives

    Alvin E. Roth. “The Economics of Matching: Stability and Incentives”. en. In: Mathematics of Operations Research 7.4 (1982), pp. 617–628

  59. [66]

    Playing the system: address manipulation and access to schools

    Andreas Bjerre-Nielsen et al. Playing the system: address manipulation and access to schools. arXiv:2305.18949 [econ, q-fin]. May 2023

  60. [67]

    Increasing Enrollment by Optimizing Schol- arship Allocations Using Machine Learning and Genetic Algorithms

    Lovenoor Aulck, Dev Nambi, and Jevin West. “Increasing Enrollment by Optimizing Schol- arship Allocations Using Machine Learning and Genetic Algorithms.” In: International Educational Data Mining Society (2020). Publisher: ERIC

  61. [68]

    The usefulness of algorithmic models in policy making

    Daan Kolkman. “The usefulness of algorithmic models in policy making”. In: Government Information Quarterly 37.3 (July 2020), p. 101488

  62. [69]

    Decision-Making with Machine Prediction: Evidence from Predictive Maintenance in Trucking

    Adam Harris and Maggie Yellen. “Decision-Making with Machine Prediction: Evidence from Predictive Maintenance in Trucking”. en. In: (Jan. 2024)

  63. [71]

    Chapter 4 - Early Warning Indicators and Inter- vention Systems: State of the Field

    Robert Balfanz and Vaughan Byrnes. “Chapter 4 - Early Warning Indicators and Inter- vention Systems: State of the Field”. In: Handbook of Student Engagement Interventions . Ed. by Jennifer A. Fredricks, Amy L. Reschly, and Sandra L. Christenson. Academic Press, Jan. 2019, pp. 45–55

  64. [72]

    Fairness and Abstraction in Sociotechnical Systems

    Andrew D. Selbst et al. “Fairness and Abstraction in Sociotechnical Systems”. en. In: Proceedings of the Conference on Fairness, Accountability, and Transparency. Atlanta GA USA: ACM, Jan. 2019, pp. 59–68

  65. [73]

    College Admissions and the Stability of Marriage

    D. Gale and L. S. Shapley. “College Admissions and the Stability of Marriage”. In: The American Mathematical Monthly 69.1 (Jan. 1962). Publisher: Taylor & Francis eprint: https://doi.org/10.1080/00029890.1962.11989827, pp. 9–15

  66. [74]

    School Choice: A Mechanism Design Ap- proach

    Atila Abdulkadiro˘ glu and Tayfun S¨ onmez. “School Choice: A Mechanism Design Ap- proach”. en. In: American Economic Review 93.3 (June 2003), pp. 729–747

  67. [75]

    Tilskud til uddannelse

    Uddannelses- og Forskningsministeriet. Tilskud til uddannelse . da. Tekst. 2023

  68. [76]

    Grading system

    Ministry of Higher Education and Science. Grading system. en. Page. 2021

  69. [77]

    A survey on educational data mining methods used for predicting students’ performance

    Wen Xiao, Ping Ji, and Juan Hu. “A survey on educational data mining methods used for predicting students’ performance”. en. In: Engineering Reports 4.5 (2022). eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/eng2.12482, e12482

  70. [78]

    GritNet: Student Performance Prediction with Deep Learning

    Byung-Hak Kim, Ethan Vizitei, and Varun Ganapathi. GritNet: Student Performance Prediction with Deep Learning. arXiv:1804.07405 [cs, stat]. Apr. 2018

  71. [79]

    Self-attention in Knowledge Tracing: Why It Works

    Shi Pu and Lee Becker. “Self-attention in Knowledge Tracing: Why It Works”. en. In: Artificial Intelligence in Education . Ed. by Maria Mercedes Rodrigo et al. Lecture Notes in Computer Science. Cham: Springer International Publishing, 2022, pp. 731–736

  72. [80]

    MOOC dropout prediction using machine learning techniques: Review and research challenges

    Fisnik Dalipi, Ali Shariq Imran, and Zenun Kastrati. “MOOC dropout prediction using machine learning techniques: Review and research challenges”. In: 2018 IEEE Global En- gineering Education Conference (EDUCON). ISSN: 2165-9567. Apr. 2018, pp. 1007–1014

  73. [81]

    Nguyen and Julian Salazar

    Toan Q. Nguyen and Julian Salazar. Transformers without Tears: Improving the Normal- ization of Self-Attention . en. arXiv:1910.05895 [cs, stat]. Dec. 2019

  74. [82]

    Decoupled Weight Decay Regularization

    Ilya Loshchilov and Frank Hutter. “Decoupled Weight Decay Regularization”. In: Inter- national Conference on Learning Representations . 2018. 19

  75. [83]

    PyTorch Lightning

    William Falcon and The PyTorch Lightning team. PyTorch Lightning. Mar. 2019

  76. [84]

    Learning Important Features Through Propagating Activation Differences

    Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. Learning Important Features Through Propagating Activation Differences. en. arXiv:1704.02685 [cs]. Oct. 2019

  77. [85]

    A Diagnostic Study of Explainability Techniques for Text Classifi- cation

    Pepa Atanasova et al. “A Diagnostic Study of Explainability Techniques for Text Classifi- cation”. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020, pp. 3256–3274

  78. [86]

    Interpretable models do not compromise accuracy or fairness in predicting college success

    Catherine Kung and Renzhe Yu. “Interpretable models do not compromise accuracy or fairness in predicting college success”. In: Proceedings of the seventh acm conference on learning@ scale. 2020, pp. 413–416

  79. [87]

    The Case for Using the MVPF in Empirical Welfare Analysis

    Nathaniel Hendren and Ben Sprung-Keyser. The Case for Using the MVPF in Empirical Welfare Analysis. en. Tech. rep. w30029. Cambridge, MA: National Bureau of Economic Research, May 2022, w30029

  80. [88]

    Returns to investment in education: a decennial review of the global literature

    George Psacharopoulos and Harry Anthony Patrinos. “Returns to investment in education: a decennial review of the global literature”. In: Education Economics 26.5 (Sept. 2018). Publisher: Routledge eprint: https://doi.org/10.1080/09645292.2018.1484426, pp. 445– 458

  81. [89]

    Chapter 4 - Returns to different postsecondary investments: Institution type, academic programs, and credentials

    Michael Lovenheim and Jonathan Smith. “Chapter 4 - Returns to different postsecondary investments: Institution type, academic programs, and credentials”. In: Handbook of the Economics of Education . Ed. by Eric A. Hanushek, Stephen Machin, and Ludger Woess- mann. Vol. 6. Elsev...

  82. [90]

    Store samfundsøkonomiske gevinster af uddannelse

    Mie Dalskov. Store samfundsøkonomiske gevinster af uddannelse . Tech. rep. 2009

  83. [91]

    Priceless: The Nonpecuniary Benefits of School- ing

    Philip Oreopoulos and Kjell G Salvanes. “Priceless: The Nonpecuniary Benefits of School- ing”. en. In: Journal of Economic Perspectives 25.1 (Feb. 2011), pp. 159–184

  84. [92]

    Økonomisk Analyse: Regneprincipper p ˚ a beskæftigelses- og overførselsomr ˚ adet

    Finansministeriet. Økonomisk Analyse: Regneprincipper p ˚ a beskæftigelses- og overførselsomr ˚ adet. da. Tech. rep. 2021

  85. [93]

    Dokumentationsnotat om opgørelse af nettoafgiftsfaktoren

    Finansministeriet. Dokumentationsnotat om opgørelse af nettoafgiftsfaktoren . da. Tech. rep. 2019

  86. [94]

    Forslag til finanslov for finans ˚ aret 2017

    Finansministeriet. Forslag til finanslov for finans ˚ aret 2017. da. 2016

  87. [95]

    A python library for confidence intervals

    Jacob Gildenblat. A python library for confidence intervals . 2023

  88. [96]

    right to explanation

    Eurostat. International Standard Classification of Education (ISCED) . en. 2023. 20 A Material and Methods In this section, we begin by describing the data sources being used in the study, the institutional setting of higher education applications in Denmark, and the specifics...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.