{"id":"994cec08-50a6-4e85-8ec2-146074b2ef07","arxiv_id":"2411.15348","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Simple machine learning models predict college completion better than GPA or human rankings in Danish admissions, and most of the benefit of AI admissions comes from models that remain interpretable.","lead":"Danish college admissions data show that even a simple machine learning model ranks applicants by predicted degree completion more accurately than high school GPA or human screening, and that deep learning adds only a small extra gain.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 86M USD revenue estimate rests on an untested no-general-equilibrium assumption: applicants and programs are assumed not to respond to the new ranking rule, yet the paper itself flags feedback and manipulation risks in the Discussion.","rationale":"The reader's CONDITIONAL verdict is appropriate. The strongest part of the paper is the out-of-sample ranking comparison: logistic regression and gradient-boosted trees reach about 68.5% AUC versus 64.6% for GPA, with standard errors around 0.26 pp, and the transformer reaches 69.6%. The contraction analysis is also internally sound for the historical cohort because both the baseline-rejected and the algorithm-rejected 10% sets are drawn from currently admitted students whose completion outcomes are observed; the selective-labels problem invoked by the reader is not what threatens this particular counterfactual. What does threaten the economic headline is the general-equilibrium and behavioral-response assumption: once the algorithm is actually used for admissions, applicants may adjust course choices and application portfolios, programs may change composition, and the completion behavior of marginal admittees may shift. The paper acknowledges this in the Introduction and Discussion, but provides no empirical test. The revenue estimate of 86M USD is directly proportional to the number of retained graduates, so even a moderate degradation of ranking performance under deployment, or a modest change in marginal-student completion rates, would erode the headline benefit. This does not undermine the core scientific claim that ML models rank currently admitted students better than current criteria; it does mean the economic and policy conclusions should be presented as conditional on a strong, untested invariance assumption. The reader's rationale also flags the abstract's fairness overstatement and the 'infinite MVPF' framing, both of which we agree are rhetorical overreach, but the central load-bearing weakness is the no-general-equilibrium assumption. Since the reader already conditions on this assumption, no verdict change is needed, but the stress-test sharpens why it is decisive for the 86M USD number rather than a peripheral caveat.","tokens_in":60172,"tokens_out":9600,"duration_ms":105378,"concrete_test":"Use Denmark's actual post-announcement cohorts (2019-2024, following the 10% intake reduction) from the same Statistics Denmark registers: retrain the logistic regression exactly as specified on data through 2016, apply it to each new cohort, and compare (i) out-of-sample AUC against the 68.5% benchmark, (ii) the actual graduation rate of the bottom decile by model score within each program against the predicted rate, and (iii) applicant behavior before versus after the policy, measured by the number and ordering of applications and by high-school course choices. If AUC and calibration hold and application portfolios are stable, the no-general-equilibrium assumption gains support; if the algorithm's edge shrinks or applicants re-sort, the 86M USD estimate should be substantially discounted or rejected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The descriptive ranking result is credible: AUC comparisons use held-out 2017 data, and the contraction analysis compares two 10% rejection sets whose outcomes are fully observed among current admittees, so the selective-labels issue does not actually threaten the within-cohort comparison. The load-bearing step is the leap from this retrospective ranking to the 377-graduate / 86M USD policy counterfactual. That leap assumes the joint distribution of grades, applications, program choice, and completion is invariant to the admission rule: no applicant changes course choices or application portfolios in response to the algorithm, no program changes composition, and no peer or resource effects alter completion for marginal admittees. The paper explicitly states this as 'an assumption of no impact on switching between study programs' in the Introduction, and the Discussion acknowledges feedback loops and manipulation risks, but no evidence is provided that the ranking is robust to these responses. Because the revenue estimate is the product of ranking gains and per-graduate tax revenue, any degradation in ranking under deployment, or any change in completion rates of marginal students, directly shrinks the headline figure. This is an acknowledged, untested, and potentially first-order threat to the economic conclusion, though not to the descriptive ranking comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses Danish registry data covering 2006-2017 to predict degree completion conditional on admission, comparing logistic regression, gradient-boosted trees, LSTM, and transformer models against the current GPA-based and human-assessment ranking systems. Models are trained on 2006-2016 and evaluated on held-out 2017 data. The main descriptive findings are that all ML models using pre-admission academic grades rank applicants by predicted completion better than current GPA or human rankings (logistic regression AUC 68.5% vs. 64.6% for GPA; transformer 69.6%), and that a 10% intake contraction guided by ML rankings rejects students with 9.4-12.7 percentage-point lower completion rates than current rankings. The paper then estimates that adopting a logistic-regression-based admission rule would increase Danish government revenue by 86 million USD annually per cohort, and argues that the policy has an infinite Marginal Value of Public Funds. Fairness comparisons using ABROCA and related metrics are also reported.","tokens_in":60362,"tokens_out":4747,"duration_ms":50847,"significance":"The descriptive ranking comparison is a valuable contribution: it uses a large national dataset, evaluates on a genuinely held-out year, and exploits the centralized admission mechanism to observe full rankings of admitted students, thereby avoiding the selective-labels problem for the within-cohort comparison. The finding that the bulk of the gain comes from replacing current rankings with any ML model, rather than from model complexity, is policy-relevant and clearly presented. The economic estimate is less secure: it depends on externally borrowed parameters, an untested invariance assumption, and is reported without uncertainty quantification. If the authors add sensitivity analysis and reframe the economic claims proportionately, the paper could make a solid contribution to the algorithmic-policy literature. I also credit the authors for explicitly acknowledging the selective-labels limitation, the general-equilibrium assumption, and the uncertainty in the override-rate cost estimate.","major_comments":[{"comment":"The 86 million USD revenue estimate is a point estimate computed as 377 additional graduates times 230,000 USD per graduate, using the lowest borrowed return estimate (2.4 million DKK from Dalskov) and Danish tax parameters. The subtraction of 15.6 million USD in running costs is driven by an 18% human override rate imported from bail decisions (ref [70]) and applied to admissions, which the authors themselves call 'very uncertain.' No confidence intervals, plausible ranges, or alternative override rates are reported. Since the headline policy conclusion depends on net revenue remaining positive, please add a systematic sensitivity analysis over the override rate (e.g., 0-50%), per-graduate revenue, implementation costs, and development delay, and state whether the 86 million USD figure is robust over that range.","section":"§2.2 and §A.9"},{"comment":"The counterfactual policy analysis assumes that outcomes observed for currently admitted students would be unchanged if admission decisions were made by the algorithm: no general-equilibrium effects on program switching, no applicant behavioral responses, and no change in program or peer composition. The Introduction states this as 'an assumption of no impact on switching between study programs,' and the Discussion acknowledges feedback loops and manipulation risks, but no evidence is provided that the ranking is robust to these responses. Because the revenue estimate is the product of ranking gains and per-graduate tax revenue, even modest degradation under deployment would directly shrink the headline figure. Please state the invariance assumption formally in A.9 and provide a sensitivity bound, for example by assuming marginal admitted students' completion rates differ by x percentage points and showing how the 86 million USD estimate changes, or by using a regression discontinuity around current GPA cutoffs to probe external validity.","section":"§1, §3, and §A.9"},{"comment":"The central descriptive claim that ML rankings outperform GPA and human rankings by 9.4-12.7 percentage points in the contracted decile is reported without uncertainty. The paper gives standard errors for the AUC comparisons (about 0.3 pp) but not for the contraction differences, which are load-bearing for the conclusion that any ML model improves on current policy. Please report standard errors or bootstrap confidence intervals for the contraction differences in Table 1 and Figure 2, clustered by study program if appropriate.","section":"Table 1 and Figure 2"},{"comment":"The statement that algorithmic admissions yield an 'infinite Marginal Value of Public Funds' is an artifact of dividing a benefit estimate by a negative net government cost; with both the numerator and denominator estimated with substantial error, an infinite ratio is not informative for policy comparison. I recommend reporting net present value under the cost scenarios in Figures 4, 6, and 7, together with uncertainty ranges, and avoiding the infinite-MVPF formulation or clearly labeling it as a limiting statement under point estimates.","section":"§2.2 and §A.9"}],"minor_comments":[{"comment":"The introduction to Appendix A says 'estimate the Marginal Value of Public Goods,' but the correct term used elsewhere is 'Marginal Value of Public Funds.'","section":"Appendix A header"},{"comment":"The phrase 'worst-best performing model (logistic regression)' is confusing; it should read 'worst-performing of the ML models' or 'best of the simple models,' as appropriate.","section":"§2.2"},{"comment":"In the infinite geometric series formula, the exponent should be k, not k-1, in the displayed expression S = sum ar^k.","section":"§A.9"},{"comment":"The figure captions contain the typo 'V alue' in place of 'Value.'","section":"Figure 4 and Figure 7 captions"},{"comment":"The sentence 'students admitted through the secondary quota being both older, having lower grade point averages and higher graduation rates' should read 'older, with lower grade point averages and higher graduation rates.'","section":"§A.2"},{"comment":"Several supplementary figure panels appear to have garbled or placeholder axis labels in the provided PDF; the authors should ensure the final version embeds readable vector text in all figure panels.","section":"Supplementary figures SI 4-SI 9"}],"recommendation":"major_revision","confidential_remarks":"The descriptive ranking result is solid and well-executed, but the economic headline is more fragile than the rest of the paper. I would advise the editor to ask for a substantive revision that reframes the 86 million USD estimate and the 'infinite MVPF' claim as illustrative, with formal sensitivity analysis over the borrowed parameters and the general-equilibrium assumption. The paper's fit with an interdisciplinary policy audience is good, and the authors' transparency about limitations is a strength; the remaining issues are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, the core descriptive result is solid and useful: on Danish national admissions data, any machine learning model ranks admitted applicants by predicted completion better than the current GPA or human screening, and the largest gain comes from switching to even a simple logistic regression, not from transformer complexity. Second, the economic headline (86M USD per year from a 10% contraction) is much softer than the ranking result and should be read as an illustrative calculation, not a policy forecast.\n\nWhat is actually new: a national-scale, held-out-year comparison of sequence models against both mandatory GPA rankings and voluntary human rankings in a centralized matching system, with full rankings observed among admittees. The contraction analysis is transparent, the AUC differences (logistic 68.5 vs GPA 64.6, transformer 69.6) come with standard errors around 0.3 pp, and the authors check initialization stability. They also clearly state the selective-labels caveat: they cannot say anything about unadmitted applicants. That honesty is earned.\n\nThe soft spots are concentrated where the reader's report says they are. The 86M USD estimate borrows external parameters (lowest per-graduate tax revenue, an 18% override rate from bail decisions) and assumes no behavioral response: no changes in application portfolios, program switching, or completion of marginal students under the new rule. The authors flag this as an assumption in the Introduction and acknowledge feedback-loop and manipulation risks in the Discussion, but they do not test it. That makes the economic counterfactual first-order fragile even though the descriptive ranking comparison is not. The \"infinite MVPF\" framing is more zeal than the evidence supports. Minor: the abstract's \"more fair\" overstates the body, where the LSTM improves fairness on some dimensions while other models roughly match current systems.\n\nWho gets value from this: anyone working on algorithmic allocation, centralized admissions, or prediction-policy problems. The paper gives a credible template for evaluating such policies and a useful benchmark result. It deserves a serious referee, not a desk reject. I would send it out, but I would push the authors to either reframe the economic section as a sensitivity analysis under a no-general-equilibrium assumption or substantially lower its prominence.","headline":"The ranking result is genuine and worth refereeing; the 86M USD economic headline is a back-of-envelope extrapolation that needs heavy caveats or removal.","tokens_in":60948,"tokens_out":1523,"would_cite":true,"duration_ms":18232,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Machine-learning models that rank Danish college applicants by predicted degree completion outperform the current GPA-based and human-assessment rankings, and even the simplest model yields a large estimated fiscal gain.","keywords":["college admissions","dropout prediction","algorithmic policy","risk scores","fairness","degree completion","human oversight","machine learning"],"falsifier":"Run a pilot admission lottery: for one cohort, fill some seats by random assignment among marginal applicants, and compare actual completion rates of the algorithm-selected and human/GPA-selected students. If the 9.4–12.7 percentage-point completion gap does not appear in the randomly assigned marginal group, the estimated 86M USD revenue gain collapses; a cheaper falsification is to test whether the model's ranking advantage persists on applicants who were not admitted under either ranking.","tokens_in":31,"feed_emoji":"🎓","tokens_out":7932,"duration_ms":134711,"temperature":0.7,"pith_summary":"Danish college admissions currently rank applicants by high-school GPA or, for those who opt in, by a human assessment. The paper claims that machine-learning models trained on pre-admission grade transcripts predict who will complete a degree better than either of those rankings, and that the biggest accuracy jump comes from replacing the GPA with any ML model rather than from choosing a deep architecture. On a nationwide dataset it reports 68.5% AUC for a simple logistic regression versus 64.6% for GPA-based ranking, with a transformer reaching 69.6%. Simulating the newly enacted 10% intake reduction, the authors estimate that the worst-performing model would cut dropout among the marginal rejected students by 9.4 to 12.7 percentage points and raise government revenue by about 86 million USD per cohort per year. They present the real price as a policy trade-off between performance, transparency, and human oversight, since deep learning buys only a small accuracy gain while complicating compliance with high-risk AI regulation.","feed_headline":"AI rankings beat Danish college admissions, adding $86M a year","feed_subtitle":"A simple logistic-regression score predicts completion better than GPA or human rankings, no black box required","key_machinery":"The load-bearing object is the risk score: a model-generated probability that an applicant completes the degree they applied for, used strictly as a ranking device. Paper-specific steps: each student's pre-admission record is converted into chronological event sequences (socio-demographic, grade, and enrollment events), then embedded and summed into fixed-dimensional inputs; the models are trained on 2006–2016 admissions and scored on 2017. Because Danish admissions uses a deferred-acceptance mechanism with two parallel rankings—GPA and human assessment—the paper can compare algorithmic rankings against both observed rankings on the same admitted population, sidestepping the selective-labels problem that usually blocks such evaluations.","core_discovery":"The central claim is that degree-completion risk scores—computed only from data available before enrollment—rank applicants more accurately than the rankings Denmark actually uses. The paper tests this on the full population of Danish students admitted between 2006 and 2017, encoding each applicant's grade history as a chronological sequence of events and feeding it to transformer, LSTM, logistic-regression, and gradient-boosted-tree models. Every model outperforms the GPA-based and human-assessment rankings; the transformer achieves 69.6% out-of-sample AUC versus 64.6% for GPA, but logistic regression already reaches 68.5%, so the marginal value of the advanced architecture is about 1.1 percentage points. The paper further claims that under a policy contracting admissions by 10%, the models identify rejected subgroups whose completion rates are 9.4 to 12.7 percentage points below those of the students currently rejected, and that using the logistic regression would add 377 graduates per cohort and an estimated 86 million USD in yearly government revenue. It concludes that simple, transparent models capture most of the benefit, while fairness—measured by sufficiency and by the ABROCA ranking metric—is not systematically worse than current practice.","pith_inferences":["The same grade-transcript features likely transfer to other centralized admissions systems, but the revenue figure is tied to Danish tax, subsidy, and funding rules; an equivalent estimate would need local parameters.","The paper's internal comparison implies a ready policy recommendation the authors only gesture at: a transparent logistic-regression score, not a deep model, is the cost-effective choice unless a context demands the extra 1.1 AUC points.","A real rollout should test for strategic response: once applicants know that grade patterns beyond the average matter, they may reshape their course choices, and the paper's counterfactual assumes no such behavioral change.","Because the evaluation only observes outcomes for admitted students, the 86M USD estimate is an extrapolation-like figure; a natural experiment or pilot admission lottery would be needed to verify it on the full applicant pool."],"forward_implications":["A 10% intake contraction ranked by a logistic-regression risk score would reject students whose completion rates are 9.4 to 12.7 percentage points lower than those rejected under current GPA or human rankings.","Using any machine learning model on pre-admission grades captures most of the ranking improvement; a transformer adds roughly 1.1 percentage points of AUC over logistic regression.","The estimated 86 million USD annual revenue gain—from 377 additional graduates per cohort under the worst-performing model—exceeds the paper's estimated implementation and running costs of about 17.6 million USD.","Algorithmic rankings satisfy a calibration-based fairness criterion (sufficiency) across sex, nativity, and socioeconomic status, and the LSTM even shows lower ABROCA disparity than current GPA and human rankings."],"supporting_citations":[{"why":"Supplies the transformer architecture used for the best-performing sequential model.","marker":"[14]"},{"why":"Supplies the LSTM architecture used as the older deep-learning baseline.","marker":"[35]"},{"why":"Supplies the gradient-boosted trees (XGBoost) tabular baseline.","marker":"[34]"},{"why":"Frames the selective-labels problem that the full-ranking design avoids.","marker":"[36]"},{"why":"Documents the human-assessment secondary quota that serves as the human-judgment baseline.","marker":"[37]"},{"why":"Describes the deferred-acceptance mechanism with voluntary information disclosure into which the rankings plug.","marker":"[38]"},{"why":"Provides the ABROCA fairness metric used to compare ranking fairness against current admission systems.","marker":"[44]"},{"why":"Provides the Marginal Value of Public Funds framework used to monetize the policy gain.","marker":"[49]"},{"why":"Is the 10% sector-dimensioning policy whose contraction scenarios the paper simulates.","marker":"[43]"},{"why":"Supplies the per-graduate public revenue estimate used to compute the 86M USD figure.","marker":"[90]"}],"fun_headline_variants":["Simple AI model beats college admissions rankings, adding $86M","Human oversight costs $86M: AI rankings win in Denmark","Danish study: simple AI beats GPA and human judgment in admissions","AI admissions: simple models beat human oversight, $86M gain","Logistic regression beats GPA for admissions, $86M benefit"],"cache_read_input_tokens":63104,"weakest_assumption_plain":"The counterfactual evaluation assumes that the students an algorithm would reject would have had exactly the same completion outcomes—and no shift in program choices, applicant behavior, or pool composition—as they do under the current admission rules, even though the paper only observes outcomes for admitted students.","fun_headline_variants_meta":{"raw":{"variants":["Simple AI model beats college admissions rankings, adding $86M","Human oversight costs $86M: AI rankings win in Denmark","Danish study: simple AI beats GPA and human judgment in admissions","AI admissions: simple models beat human oversight, $86M gain","Logistic regression beats GPA for admissions, $86M benefit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000657,"raw_usage":{"total_tokens":3026,"prompt_tokens":986,"completion_tokens":2040,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":1953}},"tokens_in":602,"tokens_out":2040,"duration_ms":15369,"temperature":1.0,"reasoning_tokens":1953,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:26:31.765607+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a pilot admission lottery: for one cohort, fill some seats by random assignment among marginal applicants, and compare actual completion rates of the algorithm-selected and human/GPA-selected students. If the 9.4–12.7 percentage-point completion gap does not appear in the randomly assigned marginal group, the estimated 86M USD revenue gain collapses; a cheaper falsification is to test whether the model's ranking advantage persists on applicants who were not admitted under either ranking.","supporting_citations":[{"cited_title":"College Admission as a Screening and Sorting Device","cited_arxiv_id":null,"evidence_quote":"Documents the human-assessment secondary quota that serves as the human-judgment baseline."},{"cited_title":"Long short-term memory","cited_arxiv_id":null,"evidence_quote":"Supplies the LSTM architecture used as the older deep-learning baseline."},{"cited_title":"XGBoost: A Scalable Tree Boosting System","cited_arxiv_id":null,"evidence_quote":"Supplies the gradient-boosted trees (XGBoost) tabular baseline."},{"cited_title":"Algorithmic Fairness","cited_arxiv_id":null,"evidence_quote":"Frames the selective-labels problem that the full-ranking design avoids."},{"cited_title":"Voluntary Information Disclosure in Centralized Matching: Efficiency Gains and Strategic Properties","cited_arxiv_id":"2206.15096","evidence_quote":"Describes the deferred-acceptance mechanism with voluntary information disclosure into which the rankings plug."},{"cited_title":"Evaluating the Fairness of Pre- dictive Student Models Through Slicing Analysis","cited_arxiv_id":null,"evidence_quote":"Provides the ABROCA fairness metric used to compare ranking fairness against current admission systems."},{"cited_title":"A Unified Welfare Analysis of Government Policies","cited_arxiv_id":null,"evidence_quote":"Provides the Marginal Value of Public Funds framework used to monetize the policy gain."},{"cited_title":"Udmøntning af sektordimensionering p ˚ a univer- siteterne","cited_arxiv_id":null,"evidence_quote":"Is the 10% sector-dimensioning policy whose contraction scenarios the paper simulates."},{"cited_title":"Store samfundsøkonomiske gevinster af uddannelse","cited_arxiv_id":null,"evidence_quote":"Supplies the per-graduate public revenue estimate used to compute the 86M USD figure."}],"review_version":1}