{"id":"a70c38b1-78f1-41a6-b1fa-cf519aeac8a6","arxiv_id":"2608.04180","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"NTK-motivated early-gradient sensitivity offers the best accuracy-stability trade-off for OUD diagnosis-code feature selection, though the benchmark's fairness is undermined by potential label leakage.","lead":"This study compares five ways to choose diagnosis codes as features for predicting opioid use disorder from electronic health records, using a BERT model to evaluate each choice. It reports that an early-gradient method works best overall, but the comparison may be biased because most methods can use the target diagnosis code as a feature.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central comparison is compromised by possible label leakage: outcome-related F11 codes are not excluded from the data-driven selectors' feature vocabulary, while the LLM prompt explicitly excludes them, so the reported NTK advantage may be inflated.","rationale":"The reader's weakest_assumption matches my reading: the outcome definition and lookback feature window create a direct leakage path that only the LLM is instructed to avoid. This is load-bearing because it affects every pairwise comparison in Table 1 and Figure 3, not just a secondary metric. The paper otherwise has strengths: a large cohort, patient-level splits, a fixed downstream architecture, multiple feature budgets, and a stability analysis. Those do not compensate for an asymmetric leakage risk, because the main 'NTK best' claim rests on the rankings being honest. The full-vocabulary AUPRC (0.458) versus K=500 NTK AUPRC (0.291) also undermines the 'moderate vocabulary / diminishing returns' claim, but that is secondary to the leakage issue; even if the vocabulary-size curve were reconciled, the method comparison would still be contaminated unless F11 codes are excluded for all methods. Therefore I concur with the reader's REJECT verdict. A focused incident-OUD re-analysis with outcome codes removed from all candidate vocabularies would settle the concern; if the NTK advantage survives that test, the paper could be revised and reconsidered. No change to the reader's verdict is needed now.","tokens_in":8409,"tokens_out":4650,"duration_ms":44216,"concrete_test":"Re-run the feature-selection pipeline on the same cohort with an incident-OUD outcome: exclude any patient whose F11 evidence predates the index opioid prescription, and remove all F11-prefixed (and, conservatively, opioid-poisoning and treatment) diagnosis categories from the candidate vocabulary for all five methods, not just the LLM. Then recompute Table 1 and Figure 3. If F11 codes appear in any current top-K ranking, the reported NTK advantage must be re-evaluated without those codes; if no F11 codes appear, the leakage concern is refuted and the ranking comparison stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the claim that NTK-motivated early gradient sensitivity is the best feature selector, the five methods must rank features under equivalent information about the outcome. They do not. OUD is defined as any F11-prefixed diagnosis in the patient's records (Methods, Data Source), and the feature vocabulary is derived from the 12 months preceding the index opioid prescription. The paper never states that F11 or OUD-consequence codes were removed from the candidate vocabulary for Recurrence Enrichment, NTK Sensitivity, LightGBM-SHAP, or Elastic Net. A patient with an F11 code in the lookback window is by construction an OUD-positive case; for the four data-driven methods that code is a perfectly informative feature, so rankings and downstream performance can be inflated by direct label leakage. The LLM method, by contrast, is explicitly prompted to exclude consequence and treatment codes for established OUD to prevent label leakage (Methods, LLM Semantic prompt). Thus the comparison is asymmetric and the central 'NTK best' conclusion is not supported as presented. The paper's own limitation section concedes temporal misalignment and calls for an incident-only sensitivity analysis, which is exactly the missing control.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a benchmark of five feature-selection methods—recurrence enrichment, NTK-motivated early gradient sensitivity, LightGBM-SHAP, Elastic Net, and LLM-guided semantic selection—for selecting ICD-10 diagnosis categories to predict opioid use disorder (OUD) from EHR data. Using a Cerner Health Facts cohort of 11.8 million patients, the authors construct diagnosis features from the 12 months preceding an index opioid prescription, rank features by each method, and evaluate top-K subsets by training a BERT classifier from scratch. The reported findings are that NTK sensitivity gives the best accuracy/stability trade-off, that performance improves up to roughly 300 features and then plateaus, and that LLM selection adds complementary but weaker standalone signal. The paper is a comparative empirical study rather than a novel methodological contribution.","tokens_in":8631,"tokens_out":8404,"duration_ms":83155,"significance":"If the comparison were valid, it would provide practical guidance for diagnosis-code feature selection in sparse EHR spaces. Strengths of the manuscript include a very large real-world cohort, a unified preprocessing and evaluation pipeline, multiple feature budgets, downstream evaluation with a consistent classifier, bootstrap stability analysis, and the inclusion of statistical, ML, and LLM-based selectors. The NTK-motivated ranking is a simple and interesting baseline. However, the central comparison is compromised by an asymmetric label-leakage setup: the outcome-defining F11 code is apparently available to the four data-driven selectors but explicitly excluded from the LLM prompt. In addition, the full-vocabulary result contradicts the diminishing-returns claim. Because the headline conclusion rests on this comparison, the manuscript in its current form does not support its main claims.","major_comments":[{"comment":"The comparison is asymmetric because the outcome-defining code F11 is not excluded from the candidate vocabulary for the four data-driven selectors, while the LLM prompt explicitly excludes 'consequences or treatment of established OUD (i.e., label leakage)'. OUD is operationalized as any F11-prefixed diagnosis in the patient's records (Methods, Data Source), and the feature window is the 12 months preceding the index opioid prescription. If an F11 code appears in that window, the patient is OUD-positive by definition, so Recurrence Enrichment (Eq. 3), NTK Sensitivity (Eqs. 5-6), LightGBM-SHAP, and Elastic Net can all rank this feature as essentially a copy of the label. The paper never states that F11 was removed from the input vocabulary for these methods. This inflates their downstream AUPRC and stability and invalidates the headline claim that NTK sensitivity is the best selector. The limitation section calling for an incident-only sensitivity analysis is exactly the missing control; it should be a primary analysis, not a deferred caveat.","section":"Methods, Data Source and LLM Semantic prompt; Table 1"},{"comment":"The 'diminishing returns around 300 features' conclusion is not supported by the reported numbers. Table 1 shows NTK Sensitivity at K=300 with AUPRC 0.247 and at K=500 with AUPRC 0.291, while the Full Vocabulary (1,908 features) reaches AUPRC 0.458, more than 50% higher than the best K=500 subset. This is not a marginal gain and is inconsistent with the abstract's claim that 'predictive performance improves rapidly up to a moderate vocabulary size (approximately 300 diagnosis categories) and then exhibits diminishing returns.' The full-vocabulary result is acknowledged in the text but only as a caveat; it should be reconciled with the conclusion, or the conclusion should be restricted to the evaluated range K=50-500.","section":"Overall Predictive Performance, Table 1"},{"comment":"The outcome definition is temporally ambiguous. The Data Source says patients are considered to have OUD if an F11 code appears 'in their diagnosis records,' with no statement that the code must occur after the feature window or after the index opioid prescription. The BERT model is instead described as predicting whether the patient 'will receive an OUD-related diagnosis at the next encounter.' If the input sequence already contains an F11 code, the model is being asked to predict a condition that is already documented. The paper should define an explicit index date, exclude F11 and related consequence/treatment codes from the feature window for all methods, and report an incident-only analysis; this is a prerequisite for the 'future OUD prediction' interpretation used throughout the abstract and conclusion.","section":"Methods, Data Source and Downstream Model"}],"minor_comments":[{"comment":"The BERT architecture and training hyperparameters (number of layers, hidden size, learning rate, batch size, training steps) are not specified; the statement that they are 'identical' across conditions is not sufficient for reproducibility.","section":"Downstream Model"},{"comment":"The grid-search ranges for the Elastic Net hyperparameters α and λ are not reported; include them so the selection procedure is reproducible.","section":"Methods, Elastic Net"},{"comment":"The method is called 'LLM Semantic' in Methods but 'LLM Semantic Similarity' in the Experiments section; please use one name consistently.","section":"Experiments and Results"},{"comment":"There is a typo in the prompt text: 'Donotbase' should be 'Do not base'.","section":"Figure 2"},{"comment":"Reference [22] cites DrugBank for the ICD-9 to ICD-10 mapping, which does not appear to be the correct source for terminology mapping; please cite the appropriate mapping resource.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The label-leakage concern is serious and aligns with the reader's report. I would not accept the paper until the comparison is rerun with F11 (and ideally all OUD-related consequence/treatment codes) removed from the candidate vocabulary for all methods, or until an incident-only cohort is used. The full-vocabulary result also needs to be reconciled with the diminishing-returns claim. If the authors cannot perform these re-analyses, the paper should be rejected. The empirical framework is valuable and the NTK baseline is interesting, but the current analysis does not support the headline conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a serious empirical benchmark: five feature-selection paradigms compared under one preprocessing pipeline, a from-scratch BERT evaluator, stability analysis under resampling, and attention to rare-code representation. The NTK-motivated early-gradient filter is a genuinely new variant, and the stability result for it (Jaccard ~0.98) is striking. I also appreciate the honesty in the limitations section, which flags temporal misalignment with true OUD onset and calls for an incident-only analysis.\n\nThat said, the main comparison is compromised. OUD is defined as any F11-prefixed code in the patient's records, and the feature vocabulary comes from the 12 months before the index opioid prescription. The paper never states that F11 codes were removed from the candidate vocabulary for Recurrence Enrichment, NTK, LightGBM-SHAP, or Elastic Net. A patient with an F11 code in the lookback window is positive by construction, so that code is a near-perfect label leak for those four methods. The LLM prompt, by contrast, explicitly instructs the model to exclude consequences or treatments of established OUD. So the data-driven rankings and their downstream AUPRC are inflated relative to the LLM, and the \"NTK best\" conclusion is not supported as presented. This is load-bearing, not a nitpick.\n\nTwo secondary issues. First, the \"diminishing returns beyond ~300 features\" claim sits uneasily with the full-vocabulary result: 1,908 features gives AUPRC 0.458 versus 0.291 for NTK at 500. That is a large jump, so the plateau claim is overstated as stated. Second, the text says bootstrap 95% confidence intervals were computed, but Table 1 lists no intervals; the reader cannot judge whether NTK's edge over LightGBM is real.\n\nThe fixes are clear and plausible: drop F11 and OUD-consequence codes from all methods' vocabularies, or define an incident-OUD outcome; report the intervals; and tone down the plateau claim. The benchmark framework is reusable, and the NTK method is worth a look once the leakage is removed.\n\nWho is this for? Anybody building EHR-based OUD prediction models and deciding how to budget a diagnosis vocabulary. It deserves a serious referee—not because the current claims hold, but because the question matters and the experimental scaffolding is strong enough that a revision could make it a useful paper. I would not cite it in its current form.","headline":"A well-built benchmark with a novel NTK-based selector, but the head-to-head comparison is not clean: F11 codes leak into the data-driven feature sets while the LLM is explicitly told to exclude them.","tokens_in":9205,"tokens_out":1494,"would_cite":false,"duration_ms":16908,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that NTK-motivated early gradient sensitivity is the best feature-selection method among five for OUD prediction from EHR diagnosis codes, with a plateau around 300 codes.","keywords":["feature selection","electronic health records","opioid use disorder","NTK gradient sensitivity","diagnosis codes","SHAP","large language models","predictive modeling"],"falsifier":"Retrain the NTK, LightGBM-SHAP, Elastic Net, and recurrence rankings after deleting every F11 and OUD-consequence code from the candidate vocabulary, then compare BERT AUPRC at $K=300$; if NTK's advantage shrinks or the selected subsets fall to LLM-only levels, the headline result is driven by label leakage.","tokens_in":8181,"feed_emoji":"📊","tokens_out":5892,"duration_ms":58919,"temperature":0.7,"pith_summary":"This paper benchmarks five ways to choose a compact set of ICD-10 diagnosis categories for predicting opioid use disorder from electronic health records, using a single preprocessing pipeline and a BERT-based downstream classifier. It claims that NTK-motivated early gradient sensitivity—ranking codes by their gradient magnitude at model initialization—achieves the best balance of predictive performance and resampling stability. It also finds that performance improves rapidly up to roughly 300 diagnosis categories and then shows diminishing returns, while large-language-model selection contributes complementary clinical codes despite lower standalone accuracy. The result matters because it identifies a reproducible, model-agnostic recipe for reducing sparse diagnosis-code spaces to a manageable clinical vocabulary.","feed_headline":"Early-gradient ranking beats four rivals for opioid-risk EHR features","feed_subtitle":"In 11.8 million patient records, 300 diagnosis codes chosen by early gradients gave the most stable, accurate OUD predictions.","key_machinery":"The central mechanism is the NTK-motivated early gradient sensitivity score. For each diagnosis $d$, the paper forms a recurrence-weighted patient feature $x'_d(p)=\\log(1+f_d(p))$, initializes logistic-regression bias $b_0=\\log(r/(1-r))$ from OUD prevalence $r$, and computes per-encounter gradient contributions $g_d(e)=(\\hat p_0-y)\\,x'_d(p)$. Codes are ranked by $S_d=\\sum_e |g_d(e)|$. The NTK argument is that this initialization-time gradient magnitude forecasts how strongly each feature will influence the trained model, and the empirical claim is that this forecast is both accurate and far more stable under resampling than SHAP or Elastic Net attributions.","core_discovery":"On the paper's own terms, the central discovery is that a feature-selection method needing no iterative training—NTK-motivated early gradient sensitivity—outperforms recurrence enrichment, LightGBM-SHAP, Elastic Net, and LLM-guided semantic selection when the selected codes feed a BERT classifier. At $K=500$ it reaches test AUPRC $0.291$, AUROC $0.901$, and optimal F1 $0.319$, with near-deterministic rankings under bootstrap resampling (Jaccard similarity $0.984 \\pm 0.006$ at $K=300$). The paper further claims that gains from larger budgets are concentrated below $K\\approx 300$ and that consensus codes across methods retain usable signal, though the full 1,908-code vocabulary still yields higher AUPRC ($0.458$).","pith_inferences":["If F11 codes were not removed from the data-driven vocabularies, the NTK advantage may shrink or vanish once exclusion is enforced; the paper's LLM prompt explicitly excludes them, so a fair comparison requires the same exclusion.","If the roughly 300-code plateau holds across outcomes and institutions, health systems could standardize on a small, interpretable diagnosis vocabulary for OUD screening, improving validation and transportability.","A hybrid pipeline that uses the LLM list to add rare clinically meaningful codes to the NTK list could exceed either method alone, but this paper does not test that combination.","The gap between the selected subsets and the full 1,908-code vocabulary suggests substantial unused signal remains, so larger budgets or hierarchy-aware grouping could capture it."],"forward_implications":["A fixed vocabulary of roughly 300 three-character ICD-10 categories can match the predictive performance of much larger diagnosis-code sets, so deployment can use compact feature budgets without major loss.","NTK sensitivity rankings are nearly deterministic under bootstrap resampling, so feature-selection results should reproduce across training-data perturbations.","LLM-selected codes overlap little with data-driven rankings, so semantic priors can be used to expand candidate sets rather than replace statistical selection.","Consensus feature sets selected by at least three methods perform between individual methods and the full vocabulary, suggesting that agreement across paradigms marks reliable signal.","BERT-based downstream models outperform logistic regression and MLP under the same feature subsets, so architecture captures temporal diagnosis patterns beyond linear or shallow nonlinear models."],"supporting_citations":[{"why":"BERT transformer encoder is the downstream model whose test AUPRC, AUROC, and F1 define the comparison.","marker":"[25]"},{"why":"SHAP values supply the importance ranking for the LightGBM-SHAP arm.","marker":"[19]"},{"why":"Motivates using LLMs to rank predictive variables without training data, the basis of the LLM Semantic arm.","marker":"[18]"},{"why":"Defines the stability framework used to assess resampling robustness of selected feature sets.","marker":"[14]"},{"why":"Supplies the Cerner Health Facts EHR database from which the 11,791,858-patient cohort is drawn.","marker":"[20]"},{"why":"Documents how OUD-related ICD codes map from ICD-9-CM to ICD-10-CM, supporting the F11.xx outcome definition.","marker":"[24]"}],"fun_headline_variants":["NTK early-gradient ranking beats 4 rivals for OUD feature selection","No-training NTK gradients top SHAP, LightGBM, LLM for opioid EHR features","300 EHR diagnosis codes via NTK gradients yield stable OUD predictions","Early-gradient feature pick wins on accuracy and stability for OUD model","Opioid risk prediction: NTK sensitivity bests four feature-selection methods"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes the 12-month lookback features are temporally separate from the outcome, but OUD is defined as any F11 code in the patient's records and the paper never states that F11 or consequence codes were removed from the input vocabulary for the four data-driven methods.","fun_headline_variants_meta":{"raw":{"variants":["NTK early-gradient ranking beats 4 rivals for OUD feature selection","No-training NTK gradients top SHAP, LightGBM, LLM for opioid EHR features","300 EHR diagnosis codes via NTK gradients yield stable OUD predictions","Early-gradient feature pick wins on accuracy and stability for OUD model","Opioid risk prediction: NTK sensitivity bests four feature-selection methods"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000237,"raw_usage":{"total_tokens":1480,"prompt_tokens":892,"completion_tokens":588,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":488}},"tokens_in":508,"tokens_out":588,"duration_ms":6715,"temperature":1.0,"reasoning_tokens":488,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T00:21:30.908856+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain the NTK, LightGBM-SHAP, Elastic Net, and recurrence rankings after deleting every F11 and OUD-consequence code from the candidate vocabulary, then compare BERT AUPRC at $K=300$; if NTK's advantage shrinks or the selected subsets fall to LLM-only levels, the headline result is driven by label leakage.","supporting_citations":[{"cited_title":"Bert: Pre-training of deep bidirectional transformers for language understanding","cited_arxiv_id":null,"evidence_quote":"BERT transformer encoder is the downstream model whose test AUPRC, AUROC, and F1 define the comparison."},{"cited_title":"Llm-select: Feature selection with large language models","cited_arxiv_id":null,"evidence_quote":"Motivates using LLMs to rank predictive variables without training data, the basis of the LLM Semantic arm."},{"cited_title":"On the stability of feature selection algorithms","cited_arxiv_id":null,"evidence_quote":"Defines the stability framework used to assess resampling robustness of selected feature sets."},{"cited_title":"Cerner Health Facts ® Data Sets - SBMI Data Service; 2018","cited_arxiv_id":null,"evidence_quote":"Supplies the Cerner Health Facts EHR database from which the 11,791,858-patient cohort is drawn."},{"cited_title":"Case Study: Exploring How Opioid-Related Diagnosis Codes Translate From ICD-9- CM to ICD-10-CM","cited_arxiv_id":null,"evidence_quote":"Documents how OUD-related ICD codes map from ICD-9-CM to ICD-10-CM, supporting the F11.xx outcome definition."}],"review_version":1}