REVIEW 3 major objections 5 minor 36 references
TEDDY, a 1.84-million-parameter transformer trained on ICD-10 diagnosis histories from 1.6 million children, anticipates first-time diagnoses across 797 conditions with a median AUC of 0.72, including rare diseases and with signal detectabl
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 03:04 UTC pith:UBW6PUHX
load-bearing objection First generative pediatric EHR model with a carefully designed evaluation, but the headline AUCs are likely inflated by not matching cases and controls on history length in the main analysis. the 3 major comments →
TEDDY: A Pediatric Foundation Model for Risk Forewarning from ICD-Coded Diagnostic Histories
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
TEDDY, a decoder-only transformer with 1.84 million parameters, is trained on about 73 million ICD-10 diagnosis records from 1.6 million children at a single pediatric institution, with each visit bracketed by special tokens and ages encoded as sinusoidal features. At the start of each visit, before any of that day's codes are revealed, the model outputs a probability distribution over the full ICD-10 vocabulary and an exponential rate for time to the next visit. Scoring only first occurrences against sex- and age-matched controls on held-out patients, the model achieves a median AUC of 0.72 across 797 conditions, and 0.793 for asthma and 0.847 for ADHD, outperforming all same-data baseline
What carries the argument
The core mechanism is a visit-boundary scoring protocol combined with a joint next-code and time-to-event head: same-day diagnosis codes are grouped into visits delimited by VISIT_START and VISIT_END tokens, predictions are read at VISIT_START before any same-day codes are visible, and the softmax of the output logits at that position gives the probability that the next event is a particular ICD-10 code. This design eliminates within-encounter leakage and puts TEDDY on the same scoring axis as baselines and large language models.
Load-bearing premise
The evaluation assumes that a child who never has a particular ICD-10 code recorded at the single pediatric institution truly never received that diagnosis; because children often receive care outside the system, some control patients may have been diagnosed elsewhere, which would inflate every reported AUC.
What would settle it
Link the same cohort to statewide or claims-based records that capture out-of-system care and re-run the per-code AUCs after removing any control who appears with the target code elsewhere. If the median AUC drops substantially (for example, toward 0.5) or the rare-condition advantage disappears, the single-site capture assumption fails; if it holds, the model's discrimination is robust.
If this is right
- Pediatric diagnostic histories alone, without lab results, vitals, or medications, may be sufficient for meaningful incident-disease prediction.
- A few million parameters can beat far larger general-purpose models on pediatric forecasting tasks when longitudinal structure is carefully represented.
- Rare diseases, often the hardest to diagnose, appear to be among the most predictable, suggesting the model captures trajectory patterns specific to low-prevalence conditions.
- Predictions remain informative years before the recorded diagnosis, motivating possible use at routine pediatric visits to flag children at elevated risk.
Where Pith is reading between the lines
- Because the evaluation treats 'never recorded at the single center' as 'never diagnosed,' the reported AUCs are likely upper-bound estimates; linking to community records or claims data could lower them and would provide a sharper test of the model's true discrimination.
- The strong rare-code performance suggests that diagnostic trajectories may carry characteristic co-occurrence patterns (e.g., referrals, genetic tests, symptom codes) that could be leveraged to generate differential diagnoses during the diagnostic odyssey, though the paper does not directly demonstrate such a decision-support tool.
- If the visit-boundary protocol matters as much as the architecture, retrofitting existing EHR models with explicit visit tokens and first-occurrence evaluation may yield larger gains than scaling parameters—an implication the paper's baseline comparisons support indirectly.
- The timing head's miscalibration of long inter-visit intervals indicates that future models may need non-exponential survival heads or continuous-time machinery to become useful for scheduling and care planning.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. TEDDY is a 1.84M-parameter decoder-only transformer trained on 73M ICD-10 diagnosis events from 1.6M pediatric patients at Texas Children’s Hospital. The model represents each patient as a sequence of same-day visit brackets, uses a joint next-code and exponential time-to-visit objective, and is evaluated by scoring each code at the visit boundary before same-day codes are revealed. On 797 first-occurrence tasks spanning 16 ICD-10 chapters, the authors report a median AUC of 0.72, superiority over DenseNet, CNN, RNN, and LSTM baselines on 96–99% of codes, and focused asthma/ADHD AUCs of 0.793/0.847. The paper also reports that discrimination remains above chance more than two years before diagnosis and that visit-timing predictions have a 3.0-day mean absolute RMST error, with acknowledged tail miscalibration.
Significance. The evaluation design is a genuine strength: visit-boundary scoring removes within-encounter leakage, first-occurrence scoring matches clinical intent, sex/age matching reduces demographic confounding, and patient-level bootstrap CIs are provided. If the headline result survives adjustment for the control-sampling issues below, it would be a valuable demonstration that a compact generative model can deliver broad and rare-disease pediatric forecasting from a single institution, contrasting with the population-scale adult models in the literature. The open-source pipeline and honest reporting of limitations also strengthen the contribution.
major comments (3)
- [Methods: Per-code AUC] Control sampling does not match history length. The text states 'a control is a VISIT_START from a patient who never has code c recorded anywhere in their trajectory' and 'Each patient contributes a single such position.' Cases are, by construction, scored at the visit immediately preceding the first recorded occurrence of c, which tends to occur later in a child's TCH record. If controls are drawn from all visits of never-c patients, their prior history is on average shorter, letting the model separate cases from controls by chart length rather than clinical signal. The focused asthma/ADHD benchmark (Fig. 3) matches history length and reports a length-only AUROC of about 0.51, but this safeguard is not applied to the 797-code pipeline, the rare/common comparison (Fig. 2D–E), or the future-horizon analysis. Please add history-length matching (or a length-only AUROC and adjusted analysis)
- [Methods: Per-code AUC] The 'never' control definition assumes that absence of code c anywhere in the TCH record means the child never received that diagnosis. The Introduction itself notes that children receive much care at community providers whose records do not reach TCH. Some 'never' controls may therefore have been diagnosed elsewhere. The direction of the resulting bias is not clear a priori, but the assumption is load-bearing for every reported AUC. Please provide a sensitivity analysis, e.g., restricting controls to children with a minimum number of TCH visits or with a later well-child visit to the center, or otherwise quantify the possible effect of incomplete out-of-network capture.
- [Future-horizon discrimination] The same control-sampling concern applies to the future-horizon analysis (Methods: Future-horizon discrimination) and to the rare-code bootstrap claim in Supplementary Fig. S1. If the per-code estimator is revised to match history length, the horizon-band results and the '202/225 rarest above chance' statement should be recomputed or the text should explain why they are insensitive.
minor comments (5)
- [Figure 1A] Caption text 'The history of a patient with rett syndrome by ICD code' is awkward and seems to duplicate the preceding line; rephrase for clarity.
- [Throughout] The paper uses both 'AUC' and 'AUROC' for the same quantity; consider standardizing the terminology.
- [Methods: Temporal calibration] The 3.0-day mean absolute RMST error is a decile-weighted average, not a per-position average; the abstract could make this explicit.
- [Table 1] Verify the 'ReClaim' row and that the model name is consistent with the reference and with the main text.
- [Abstract / Discussion] The term 'foundation model' is used; the paper might briefly justify why a 1.84M-parameter single-institution model qualifies as a foundation model under the usual definition.
Circularity Check
No significant circularity: TEDDY's predictions are genuine out-of-sample scores and the central derivation is self-contained.
full rationale
The paper's central derivation is self-contained and does not reduce to its inputs by construction. The model is trained on a 90/5/5 patient-level split, and all headline AUCs are computed on the held-out test split. The per-code AUC reads the model's next-token probability at a VISIT_START position before any same-day codes are revealed, with cases defined as first-occurrence visits and controls as patients who never carry the code, reweighted for sex and age. No test labels are used to set parameters, and the same-data baselines are trained on the same canonical dataset and evaluated through the same pipeline. The NO_EVENT gap-fill cadence and the exponential time head are modeling choices that shape the training objective, not fitted to test outcomes. The paper explicitly acknowledges its limitations—single-institution data, community-care truncation, and the fact that history-length matching is applied only in the focused benchmark—but these are evaluation-validity concerns, not circular reductions. The Delphi-based architecture is an external, openly credited framework, not a self-citation chain. No step in the derivation is equivalent to its inputs by definition or by fitting, so the appropriate circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (3)
- Rare/common prevalence cutoff =
5e-4 (5 in 10,000 children)
- Minimum case/control threshold for including a code =
20 cases and 20 controls
- AAP gap-fill intervals =
15, 45, 135, 270, 365 days by age bucket
axioms (4)
- domain assumption Children who never have a diagnosis code in the TCH record are true non-cases for that diagnosis.
- domain assumption Same-day ICD-10 codes carry no internal order, so scoring at VISIT_START is leakage-free.
- domain assumption An exponential time-to-event distribution with one rate per position is an adequate model for visit timing.
- domain assumption ICD-10 diagnosis codes as recorded in the EHR adequately reflect the child's true disease state.
invented entities (2)
-
NO_EVENT token
no independent evidence
-
VISIT_START and VISIT_END bracket tokens
no independent evidence
read the original abstract
Pediatric electronic health records capture developmentally structured clinical trajectories, yet their potential for generative healthcare foundation models remains largely unexplored. Here we present TEDDY (Temporal Event Decoder for Disease in Youth), a 1.84-million-parameter decoder transformer trained on approximately 73 million ICD-10 diagnoses from 1.6 million children at a single pediatric institution. TEDDY models longitudinal diagnosis trajectories and visit timing. Predictions were made before visit codes were revealed, limited to first occurrences, and evaluated against sex- and age-matched controls. Across 797 disease-onset prediction tasks spanning 16 ICD-10 chapters, TEDDY achieved a median AUC of 72.0%, outperforming same-data DenseNet (50.0%), CNN (57.2%), RNN (60.1%), and LSTM (62.7%) baselines on 96-99% of tasks. Performance held across sex and age and was strongest among lower-prevalence diagnoses; 202 of the 225 rarest conditions (90%) had 95% confidence intervals above chance. Predictive signal remained detectable more than two years before first recorded diagnosis, with median AUCs of 59.7% in the unrestricted analysis and 64.4% in a fixed-cohort sensitivity analysis. In asthma and attention-deficit/hyperactivity disorder benchmarks, AUCs were 79.3% and 84.7%, compared with 62.7% and 71.7% for the strongest comparators, including a general-purpose language model three orders of magnitude larger. Visit-timing predictions had a 3.0-day mean absolute restricted mean survival-time error over 365 days, although median and long-tail return intervals remained miscalibrated. Together, these results establish pediatric diagnostic histories as a substrate for compact generative models supporting broad, rare-disease, and long-horizon risk forecasting without population-scale data or billion-parameter models.
Figures
Reference graph
Works this paper leans on
-
[1]
Developmental pharmacology: neonates are not just small adults
Karel Allegaert, René Verbesselt, Gunnar Naulaers, JN Van Den Anker, Maissa Rayyan, Anne Debeer, and Jan de Hoon. Developmental pharmacology: neonates are not just small adults. . . .Acta Clinica Belgica, 63(1):16–24, 2008
2008
-
[2]
Ruisong Wang, Xiaoman Ding, Wanyue Zhang, and Tieliu Shi. The transformative potential of artificial intelli- gence in pediatric medicine: Current applications, methodological challenges, and future directions.Pediatric Investigation, 2026
2026
-
[3]
Off-label and unapproved pediatric drug utilization: A meta-analysis.Experimental and Therapeutic Medicine, 28(5):412, 2024
Xingxing Yuan, Jiawei Gao, Liuxin Yang, Yurong Tan, and Ousman Bajinka. Off-label and unapproved pediatric drug utilization: A meta-analysis.Experimental and Therapeutic Medicine, 28(5):412, 2024
2024
-
[4]
Advancing pediatric medication safety using real-world data: current problems and potential solutions.Journal of hospital medicine, 18(9):865, 2023
James W Antoon, James A Feinstein, Jennifer L Goldman, Kathryn E Kyler, Samir S Shah, Children’s Hos- pital Association Pharmacoepidemiology, and Research Group. Advancing pediatric medication safety using real-world data: current problems and potential solutions.Journal of hospital medicine, 18(9):865, 2023
2023
-
[5]
Lack of children in public medical imaging data points to growing age bias in biomedical ai.medRxiv, 2025
Stanley Bryan Zamora Hua, Nicholas Heller, Ping He, Alexander J Towbin, Irene Y Chen, Alex X Lu, and Lauren Erdman. Lack of children in public medical imaging data points to growing age bias in biomedical ai.medRxiv, 2025
2025
-
[6]
Us fda approval of pediatric artificial intelligence and machine learning–enabled medical devices.JAMA pediatrics, 179(2):212–214, 2025
Ryan CL Brewster, Matthew Nagy, Susmitha Wunnava, and Florence T Bourgeois. Us fda approval of pediatric artificial intelligence and machine learning–enabled medical devices.JAMA pediatrics, 179(2):212–214, 2025
2025
-
[7]
Marjan van den Akker, Mirjam Dieckelmann, Mohammad Akhtar Hussain, Daniela Bond-Smith, Christiane Muth, Sanghamitra Pati, Sonia Saxena, Desiree Silva, Rachel Skoss, Leon Straker, et al. Children and adolescents are 15 APREPRINT- JULY17, 2026 not small adults: toward a better understanding of multimorbidity in younger populations.Journal of Clinical Epidem...
2026
-
[8]
The magnitude of multimorbidity in childhood: a global systematic review.Systematic Reviews, 2026
Somen Kumar Pradhan, Jogesh Murmu, Upasana Nayak, Abhinav Sinha, Marjan van den Akker, Moham- mad Akhtar Hussain, Krushna Chandra Sahoo, Debdutta Bhattacharya, Jaya Singh Kshatri, Durga Madhab Satapathy, et al. The magnitude of multimorbidity in childhood: a global systematic review.Systematic Reviews, 2026
2026
-
[9]
Lambert, Annie Olry, Charlotte Rodwell, Charlotte Gueydan, Valérie Lanneau, Daniel Murphy, Yann Le Cam, and Ana Rath
Stéphanie Nguengang Wakap, Deborah M. Lambert, Annie Olry, Charlotte Rodwell, Charlotte Gueydan, Valérie Lanneau, Daniel Murphy, Yann Le Cam, and Ana Rath. Estimating cumulative point prevalence of rare diseases: analysis of the Orphanet database.European Journal of Human Genetics, 28(2):165–173, 2020
2020
-
[10]
Hendriksz
Christian J. Hendriksz. Rare disease impact report: Insights from patients and the medical community. Technical report, Shire Human Genetic Therapies, April 2013
2013
-
[11]
Diagnostic delay in rare diseases: Data from the spanish rare diseases patient registry.Orphanet Journal of Rare Diseases, 17(1):418, 2022
Juan Benito-Lozano, Blanca López-Villalba, Greta Arias-Merino, Manuel Posada de la Paz, and Verónica Alonso- Ferreira. Diagnostic delay in rare diseases: Data from the spanish rare diseases patient registry.Orphanet Journal of Rare Diseases, 17(1):418, 2022
2022
-
[12]
The national economic burden of rare disease in the united states in 2019.Orphanet Journal of Rare Diseases, 17(1):163, 2022
Grace Yang, Inna Cintina, Anne Pariser, Elisabeth Oehrlein, Jamie Sullivan, and Annie Kennedy. The national economic burden of rare disease in the united states in 2019.Orphanet Journal of Rare Diseases, 17(1):163, 2022
2019
-
[13]
The cost of delayed diagnosis in rare disease: A health economic study: Full study report
The Lewin Group, Inc. The cost of delayed diagnosis in rare disease: A health economic study: Full study report. Technical report, EveryLife Foundation for Rare Diseases, September 2023
2023
-
[14]
Timeliness of diagnosis and treatment: the challenge of childhood cancers.British journal of cancer, 125(12):1612–1620, 2021
Callum JR Mullen, Ronald D Barr, and Eduardo L Franco. Timeliness of diagnosis and treatment: the challenge of childhood cancers.British journal of cancer, 125(12):1612–1620, 2021
2021
-
[15]
BEHRT: Transformer for electronic health records
Yikuan Li, Shishir Rao, Jose Roberto Ayala Solares, Abdelaali Hassaine, Dexter Canoy, Yajie Zhu, Farideh Rahimian, Gholamreza Salimi-Khorshidi, Kazem Khaleeli, Diletta Taddei, Davide Luciano, Azeem Majeed, Madelon Seewah Mah, Rohit Joshi, and Kazem Rahimi. BEHRT: Transformer for electronic health records. Scientific Reports, 10:7155, 2020
2020
-
[16]
Med-BERT: Pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction.npj Digital Medicine, 4:86, 2021
Laila Rasmy, Yang Xiang, Ziqian Xie, Cui Tao, and Degui Zhi. Med-BERT: Pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction.npj Digital Medicine, 4:86, 2021
2021
-
[17]
Kalluri, Matthew Spotnitz, Ruijun Chen, Erica A
Chao Pang, Xinzhuo Jiang, Krishna S. Kalluri, Matthew Spotnitz, Ruijun Chen, Erica A. V oss, and Karthik Natarajan. CEHR-BERT: Incorporating temporal information from structured EHR data to improve prediction tasks.arXiv preprint arXiv:2111.08585, 2021
Pith/arXiv arXiv 2021
-
[18]
Yikuan Li, Mohammad Mamouei, Gholamreza Salimi-Khorshidi, Shishir Rao, Abdelaali Hassaine, Dexter Canoy, Thomas Lukasiewicz, and Kazem Rahimi. Hi-BEHRT: Hierarchical transformer-based model for accurate prediction of clinical events using multimodal longitudinal electronic health records.IEEE Journal of Biomedical and Health Informatics, 27:1106–1117, 2023
2023
-
[19]
Learning the natural history of human disease with generative transformers
Artem Shmatko, Alexander Wolfgang Jung, Kumar Gaurav, Søren Brunak, Laust Hvas Mortensen, Ewan Birney, Tom Fitzgerald, and Moritz Gerstung. Learning the natural history of human disease with generative transformers. Nature, 647(8088):248–256, 2025
2025
-
[20]
Fan Ma, Yuntian Liu, Xiang Lan, Weipeng Zhou, Jun Ni, Mauro Giuffrè, Lingfei Qian, Xueqing Peng, Yujia Zhou, Ruey-Ling Weng, Huan He, Lu Li, Huiyuan Wang, Qingyu Chen, Andrew Loza, Laila Rasmy, Degui Zhi, Yuan Lu, Chenjie Zeng, Joshua C Denny, Lee Schwamm, Daniella Meeker, Lucila Ohno-Machado, Yong Chen, and Hua Xu. Foundation models to unlock real-world ...
Pith/arXiv arXiv 2026
-
[21]
Generative medical event models improve with scale.arXiv preprint arXiv:2508.12104, 2025
Shane Waxler, Paul Blazek, Davis White, Daniel Sneider, Kevin Chung, Mani Nagarathnam, Patrick Williams, Hank V oeller, Karen Wong, Matthew Swanhorst, Sheng Zhang, Naoto Usuyama, Cliff Wong, Tristan Naumann, Hoifung Poon, Andrew Loza, Daniella Meeker, Seth Hain, and Rahul Shah. Generative medical event models improve with scale.arXiv preprint arXiv:2508.1...
arXiv 2025
-
[22]
Yeung, Daniel Bean, James Teo, and Richard J
Zeljko Kraljevic, Joe A. Yeung, Daniel Bean, James Teo, and Richard J. Dobson. Foresight—a generative pretrained transformer for modelling of patient timelines using electronic health records: A retrospective modelling study.The Lancet Digital Health, 6:e281–e290, 2024
2024
-
[23]
MOTOR: A time-to-event foundation model for structured medical records
Ethan Steinberg, Jason Alan Fries, Yizhe Xu, and Nigam Shah. MOTOR: A time-to-event foundation model for structured medical records. InInternational Conference on Learning Representations (ICLR), 2024
2024
-
[24]
TransformEHR: Transformer-based encoder-decoder generative model to enhance prediction of disease outcomes using electronic health records
Zhichao Yang, Avijit Mitra, Weijie Liu, Dan Berlowitz, and Hong Yu. TransformEHR: Transformer-based encoder-decoder generative model to enhance prediction of disease outcomes using electronic health records. Nature Communications, 14:7857, 2023. 16 APREPRINT- JULY17, 2026
2023
-
[25]
Using sequences of life-events to predict human lives.Nature Computational Science, 4:43–56, 2024
Germans Savcisens, Tina Eliassi-Rad, Lars Kai Hansen, Laust Hvas Mortensen, Lau Lilleholt, Anna Rogers, Ingo Zettler, and Sune Lehmann. Using sequences of life-events to predict human lives.Nature Computational Science, 4:43–56, 2024
2024
-
[26]
Linwood, and Chang Liu
Xianlong Zeng, Simon L. Linwood, and Chang Liu. Pretrained transformer framework on pediatric claims data for population specific tasks.Scientific Reports, 12:3651, 2022
2022
-
[27]
Recommendations for preventive pediatric health care (pe- riodicity schedule)
Bright Futures/American Academy of Pediatrics. Recommendations for preventive pediatric health care (pe- riodicity schedule). Technical report, American Academy of Pediatrics, February 2025. Schedule reflects recommendations approved December 2024 and published February 2025
2025
-
[28]
Lipkin, Michelle M
Paul H. Lipkin, Michelle M. Macias, AAP Council on Children with Disabilities, and Section on Developmen- tal and Behavioral Pediatrics. Promoting optimal development: Identifying infants and young children with developmental disorders through developmental surveillance and screening.Pediatrics, 145(1):e20193449, 2020
2020
-
[29]
Wisk and Niraj Sharma
Lauren E. Wisk and Niraj Sharma. Prevalence and trends in pediatric-onset chronic conditions in the united states, 1999–2018.Academic Pediatrics, 25(4):102810, 2025
1999
-
[30]
UK Biobank: An open access resource for identifying the causes of a wide range of complex diseases of middle and old age.PLoS Medicine, 12(3):e1001779, 2015
Cathie Sudlow, John Gallacher, Naomi Allen, Valerie Beral, Paul Burton, John Danesh, Paul Downey, Paul Elliott, Jane Green, Martin Landray, Bette Liu, Paul Matthews, Giok Ong, Jill Pell, Alan Silman, Alan Young, Tim Sprosen, Tim Peakman, and Rory Collins. UK Biobank: An open access resource for identifying the causes of a wide range of complex diseases of...
2015
-
[31]
Alistair E. W. Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J. Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, Li-wei H. Lehman, Leo A. Celi, and Roger G. Mark. MIMIC-IV, a freely accessible electronic health record dataset.Scientific Data, 10(1):1, 2023
2023
-
[32]
Forrest, Peter A
Christopher B. Forrest, Peter A. Margolis, L. Charles Bailey, Keith Marsolo, Mark A. Del Beccaro, Jonathan A. Finkelstein, David E. Milov, Veronica J. Vieland, Bryan A. Wolf, Feliciano B. Yu, and Michael G. Kahn. PEDSnet: A national pediatric learning health system.Journal of the American Medical Informatics Association, 21(4):602–606, 2014
2014
-
[33]
Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, Gaël Liu, Francesco Visin, Kathleen Kenealy, Lucas Bey...
Pith/arXiv arXiv 2025
-
[34]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems, volume 30, pages 5998–6008. Curran Associates, Inc., 2017
2017
-
[35]
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. Technical report, OpenAI, 2018
2018
-
[36]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. InInternational Conference on Learning Representations (ICLR), 2019. 18 APREPRINT- JULY17, 2026 Supplementary Figures and Tables 0 50 100 150 200 Rare conditions, ranked by AUC 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0AUC (first occurrence) chance common-stratum median (0.71) 202/225 (90%) cle...
2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.