{"schema":"pith.reference-change-event.v1","doi":"10.1038/s41591-022-01772-9","canonical_url":"https://pith.science/event/10.1038/s41591-022-01772-9","json_url":"https://pith.science/event/10.1038/s41591-022-01772-9.json","not_a_judgment":"This page records that a citing paper's bibliography includes a work with a published notice. It is not a judgment on the citing paper.","primary":{"event_id":347141,"doi":"10.1038/s41591-022-01772-9","event_type":"correction","event_type_label":"Correction","source":"crossref","source_label":"Crossref","event_date":"2022-08-12","title":"Publisher Correction: Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI","work_title":"Nature Medicine 28, 924–933","work_doi":"10.1038/s41591-022-01772-9","work_arxiv_id":null,"notice_doi":"10.1038/s41591-022-01951-8","flag_count":0,"flags_open":0,"flags_disputed":0,"latest_flag_at":null,"human_href":"/event/10.1038/s41591-022-01772-9","json_href":"/event/10.1038/s41591-022-01772-9.json"},"events":[{"event_id":347141,"doi":"10.1038/s41591-022-01772-9","event_type":"correction","event_type_label":"Correction","source":"crossref","source_label":"Crossref","event_date":"2022-08-12","title":"Publisher Correction: Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI","work_title":"Nature Medicine 28, 924–933","work_doi":"10.1038/s41591-022-01772-9","work_arxiv_id":null,"notice_doi":"10.1038/s41591-022-01951-8","flag_count":0,"flags_open":0,"flags_disputed":0,"latest_flag_at":null,"human_href":"/event/10.1038/s41591-022-01772-9","json_href":"/event/10.1038/s41591-022-01772-9.json"}],"flags":[{"id":5628,"status":"open","status_label":"Open","citing_arxiv_id":"2605.04135","citing_title":"Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation","ref_index":5,"evidence_raw":"URL https://www.aisi.gov.uk/frontier-ai-trends-report. First public evidence-based assessment aggregating two years of AISI’s frontier model testing (November 2023 through October 2025); cited for the frontier-trajectory reframe of capability evaluation. Baptiste Vasey, Myura Nagendran, others, and DECIDE-AI Expert Group. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI.Nature Medicine, 28(5):924–933, 2022. doi: 10.1038/s41591-022-01772-9. Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In International Conference on Learning Representations (ICLR), 2023. Self-consistency gains of+6.4– +17.9pp on math / reasoning benchmarks; used to calibrate the sampling-axis chip conservatively for SWE-Bench-Verified pass@1. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems (NeurIPS), 2022. Companio","evidence_cleaned":null,"evidence_source_label":"bibliography line","event_type":"correction","event_type_label":"Correction","source_label":"Crossref","event_date":"2022-08-12","work_title":"Nature Medicine 28, 924–933","work_doi":"10.1038/s41591-022-01772-9","event_doi":"10.1038/s41591-022-01772-9","flag_href":"/flags/5628","event_href":"/event/10.1038/s41591-022-01772-9","paper_href":"/paper/2605.04135","created_at":"2026-07-11T03:19:11.601274Z","dispute_note":null,"disputed_at":null,"disputed_by":null},{"id":5630,"status":"open","status_label":"Open","citing_arxiv_id":"2605.04135","citing_title":"Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation","ref_index":5,"evidence_raw":"URL https://www.aisi.gov.uk/frontier-ai-trends-report. First public evidence-based assessment aggregating two years of AISI’s frontier model testing (November 2023 through October 2025); cited for the frontier-trajectory reframe of capability evaluation. Baptiste Vasey, Myura Nagendran, others, and DECIDE-AI Expert Group. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI.Nature Medicine, 28(5):924–933, 2022. doi: 10.1038/s41591-022-01772-9. Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In 43 International Conference on Learning Representations (ICLR), 2023. Self-consistency gains of+6.4– +17.9pp on math / reasoning benchmarks; used to calibrate the sampling-axis chip conservatively for SWE-Bench-Verified pass@1. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems (NeurIPS), 2022. Compa","evidence_cleaned":null,"evidence_source_label":"bibliography line","event_type":"correction","event_type_label":"Correction","source_label":"Crossref","event_date":"2022-08-12","work_title":"Nature Medicine 28, 924–933","work_doi":"10.1038/s41591-022-01772-9","event_doi":"10.1038/s41591-022-01772-9","flag_href":"/flags/5630","event_href":"/event/10.1038/s41591-022-01772-9","paper_href":"/paper/2605.04135","created_at":"2026-07-11T03:19:11.601274Z","dispute_note":null,"disputed_at":null,"disputed_by":null},{"id":5629,"status":"open","status_label":"Open","citing_arxiv_id":"2605.10601","citing_title":"The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime","ref_index":37,"evidence_raw":"Baptiste Vasey, Myura Nagendran, Bruce Campbell, David A. Clifton, Gary S. Collins, Spiros Denaxas, Alastair K. Denniston, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI.Nature Medicine, 28:924–933, 2022. doi: 10.1038/s41591-022-01772-9","evidence_cleaned":null,"evidence_source_label":"bibliography line","event_type":"correction","event_type_label":"Correction","source_label":"Crossref","event_date":"2022-08-12","work_title":"Nature Medicine 28, 924–933","work_doi":"10.1038/s41591-022-01772-9","event_doi":"10.1038/s41591-022-01772-9","flag_href":"/flags/5629","event_href":"/event/10.1038/s41591-022-01772-9","paper_href":"/paper/2605.10601","created_at":"2026-07-11T03:19:11.601274Z","dispute_note":null,"disputed_at":null,"disputed_by":null},{"id":5631,"status":"open","status_label":"Open","citing_arxiv_id":"2607.00019","citing_title":"LLMs in the Real World: Evaluating \"AI\" in Emergency Contexts","ref_index":80,"evidence_raw":"Baptiste Vasey, Myura Nagendran, Bruce Campbell, David A Clifton, Gary S Collins, Spiros Denaxas, Alastair K Denniston, Livia Faes, Bart Geerts, Mudathir Ibrahim, Xiaoxuan Liu, Bilal A Mateen, Piyush Mathur, Melissa D McCradden, Lauren Morgan, Johan Ordish, Chris Rogers, Suchi Saria, Daniel Shu Wei Ting, and 4 others. 2022. https://doi.org/10.1038/s41591-022-01772-9 Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI . Nature Medicine, 28(5):924--933","evidence_cleaned":null,"evidence_source_label":"bibliography line","event_type":"correction","event_type_label":"Correction","source_label":"Crossref","event_date":"2022-08-12","work_title":"Nature Medicine 28, 924–933","work_doi":"10.1038/s41591-022-01772-9","event_doi":"10.1038/s41591-022-01772-9","flag_href":"/flags/5631","event_href":"/event/10.1038/s41591-022-01772-9","paper_href":"/paper/2607.00019","created_at":"2026-07-11T03:19:11.601274Z","dispute_note":null,"disputed_at":null,"disputed_by":null},{"id":5632,"status":"open","status_label":"Open","citing_arxiv_id":"2607.05311","citing_title":"Deep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical Translation","ref_index":109,"evidence_raw":"Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: Decide-ai. Nature Medicine 28, 924–933. doi:10.1038/s41591-022-01772-9","evidence_cleaned":null,"evidence_source_label":"bibliography line","event_type":"correction","event_type_label":"Correction","source_label":"Crossref","event_date":"2022-08-12","work_title":"Nature Medicine 28, 924–933","work_doi":"10.1038/s41591-022-01772-9","event_doi":"10.1038/s41591-022-01772-9","flag_href":"/flags/5632","event_href":"/event/10.1038/s41591-022-01772-9","paper_href":"/paper/2607.05311","created_at":"2026-07-11T03:19:11.601274Z","dispute_note":null,"disputed_at":null,"disputed_by":null}],"flag_count":5,"flags_open":5,"flags_disputed":0,"desk_url":"https://pith.science/flags"}