{"paper":{"title":"Underspecification Presents Challenges for Credibility in Modern Machine Learning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["stat.ML"],"primary_cat":"cs.LG","authors_text":"Akinori Mitani, Alan Karthikesalingam, Alexander D'Amour, Alex Beutel, Andrea Montanari, Babak Alipanahi, Ben Adlam, Christina Chen, Christopher Nielson, Cory McLean, Dan Moldovan, Diana Mincu, D. Sculley, Farhad Hormozdiari, Ghassen Jerfel, Harini Suresh, Jacob Eisenstein, Jessica Schrouff, Jonathan Deaton, Katherine Heller, Kellie Webster, Kim Ramasamy, Mario Lucic, Martin Seneviratne, Matthew D. Hoffman, Max Vladymyrov, Neil Houlsby, Rajiv Raman, Rory Sayres, Shannon Sequeira, Shaobo Hou, Steve Yadlowsky, Taedong Yun, Thomas F. Osborne, Victor Veitch, Vivek Natarajan, Xiaohua Zhai, Xuezhi Wang, Yian Ma, Zachary Nado","submitted_at":"2020-11-06T14:53:13Z","abstract_excerpt":"ML models often exhibit unexpectedly poor behavior when they are deployed in real-world domains. We identify underspecification as a key reason for these failures. An ML pipeline is underspecified when it can return many predictors with equivalently strong held-out performance in the training domain. Underspecification is common in modern ML pipelines, such as those based on deep learning. Predictors returned by underspecified pipelines are often treated as equivalent based on their training domain performance, but we show here that such predictors can behave very differently in deployment dom"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2011.03395","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2011.03395/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}