{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2020:TKIRAZBVTYELJNONKSDX3JDTLT","short_pith_number":"pith:TKIRAZBV","schema_version":"1.0","canonical_sha256":"9a911064359e08b4b5cd54877da4735ce388602f339e52db330a01e098429ed4","source":{"kind":"arxiv","id":"2011.03395","version":2},"attestation_state":"computed","paper":{"title":"Underspecification Presents Challenges for Credibility in Modern Machine Learning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["stat.ML"],"primary_cat":"cs.LG","authors_text":"Akinori Mitani, Alan Karthikesalingam, Alexander D'Amour, Alex Beutel, Andrea Montanari, Babak Alipanahi, Ben Adlam, Christina Chen, Christopher Nielson, Cory McLean, Dan Moldovan, Diana Mincu, D. Sculley, Farhad Hormozdiari, Ghassen Jerfel, Harini Suresh, Jacob Eisenstein, Jessica Schrouff, Jonathan Deaton, Katherine Heller, Kellie Webster, Kim Ramasamy, Mario Lucic, Martin Seneviratne, Matthew D. Hoffman, Max Vladymyrov, Neil Houlsby, Rajiv Raman, Rory Sayres, Shannon Sequeira, Shaobo Hou, Steve Yadlowsky, Taedong Yun, Thomas F. Osborne, Victor Veitch, Vivek Natarajan, Xiaohua Zhai, Xuezhi Wang, Yian Ma, Zachary Nado","submitted_at":"2020-11-06T14:53:13Z","abstract_excerpt":"ML models often exhibit unexpectedly poor behavior when they are deployed in real-world domains. We identify underspecification as a key reason for these failures. An ML pipeline is underspecified when it can return many predictors with equivalently strong held-out performance in the training domain. Underspecification is common in modern ML pipelines, such as those based on deep learning. Predictors returned by underspecified pipelines are often treated as equivalent based on their training domain performance, but we show here that such predictors can behave very differently in deployment dom"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2011.03395","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2020-11-06T14:53:13Z","cross_cats_sorted":["stat.ML"],"title_canon_sha256":"70018e6fc56f6cf01fae71bbceec9f81941e292fbfd211899e4c056d4b4037fb","abstract_canon_sha256":"952f73427057b3c42080e341dfe405fbd1ebcfe07204649779f9a430555a61bd"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T01:54:31.995699Z","signature_b64":"tuH0Pl+jayoDxWr72WIvqSAC7iS9I3IaH5kpWsafJrCyPWV7rQ5y4qng7ao/WpEkUMycOQJQ6F/ZnbYHIPlhCg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"9a911064359e08b4b5cd54877da4735ce388602f339e52db330a01e098429ed4","last_reissued_at":"2026-07-05T01:54:31.995233Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T01:54:31.995233Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Underspecification Presents Challenges for Credibility in Modern Machine Learning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["stat.ML"],"primary_cat":"cs.LG","authors_text":"Akinori Mitani, Alan Karthikesalingam, Alexander D'Amour, Alex Beutel, Andrea Montanari, Babak Alipanahi, Ben Adlam, Christina Chen, Christopher Nielson, Cory McLean, Dan Moldovan, Diana Mincu, D. Sculley, Farhad Hormozdiari, Ghassen Jerfel, Harini Suresh, Jacob Eisenstein, Jessica Schrouff, Jonathan Deaton, Katherine Heller, Kellie Webster, Kim Ramasamy, Mario Lucic, Martin Seneviratne, Matthew D. Hoffman, Max Vladymyrov, Neil Houlsby, Rajiv Raman, Rory Sayres, Shannon Sequeira, Shaobo Hou, Steve Yadlowsky, Taedong Yun, Thomas F. Osborne, Victor Veitch, Vivek Natarajan, Xiaohua Zhai, Xuezhi Wang, Yian Ma, Zachary Nado","submitted_at":"2020-11-06T14:53:13Z","abstract_excerpt":"ML models often exhibit unexpectedly poor behavior when they are deployed in real-world domains. We identify underspecification as a key reason for these failures. An ML pipeline is underspecified when it can return many predictors with equivalently strong held-out performance in the training domain. Underspecification is common in modern ML pipelines, such as those based on deep learning. Predictors returned by underspecified pipelines are often treated as equivalent based on their training domain performance, but we show here that such predictors can behave very differently in deployment dom"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2011.03395","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2011.03395/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2011.03395","created_at":"2026-07-05T01:54:31.995291+00:00"},{"alias_kind":"arxiv_version","alias_value":"2011.03395v2","created_at":"2026-07-05T01:54:31.995291+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2011.03395","created_at":"2026-07-05T01:54:31.995291+00:00"},{"alias_kind":"pith_short_12","alias_value":"TKIRAZBVTYEL","created_at":"2026-07-05T01:54:31.995291+00:00"},{"alias_kind":"pith_short_16","alias_value":"TKIRAZBVTYELJNON","created_at":"2026-07-05T01:54:31.995291+00:00"},{"alias_kind":"pith_short_8","alias_value":"TKIRAZBV","created_at":"2026-07-05T01:54:31.995291+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":9,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.17582","citing_title":"Collaborative Large and Small Language Models for Accurate and Scalable Data Repair","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2606.12277","citing_title":"Finding Multiple Interpretations in Datasets","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2606.31589","citing_title":"From Failure to Alignment: A Requirements Engineering Framework for Machine Learning Systems","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09881","citing_title":"Toward Calibrated, Fair, and accurate Deepfake Detection","ref_index":121,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15183","citing_title":"When Are Two Networks the Same? Tensor Similarity for Mechanistic Interpretability","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13826","citing_title":"Reducing cross-sample prediction churn in scientific machine learning","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2605.04905","citing_title":"Cross-Model Consistency of Feature Importance in Electrospinning: Separating Robust from Model-Dependent Features","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08192","citing_title":"Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2605.04905","citing_title":"Cross-Model Consistency of Feature Importance in Electrospinning: Separating Robust from Model-Dependent Features","ref_index":14,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/TKIRAZBVTYELJNONKSDX3JDTLT","json":"https://pith.science/pith/TKIRAZBVTYELJNONKSDX3JDTLT.json","graph_json":"https://pith.science/api/pith-number/TKIRAZBVTYELJNONKSDX3JDTLT/graph.json","events_json":"https://pith.science/api/pith-number/TKIRAZBVTYELJNONKSDX3JDTLT/events.json","paper":"https://pith.science/paper/TKIRAZBV"},"agent_actions":{"view_html":"https://pith.science/pith/TKIRAZBVTYELJNONKSDX3JDTLT","download_json":"https://pith.science/pith/TKIRAZBVTYELJNONKSDX3JDTLT.json","view_paper":"https://pith.science/paper/TKIRAZBV","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2011.03395&json=true","fetch_graph":"https://pith.science/api/pith-number/TKIRAZBVTYELJNONKSDX3JDTLT/graph.json","fetch_events":"https://pith.science/api/pith-number/TKIRAZBVTYELJNONKSDX3JDTLT/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/TKIRAZBVTYELJNONKSDX3JDTLT/action/timestamp_anchor","attest_storage":"https://pith.science/pith/TKIRAZBVTYELJNONKSDX3JDTLT/action/storage_attestation","attest_author":"https://pith.science/pith/TKIRAZBVTYELJNONKSDX3JDTLT/action/author_attestation","sign_citation":"https://pith.science/pith/TKIRAZBVTYELJNONKSDX3JDTLT/action/citation_signature","submit_replication":"https://pith.science/pith/TKIRAZBVTYELJNONKSDX3JDTLT/action/replication_record"}},"created_at":"2026-07-05T01:54:31.995291+00:00","updated_at":"2026-07-05T01:54:31.995291+00:00"}