{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2022:4PQYXXNT3XJLFDYBVQC2WNCDH6","short_pith_number":"pith:4PQYXXNT","schema_version":"1.0","canonical_sha256":"e3e18bddb3ddd2b28f01ac05ab34433faca427f4e0532cbe6708657207dca654","source":{"kind":"arxiv","id":"2211.09110","version":2},"attestation_state":"computed","paper":{"title":"Holistic Evaluation of Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Language models are now densely benchmarked on the same 42 scenarios and 7 metrics under standardized conditions for all 30 models evaluated.","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CL","authors_text":"Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher R\\'e, Deepak Narayanan, Diana Acosta-Navas, Dilara Soylu, Dimitris Tsipras, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Michihiro Yasunaga, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri Chatterji, Omar Khattab, Percy Liang, Peter Henderson, Qian Huang, Rishi Bommasani, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Tony Lee, Vishrav Chaudhary, William Wang, Xuechen Li, Yian Zhang, Yifan Mai, Yuhuai Wu, Yuhui Zhang, Yuta Koreeda","submitted_at":"2022-11-16T18:51:34Z","abstract_excerpt":"Language models (LMs) are becoming the foundation for almost all major language technologies, but their capabilities, limitations, and risks are not well understood. We present Holistic Evaluation of Language Models (HELM) to improve the transparency of language models. First, we taxonomize the vast space of potential scenarios (i.e. use cases) and metrics (i.e. desiderata) that are of interest for LMs. Then we select a broad subset based on coverage and feasibility, noting what's missing or underrepresented (e.g. question answering for neglected English dialects, metrics for trustworthiness)."},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":false},"canonical_record":{"source":{"id":"2211.09110","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2022-11-16T18:51:34Z","cross_cats_sorted":["cs.AI","cs.LG"],"title_canon_sha256":"8b61ecb21f40f5050d219aa6879d32a1f08e0e2a55ff5e7ad0fa2e141ccbc56c","abstract_canon_sha256":"5e8efe20a353e260589415162a56e5c0872f9e62d4f58f9e1aa97f12c642bd32"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T06:55:50.810673Z","signature_b64":"0S+8V/4T0dxSNxOqNMQpkNmJUpTEorOjsR0X/BFyqmzVvw/jwfovK0470Vbvf207nSpvrou+9ozZKhapENfKAw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"e3e18bddb3ddd2b28f01ac05ab34433faca427f4e0532cbe6708657207dca654","last_reissued_at":"2026-07-05T06:55:50.810187Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T06:55:50.810187Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Holistic Evaluation of Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Language models are now densely benchmarked on the same 42 scenarios and 7 metrics under standardized conditions for all 30 models evaluated.","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CL","authors_text":"Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher R\\'e, Deepak Narayanan, Diana Acosta-Navas, Dilara Soylu, Dimitris Tsipras, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Michihiro Yasunaga, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri Chatterji, Omar Khattab, Percy Liang, Peter Henderson, Qian Huang, Rishi Bommasani, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Tony Lee, Vishrav Chaudhary, William Wang, Xuechen Li, Yian Zhang, Yifan Mai, Yuhuai Wu, Yuhui Zhang, Yuta Koreeda","submitted_at":"2022-11-16T18:51:34Z","abstract_excerpt":"Language models (LMs) are becoming the foundation for almost all major language technologies, but their capabilities, limitations, and risks are not well understood. We present Holistic Evaluation of Language Models (HELM) to improve the transparency of language models. First, we taxonomize the vast space of potential scenarios (i.e. use cases) and metrics (i.e. desiderata) that are of interest for LMs. Then we select a broad subset based on coverage and feasibility, noting what's missing or underrepresented (e.g. question answering for neglected English dialects, metrics for trustworthiness)."},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"We improve this to 96.0%: now all 30 models have been densely benchmarked on the same core scenarios and metrics under standardized conditions. Our evaluation surfaces 25 top-level findings.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"The selection of a broad but feasible subset of scenarios and metrics from the full taxonomy is sufficient to deliver a holistic view, even while the paper explicitly notes missing or underrepresented areas such as question answering for neglected English dialects and metrics for trustworthiness.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"HELM establishes a multi-metric evaluation covering 30 language models on 42 scenarios (16 core) to raise average scenario coverage from 17.9% to 96% under uniform conditions while releasing all prompts, completions, and a toolkit.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Language models are now densely benchmarked on the same 42 scenarios and 7 metrics under standardized conditions for all 30 models evaluated.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"6a70928f421528a515b1171351ce455e739868126c73acce62221e18b8b435da"},"source":{"id":"2211.09110","kind":"arxiv","version":2},"verdict":{"id":"f9f25deb-e0b1-4da7-be48-d5403e3b1731","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-24T10:04:27.979965Z","strongest_claim":"We improve this to 96.0%: now all 30 models have been densely benchmarked on the same core scenarios and metrics under standardized conditions. Our evaluation surfaces 25 top-level findings.","one_line_summary":"HELM establishes a multi-metric evaluation covering 30 language models on 42 scenarios (16 core) to raise average scenario coverage from 17.9% to 96% under uniform conditions while releasing all prompts, completions, and a toolkit.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"The selection of a broad but feasible subset of scenarios and metrics from the full taxonomy is sufficient to deliver a holistic view, even while the paper explicitly notes missing or underrepresented areas such as question answering for neglected English dialects and metrics for trustworthiness.","pith_extraction_headline":"Language models are now densely benchmarked on the same 42 scenarios and 7 metrics under standardized conditions for all 30 models evaluated."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2211.09110/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":21,"sample":[{"doi":"10.18653/v1/2021.naacl-main.385","year":2021,"title":"Language Models are Few-Shot Learners","work_id":"214732c0-2edd-44a0-af9e-28184a2b8279","ref_index":1,"cited_arxiv_id":"2005.14165","is_internal_anchor":true},{"doi":"10.18653/v1/2021.acl-long.150","year":2021,"title":"doi: 10.18653/v1/2021.acl-long.150","work_id":"28ca0026-3906-4281-9006-088b556137d2","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"10.5281/zenodo.4761960","year":2018,"title":"URLhttps://glottolog.org/accessed2021-08-08","work_id":"b03a94d5-21f1-4c7f-8e70-54e70f8cd4db","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"10.18653/v1/2021.eacl-main.225","year":2021,"title":"Measuring Coding Challenge Competence With APPS","work_id":"c014c12f-1080-4cb2-ae03-ab6b7c09445c","ref_index":4,"cited_arxiv_id":"2105.09938","is_internal_anchor":true},{"doi":"10.1093/oxfordhb/9780199286546.001.0001/","year":2021,"title":"In Christopher Hitchcock & Alan Hajek, edi- tors: Oxford Handbook of Probability and Philosophy , Oxford University Press, pp","work_id":"14395009-699b-40c7-94de-6004ef131037","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":21,"snapshot_sha256":"acf92c8ccb24de906fa513991324cadf28c42956427d1f31fc748d008597ba99","internal_anchors":7},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2211.09110","created_at":"2026-07-05T06:55:50.810272+00:00"},{"alias_kind":"arxiv_version","alias_value":"2211.09110v2","created_at":"2026-07-05T06:55:50.810272+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2211.09110","created_at":"2026-07-05T06:55:50.810272+00:00"},{"alias_kind":"pith_short_12","alias_value":"4PQYXXNT3XJL","created_at":"2026-07-05T06:55:50.810272+00:00"},{"alias_kind":"pith_short_16","alias_value":"4PQYXXNT3XJLFDYB","created_at":"2026-07-05T06:55:50.810272+00:00"},{"alias_kind":"pith_short_8","alias_value":"4PQYXXNT","created_at":"2026-07-05T06:55:50.810272+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":185,"internal_anchor_count":181,"sample":[{"citing_arxiv_id":"2607.08028","citing_title":"From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents","ref_index":62,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25476","citing_title":"A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation","ref_index":82,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24391","citing_title":"Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War","ref_index":4,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24888","citing_title":"DiffusionBench: On Holistic Evaluation of Diffusion Transformers","ref_index":206,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26429","citing_title":"DualEval: Joint Model-Item Calibration for Unified LLM Evaluation","ref_index":19,"is_internal_anchor":true},{"citing_arxiv_id":"2606.27047","citing_title":"NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models","ref_index":17,"is_internal_anchor":true},{"citing_arxiv_id":"2606.27291","citing_title":"Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search","ref_index":18,"is_internal_anchor":true},{"citing_arxiv_id":"2606.27274","citing_title":"BetXplain: An Explanation-Annotated Dataset for Detecting Manipulative Betting Advertisements on Social Media","ref_index":158,"is_internal_anchor":true},{"citing_arxiv_id":"2606.23654","citing_title":"EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions","ref_index":6,"is_internal_anchor":true},{"citing_arxiv_id":"2606.22610","citing_title":"PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement","ref_index":74,"is_internal_anchor":true},{"citing_arxiv_id":"2606.21937","citing_title":"Latent Confidence Alignment for LLM Self-Assessment","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2606.21123","citing_title":"A Multi-Agent Audit Framework for High-Stakes Reasoning: Evaluation and Interpretability in Clinical Mental Health Screening","ref_index":24,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19704","citing_title":"Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents","ref_index":25,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19057","citing_title":"Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning","ref_index":9,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19501","citing_title":"DeXposure-Claw: An Agentic System for DeFi Risk Supervision","ref_index":65,"is_internal_anchor":true},{"citing_arxiv_id":"2607.01254","citing_title":"The Benchmark Ceiling: Human Judgment, Evaluation Scarcity, and the Political Economy of AI Capability Measurement","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"2606.17383","citing_title":"Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation","ref_index":23,"is_internal_anchor":true},{"citing_arxiv_id":"2606.18532","citing_title":"AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework","ref_index":79,"is_internal_anchor":true},{"citing_arxiv_id":"2607.01152","citing_title":"AGC-Bench: Measuring Artificial General Creativity","ref_index":40,"is_internal_anchor":true},{"citing_arxiv_id":"2606.17574","citing_title":"DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack","ref_index":21,"is_internal_anchor":true},{"citing_arxiv_id":"2606.13111","citing_title":"M\\\"OVE: A Holistic LLM Benchmark for the German Public Sector","ref_index":10,"is_internal_anchor":true},{"citing_arxiv_id":"2607.01740","citing_title":"Meta-Benchmarks for Financial-Services LLM Evaluation","ref_index":13,"is_internal_anchor":true},{"citing_arxiv_id":"2606.18285","citing_title":"RELIANCE: Curating and Evaluating Reproductive Health Information on Social Media","ref_index":27,"is_internal_anchor":true},{"citing_arxiv_id":"2606.09809","citing_title":"Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting","ref_index":64,"is_internal_anchor":true},{"citing_arxiv_id":"2606.28369","citing_title":"Multimodal and Multiscale Spatial-Temporal Semantic Search and Recommendation with AI Foundation Models","ref_index":29,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/4PQYXXNT3XJLFDYBVQC2WNCDH6","json":"https://pith.science/pith/4PQYXXNT3XJLFDYBVQC2WNCDH6.json","graph_json":"https://pith.science/api/pith-number/4PQYXXNT3XJLFDYBVQC2WNCDH6/graph.json","events_json":"https://pith.science/api/pith-number/4PQYXXNT3XJLFDYBVQC2WNCDH6/events.json","paper":"https://pith.science/paper/4PQYXXNT"},"agent_actions":{"view_html":"https://pith.science/pith/4PQYXXNT3XJLFDYBVQC2WNCDH6","download_json":"https://pith.science/pith/4PQYXXNT3XJLFDYBVQC2WNCDH6.json","view_paper":"https://pith.science/paper/4PQYXXNT","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2211.09110&json=true","fetch_graph":"https://pith.science/api/pith-number/4PQYXXNT3XJLFDYBVQC2WNCDH6/graph.json","fetch_events":"https://pith.science/api/pith-number/4PQYXXNT3XJLFDYBVQC2WNCDH6/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/4PQYXXNT3XJLFDYBVQC2WNCDH6/action/timestamp_anchor","attest_storage":"https://pith.science/pith/4PQYXXNT3XJLFDYBVQC2WNCDH6/action/storage_attestation","attest_author":"https://pith.science/pith/4PQYXXNT3XJLFDYBVQC2WNCDH6/action/author_attestation","sign_citation":"https://pith.science/pith/4PQYXXNT3XJLFDYBVQC2WNCDH6/action/citation_signature","submit_replication":"https://pith.science/pith/4PQYXXNT3XJLFDYBVQC2WNCDH6/action/replication_record"}},"created_at":"2026-07-05T06:55:50.810272+00:00","updated_at":"2026-07-05T06:55:50.810272+00:00"}