{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:DXOIU5VTEFEPZCHODNQ5RNYOG2","short_pith_number":"pith:DXOIU5VT","schema_version":"1.0","canonical_sha256":"1ddc8a76b32148fc88ee1b61d8b70e368740a3c8196e216b97f215013e863ba9","source":{"kind":"arxiv","id":"2501.10711","version":5},"attestation_state":"computed","paper":{"title":"Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.SE","authors_text":"Boxi Yu, Chaozheng Wang, Jialun Cao, Michael R. Lyu, Mingwei Liu, Pinjia He, Ruixi Qiao, Shing-Chi Cheung, Shuai Wang, Shuqing Li, Wenxuan Wang, Yuk-Kit Chan, Yuting Han, Zibin Zheng, Zixuan Ling","submitted_at":"2025-01-18T09:51:57Z","abstract_excerpt":"Code-related benchmarks play a critical role in evaluating large language models (LLMs), yet their quality fundamentally shapes how the community interprets model capabilities. In the past few years, awareness of benchmark quality has grown. Yet, after a decade-scale (2014-2025) survey over 672 code benchmarks, we observed a lag between growing awareness and actual practice. For example, in 2025 alone, the number of benchmarks that ignore code coverage when providing test cases nearly matches the total count accumulated across the previous ten years. In response, we take a clear position: Code"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2501.10711","kind":"arxiv","version":5},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.SE","submitted_at":"2025-01-18T09:51:57Z","cross_cats_sorted":["cs.AI","cs.CL"],"title_canon_sha256":"0cc7a9eecf6bc0db96d6167e710f1bf6df3958a9de179fc95c3826f1381fd20b","abstract_canon_sha256":"7db5601ed86e13add943dad5d203c8bd59e182fea143bc291ca494df015ced74"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-07T02:19:37.250235Z","signature_b64":"+AorFA71Sj21iFRhKEiD0CXa5BLm/CUhRaen7S/xy9hllAqTjLqGyXR6HjGyaqDLR4TYvyjZVgx8P9VCkPNADw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"1ddc8a76b32148fc88ee1b61d8b70e368740a3c8196e216b97f215013e863ba9","last_reissued_at":"2026-07-07T02:19:37.249459Z","signature_status":"signed_v1","first_computed_at":"2026-07-07T02:19:37.249459Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.SE","authors_text":"Boxi Yu, Chaozheng Wang, Jialun Cao, Michael R. Lyu, Mingwei Liu, Pinjia He, Ruixi Qiao, Shing-Chi Cheung, Shuai Wang, Shuqing Li, Wenxuan Wang, Yuk-Kit Chan, Yuting Han, Zibin Zheng, Zixuan Ling","submitted_at":"2025-01-18T09:51:57Z","abstract_excerpt":"Code-related benchmarks play a critical role in evaluating large language models (LLMs), yet their quality fundamentally shapes how the community interprets model capabilities. In the past few years, awareness of benchmark quality has grown. Yet, after a decade-scale (2014-2025) survey over 672 code benchmarks, we observed a lag between growing awareness and actual practice. For example, in 2025 alone, the number of benchmarks that ignore code coverage when providing test cases nearly matches the total count accumulated across the previous ten years. In response, we take a clear position: Code"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2501.10711","kind":"arxiv","version":5},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2501.10711/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2501.10711","created_at":"2026-07-07T02:19:37.249563+00:00"},{"alias_kind":"arxiv_version","alias_value":"2501.10711v5","created_at":"2026-07-07T02:19:37.249563+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2501.10711","created_at":"2026-07-07T02:19:37.249563+00:00"},{"alias_kind":"pith_short_12","alias_value":"DXOIU5VTEFEP","created_at":"2026-07-07T02:19:37.249563+00:00"},{"alias_kind":"pith_short_16","alias_value":"DXOIU5VTEFEPZCHO","created_at":"2026-07-07T02:19:37.249563+00:00"},{"alias_kind":"pith_short_8","alias_value":"DXOIU5VT","created_at":"2026-07-07T02:19:37.249563+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":8,"internal_anchor_count":8,"sample":[{"citing_arxiv_id":"2606.11166","citing_title":"Flaws in the LLM Automation Narrative","ref_index":24,"is_internal_anchor":true},{"citing_arxiv_id":"2607.00062","citing_title":"AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation","ref_index":4,"is_internal_anchor":true},{"citing_arxiv_id":"2508.15503","citing_title":"Guidelines for Empirical Studies in Software Engineering involving Large Language Models","ref_index":19,"is_internal_anchor":true},{"citing_arxiv_id":"2506.03535","citing_title":"Across Programming Language Silos: A Study on Cross-Lingual Retrieval-augmented Code Generation","ref_index":23,"is_internal_anchor":true},{"citing_arxiv_id":"2508.15503","citing_title":"Guidelines for Empirical Studies in Software Engineering involving Large Language Models","ref_index":19,"is_internal_anchor":true},{"citing_arxiv_id":"2512.00380","citing_title":"Knowledge-Graph-Driven Data Synthesis for Low-Resource Software Development: A HarmonyOS Case Study","ref_index":5,"is_internal_anchor":true},{"citing_arxiv_id":"2604.05100","citing_title":"Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks","ref_index":4,"is_internal_anchor":true},{"citing_arxiv_id":"2508.04325","citing_title":"Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models","ref_index":2,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/DXOIU5VTEFEPZCHODNQ5RNYOG2","json":"https://pith.science/pith/DXOIU5VTEFEPZCHODNQ5RNYOG2.json","graph_json":"https://pith.science/api/pith-number/DXOIU5VTEFEPZCHODNQ5RNYOG2/graph.json","events_json":"https://pith.science/api/pith-number/DXOIU5VTEFEPZCHODNQ5RNYOG2/events.json","paper":"https://pith.science/paper/DXOIU5VT"},"agent_actions":{"view_html":"https://pith.science/pith/DXOIU5VTEFEPZCHODNQ5RNYOG2","download_json":"https://pith.science/pith/DXOIU5VTEFEPZCHODNQ5RNYOG2.json","view_paper":"https://pith.science/paper/DXOIU5VT","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2501.10711&json=true","fetch_graph":"https://pith.science/api/pith-number/DXOIU5VTEFEPZCHODNQ5RNYOG2/graph.json","fetch_events":"https://pith.science/api/pith-number/DXOIU5VTEFEPZCHODNQ5RNYOG2/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/DXOIU5VTEFEPZCHODNQ5RNYOG2/action/timestamp_anchor","attest_storage":"https://pith.science/pith/DXOIU5VTEFEPZCHODNQ5RNYOG2/action/storage_attestation","attest_author":"https://pith.science/pith/DXOIU5VTEFEPZCHODNQ5RNYOG2/action/author_attestation","sign_citation":"https://pith.science/pith/DXOIU5VTEFEPZCHODNQ5RNYOG2/action/citation_signature","submit_replication":"https://pith.science/pith/DXOIU5VTEFEPZCHODNQ5RNYOG2/action/replication_record"}},"created_at":"2026-07-07T02:19:37.249563+00:00","updated_at":"2026-07-07T02:19:37.249563+00:00"}