{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:76CK2P3OSRVGTSZ5JFVIZ6XWJG","short_pith_number":"pith:76CK2P3O","schema_version":"1.0","canonical_sha256":"ff84ad3f6e946a69cb3d496a8cfaf6499722ef31fd83a0636be80ac2eee010e8","source":{"kind":"arxiv","id":"2403.19114","version":1},"attestation_state":"computed","paper":{"title":"Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL","cs.LG","cs.PL"],"primary_cat":"cs.SE","authors_text":"Chunqiu Steven Xia, Lingming Zhang, Yinlin Deng","submitted_at":"2024-03-28T03:10:39Z","abstract_excerpt":"LLMs have become the go-to choice for code generation tasks, with an exponential increase in the training, development, and usage of LLMs specifically for code generation. To evaluate the ability of LLMs on code, both academic and industry practitioners rely on popular handcrafted benchmarks. However, prior benchmarks contain only a very limited set of problems, both in quantity and variety. Further, due to popularity and age, many benchmarks are prone to data leakage where example solutions can be readily found on the web and thus potentially in training data. Such limitations inevitably lead"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2403.19114","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.SE","submitted_at":"2024-03-28T03:10:39Z","cross_cats_sorted":["cs.CL","cs.LG","cs.PL"],"title_canon_sha256":"2b918a89a4c5ad2ad9d7eb18b47c11fb21af7ba1f567caf6a1a0a1165b12c6a9","abstract_canon_sha256":"ba320b0c33f3b9091c66df555bd6f356509f431fc339edeb2e29ad9eb694372b"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:01:44.707990Z","signature_b64":"PVipEWDycgWeWYFIz6X200lFUhMJ+SMBbGiN0EQutHaZFmvxoGkF3Edm798zxNkoJQW6ZwM7i5p0KfhpwUtzCw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"ff84ad3f6e946a69cb3d496a8cfaf6499722ef31fd83a0636be80ac2eee010e8","last_reissued_at":"2026-07-05T08:01:44.707468Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:01:44.707468Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL","cs.LG","cs.PL"],"primary_cat":"cs.SE","authors_text":"Chunqiu Steven Xia, Lingming Zhang, Yinlin Deng","submitted_at":"2024-03-28T03:10:39Z","abstract_excerpt":"LLMs have become the go-to choice for code generation tasks, with an exponential increase in the training, development, and usage of LLMs specifically for code generation. To evaluate the ability of LLMs on code, both academic and industry practitioners rely on popular handcrafted benchmarks. However, prior benchmarks contain only a very limited set of problems, both in quantity and variety. Further, due to popularity and age, many benchmarks are prone to data leakage where example solutions can be readily found on the web and thus potentially in training data. Such limitations inevitably lead"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2403.19114","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2403.19114/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2403.19114","created_at":"2026-07-05T08:01:44.707539+00:00"},{"alias_kind":"arxiv_version","alias_value":"2403.19114v1","created_at":"2026-07-05T08:01:44.707539+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2403.19114","created_at":"2026-07-05T08:01:44.707539+00:00"},{"alias_kind":"pith_short_12","alias_value":"76CK2P3OSRVG","created_at":"2026-07-05T08:01:44.707539+00:00"},{"alias_kind":"pith_short_16","alias_value":"76CK2P3OSRVGTSZ5","created_at":"2026-07-05T08:01:44.707539+00:00"},{"alias_kind":"pith_short_8","alias_value":"76CK2P3O","created_at":"2026-07-05T08:01:44.707539+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.28998","citing_title":"Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2406.04244","citing_title":"Benchmark Data Contamination of Large Language Models: A Survey","ref_index":166,"is_internal_anchor":false},{"citing_arxiv_id":"2603.00989","citing_title":"Sustainable Code Generation Using Large Language Models: A Systematic Literature Review","ref_index":144,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14445","citing_title":"FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2604.17308","citing_title":"SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents","ref_index":36,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/76CK2P3OSRVGTSZ5JFVIZ6XWJG","json":"https://pith.science/pith/76CK2P3OSRVGTSZ5JFVIZ6XWJG.json","graph_json":"https://pith.science/api/pith-number/76CK2P3OSRVGTSZ5JFVIZ6XWJG/graph.json","events_json":"https://pith.science/api/pith-number/76CK2P3OSRVGTSZ5JFVIZ6XWJG/events.json","paper":"https://pith.science/paper/76CK2P3O"},"agent_actions":{"view_html":"https://pith.science/pith/76CK2P3OSRVGTSZ5JFVIZ6XWJG","download_json":"https://pith.science/pith/76CK2P3OSRVGTSZ5JFVIZ6XWJG.json","view_paper":"https://pith.science/paper/76CK2P3O","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2403.19114&json=true","fetch_graph":"https://pith.science/api/pith-number/76CK2P3OSRVGTSZ5JFVIZ6XWJG/graph.json","fetch_events":"https://pith.science/api/pith-number/76CK2P3OSRVGTSZ5JFVIZ6XWJG/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/76CK2P3OSRVGTSZ5JFVIZ6XWJG/action/timestamp_anchor","attest_storage":"https://pith.science/pith/76CK2P3OSRVGTSZ5JFVIZ6XWJG/action/storage_attestation","attest_author":"https://pith.science/pith/76CK2P3OSRVGTSZ5JFVIZ6XWJG/action/author_attestation","sign_citation":"https://pith.science/pith/76CK2P3OSRVGTSZ5JFVIZ6XWJG/action/citation_signature","submit_replication":"https://pith.science/pith/76CK2P3OSRVGTSZ5JFVIZ6XWJG/action/replication_record"}},"created_at":"2026-07-05T08:01:44.707539+00:00","updated_at":"2026-07-05T08:01:44.707539+00:00"}