{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:C7KH3B52IEHJVXZBBOVM5BED3Y","short_pith_number":"pith:C7KH3B52","schema_version":"1.0","canonical_sha256":"17d47d87ba410e9adf210baace8483de0e1daf44703876df741ab425810daffd","source":{"kind":"arxiv","id":"2506.11928","version":1},"attestation_state":"computed","paper":{"title":"LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL","cs.LG"],"primary_cat":"cs.SE","authors_text":"Aleksandra Korolova, Dongruixuan Li, Hangyi Hao, Hansen He, Jianzhu Yao, Jingbo Shang, Kaiyuan Liu, Peiyao Sheng, Peter Henderson, Pramod Viswanath, Saining Xie, Sanjeev Arora, Shang Zhou, Stanley Wei, Wenhao Chai, Zerui Cheng, Zeyu Shen, Zihan Zheng, Zixuan Wang","submitted_at":"2025-06-13T16:29:09Z","abstract_excerpt":"Recent reports claim that large language models (LLMs) now outperform elite humans in competitive programming. Drawing on knowledge from a group of medalists in international algorithmic contests, we revisit this claim, examining how LLMs differ from human experts and where limitations still remain. We introduce LiveCodeBench Pro, a benchmark composed of problems from Codeforces, ICPC, and IOI that are continuously updated to reduce the likelihood of data contamination. A team of Olympiad medalists annotates every problem for algorithmic categories and conducts a line-by-line analysis of faile"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2506.11928","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.SE","submitted_at":"2025-06-13T16:29:09Z","cross_cats_sorted":["cs.AI","cs.CL","cs.LG"],"title_canon_sha256":"1ca327ccc44be8efaa03c71d5af5c9046129ae010775fe4de4a9f547d8523db4","abstract_canon_sha256":"ccbe1ef3da8211938474b9d8b576e7459cf56f6e04bc86a29ae117211f613ab5"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:21:13.002763Z","signature_b64":"c7n/XBjzHe3Lgrl9lqRftBVel1l7oc0/USLl7/knghu4hBYpoe/tEaJ6a5Vi1n9sRo6avFQzbU3Znz7jrLROAw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"17d47d87ba410e9adf210baace8483de0e1daf44703876df741ab425810daffd","last_reissued_at":"2026-07-05T11:21:13.002273Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:21:13.002273Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL","cs.LG"],"primary_cat":"cs.SE","authors_text":"Aleksandra Korolova, Dongruixuan Li, Hangyi Hao, Hansen He, Jianzhu Yao, Jingbo Shang, Kaiyuan Liu, Peiyao Sheng, Peter Henderson, Pramod Viswanath, Saining Xie, Sanjeev Arora, Shang Zhou, Stanley Wei, Wenhao Chai, Zerui Cheng, Zeyu Shen, Zihan Zheng, Zixuan Wang","submitted_at":"2025-06-13T16:29:09Z","abstract_excerpt":"Recent reports claim that large language models (LLMs) now outperform elite humans in competitive programming. Drawing on knowledge from a group of medalists in international algorithmic contests, we revisit this claim, examining how LLMs differ from human experts and where limitations still remain. We introduce LiveCodeBench Pro, a benchmark composed of problems from Codeforces, ICPC, and IOI that are continuously updated to reduce the likelihood of data contamination. A team of Olympiad medalists annotates every problem for algorithmic categories and conducts a line-by-line analysis of faile"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2506.11928","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2506.11928/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2506.11928","created_at":"2026-07-05T11:21:13.002344+00:00"},{"alias_kind":"arxiv_version","alias_value":"2506.11928v1","created_at":"2026-07-05T11:21:13.002344+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2506.11928","created_at":"2026-07-05T11:21:13.002344+00:00"},{"alias_kind":"pith_short_12","alias_value":"C7KH3B52IEHJ","created_at":"2026-07-05T11:21:13.002344+00:00"},{"alias_kind":"pith_short_16","alias_value":"C7KH3B52IEHJVXZB","created_at":"2026-07-05T11:21:13.002344+00:00"},{"alias_kind":"pith_short_8","alias_value":"C7KH3B52","created_at":"2026-07-05T11:21:13.002344+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":10,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.20517","citing_title":"Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2607.00248","citing_title":"Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity","ref_index":143,"is_internal_anchor":false},{"citing_arxiv_id":"2606.32007","citing_title":"AxDafny: Agentic Verified Code Generation in Dafny","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15177","citing_title":"OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15607","citing_title":"Syntax Without Semantics: Teaching Large Language Models to Code in an Unseen Language","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2603.20633","citing_title":"Seed1.8 Model Card: Towards Generalized Real-World Agency","ref_index":93,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15177","citing_title":"OpenDeepThink: Parallel Reasoning via Bradley-Terry Aggregation","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2604.14164","citing_title":"How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08478","citing_title":"When Independent Sampling Outperforms Agentic Reasoning","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2604.12268","citing_title":"CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation","ref_index":42,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/C7KH3B52IEHJVXZBBOVM5BED3Y","json":"https://pith.science/pith/C7KH3B52IEHJVXZBBOVM5BED3Y.json","graph_json":"https://pith.science/api/pith-number/C7KH3B52IEHJVXZBBOVM5BED3Y/graph.json","events_json":"https://pith.science/api/pith-number/C7KH3B52IEHJVXZBBOVM5BED3Y/events.json","paper":"https://pith.science/paper/C7KH3B52"},"agent_actions":{"view_html":"https://pith.science/pith/C7KH3B52IEHJVXZBBOVM5BED3Y","download_json":"https://pith.science/pith/C7KH3B52IEHJVXZBBOVM5BED3Y.json","view_paper":"https://pith.science/paper/C7KH3B52","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2506.11928&json=true","fetch_graph":"https://pith.science/api/pith-number/C7KH3B52IEHJVXZBBOVM5BED3Y/graph.json","fetch_events":"https://pith.science/api/pith-number/C7KH3B52IEHJVXZBBOVM5BED3Y/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/C7KH3B52IEHJVXZBBOVM5BED3Y/action/timestamp_anchor","attest_storage":"https://pith.science/pith/C7KH3B52IEHJVXZBBOVM5BED3Y/action/storage_attestation","attest_author":"https://pith.science/pith/C7KH3B52IEHJVXZBBOVM5BED3Y/action/author_attestation","sign_citation":"https://pith.science/pith/C7KH3B52IEHJVXZBBOVM5BED3Y/action/citation_signature","submit_replication":"https://pith.science/pith/C7KH3B52IEHJVXZBBOVM5BED3Y/action/replication_record"}},"created_at":"2026-07-05T11:21:13.002344+00:00","updated_at":"2026-07-05T11:21:13.002344+00:00"}