Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:37:58.494276Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 50 inbound Pith citation observations for arXiv:2501.01257.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:37:58.494276Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:47:34.593855Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
27 of 27 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
Observation b772444c-d80b-42a2-a4d5-6e66bf30def7 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Program Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88158010-b2c5-493d-ba48-a9e991195b3a · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings The Llama 3 Herd of Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17dd5632-809a-4f61-b35c-7578a9275979 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings (2024) - o1-mini OpenAI (2024a) - Qwen2.5-Coder-1.5B-Instruct Hui et al
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6dd21d7b-f2a0-4048-9420-f08e9fb966e3 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e54c3c-dcf0-48a2-96ed-46ab41d35371 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Understanding HTML with Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83a150fd-a97b-42b2-8d1a-7e984b875e57 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings abc", acceptable outputs could include
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2e877dc5-c25c-40d0-82fa-a26b61ba6695 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 321408de-878e-48b8-8417-a416e6f32739 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Qwen2.5-Coder Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0001b78-c5fb-48dc-aea4-31ebea6e08f7 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings GPT-4o System Card
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff8bec78-c878-4303-8ae5-41bfe2803048 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1aeb3b1-fc3b-4341-888c-e070d678e16c · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Mistral 7B
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5f18db5-4279-4d3e-ae74-69077925e2ad · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings xCodeEval: A Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and Retrieval
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0704a626-34ac-4c4f-a57f-44a51bcae9e7 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings TACO: Topics in Algorithmic COde generation dataset
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23b738d-afc2-495b-97df-a64d1b3f2a41 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a25242d0-112a-4156-a918-5d7d2e5af09c · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings American invitational mathematics examination - aime
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 05901e81-50bc-4c0f-9b2e-a587bfe1aac1 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 410498c9-dafa-4952-9c4e-0f48923cabe9 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Can Language Models Solve Olympiad Programming?
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf1502dd-0247-4114-9829-c97e27dc20f0 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings SelfCodeAlign: Self-Alignment for Code Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e65110a8-14fb-4ebc-b65e-67560475f45a · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Qwen2 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf128bdd-9cf7-4ea2-b372-48d974124012 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69cfdfb7-2358-43f3-8ebc-89730dd5a836 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4e100ea-b366-4e14-b4e7-5be57b7e319b · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Namely, the new rating can be calculated using the formula ri = ri−1+E(ri) 2
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 30800675-fe33-4030-8767-e9e2aca13a71 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 1978
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a39333c0-5da3-469d-ad83-5222024ccd1d · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Evaluating Large Language Models Trained on Code
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41b15a40-e11d-47b6-9cbe-12c1231a312f · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Measuring Coding Challenge Competence With APPS
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2364059a-7b44-492a-a8ee-61690f8fd4b7 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Mixtral of Experts
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc2795d0-aa38-481f-b493-40bf8d7571c1 · outbound
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings io/blog.html?post=en/2024-09-05-A-Small-but-Mighty-LLM-for-Code.md
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7f08bf2f-867b-428e-a3d1-997c413519e5 · inbound
LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a25498-4e45-4f51-b8e2-dfb4120f10bb · inbound
Evaluating LLM Metrics Through Real-World Capabilities CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3ac7d4f-3fd2-4b35-abfa-1568aaa272c6 · inbound
Qwen3 Technical Report CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0073fe15-f24f-46b0-a18d-b312d982eae9 · inbound
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd8354c2-1db0-4c14-afcf-87001b2a3695 · inbound
How Programming Concepts and Neurons Are Shared in Code Language Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6dca7f2-f5cf-484a-bac9-ba33ab5a1375 · inbound
Seed-Coder: Let the Code Model Curate Data for Itself CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d8cdb37-b2b4-48b3-9ac1-14879d128c64 · inbound
ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39d456e6-5449-4eaa-8845-ed2cdc25096c · inbound
SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89daff14-5780-49b5-be69-a6f71cfab640 · inbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 307f5057-6961-4614-a47e-f71a8c6133cb · inbound
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 90499ba3-6223-499a-a37c-87532b727c53 · inbound
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming? CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 984bb649-2cce-46ed-8cac-8decff31f390 · inbound
Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4973fe17-3e77-4718-b797-ff733544a9dc · inbound
BACTA-GPT: An AI-Based Bayesian Adaptive Clinical Trial Architect CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99e58958-ce9a-4adf-b53e-55b5051e5786 · inbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b43dddd9-fb6d-42ab-b506-02f13e1b990b · inbound
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce54451-14c6-4b2c-aa51-696bf63c5a03 · inbound
REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75acb1a9-c225-43a2-9ff5-3e318ecf04b2 · inbound
Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e07aeb6-7577-4eec-a152-401656ddacac · inbound
IFEvalCode: Controlled Code Generation CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17e42d19-1e94-48f3-aa19-8c9786923d78 · inbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed960569-c49f-4db2-8a2c-55242f2ee0fc · inbound
AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf84de14-32ec-4fa4-bdbf-4020e10daa64 · inbound
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e943bfea-6c25-45fa-bd73-6cb01e8f4444 · inbound
ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cdb9ec5f-aa9d-478d-b68d-d7e2d4219d49 · inbound
XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5abadc6a-e85d-47df-b5c9-e547f06fe00d · inbound
Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10933620-c4ea-48c7-a354-87fca0c14cac · inbound
Parameter-Efficient Multi-Task Fine-Tuning in Code-Related Tasks CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 657e1a74-7e37-462b-8b5f-a8869b0d60ed · inbound
Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67447ea7-db89-40a5-8f3f-abca1710acf9 · inbound
GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 764cb704-d7d5-4a82-b4aa-106689f8aed8 · inbound
A meta-analysis of the effect of generative AI on productivity and learning in programming CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 64b3b35e-2794-4a68-964c-1e0659eff5e9 · inbound
Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7ed6bb66-9a37-4efe-9dec-fe46e2c80605 · inbound
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0eaf367b-58f3-452c-a215-ced211e89e39 · inbound
Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 33283815-219a-4ad1-af20-da60503ea4e9 · inbound
When Independent Sampling Outperforms Agentic Reasoning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8fcaa0af-3e5c-4808-8f18-4c8ca527c9d8 · inbound
Scalable Token-Level Hallucination Detection in Large Language Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c269d1fe-9e67-4a66-ba22-0081522b6f8c · inbound
Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0cfd1dbb-04aa-42a4-a833-1b170a907e2c · inbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d0b9eb4a-8f38-4e2e-b05c-42ed6bf2bafe · inbound
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4e2da0ee-e08a-4f50-bc67-3d540bdada9c · inbound
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 38913d1a-1d2f-4d63-be86-4cc366c79fe7 · inbound
Sakana Fugu Technical Report CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 157
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 50834b1b-e859-45ec-91b6-7194cab1928d · inbound
The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fc0d15ef-6e2d-4e85-b267-a5672e39e6f3 · inbound
The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8a0297f1-5296-4827-865e-a8e650d38072 · inbound
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 673b4857-0913-4c36-9044-6e0828a42c0f · inbound
Selective Left-Shift: Turning Test-Time Compute and Difficulty-based Curation into Training Data for Low-Resource Code Generation CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation afd13c9d-edc2-4bca-8b40-b167ed868f64 · inbound
From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6e143a64-a9d3-4332-8676-c5b4560bf0d4 · inbound
Quantize with Confidence? An Empirical Study of Quantization for Code Generation CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c609dd5c-a15f-49bc-a80a-45d142a36396 · inbound
PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55439718-bb99-4fd1-9316-b29d5bf06fa2 · inbound
Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42abb11-d600-4e65-87ea-38efbb551916 · inbound
DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 956cc784-859b-4068-b5b5-4853613f2aa1 · inbound
DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b0e0aff-407c-4cf3-8c50-52c0e21cf427 · inbound
CurveShift: Is Agent Progress Scalar? Separating Level from Shape CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67a66d07-2ecf-47b6-9302-3ae9a415e5a6 · inbound
Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.