Pith. sign in

Paper Citation Record · LEDGER

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

As of 21 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 50 inbound Pith citation observations for arXiv:2501.01257.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01257 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:37:58.494276Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:47:34.593855Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b772444c-d80b-42a2-a4d5-6e66bf30def7 · outbound

This paper cites Program Synthesis with Large Language Models.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.399925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.399925Z digest=sha256:84a7d449b8f064f9770658d91885f22496c138ec26328a0f4965c9e92e4aefae

Observation 88158010-b2c5-493d-ba48-a9e991195b3a · outbound

This paper cites The Llama 3 Herd of Models.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.408634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.408634Z digest=sha256:3acf7ab1aeea42a2913966375aaca046e0f4aae58b097c231527366b6abb9793

Observation 17dd5632-809a-4f61-b35c-7578a9275979 · outbound

This paper cites (2024) - o1-mini OpenAI (2024a) - Qwen2.5-Coder-1.5B-Instruct Hui et al.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings (2024) - o1-mini OpenAI (2024a) - Qwen2.5-Coder-1.5B-Instruct Hui et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:37:58.863972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:37:58.485989Z digest=sha256:db53166f39122b8c1079becbc09b7219416b969cad52261530e55605401b4c80

Observation 6dd21d7b-f2a0-4048-9420-f08e9fb966e3 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.417005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.417005Z digest=sha256:50a5fa1ab2b674bbdb1131dcfdc3dd4b772b24eb0fdf81ec22152e3febf11df9

Observation 26e54c3c-dcf0-48a2-96ed-46ab41d35371 · outbound

This paper cites Understanding HTML with Large Language Models.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Understanding HTML with Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.421209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.421209Z digest=sha256:3cdfc92234e798acd52065b1c923231c1f803c762c7ec86069ed4a5766222c90

Observation 83a150fd-a97b-42b2-8d1a-7e984b875e57 · outbound

This paper cites abc", acceptable outputs could include.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings abc", acceptable outputs could include

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:37:58.840939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:37:58.494276Z digest=sha256:dbb350d57429676f3cea831fe3d1bf082e0c3826e77d163454c1c43df88d59e9

Observation 2e877dc5-c25c-40d0-82fa-a26b61ba6695 · outbound

This paper cites OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.429313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.429313Z digest=sha256:c5979fc8d315a117b890d617aeac37d04eeb2ab0c55e77bbeb2de2e29ca41ba0

Observation 321408de-878e-48b8-8417-a416e6f32739 · outbound

This paper cites Qwen2.5-Coder Technical Report.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Qwen2.5-Coder Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.433146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.433146Z digest=sha256:6df0e453d1fe383d84dc19e6dc62e84973c95acc566d20e5c26e17dfc0235c66

Observation d0001b78-c5fb-48dc-aea4-31ebea6e08f7 · outbound

This paper cites GPT-4o System Card.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings GPT-4o System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.437135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.437135Z digest=sha256:8d53ff2ff51f83669f232b5fb7dd76e351f992fa07ff10f34262878341450abb

Observation ff8bec78-c878-4303-8ae5-41bfe2803048 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.440684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.440684Z digest=sha256:cc4e76c5a100b805622a62f1b417bb49e96aa951d248ae2e7688f80612da921b

Observation c1aeb3b1-fc3b-4341-888c-e070d678e16c · outbound

This paper cites Mistral 7B.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Mistral 7B

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.444503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.444503Z digest=sha256:31e3af68894ea3e7e29a3376e4d36713e0d5d6d51637c676d99d75143e207239

Observation a5f18db5-4279-4d3e-ae74-69077925e2ad · outbound

This paper cites xCodeEval: A Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and Retrieval.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings xCodeEval: A Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and Retrieval

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.451483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.451483Z digest=sha256:b72299fdd7c156c690dda21fd54b24eab7cf45ef0eecb3493962d6a771231df2

Observation 0704a626-34ac-4c4f-a57f-44a51bcae9e7 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings TACO: Topics in Algorithmic COde generation dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.454843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.454843Z digest=sha256:91477d164040fbb80d07ee475aec3eff4f0d0b947af03972289b85d0a83b8650

Observation c23b738d-afc2-495b-97df-a64d1b3f2a41 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.458249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.458249Z digest=sha256:112277e27833dbd26dd1dc001beedec6cd8c4943ea017ae4dcc83bcf83a20b31

Observation a25242d0-112a-4156-a918-5d7d2e5af09c · outbound

This paper cites American invitational mathematics examination - aime.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings American invitational mathematics examination - aime

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:37:58.874826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:37:58.461544Z digest=sha256:798b46e61371cae65821ce123159d63084799867285f71d9b955baa69295b503

Observation 05901e81-50bc-4c0f-9b2e-a587bfe1aac1 · outbound

This paper cites an unresolved cited work.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Unresolved cited work

Reference 19

Resolution
verified exact
raw_fallback, observed 2026-08-10T22:37:58.659524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:37:58.464950Z digest=sha256:cf788b85fae7ec1efb9aa09126e079d406bc6a8cfb013c093513b9dfbfa4d481

Observation 410498c9-dafa-4952-9c4e-0f48923cabe9 · outbound

This paper cites Can Language Models Solve Olympiad Programming?.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Can Language Models Solve Olympiad Programming?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.468295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.468295Z digest=sha256:5d278aa230334e325f22dbbb67277745c4d738fc249b2a5e37e7cc40f7f59f64

Observation cf1502dd-0247-4114-9829-c97e27dc20f0 · outbound

This paper cites SelfCodeAlign: Self-Alignment for Code Generation.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings SelfCodeAlign: Self-Alignment for Code Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.471707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.471707Z digest=sha256:90927c304b5d127e08ab6d6ee1cc953659332c681cd3bb0ce60ce6bea5b7d1ff

Observation e65110a8-14fb-4ebc-b65e-67560475f45a · outbound

This paper cites Qwen2 Technical Report.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Qwen2 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.475033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.475033Z digest=sha256:f1e882d021229e93378d7478cd1fe1435212f5eeb79147e787c33d8b86abd9e4

Observation bf128bdd-9cf7-4ea2-b372-48d974124012 · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.478693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.478693Z digest=sha256:f64d2222c8832cef655ade4119e4db27266e0432056cd6f07ffab44bc3db3b97

Observation 69cfdfb7-2358-43f3-8ebc-89730dd5a836 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.482301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.482301Z digest=sha256:7357bb25e429bd4fff5b0a1d3b33a38e2668a0c771a48db705abd71123e01e7a

Observation f4e100ea-b366-4e14-b4e7-5be57b7e319b · outbound

This paper cites Namely, the new rating can be calculated using the formula ri = ri−1+E(ri) 2.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Namely, the new rating can be calculated using the formula ri = ri−1+E(ri) 2

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:37:58.853078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:37:58.490026Z digest=sha256:b41f01af47dd87545fb444593bcc484c6cc5a0424a313472394b2fb801852b04

Observation 30800675-fe33-4030-8767-e9e2aca13a71 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.412771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.412771Z digest=sha256:49941e5ef17b77cb27777253c7e5dd1ceb75e81af036ff665ae444624f8dc619

Observation a39333c0-5da3-469d-ad83-5222024ccd1d · outbound

This paper cites Evaluating Large Language Models Trained on Code.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.404474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.404474Z digest=sha256:5aceb4349b79a9d7bafc79363eedb045fca779ee07be264376c338d4a9fc60eb

Observation 41b15a40-e11d-47b6-9cbe-12c1231a312f · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Measuring Coding Challenge Competence With APPS

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.425527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.425527Z digest=sha256:97af28177419587c701e81a6de32850f9da6302481ff3afc8f308e9dcdee17df

Observation 2364059a-7b44-492a-a8ee-61690f8fd4b7 · outbound

This paper cites Mixtral of Experts.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings Mixtral of Experts

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:37:58.448033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:37:58.448033Z digest=sha256:8c2615a8b8668b6ed1d839e58d59bc62ee58c1bff42ac3651c51660eccf94806

Observation fc2795d0-aa38-481f-b493-40bf8d7571c1 · outbound

This paper cites io/blog.html?post=en/2024-09-05-A-Small-but-Mighty-LLM-for-Code.md.

CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings io/blog.html?post=en/2024-09-05-A-Small-but-Mighty-LLM-for-Code.md

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:37:58.886501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-10T22:37:58.395499Z digest=sha256:9531f7e47834fdf599bffd70d6814f1d234ba61c17dd431bfeadbf65dcac9bdb

Pith citing papers

Observation 7f08bf2f-867b-428e-a3d1-997c413519e5 · inbound

LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs cites this paper.

LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:47:34.593855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:47:34.593855Z digest=sha256:80af0926125c5a2ab3e6e2e36edc023e8f5b227da3b997d9bcfc45aef1963aaf

Observation f7a25498-4e45-4f51-b8e2-dfb4120f10bb · inbound

Evaluating LLM Metrics Through Real-World Capabilities cites this paper.

Evaluating LLM Metrics Through Real-World Capabilities CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T22:03:20.437616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:03:20.437616Z digest=sha256:afd4c415a3441a8464c6b6fecdccd2b1ead6ce7b85680de38e0faa1af7857534

Observation d3ac7d4f-3fd2-4b35-abfa-1568aaa272c6 · inbound

Qwen3 Technical Report cites this paper.

Qwen3 Technical Report CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:28.581174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T06:35:27.813995Z digest=sha256:148902b8df74bd02f967524fc8512db336fdaea7940bcfa0d27f275dbb1404d0

Observation 0073fe15-f24f-46b0-a18d-b312d982eae9 · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:05.108016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:05.108016Z digest=sha256:1bc08b90a584b50bcbfa7f75032caa72810b40ee80b46104837f8f618c08aa67

Observation fd8354c2-1db0-4c14-afcf-87001b2a3695 · inbound

How Programming Concepts and Neurons Are Shared in Code Language Models cites this paper.

How Programming Concepts and Neurons Are Shared in Code Language Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:18.731030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:56:18.731030Z digest=sha256:7697de5be1b41e06b3d9ff20887578ae55e9213c978db2d6b6e29eaf0ffa68da

Observation d6dca7f2-f5cf-484a-bac9-ba33ab5a1375 · inbound

Seed-Coder: Let the Code Model Curate Data for Itself cites this paper.

Seed-Coder: Let the Code Model Curate Data for Itself CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:56.850042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:56.850042Z digest=sha256:fd7c2a9b2079710247df0d13399abcc615dfc73d3ed357702372111079fa75f3

Observation 7d8cdb37-b2b4-48b3-9ac1-14879d128c64 · inbound

ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests cites this paper.

ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:02.087340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:02.087340Z digest=sha256:413a3a0f2c7bc987ef7143f4e7408edba7d8c52f189fba7fb095c880f707781c

Observation 39d456e6-5449-4eaa-8845-ed2cdc25096c · inbound

SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows cites this paper.

SCGAgent: Recreating the Benefits of Reasoning Models for Secure Code Generation with Agentic Workflows CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:54.942793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:54.942793Z digest=sha256:0baff394a8cba6330d100d6cb0366604703df14ff1a92d8d383431b608e9a923

Observation 89daff14-5780-49b5-be69-a6f71cfab640 · inbound

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics cites this paper.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:08.345343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:08.345343Z digest=sha256:78a1a7fea985fbbd528b5dfa472df70fb5cc4da49e9e5a986ac3b9f88d19b58e

Observation 307f5057-6961-4614-a47e-f71a8c6133cb · inbound

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents cites this paper.

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:07:39.497630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T08:07:39.384613Z digest=sha256:955b5fe09a9d4f0c4a831ee9983ec67f64526e29bd930517a140a9f4f6292120

Observation 90499ba3-6223-499a-a37c-87532b727c53 · inbound

LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming? cites this paper.

LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming? CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:00.872714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:08:00.872714Z digest=sha256:6ee1bcdff8603fd65864409655cd749bdce94fb7b57a04b19f7cb96beb95b383

Observation 984bb649-2cce-46ed-8cac-8decff31f390 · inbound

Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning cites this paper.

Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:54:39.321936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:54:39.321936Z digest=sha256:5a672f2184d1e732b36af0e8473be44ccdd1f74a752718fadefc7b1452a6cf92

Observation 4973fe17-3e77-4718-b797-ff733544a9dc · inbound

BACTA-GPT: An AI-Based Bayesian Adaptive Clinical Trial Architect cites this paper.

BACTA-GPT: An AI-Based Bayesian Adaptive Clinical Trial Architect CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:42:55.226122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:42:55.226122Z digest=sha256:e29209e3a22d813c38de8348ebe905d1e0fe0907e26d9c249ea117327ce39b21

Observation 99e58958-ce9a-4adf-b53e-55b5051e5786 · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:25.910315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:25.910315Z digest=sha256:02ce786ad2e12295c1e3a4989b8e27b0d47b01883e1442a2674357e952470947

Observation b43dddd9-fb6d-42ab-b506-02f13e1b990b · inbound

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity cites this paper.

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:29.982589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:29.982589Z digest=sha256:df4b89f25be6fa8545668fd13565583819050972ac70aa98975ec666cac18367

Observation bce54451-14c6-4b2c-aa51-696bf63c5a03 · inbound

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once cites this paper.

REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:31.589921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:35:31.589921Z digest=sha256:4bf5d43ecd869bb4fa2ad615bf52bfa4272d19ee7edc6aa86847b80d3b44d127

Observation 75acb1a9-c225-43a2-9ff5-3e318ecf04b2 · inbound

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench cites this paper.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.423730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.423730Z digest=sha256:a30ba053cf3f91b81a564f05d82640cabb75b186bcd25f3dd917087db0ccb8bf

Observation 1e07aeb6-7577-4eec-a152-401656ddacac · inbound

IFEvalCode: Controlled Code Generation cites this paper.

IFEvalCode: Controlled Code Generation CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:34.051268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:44:34.051268Z digest=sha256:024517a6131d3d56e89a51e1a893b8920df54718add63db122cae0cb368c568e

Observation 17e42d19-1e94-48f3-aa19-8c9786923d78 · inbound

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators cites this paper.

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:27.805253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:18:27.805253Z digest=sha256:5b95aba83db292a53788c1799b3936a31973c971376cbe19d1069090422b2da1

Observation ed960569-c49f-4db2-8a2c-55242f2ee0fc · inbound

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications cites this paper.

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T17:26:28.894359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:26:28.894359Z digest=sha256:ea9e82e6c8a949617dc28d12fe80fddc0d8965e1e982551ff6617b1717ed3c2c

Observation bf84de14-32ec-4fa4-bdbf-4020e10daa64 · inbound

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models cites this paper.

AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:13.786271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:17:13.786271Z digest=sha256:b451b22a1b29370b594f19208bf16196b9f96b2acf9bc9ccdfd49945f22034bb

Observation e943bfea-6c25-45fa-bd73-6cb01e8f4444 · inbound

ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution cites this paper.

ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 121

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:58:58.872793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-16T13:58:58.627748Z digest=sha256:e3c14c1ff2ede22752003fabae6702f6c7081b10c34c6d9184800abc7dd59da5

Observation cdb9ec5f-aa9d-478d-b68d-d7e2d4219d49 · inbound

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation cites this paper.

XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T11:11:12.184681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:11:12.184681Z digest=sha256:ad2276946f78e56bfd44ba9dc6bca2a54010256f68d0999ba785c13b9663f988

Observation 5abadc6a-e85d-47df-b5c9-e547f06fe00d · inbound

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation cites this paper.

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T12:19:33.502600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:19:33.502600Z digest=sha256:549aa6c80beea9a580e461dabedfdd7cf96bc74ed5f6c63b7b87c7adc1c17d8f

Observation 10933620-c4ea-48c7-a354-87fca0c14cac · inbound

Parameter-Efficient Multi-Task Fine-Tuning in Code-Related Tasks cites this paper.

Parameter-Efficient Multi-Task Fine-Tuning in Code-Related Tasks CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T09:01:32.227320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:01:32.227320Z digest=sha256:c6dfbfe9f68233873994c7fa6e0780ef3f33919aebaf2ec3edb8adc304af2b99

Observation 657e1a74-7e37-462b-8b5f-a8869b0d60ed · inbound

Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention cites this paper.

Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:43:29.252596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:43:29.252596Z digest=sha256:30353affde9691c8a8eba3b4e959409047eb01248187c434cda4ef5fb096325f

Observation 67447ea7-db89-40a5-8f3f-abca1710acf9 · inbound

GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning cites this paper.

GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:16.679979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T20:30:17.791973Z digest=sha256:03c4fe8daf37d10ac2d7832ae9e2f53bd0c9a737386a51c92d0d263349cf8dc4

Observation 764cb704-d7d5-4a82-b4aa-106689f8aed8 · inbound

A meta-analysis of the effect of generative AI on productivity and learning in programming cites this paper.

A meta-analysis of the effect of generative AI on productivity and learning in programming CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:05.646008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T17:22:26.933712Z digest=sha256:8e0c79d12c26802157f35aad257c0442338e9656437d9ac9d0d450a99966a9ad

Observation 64b3b35e-2794-4a68-964c-1e0659eff5e9 · inbound

Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts cites this paper.

Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:50:57.504801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T02:11:19.295354Z digest=sha256:9aa77fd4d17fe1d00036acebebefcaff4830f1de50ecff05c7519787dab6473f

Observation 7ed6bb66-9a37-4efe-9dec-fe46e2c80605 · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:53.378167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T03:40:04.692279Z digest=sha256:750831e59e4d4259ce50ce5849915c91b0f62029a45a0c70c27e4faf995344dc

Observation 0eaf367b-58f3-452c-a215-ced211e89e39 · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.924499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T08:14:55.858466Z digest=sha256:4ecf7b9cc7b44dab4b7b4131edb368a3ba552c991d6b976f6f5732e05b23a841

Observation 33283815-219a-4ad1-af20-da60503ea4e9 · inbound

When Independent Sampling Outperforms Agentic Reasoning cites this paper.

When Independent Sampling Outperforms Agentic Reasoning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:31:24.569785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T02:46:48.375901Z digest=sha256:d51d6e95fdb9b1c6ba2c0bd173b31bbd1d058021067827f7f4cda605344f803a

Observation 8fcaa0af-3e5c-4808-8f18-4c8ca527c9d8 · inbound

Scalable Token-Level Hallucination Detection in Large Language Models cites this paper.

Scalable Token-Level Hallucination Detection in Large Language Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.415561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T05:49:23.534294Z digest=sha256:72ad19d0741bb5325c8493b86231e00e920daba3b8848d909c3ba4f3cde55ccd

Observation c269d1fe-9e67-4a66-ba22-0081522b6f8c · inbound

Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution cites this paper.

Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:32:39.503059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T16:28:23.113976Z digest=sha256:1269ad0357b5dc8e6c96492e805e849cf1ab28c7cb945ba403735ed4f528d12c

Observation 0cfd1dbb-04aa-42a4-a833-1b170a907e2c · inbound

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning cites this paper.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.261829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T14:52:33.141693Z digest=sha256:fda6a1a9ee5dc3432e18e80680692abd015fdb7a90a28e2288885a1624e8d0b4

Observation d0b9eb4a-8f38-4e2e-b05c-42ed6bf2bafe · inbound

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning cites this paper.

Leveraging Error Diversity in Group Rollouts for Reinforcement Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.422916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T19:03:10.268534Z digest=sha256:c9856acb9b5bcf1d362de6e2a8b0ec2a48842c30fc9eadf9e69d3140040e6e94

Observation 4e2da0ee-e08a-4f50-bc67-3d540bdada9c · inbound

CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models cites this paper.

CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:33:31.675895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T06:25:56.404246Z digest=sha256:78c68fc29ccc5fcc218fc07d682ecad937ff32c84fde932cb56f13237c6a8a70

Observation 38913d1a-1d2f-4d63-be86-4cc366c79fe7 · inbound

Sakana Fugu Technical Report cites this paper.

Sakana Fugu Technical Report CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 157

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:29:38.282952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T14:22:37.596720Z digest=sha256:01368e5b5581b85d6983a0065f28b7c092afc14e2aab01bc81fa6fcac87956a1

Observation 50834b1b-e859-45ec-91b6-7194cab1928d · inbound

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms cites this paper.

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.073584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-25T20:55:15.784610Z digest=sha256:0e6217af84c6534abb2ab3437c41cb55fb211c8349fe5efde87a492378310a80

Observation fc0d15ef-6e2d-4e85-b267-a5672e39e6f3 · inbound

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms cites this paper.

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.536539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T05:29:21.598397Z digest=sha256:9c0fc36157dfcf3042b76343f0dbd52425e8ff0b2ff682426cb12034c6abea7d

Observation 8a0297f1-5296-4827-865e-a8e650d38072 · inbound

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity cites this paper.

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:07:17.121388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-02T18:57:46.841456Z digest=sha256:a5067865217f993dfcdd6e4152de00c71031d615a487d6a6aa48fa150d21a63b

Observation 673b4857-0913-4c36-9044-6e0828a42c0f · inbound

Selective Left-Shift: Turning Test-Time Compute and Difficulty-based Curation into Training Data for Low-Resource Code Generation cites this paper.

Selective Left-Shift: Turning Test-Time Compute and Difficulty-based Curation into Training Data for Low-Resource Code Generation CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-10T19:57:34.142769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T19:48:27.137130Z digest=sha256:18aa37045f0ce841e7b11e70be2c58de4ac097e30b0a20f992324c759bfc5bc7

Observation afd13c9d-edc2-4bca-8b40-b167ed868f64 · inbound

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs cites this paper.

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 145

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T13:57:06.683592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-10T13:49:17.343893Z digest=sha256:2dc2c8e61fdc3e058d5344290d2213f1a0674a12c46cfaca6339fbb3e1b5784d

Observation 6e143a64-a9d3-4332-8676-c5b4560bf0d4 · inbound

Quantize with Confidence? An Empirical Study of Quantization for Code Generation cites this paper.

Quantize with Confidence? An Empirical Study of Quantization for Code Generation CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T03:32:56.869485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:32:56.869485Z digest=sha256:54ea831bb0f01786908b2ff6b20c4d31c7e47fc51188d68803389e31ee1905a7

Observation c609dd5c-a15f-49bc-a80a-45d142a36396 · inbound

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents cites this paper.

PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T14:11:55.543471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:11:55.543471Z digest=sha256:e30d0a3fbf1485fc55515f8b80c3187268498762cb3728801d5408648de94fad

Observation 55439718-bb99-4fd1-9316-b29d5bf06fa2 · inbound

Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models cites this paper.

Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T05:52:42.177022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:52:42.177022Z digest=sha256:a8f91701ab3ad7b35c55be32e04ea6118d72006a20915885515ba1c4f66be7bb

Observation c42abb11-d600-4e65-87ea-38efbb551916 · inbound

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning cites this paper.

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T15:29:54.308202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:29:54.308202Z digest=sha256:0f7814b703a327f8c66122f89fb1a42ed11adbc5350db3fefcf292b447435be2

Observation 956cc784-859b-4068-b5b5-4853613f2aa1 · inbound

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning cites this paper.

DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:42.973857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:27:42.973857Z digest=sha256:29dd454f935842ccffcb9da0a8746cdfb4a3f37270b7fef0737b7dddb0fc3e85

Observation 7b0e0aff-407c-4cf3-8c50-52c0e21cf427 · inbound

CurveShift: Is Agent Progress Scalar? Separating Level from Shape cites this paper.

CurveShift: Is Agent Progress Scalar? Separating Level from Shape CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T00:43:21.537277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:43:21.537277Z digest=sha256:2e8d7a3860d411c5f5993e95e0002536c8a48c6a0c107c16fc10c2bebc510b25

Observation 67a66d07-2ecf-47b6-9302-3ae9a415e5a6 · inbound

Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation cites this paper.

Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:30:09.691906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:30:09.691906Z digest=sha256:f84e082e32e60f655777845ff322a2c6ff7e4cd01dea883bb7b5934798033832