Pith. sign in

Paper Citation Record · LEDGER

Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2403.19114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.19114 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:24:17.725599Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T08:34:27.439389Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e90f8d23-761a-4087-99c1-79c4c7dbde82 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.129426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:92cdf9d4df7bb3e71a4bce8e991d4723bc25596b411c26f2794b52e2415d542a

Observation ee0eebcf-5173-4998-9590-78a7c320bf4e · inbound

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities cites this paper.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.725599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.725599Z digest=sha256:0cc29751963165768a2034a906e7c64fc7b199129d33d7ec75f84ca69e98b1b9

Observation 96d3e6c8-3e6f-4a8a-82ff-6496f9855eee · inbound

Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar cites this paper.

Unseen Horizons: Unveiling the Real Capability of LLM Code Generation Beyond the Familiar Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.670461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.670461Z digest=sha256:0150f12339f3c098dbfc4618e8109843f3e6ff024c03715b369937805272ddf6

Observation 774f1bd7-85c3-42f8-b5e7-8a24ba23a3bd · inbound

Sustainable Code Generation Using Large Language Models: A Systematic Literature Review cites this paper.

Sustainable Code Generation Using Large Language Models: A Systematic Literature Review Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:50:16.506793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T18:49:01.097179Z digest=sha256:81ad36af0ca785e95a1c6da7979c23ec6295597bf3d8a279701dee03be03e8f4

Observation 4effb991-0933-4ee7-a21f-50f16cdbac37 · inbound

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents cites this paper.

SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:27.205216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T06:13:32.434201Z digest=sha256:a3b855cea6996091ed0d7c5f1fb6ab5ced1667baa689bf59df39a526b11ad846

Observation 6824477f-af6f-4a10-ad7f-2b49a9c628a8 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:03:29.013236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:257cd13c64be9f2a5eb2ff014be5fc203a6ba07b180b531fca452f0c92b0d90c

Observation b7683286-aac9-48aa-a80a-85bb01283e22 · inbound

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation cites this paper.

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 56

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T08:34:27.441561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T08:30:37.016856Z digest=sha256:7200ccf5aadd6dd9a8cddb6fdfa213f1e29f1e9789047ed06f746e9563113c43

Observation 6945d7be-c82c-4611-9500-12bf084d04fc · inbound

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations cites this paper.

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T04:56:20.002453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T04:56:20.002453Z digest=sha256:a7eb2aca783f9822ff1ab08dbe94a53c86b33b1ea7cced8c806905bc2293e40b

Observation 8440dad5-8fea-4ce1-9b80-024b1dfbed09 · inbound

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations cites this paper.

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T07:43:20.935793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:43:20.935793Z digest=sha256:d5137aace8e5e5d8f7092891f75e44b7cc4389f5ff709d39c215d96193e26209