Pith. sign in

Paper Citation Record · LEDGER

Competition-Level Problems are Effective LLM Evaluators

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2312.02143.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.02143 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:07:59.838876Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:10:40.976975Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ef6d5019-f883-453a-9dc3-45ec69b674ba · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Competition-Level Problems are Effective LLM Evaluators

Reference 259

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:34:42.915321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:606198912d607db0617f6771e2a7b6b1cad4eb944202cffefad086b3d679e312

Observation c7a6f056-30ae-4c52-9b02-c7a49d20297a · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Competition-Level Problems are Effective LLM Evaluators

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:40.979836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:b7df04342614a23e7b1789df6160ea5d033a0e22e149805cf2f08329a31c2ccb

Observation 5637c6f3-e880-4a07-a26c-5372d634466d · inbound

LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming? cites this paper.

LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming? Competition-Level Problems are Effective LLM Evaluators

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T01:07:59.838876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:07:59.838876Z digest=sha256:52f8a1f6770559935eb470ed5a1aae8759f1793dec4799eceff833f739d37b30

Observation 380c017e-2dc7-48cc-a404-6453a238f449 · inbound

Evaluating and Improving Large Language Models for Competitive Program Generation cites this paper.

Evaluating and Improving Large Language Models for Competitive Program Generation Competition-Level Problems are Effective LLM Evaluators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:28.782849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:28.782849Z digest=sha256:d5cab739c63400e782041a85ca443cc70a584d7df2a5ac925378339b42279231

Observation 66a5db3d-56cb-4e92-97d9-0cc2189d4012 · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Competition-Level Problems are Effective LLM Evaluators

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.360747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:e8c3a93f54bebd61ec6945dc63760991f789c922215dd4dbe5dacf05087a58cf

Observation 6ecc6ae9-09b0-427f-a1f9-101da0c6a33f · inbound

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving cites this paper.

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving Competition-Level Problems are Effective LLM Evaluators

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T15:11:32.713054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T15:08:03.793304Z digest=sha256:43f5f4319c7d39364e39090127718bde0bc912c45f84fcb56897f3c39656a68c

Observation 7cb9072a-e232-44bf-934b-86ca93832e31 · inbound

AutoBaxBuilder: Bootstrapping Code Security Benchmarking cites this paper.

AutoBaxBuilder: Bootstrapping Code Security Benchmarking Competition-Level Problems are Effective LLM Evaluators

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:16:31.800716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T12:15:45.967861Z digest=sha256:06ac2add89f604a9a76db4c0977344cfd20f469a7ecaf743a9455dc74d6481b3