Pith. sign in

Paper Citation Record · LEDGER

Competition-Level Problems are Effective LLM Evaluators

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2312.02143.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.02143 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T23:10:40.420241Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:10:40.976975Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ef6d5019-f883-453a-9dc3-45ec69b674ba · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Competition-Level Problems are Effective LLM Evaluators

Reference 259

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:34:42.915321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:e1160cb4bfa52f954723016eddda116c14aeb0eb6ac34995853003b418596d62

Observation c7a6f056-30ae-4c52-9b02-c7a49d20297a · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Competition-Level Problems are Effective LLM Evaluators

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:40.979836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:1c2d05f6d4849e73e633e17c3f962e350d5b62165d61b1da5dbaaa616e7a3d22

Observation 66a5db3d-56cb-4e92-97d9-0cc2189d4012 · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Competition-Level Problems are Effective LLM Evaluators

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.360747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:fdfc68f673fad199179882567c04ed0861db39d6db33c6b0383171aa84d718ba

Observation 6ecc6ae9-09b0-427f-a1f9-101da0c6a33f · inbound

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving cites this paper.

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving Competition-Level Problems are Effective LLM Evaluators

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T15:11:32.713054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T15:08:03.793304Z digest=sha256:33f2eba411df1d30d8f9d061c7148fa80d12f860f67231b3ec8701b511db75b8

Observation 7cb9072a-e232-44bf-934b-86ca93832e31 · inbound

AutoBaxBuilder: Bootstrapping Code Security Benchmarking cites this paper.

AutoBaxBuilder: Bootstrapping Code Security Benchmarking Competition-Level Problems are Effective LLM Evaluators

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:16:31.800716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T12:15:45.967861Z digest=sha256:708360f78cf0dfde2099cb1aba4b3c4d9dbde34291b33d745b883c6f22faf992