Pith. sign in

Paper Citation Record · LEDGER

NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2405.04520.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.04520 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:56.926699Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:42:36.222980Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aa7df653-a6fc-4082-bd95-810bf5c01e09 · inbound

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools cites this paper.

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:08:09.715752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T08:08:09.444352Z digest=sha256:7595d78c39b9f4cfdf98855bfe73cf65895b113a07ecfcd76e0c84bff53391e6

Observation feece394-c6f3-426e-bf88-a55688ba4792 · inbound

Seed-Coder: Let the Code Model Curate Data for Itself cites this paper.

Seed-Coder: Let the Code Model Curate Data for Itself NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:56.926699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:56.926699Z digest=sha256:09fd8d4afc8b2d1f44fa2083ed431dcaafab99012fa57febbe252ab244825408

Observation 852763eb-388f-4f56-a45b-c50e763b21d6 · inbound

IFEvalCode: Controlled Code Generation cites this paper.

IFEvalCode: Controlled Code Generation NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:34.139423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:44:34.139423Z digest=sha256:41845567a59372abe4f4651889f2b1e2c13bf39407172be3710fa4645ebf15cf

Observation 683694ca-d5c2-46a4-ac32-93202e1c86ed · inbound

Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference cites this paper.

Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:17:50.333004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T16:17:50.215935Z digest=sha256:9d8b37f26fb7ab8d6a53111b5f087d87276a9c6e8de5063db7583fbd76fd88ce

Observation 71b73a86-1486-4778-a990-b970b84c1c16 · inbound

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators cites this paper.

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:28.467608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:18:28.467608Z digest=sha256:4f20b8c2528ccc22d1fb4862a5280d4c0467bc6203137b2b445d22f6803fe693

Observation 04c5bdb8-117f-4e82-9696-6c2fc7fe2736 · inbound

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications cites this paper.

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:42:36.224566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T18:53:18.645984Z digest=sha256:9128c74eab1edfac6d936dd205baaad838cf025da7e434d895632ee7b33e3246