Pith. sign in

Paper Citation Record · LEDGER

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks

As of 4 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 1 inbound Pith citation observation for arXiv:2511.04355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.04355 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T23:43:15.246548Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T07:16:30.725358Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:56:44.727065Z

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cad37e20-552c-414a-859e-82199c07ed36 · outbound

This paper cites an unresolved cited work.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:14.703226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:14.703226Z digest=sha256:f65da3833639becde98f76fe31d9accdf66b71b84aa2f6d38d1db65e43934fa0

Observation b212e767-b003-45ec-9d31-97f35e21915d · outbound

This paper cites an unresolved cited work.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:14.735306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:14.735306Z digest=sha256:360fcbc684e3c1c38774098a46984c45164c219a1385d8b61fbd4fae693f11fa

Observation d89f27f7-d5a2-4b29-b2a9-4479aab96e78 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Evaluating Large Language Models Trained on Code

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:14.786547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:14.786547Z digest=sha256:85332b6fae8b69ac290315cf0ee07093b7e78b917dc2762caa81d33c31317077

Observation a8f4ff59-3277-4774-950d-ec6952061c44 · outbound

This paper cites Program Synthesis with Large Language Models.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:14.853124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:14.853124Z digest=sha256:1333bcee695b26e19ed3f1ec2bc2537fa61b549c9cf25bdbd5590a8462a599c1

Observation b511eb66-97da-4e5a-9526-7986e81dd6df · outbound

This paper cites A Comprehensive Survey of Contamination Detection Methods in Large Language Models.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks A Comprehensive Survey of Contamination Detection Methods in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:14.882558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:14.882558Z digest=sha256:e27d5fcdf8300a00c2ec4e0e709b7b97554f7030cf8768f4831a140d5c3d5aa6

Observation 0f02549c-e2cb-47f2-ad33-7dd4011b3610 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:14.940179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:14.940179Z digest=sha256:8d797ac37193c8354448cb55d535ec9da257bf9d9ffa956b4ea511a0bbfa21d4

Observation 56be89f8-b821-4f69-b6f4-d5f5a7191cdb · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:14.983161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:14.983161Z digest=sha256:c840949eca689b2f1cec411b657bc547de1fe5afe5a7d0036d97119569f87751

Observation e4ee9963-cdfa-4185-a2a3-c30eee9424e3 · outbound

This paper cites an unresolved cited work.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:15.074745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:15.074745Z digest=sha256:90380c79adb3ffae1bbff0cfc802c401ede190ba5c01ec377f505caf95ea8d93

Observation 3b213d7f-3344-4e7e-9e2a-b143a710c997 · outbound

This paper cites HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:15.114501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:15.114501Z digest=sha256:71f7357e3635a871e5a437567d6091c8b5969814a1e8ffd3742bec1f2a56f6cd

Observation 9f913bb4-fb57-465d-95e6-93f62a46e99f · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Measuring Coding Challenge Competence With APPS

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:15.170191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:15.170191Z digest=sha256:81028cec5c7150d41d5fa785fa4503a4c8495aada1d67d8634a6d10028bae1ba

Observation ab594daf-a75d-4197-8369-128f0ea6b1d5 · outbound

This paper cites an unresolved cited work.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:15.246548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:15.246548Z digest=sha256:1ff7a165d56bb014c90f67045326419a1804fa6b76d197e702e34165df696325

Pith citing papers

Observation a62c1ff0-61b9-4a91-91ea-ba6bd33f0908 · inbound

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses cites this paper.

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:14.138901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T07:16:30.725358Z digest=sha256:372a09dd80b6e5fdc127541d1c4fc7c13b3e2380aa7e045b036de22f4c458cc1