Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T23:43:15.246548Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 1 inbound Pith citation observation for arXiv:2511.04355.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T23:43:15.246548Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T07:16:30.725358Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T06:56:44.727065Z
11 of 11 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cad37e20-552c-414a-859e-82199c07ed36 · outbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b212e767-b003-45ec-9d31-97f35e21915d · outbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d89f27f7-d5a2-4b29-b2a9-4479aab96e78 · outbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Evaluating Large Language Models Trained on Code
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8f4ff59-3277-4774-950d-ec6952061c44 · outbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Program Synthesis with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b511eb66-97da-4e5a-9526-7986e81dd6df · outbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks A Comprehensive Survey of Contamination Detection Methods in Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f02549c-e2cb-47f2-ad33-7dd4011b3610 · outbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56be89f8-b821-4f69-b6f4-d5f5a7191cdb · outbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4ee9963-cdfa-4185-a2a3-c30eee9424e3 · outbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b213d7f-3344-4e7e-9e2a-b143a710c997 · outbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f913bb4-fb57-465d-95e6-93f62a46e99f · outbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Measuring Coding Challenge Competence With APPS
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab594daf-a75d-4197-8369-128f0ea6b1d5 · outbound
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a62c1ff0-61b9-4a91-91ea-ba6bd33f0908 · inbound
FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.