Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:46:23.635504Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 8 of 8 outbound references and 0 inbound Pith citation observations for arXiv:2502.00226.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:46:23.635504Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
8 of 8 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c0188b4d-f77f-400f-b809-c542f336ef2e · outbound
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5472529-8f3f-4855-a760-60bbce9b0ebd · outbound
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45939f3d-620a-46b6-85e1-1f2ce27f1231 · outbound
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d1a30ce-b9ac-4416-b692-f81861dea393 · outbound
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b2843a-be17-4e07-a091-5aa9935f771a · outbound
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems Enhancing Large Language Models in Coding Through Multi-Perspective Self-Consistency
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abfbb0e0-656b-4b1e-92a5-d5f25c2dccbd · outbound
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0156aa4-85de-49ac-8a05-1faef45343f4 · outbound
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems AgentBench: Evaluating LLMs as Agents
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b1e7f17-7f48-4c2c-952d-dd87aa842518 · outbound
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.