Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:14:57.176868Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 1 inbound Pith citation observation for arXiv:2508.13757.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:14:57.176868Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-08T17:58:54.689999Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-09T06:55:42.738972Z
9 of 9 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 00a574af-c152-449d-b83b-80f9c7d668c1 · outbound
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models Evaluating Large Language Models Trained on Code
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dca436d3-46f3-401f-b525-01cfac0e62f2 · outbound
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models Program Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff5ac8a7-4678-434c-b26c-f595df2f3d8d · outbound
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ce98c4a-db03-4b39-8122-7ebc8fade0fc · outbound
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models Hackerrank-astra: A benchmark for evaluating llms in coding competitions,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 61a6db8a-f570-4ad7-96fa-a829041d8ecc · outbound
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models Managing technical debt with the sqale method,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 418899ed-cb2e-4fba-b96d-7021727ba079 · outbound
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models A metrics suite for object oriented design,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 982d7808-e294-4979-8dd3-d48ac183144f · outbound
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models Measuring Coding Challenge Competence With APPS
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a5c496-06e0-4d3e-bb3e-4b09d06320ae · outbound
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001f5b98-4def-4c60-9095-fbcb3ea4ece0 · outbound
COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models Modeling the performance prediction problem in industrial and organizational psychology,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b94b25d9-0201-41c3-bc78-75150cac7030 · inbound
ARIADNE: Agentic Reward-Informed Adaptive Decision Exploration via Blackboard-Driven MCTS for Competitive Program Generation COMPASS: A Multi-Dimensional Benchmark for Evaluating Code Generation in Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.