Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:18:29.153977Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 11 inbound Pith citation observations for arXiv:2508.09101.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:18:29.153977Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T15:13:50.935010Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
30 of 30 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 404202ab-2ac9-4e89-8da3-dda9575e996a · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f4636e5-23dd-41ba-983f-f6bc70d50431 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators Evaluating Large Language Models Trained on Code
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88c0a9ab-50c2-47c8-baa9-c206e77a5ff7 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators DeepSeek-V3 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc1b5865-60ff-4b34-ace1-859c6c1ec486 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e425ee9-5964-4615-9b3b-667575a21e79 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa5a5a28-0ac1-4775-851b-8c5c793c757c · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00d38d5b-e319-42f8-aae5-5374c6abcd54 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6975957d-e735-491d-823a-f8a3aa8a3483 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators A Survey on Large Language Models for Code Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67f9ac42-4e66-4c4a-9d08-e37b918d8067 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 374da8da-6e88-42f5-a76b-56d2633b4d27 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators Kimi K2: Open Agentic Intelligence
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 286b0d4e-c561-4888-9959-9117a668303d · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators doi: 10.1126/science.abq1158
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6400a4c-50e6-481e-bee3-ff5cc6e04de4 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17e42d19-1e94-48f3-aa19-8c9786923d78 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89d7ccf0-3ccd-4544-97a7-104087f96718 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators Qwen2.5 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f91e24b2-3419-459c-81de-f8f935449a3f · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators Seed-Coder: Let the Code Model Curate Data for Itself
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 198c1050-280b-453d-ac36-ed6b2c83b42c · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators Kaixin Wang, Tianlin Li, Xiaoyu Zhang, Chong Wang, Weisong Sun, Yang Liu, and Bin Shi
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f29a7ea-cecd-4d38-b937-1b72fb680195 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95dcbbed-88af-4ba4-8607-d4235f3ba155 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71b73a86-1486-4778-a990-b970b84c1c16 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd3f18c5-5e10-43a6-ad1f-60500951a413 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators doi: 10.18653/v1/2024.findings-acl.762
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b367fcbe-3f42-4b12-972c-c2789e8c36da · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d73353c-3220-4953-9140-324f45194dad · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators RefineCoder: Iterative Improving of Large Language Models via Adaptive Critique Refinement for Code Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13979a02-a595-4766-af05-c8b0caeed694 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 96c61cb4-bdfe-4e48-a81e-6586fbc8ad55 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators __main__
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8eefd89b-f45a-4c78-80ed-437b17880e0f · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators 22 # Code Benchmark Construction TaskIn order to build the code benchmark, I need you to help me create a Python function, as well as two test functions
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 93d537fe-9023-497f-ac4f-9767c15470aa · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators Qwen3 Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcc817ce-4445-42ca-aa1a-d741bb9ca7ad · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators Program Synthesis with Large Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97742432-eee2-45d9-8f71-b792fa671509 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70e884f0-4d1e-45ba-8361-a00a8099eb0a · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5114ecfe-fb52-4e44-8466-be759840ff37 · outbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7371962a-864c-4fe9-b8c9-13480715b815 · inbound
From Prompt Optimization to Multi-Dimensional Credibility Evaluation: Enhancing Trustworthiness of Chinese LLM-Generated Liver MRI Reports -- with Preliminary Extension to Lung Cancer AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c58ae96-276a-4872-823c-8b0dc95a06ac · inbound
SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 047e65da-4768-4396-9fd1-9824a011065d · inbound
BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 73863823-1d47-4baa-98a7-64b200170ad5 · inbound
BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b6d6a3cd-23bd-4de5-b7da-61bbe650cd73 · inbound
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 20d73053-a8b1-4e43-9d1c-3281bf908ec5 · inbound
PBT-Bench: Benchmarking AI Agents on Property-Based Testing AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6ce5e85f-b918-4633-a49a-516fb707b71b · inbound
PBT-Bench: Benchmarking AI Agents on Property-Based Testing AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 497e3519-f1de-4a52-9bd7-b0c5b14b951f · inbound
PBT-Bench: Benchmarking AI Agents on Property-Based Testing AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8bed6de6-6e17-4131-9582-0a657bbfb095 · inbound
Code-QA-Bench: Separating Code Reasoning from Documentation Memorization in Repository-Level QA AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3bdce69e-548b-480a-8975-a62cb8b0a0b5 · inbound
PERFOPT-Bench: Evaluating Coding Agents on Software Performance Optimization AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 452afdc8-54b1-40b5-894c-937eddea0652 · inbound
Cross-Domain Hybrid OPD for Generalizable Search Agents AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.