Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:47:07.814079Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 3 inbound Pith citation observations for arXiv:2505.09027.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:47:07.814079Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:49.870701Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T21:47:27.669855Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0462b0fe-f9b9-449a-876d-c752e7a6d3b4 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation https://fireship.io/, 2017
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3650eac4-9627-493e-9f51-7021453ed6d3 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation https://huggingface.co/spaces/onekq-ai/WebApp1K-models-leaderboard, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dd21de74-1115-4612-92d6-aae255055870 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Fullstack React: The Complete Guide to ReactJS and Friends
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0d262345-6a76-4e2c-a032-ac6fb2a700c5 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation W., Tian, Z., and Barber, D
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3ff65f97-89f7-4589-ac24-76abb06bcb5f · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Program Synthesis with Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 720d92f7-4a2a-4aaa-9586-9164d968b40d · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Test Driven Development: By Example
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 76ea2de5-545f-4994-994b-648d53a26076 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Evaluating Large Language Models Trained on Code
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0584331-f6f9-4d53-872d-dae2d6bb9602 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Batch Prompting: Efficient Inference with Large Language Model APIs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04b4a062-6245-4b85-926f-29b4690499bb · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation INSTRUCTEVAL: Towards Holistic Evaluation of Instruction-Tuned Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 719c4ffa-361f-4604-a5ce-d629d2676e6d · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a0f1859-90db-4a62-83c0-739ebb42685c · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation A Survey on In-context Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 461be28f-c97b-4641-9370-7ef4f096f8d3 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd1fbb5f-f044-4772-8892-8889ee8514fc · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c5cb4a5-f725-440f-a4cf-25c7d7ffcea3 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Instruction Following without Instruction Tuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74e871a1-6597-4553-9e42-b263a33f2bae · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87ec06e8-8f97-4d0a-ac33-31ea501dc86e · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Qwen2.5-Coder Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e64c5eb5-2b86-49d0-8479-78cff52564e6 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Code Security Vulnerability Repair Using Reinforcement Learning with Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9525066-813c-4d68-9fb3-4c8eb248a001 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation LLM-Powered Code Vulnerability Repair with Reinforcement Learning and Semantic Reward
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation defd3581-e110-4905-9dc0-8d4a0f74517c · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Coarse-Tuning Models of Code with Reinforcement Learning Feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c69ee599-1f0a-498a-a851-125011ea847d · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03882f12-0034-442f-8a59-4e8313df98a2 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4adbef44-8446-42d9-8d7a-f93b43b868dd · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 915a6cd4-65f3-4ef4-a628-341ea80bc722 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation StarCoder: may the source be with you!
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fa98f54-f811-429b-a87d-1c8a57f09ffe · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Implicit In-context Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed84cfbb-69f1-41ea-ab58-a338c2f77e91 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Let's Verify Step by Step
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ea5523-8453-4789-bad1-2f6484f8211b · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c97fe7e4-7e51-47c6-9df7-1b3c0cff747c · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Large Language Model Instruction Following: A Survey of Progresses and Challenges
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b2882a-a8e7-4d3a-8d5b-abc2d9b55254 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation StarCoder 2 and The Stack v2: The Next Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df2d2159-68ab-4370-a72f-070494b27d7b · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Test-Driven Development for Code Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 271b15e7-e93d-4fdf-a334-d587f66cd78e · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation React framework
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3a550d67-7c0d-4c16-a61d-2ce2f3e3ca1c · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Mdn web docs
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5a25c69c-92ef-484e-8d02-4068ba50ab84 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Testing LLMs on Code Generation with Varying Levels of Prompt Specificity
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb585c0-1a04-4db3-b8ec-b7f23788334e · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Introducing swe-bench verified
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c5ec8b29-f0af-4db3-9c5c-a70c71903b42 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation LLM4TDD: Best Practices for Test Driven Development Using Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc88c425-18c2-4da5-b1dd-5dbbe5c68ea1 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation InFoBench: Evaluating Instruction Following Ability in Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d6b7c9b-327d-494d-8a9f-0cb3537a75e6 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb5ffa12-f980-4474-b14d-ee50c8e0538a · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 247971a5-568a-418d-99b1-f7ce0686afbf · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69f1cd5d-4d97-47f8-bf9a-4bac9cd57cc5 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d573ca1f-aeb5-4510-b5e1-5736206e424c · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Hypothesis search: Inductive reasoning with language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 19ed2752-1f40-477b-9d16-e94153571148 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26dfcfc7-7d51-424f-beb6-92c6198698e4 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Finetuned Language Models Are Zero-Shot Learners
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1636f878-e7bc-42ae-b828-47a5a3305420 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation An Explanation of In-context Learning as Implicit Bayesian Inference
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a8ee398-e969-410e-a2c8-3c5ff1b3792f · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation C.-J., Zhang, T., Patil, S
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a124e88e-1682-4684-a24d-b62c2944c9b1 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Codereval: A benchmark of pragmatic code generation with generative pre-trained models, 2023
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e55de092-178a-497a-9eb0-142bbdfa4e10 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 15bc3cae-9ad1-4898-a751-1b50d544cbeb · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b29a5d1f-77b3-4244-9464-5dc0aa06eeb8 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation A survey on self-play methods in reinforcement learning, 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d27b172-205f-4406-bbb0-0f5b22181f5a · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation write newline
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 234bdd77-f35f-4f1d-ac78-5dfb1b12e841 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation @esa (Ref
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 906089b1-26ce-4101-9caa-fd1507ce5f58 · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 009752c6-dfa6-4d46-8b19-e93ff9b46d5e · outbound
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9739363-1866-4844-9794-a79322847282 · inbound
Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
Reference 191
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a7f6292-c31c-4d4b-a303-b8ef39fabb6b · inbound
TDD Governance for Multi-Agent Code Generation via Prompt Engineering Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 599f82cd-22cb-4a15-ae19-d5bf46f4818e · inbound
TICoder: A Repository-Level Code Generation Framework with Test-Driven Planning and Implementation-Aware Reuse Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.