Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:57:13.152676Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2507.06980.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:57:13.152676Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fd190d38-f3f5-40f7-9983-5ec715a5a2e7 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation https://deepmind
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e420d59f-78dc-4941-9f59-94a7e5a1999c · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation https://platform.openai.com/docs/models/o1/
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0dd13243-0a6e-4ccb-b74b-5f82f94bdd86 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Let the llms talk: Simu- lating human-to-human conversational qa via zero-shot llm-to-llm interactions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dddb8680-2d31-4975-acb9-d35b4364e831 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 090910c8-f8c0-45e2-b788-b55841329f11 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Is github’s copilot as bad as humans at introducing vulnerabilities in code? Empirical Software Engineering , 28(6):129, 2023
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05f384a1-10a6-4683-be98-6259d406079d · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Language Models are Few-Shot Learners
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 045be678-fa73-44e7-a757-e3b6f3933bdc · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeT: Code Generation with Generated Tests
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 148742a7-2c10-4e69-8977-1e0c1515a2ed · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Evaluating Large Language Models Trained on Code
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 225d0a00-59af-4c61-91b4-d43e484106a3 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation LocAgent: Graph-Guided LLM Agents for Code Localization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80cf230c-d9ba-41d1-90f5-b62e0adc921f · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation TestART: Improving LLM-based Unit Testing via Co-evolution of Automated Generation and Repair Iteration
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32968c5c-115a-42f3-b597-5eced2799e7b · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e300668-1bd3-483e-8cd5-14f4e1a3a8cf · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Reasoning with Language Model is Planning with World Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0051cb81-9ce3-476a-8660-fb9c4255a904 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb95e5f-90fc-4f40-9179-e16ed179174e · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Adaptive mixtures of local experts
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ba85e74d-42fd-45f5-948b-472c83fb2b4a · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Devanbu, and Emily Morgan
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9698e8eb-768d-4d64-a7ba-e16e2bd72a79 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Self- planning code generation with large language models,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c72dfe49-e09c-4735-9f4c-533e027ca233 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9976e03-e007-4c23-ab53-7cb684635acf · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Grace: Discriminator- guided chain-of-thought reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cc63a596-dd09-4fde-9c68-dca2c4fcba1e · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Large language models are zero-shot reasoners
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3beeed5-0e39-4cf6-b8c5-fd66db7ff908 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65298404-a8eb-4a09-a9d3-95c87983cf7a · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0695e282-24b3-438e-9c2b-e81d33b5cdbe · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Structured chain-of-thought prompting for code generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 882c4230-c241-4e77-bd71-ce8f3d35ecb4 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19e30428-4e58-4c6d-a4bd-a0a367e5d86d · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bfbd08e-aa7c-45d3-8557-4f96bc069445 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Fastfixer: An efficient and effective approach for repair- ing programming assignments
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6c66ca81-1f1b-49cd-8d84-7eaa522da989 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e83878d-5eec-411e-9108-24a62f1728de · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b52eba2d-4f23-4672-88a9-095b01315893 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Refining chatgpt-generated code: Characterizing and mitigating code quality issues
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1257cabc-3675-4503-869c-657aec8b956a · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation No Need to Lift a Finger Anymore? Assessing the Quality of Code Generation by ChatGPT
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 799725ee-bf4a-4800-aec3-7672cdaf2d4e · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Bridging Code Semantic and LLMs: Semantic Chain-of-Thought Prompting for Code Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 625000bc-dd32-486c-9730-b64f0772c4aa · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Clarifygpt: A framework for enhancing llm-based code generation via requirements clarification
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb89cb4-0d81-4e13-aba0-cd6f81df5a2b · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ab23470-2b95-42f2-8ac8-d0fe97ba9b41 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dedb6c62-79e1-4381-af05-d974071d4622 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Code Llama: Open Foundation Models for Code
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db44968b-3743-45c5-aaa2-7eae9b221c1d · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Amazon codewhisperer, 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 74954589-49b8-41f1-8aa4-c9aae5fc6359 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Calibration and Correctness of Language Models for Code
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f73a6b6-88ab-49be-9bff-5e2aa94ad9c6 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2cd16ae-52fb-4962-b53a-3e929e395b55 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Bugs in Large Language Models Generated Code: An Empirical Study
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38217dfd-d3c4-471e-8674-c43996863060 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Code repair with llms gives an exploration-exploitation tradeoff
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 742f76d4-a924-4161-b7e2-6746303a1dee · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6faa49c-bc50-40aa-a7a2-05acbcfa9af2 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa0f77b2-7ff5-4ae6-929f-ed9c8c3642ef · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Expectation vs
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44cb0885-084c-434e-b6cc-0b8d01df9bad · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Chain-of-thought prompting elicits reasoning in large language models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9478c76-6c3f-4f23-b454-ca30fe6e3452 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Using github copilot to solve simple programming problems
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8ec48f14-0c06-4991-a1aa-9c8606220ce2 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Agentless: Demystifying LLM-based Software Engineering Agents
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c5a2ce-f01a-44d3-961e-dbda073ef8af · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Demystifying llm-based software en- gineering agents
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 189cd40d-6060-4559-9255-5475497315fb · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Self- evaluation guided beam search for reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6216f725-8dfc-4c42-a584-667df91f637a · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4486a841-2aad-405e-b4ff-bfb4a2205713 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Tree of thoughts: Deliberate problem solving with large language models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0213a368-cc4d-48a4-a603-1554521f1d9c · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Framework for evalu- ating code generation ability of large language models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8dbffacd-d2e2-48b2-93b9-8ecc1bef1d57 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27c36329-ccc9-4fb1-a4f2-2a3602d51c47 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Learning-based widget matching for migrating gui test cases
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd830f19-23cf-4218-b2d1-0d95568360ef · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation ToolChain*: Efficient Action Space Navigation in Large Language Models with A* Search
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3998ee99-e134-430a-a740-63a4965287f0 · outbound
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.