Pith. sign in

Paper Citation Record · LEDGER

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation

As of 7 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2507.06980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06980 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:57:13.152676Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd190d38-f3f5-40f7-9983-5ec715a5a2e7 · outbound

This paper cites https://deepmind.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation https://deepmind

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.934083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:12.986869Z digest=sha256:3fe15c5e725df12f4bce79ec9d1642167268ea6595308d1e7e7e5e194b9f4f0c

Observation e420d59f-78dc-4941-9f59-94a7e5a1999c · outbound

This paper cites https://platform.openai.com/docs/models/o1/.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation https://platform.openai.com/docs/models/o1/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.925186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:12.990649Z digest=sha256:669adde25661a7a448dd5b36e507cced6c216cdafa60ff8abe4ea5bf37d6d6f0

Observation 0dd13243-0a6e-4ccb-b74b-5f82f94bdd86 · outbound

This paper cites Let the llms talk: Simu- lating human-to-human conversational qa via zero-shot llm-to-llm interactions.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Let the llms talk: Simu- lating human-to-human conversational qa via zero-shot llm-to-llm interactions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.916708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:12.994078Z digest=sha256:288acf9f1e7f7b759260c56c4ce3d53391fef47f13ca62ec3acae889cc8b35fb

Observation dddb8680-2d31-4975-acb9-d35b4364e831 · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:12.997242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:12.997242Z digest=sha256:73662106c2944f747d3dd7faa0a54fd4cd83bd31c1869d4cd67ea5c43a6fd2cd

Observation 090910c8-f8c0-45e2-b788-b55841329f11 · outbound

This paper cites Is github’s copilot as bad as humans at introducing vulnerabilities in code? Empirical Software Engineering , 28(6):129, 2023.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Is github’s copilot as bad as humans at introducing vulnerabilities in code? Empirical Software Engineering , 28(6):129, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.908021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.000811Z digest=sha256:df6db19fc4933fd3fdf3e20f6a0927135ce9ca2c46132b4905e663be13c50ff7

Observation 05f384a1-10a6-4683-be98-6259d406079d · outbound

This paper cites Language Models are Few-Shot Learners.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Language Models are Few-Shot Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.003900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.003900Z digest=sha256:3adb54f84399bcfa0dd0d21c59b27748f4f3b1d88f8fb72a8f0a8552e8b3c8a0

Observation 045be678-fa73-44e7-a757-e3b6f3933bdc · outbound

This paper cites CodeT: Code Generation with Generated Tests.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeT: Code Generation with Generated Tests

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.007465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.007465Z digest=sha256:fb091ec27626b9a17ad62c7fca5265b69eb9626ae5b3ab7dece5e8cbcbe86587

Observation 148742a7-2c10-4e69-8977-1e0c1515a2ed · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.010764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.010764Z digest=sha256:6c3df31127dbe210d93c174f7b635e2617a14442a04daea51c79aadc31455d06

Observation 225d0a00-59af-4c61-91b4-d43e484106a3 · outbound

This paper cites LocAgent: Graph-Guided LLM Agents for Code Localization.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation LocAgent: Graph-Guided LLM Agents for Code Localization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.014157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.014157Z digest=sha256:b2c2f8b11c32a7adb879366e387c9cfd15cab5a6dc3a307ec5b8941d4c6e86b3

Observation 80cf230c-d9ba-41d1-90f5-b62e0adc921f · outbound

This paper cites TestART: Improving LLM-based Unit Testing via Co-evolution of Automated Generation and Repair Iteration.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation TestART: Improving LLM-based Unit Testing via Co-evolution of Automated Generation and Repair Iteration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.017270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.017270Z digest=sha256:60d6f193581c9b830b0f6ef0d255af6f1484587f3726561f4750844375f6e918

Observation 32968c5c-115a-42f3-b597-5eced2799e7b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.020754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.020754Z digest=sha256:b7c68a7a6087d6763d9f525d72849941bb44e7dc33e56b1145eed31508b169b5

Observation 9e300668-1bd3-483e-8cd5-14f4e1a3a8cf · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Reasoning with Language Model is Planning with World Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.023840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.023840Z digest=sha256:1f6e03b639fcf6609af73eb1c67b298af27926109bff620c5daa4bb79f52f5a1

Observation 0051cb81-9ce3-476a-8660-fb9c4255a904 · outbound

This paper cites CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.026859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.026859Z digest=sha256:f3622f18c893af16b8e4aff6b7fadebdf8c86b911b90289049430ef8022ce78f

Observation ffb95e5f-90fc-4f40-9179-e16ed179174e · outbound

This paper cites Adaptive mixtures of local experts.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Adaptive mixtures of local experts

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.899144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.029804Z digest=sha256:ca23a4b31ee2610e0f45c4fcf1c28b158d47479f44eaee49fc374b68f0bc812b

Observation ba85e74d-42fd-45f5-948b-472c83fb2b4a · outbound

This paper cites Devanbu, and Emily Morgan.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Devanbu, and Emily Morgan

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.032645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.032645Z digest=sha256:0d213a8c7a2c0134bafde1bc394df333aea42a5fe645c55ba0e8754523f8c63d

Observation 9698e8eb-768d-4d64-a7ba-e16e2bd72a79 · outbound

This paper cites Self- planning code generation with large language models,.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Self- planning code generation with large language models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.890614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.035407Z digest=sha256:39c6b7dc14a1112c041bc4ae54b5a6cd0e4cbad4f00b6d3f5ee94bc6e5c27044

Observation c72dfe49-e09c-4735-9f4c-533e027ca233 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.041485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.041485Z digest=sha256:e5242e7e1f2b2d03f81bd860920791eb580cfc3002888ed4e18aac54bad52a58

Observation d9976e03-e007-4c23-ab53-7cb684635acf · outbound

This paper cites Grace: Discriminator- guided chain-of-thought reasoning.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Grace: Discriminator- guided chain-of-thought reasoning

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-06T18:57:13.499467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.044893Z digest=sha256:26f0a60a9cccfb69454705ba2afd3db8bdb27cac0a8c77070f2fd037c0e8f5b3

Observation cc63a596-dd09-4fde-9c68-dca2c4fcba1e · outbound

This paper cites Large language models are zero-shot reasoners.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Large language models are zero-shot reasoners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.881690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.047794Z digest=sha256:caa2feebcb60ebdf2c46640c9692d2827b654ab62d58c17c200f246f32dea65a

Observation c3beeed5-0e39-4cf6-b8c5-fd66db7ff908 · outbound

This paper cites CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.050671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.050671Z digest=sha256:cbf24c5b87ea544f672a3a25774971b1ee2fd474b5e103b2e86f91c06f566e41

Observation 65298404-a8eb-4a09-a9d3-95c87983cf7a · outbound

This paper cites CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.053921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.053921Z digest=sha256:1d3c35419f71d33af6cb9cc337333d8a633fa8922fa4c9b7231e9d05f37081b9

Observation 0695e282-24b3-438e-9c2b-e81d33b5cdbe · outbound

This paper cites Structured chain-of-thought prompting for code generation.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Structured chain-of-thought prompting for code generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.872380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.056940Z digest=sha256:e7591e8828c16621f091eb1928ff891678e74d20949d61c28b8d1873132740d7

Observation 882c4230-c241-4e77-bd71-ce8f3d35ecb4 · outbound

This paper cites CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeTree: Agent-guided Tree Search for Code Generation with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.060665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.060665Z digest=sha256:e39546150d700e9e596e07fd753b264b406d82bbecc982b8e16e1b09154394d5

Observation 19e30428-4e58-4c6d-a4bd-a0a367e5d86d · outbound

This paper cites Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.064699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.064699Z digest=sha256:3e8eb4cbc007a959634d95ed8e29c1d5d3061a9cace69eb3cd163ef253e260fc

Observation 9bfbd08e-aa7c-45d3-8557-4f96bc069445 · outbound

This paper cites Fastfixer: An efficient and effective approach for repair- ing programming assignments.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Fastfixer: An efficient and effective approach for repair- ing programming assignments

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.863794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.068804Z digest=sha256:6b2f1543f8c32ae4edb0836e6a16db64f2e4273687cb0df9f9039c91968b1c0d

Observation 6c66ca81-1f1b-49cd-8d84-7eaa522da989 · outbound

This paper cites Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.854116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.071864Z digest=sha256:42ec4a5572ce3e5edbabecd6475f81958201853cf84ca207aa54db9b05c5e393

Observation 8e83878d-5eec-411e-9108-24a62f1728de · outbound

This paper cites Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.843107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.074817Z digest=sha256:6b666725483d60d928af03c7e9bb074cc6006ff4ed8befec487a8020e621fed0

Observation b52eba2d-4f23-4672-88a9-095b01315893 · outbound

This paper cites Refining chatgpt-generated code: Characterizing and mitigating code quality issues.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Refining chatgpt-generated code: Characterizing and mitigating code quality issues

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.834059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.077529Z digest=sha256:5286baef9d13c92c54627702b89fea296a924d650fd62d1199e8f3a30301c517

Observation 1257cabc-3675-4503-869c-657aec8b956a · outbound

This paper cites No Need to Lift a Finger Anymore? Assessing the Quality of Code Generation by ChatGPT.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation No Need to Lift a Finger Anymore? Assessing the Quality of Code Generation by ChatGPT

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.080535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.080535Z digest=sha256:0927d48ab6516ebad0d37cc3a0b4d985705f66fa654051dc824d3bd232221076

Observation 799725ee-bf4a-4800-aec3-7672cdaf2d4e · outbound

This paper cites Bridging Code Semantic and LLMs: Semantic Chain-of-Thought Prompting for Code Generation.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Bridging Code Semantic and LLMs: Semantic Chain-of-Thought Prompting for Code Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.083644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.083644Z digest=sha256:f5ae31265c8aaa373435dc62ed9ab32c782704d07aedcc2d1e15673467023655

Observation 625000bc-dd32-486c-9730-b64f0772c4aa · outbound

This paper cites Clarifygpt: A framework for enhancing llm-based code generation via requirements clarification.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Clarifygpt: A framework for enhancing llm-based code generation via requirements clarification

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.086901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.086901Z digest=sha256:7e4b2d1504fa20c272a7320d8afe242d188674f077d852003e236c5a959d7503

Observation cdb89cb4-0d81-4e13-aba0-cd6f81df5a2b · outbound

This paper cites CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.089866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.089866Z digest=sha256:76f4414f21012906e25b45e7c5d823f745b5d7f47967312d4781b87d3fe6ef82

Observation 1ab23470-2b95-42f2-8ac8-d0fe97ba9b41 · outbound

This paper cites Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.092944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.092944Z digest=sha256:06d6b100b51385486fdadf404d95a1d601259a53ec8b6014b6896b86cb80c98e

Observation dedb6c62-79e1-4381-af05-d974071d4622 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Code Llama: Open Foundation Models for Code

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.096040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.096040Z digest=sha256:f6319da40b0574da97bf943bf1b9413a51ab29d43a80599dc2460d5195855e78

Observation db44968b-3743-45c5-aaa2-7eae9b221c1d · outbound

This paper cites Amazon codewhisperer, 2024.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Amazon codewhisperer, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.824385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.099281Z digest=sha256:3fe52c39aca66f6307ebfb922b0f719075772d28821180103a1b9e110a8d7498

Observation 74954589-49b8-41f1-8aa4-c9aae5fc6359 · outbound

This paper cites Calibration and Correctness of Language Models for Code.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Calibration and Correctness of Language Models for Code

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.102050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.102050Z digest=sha256:8231f7dcb8b86b4680f8d3f637adb94a37c5749437ae28205258a227ab396972

Observation 0f73a6b6-88ab-49be-9bff-5e2aa94ad9c6 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.105318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.105318Z digest=sha256:8f6cca56f2d55e4a055b3d4c226cfe790f9cf7cbc2a32d95b5d7ce2ddf29d3bf

Observation c2cd16ae-52fb-4962-b53a-3e929e395b55 · outbound

This paper cites Bugs in Large Language Models Generated Code: An Empirical Study.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Bugs in Large Language Models Generated Code: An Empirical Study

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.108562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.108562Z digest=sha256:e6a05d3df91ee37d935a2e93d96b2b4a3b27d669775674565ae38ed4505ea431

Observation 38217dfd-d3c4-471e-8674-c43996863060 · outbound

This paper cites Code repair with llms gives an exploration-exploitation tradeoff.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Code repair with llms gives an exploration-exploitation tradeoff

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.815606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.111596Z digest=sha256:e7717accb433c6316bb66c69d4e4bb17f877abe877945bdbe5bf216fcbed56a9

Observation 742f76d4-a924-4161-b7e2-6746303a1dee · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.114357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.114357Z digest=sha256:8522ff7a22d2cbb5ef7deb1638d6bfccdf73ad5f5b65ae82109355894b6c1266

Observation b6faa49c-bc50-40aa-a7a2-05acbcfa9af2 · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.117461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.117461Z digest=sha256:5e3d106aa5a2cd2a29d20ed96acafb707047e35229b4fa65ac3b1a51508f4d35

Observation fa0f77b2-7ff5-4ae6-929f-ed9c8c3642ef · outbound

This paper cites Expectation vs.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Expectation vs

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.801421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.120602Z digest=sha256:2755b2e7b7b5514f85353f48064b19816c631d93fb3137ec01f5a54a68685d28

Observation 44cb0885-084c-434e-b6cc-0b8d01df9bad · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Chain-of-thought prompting elicits reasoning in large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.123452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.123452Z digest=sha256:765b7a8abd5e28032bf84f250fd76961200cc18fb7fa22f0b69594fc684dc56f

Observation b9478c76-6c3f-4f23-b454-ca30fe6e3452 · outbound

This paper cites Using github copilot to solve simple programming problems.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Using github copilot to solve simple programming problems

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.786801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.126255Z digest=sha256:11296d9e4175dc14e62950cacd3faf6a6d002cc752f29fccc53208254bd270ca

Observation 8ec48f14-0c06-4991-a1aa-9c8606220ce2 · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Agentless: Demystifying LLM-based Software Engineering Agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.129076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.129076Z digest=sha256:3f7096a0c4d33dfcc5254cfb93f567b0462968cd073381be404fc2fcb928bd1b

Observation 21c5a2ce-f01a-44d3-961e-dbda073ef8af · outbound

This paper cites Demystifying llm-based software en- gineering agents.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Demystifying llm-based software en- gineering agents

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.777752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.132297Z digest=sha256:0f64e5935faccf625e4fb29566e189cea7f00922085a3434d6192f32d3efdac3

Observation 189cd40d-6060-4559-9255-5475497315fb · outbound

This paper cites Self- evaluation guided beam search for reasoning.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Self- evaluation guided beam search for reasoning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.769096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.134962Z digest=sha256:3d3ffbb47a6f9ded595604f3cb61f30a481e4f3bdf06484ebe94f14b9b6bb3f4

Observation 6216f725-8dfc-4c42-a584-667df91f637a · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.137811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.137811Z digest=sha256:1fa0cd09e2b1b2da221db8af83f7b0c545719e15822a101fbe7d46ef4491ccca

Observation 4486a841-2aad-405e-b4ff-bfb4a2205713 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Tree of thoughts: Deliberate problem solving with large language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.760345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.140899Z digest=sha256:def5b5effb2a06bef4824101e6900d5b9a5fddc2b2ed524f829117d5df6cf4c2

Observation 0213a368-cc4d-48a4-a603-1554521f1d9c · outbound

This paper cites Framework for evalu- ating code generation ability of large language models.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Framework for evalu- ating code generation ability of large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:57:13.751610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:57:13.143896Z digest=sha256:977b558eb85d5b09cdb0f533c4fb038bac25b3ba99e3549c19fb869fda43e147

Observation 8dbffacd-d2e2-48b2-93b9-8ecc1bef1d57 · outbound

This paper cites Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.146831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.146831Z digest=sha256:d2e46d621409642098a91ef9ade2aadd239919b797155edb3c69da34280796c2

Observation 27c36329-ccc9-4fb1-a4f2-2a3602d51c47 · outbound

This paper cites Learning-based widget matching for migrating gui test cases.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Learning-based widget matching for migrating gui test cases

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.149843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.149843Z digest=sha256:d9c1c484d3ff19e50c5852f8342c19940f6e4197bd83ad1ca61e661e3188ceb0

Observation cd830f19-23cf-4218-b2d1-0d95568360ef · outbound

This paper cites ToolChain*: Efficient Action Space Navigation in Large Language Models with A* Search.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation ToolChain*: Efficient Action Space Navigation in Large Language Models with A* Search

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.152676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.152676Z digest=sha256:eefb9c57d5b1108bde8afeec15c6b42b4cfe0ccf8b1548c1f3767617aaf14cc5

Observation 3998ee99-e134-430a-a740-63a4965287f0 · outbound

This paper cites an unresolved cited work.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:13.038475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:13.038475Z digest=sha256:491653c3abe494dcfa013553e7f4215419c5fb7cebd29756f72661b5a0055316

Pith citing papers

No inbound Pith citation observations are available.