Pith. sign in

Paper Citation Record · LEDGER

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency

As of 20 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 7 inbound Pith citation observations for arXiv:2505.02133.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02133 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T01:04:43.459105Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:48.502829Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:26:12.298078Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 63bb2c79-f556-49ea-a6f7-8a66d6ddcd94 · outbound

This paper cites GPT-4o System Card,.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency GPT-4o System Card,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:44.124310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.266413Z digest=sha256:18c5ebdfba81714a771ca4d8ff9118fa4444cb69cfa15dbaff4c460a67b928fa

Observation 891f2155-9151-4dc2-ab21-4d87868b876f · outbound

This paper cites The Claude 3 Model Family: Opus, Sonnet, Haiku Anthropic.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency The Claude 3 Model Family: Opus, Sonnet, Haiku Anthropic

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:44.108135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.271609Z digest=sha256:1f412b335cc7c5c3d4db88c67a7b87e904fd677cc705407d9d27b1c432254f25

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.276300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.276300Z digest=sha256:a02e0d1562a7cd3b195268215db0d4ff1b1c2706d2a92f23d2ee39f207e3f2fb

Observation 34f46a74-908c-4b6c-8ef4-e2089dbbd57c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Gemini: A Family of Highly Capable Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.281067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.281067Z digest=sha256:1946935c123bcb5df0100eea0e9ef615f6872636fc268167ea6be2f3414b2c01

Observation d6cc477e-051d-41b4-a8a8-a5972107493f · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size,.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Gemma 2: Improving Open Language Models at a Practical Size,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:44.092914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.285984Z digest=sha256:c47cf6612da9f45a6821068231233f1c3b9874f586636d9c66eecb79d15ba1c8

Observation 3cf27a4f-3245-4c05-be8a-9e65418cc5e5 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency LLaMA: Open and Efficient Foundation Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.290623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.290623Z digest=sha256:cef083d9852935c95d694a569f6381bd77cf2bc329550c688c755cf633300e4c

Observation 356a5fc6-1917-4645-aa4d-33c03899eb1f · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.295716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.295716Z digest=sha256:00a25b9808f9ecc1596b0667985103407306e284deb5395aadbe1a07adeb7e6b

Observation 2f83eb6b-2293-489b-8372-f75c999c899d · outbound

This paper cites Qwen2 Technical Report,.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Qwen2 Technical Report,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:44.078464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.301478Z digest=sha256:6e82f0c2dd66a360127dc0788925d0fece4f5c1fa3655612ffdd554e89aaf36b

Observation ed9ec64f-bb83-450b-a6ce-dc0377c002c2 · outbound

This paper cites Mixtral of Experts,.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Mixtral of Experts,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:44.063122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.305884Z digest=sha256:04632bed9ae2d077882d4722c38d76e71e709ac327dca95aef8bc4daf5f95221

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.310706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.310706Z digest=sha256:d2559a178a1219aa69c5c69487baced966829c198153735fab2e2943d0aa50ec

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.315027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.315027Z digest=sha256:fec88a4688ebcb85ffa8edc3cce0ba089729bb09519a2c9cdf7fa87bbc04ae34

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.319213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.319213Z digest=sha256:811801d4a59757d0c5ef01d5b563bda5497ec72c7c1a1c001a443ae3d921003d

Observation b8314428-cd69-4830-ab02-a5e74bfdba4f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Evaluating Large Language Models Trained on Code

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.323804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.323804Z digest=sha256:df2becd8a44f78c0c244341d7db2d365d8508ce74f215cd7b624f362a8544e35

Observation 8efe2e37-41fc-4496-80eb-b096aca127c8 · outbound

This paper cites From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.327578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.327578Z digest=sha256:a987f66261282d6d8702dbd9cb823d55c97be9e90835fc1e3c700bc64e77373d

Observation 13c4f10b-ffe2-4cc2-9b95-ee16f70fd48a · outbound

This paper cites Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.331734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.331734Z digest=sha256:6cefc58ce4a81defdc4ff8baffc8039eed65e41a658d6e899698bfb3ab7cfaf6

Observation ed483d6e-8357-407d-b9d9-70472e7a24f4 · outbound

This paper cites AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.336200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.336200Z digest=sha256:8eeff6aa31ccb914486d2e297b97499c6d6eb0f2732d05c76ec21eb0441506fe

Observation c3152481-f5b3-4957-8eb9-9c49009869ff · outbound

This paper cites MapCoder: Multi-Agent Code Generation for Competitive Problem Solving.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency MapCoder: Multi-Agent Code Generation for Competitive Problem Solving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.340868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.340868Z digest=sha256:ebf3a61d6e4ceea9fe48209c3cccbbd4645bd4d2a5d610a2a8028ddaa04bec26

Observation 397b354b-bcfc-42e0-a65e-f567584cbd13 · outbound

This paper cites Scaling Large Language Model-based Multi-Agent Collaboration.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Scaling Large Language Model-based Multi-Agent Collaboration

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.345621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.345621Z digest=sha256:f01aee0167878ef733bd60cf501d64fea4991ab612203908be241e9c82bf0dc4

Observation f205114c-7806-4330-9011-b1b6b66be203 · outbound

This paper cites ChatDev: Communicative Agents for Software Development.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency ChatDev: Communicative Agents for Software Development

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.350479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.350479Z digest=sha256:d3213ae757a3b99f3ccb661f492db5325cebd1f4e2edbdead315234c01c2d7e9

Observation 76be9be7-25f4-47af-a821-01f241a3546b · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework,.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:44.048565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.355059Z digest=sha256:e43b4a53298db6a03032fcbecce612c37af008708d71b591de9ff3a3a9a740f6

Observation a21010ec-3fb1-4770-8208-e1fcef4692ce · outbound

This paper cites Self-collaboration Code Generation via ChatGPT.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Self-collaboration Code Generation via ChatGPT

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.359761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.359761Z digest=sha256:8add698e03eb7ae4d0babf5c18ef1278164e5ffc85620505365a281ce74703e6

Observation 6c38a6e5-4dc2-4946-8027-ad502f73829b · outbound

This paper cites CY- CLE: Learning to Self-Refine the Code Genera- tion,.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency CY- CLE: Learning to Self-Refine the Code Genera- tion,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.364308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.364308Z digest=sha256:9b7d2a9b8194399cb0cc0fc5f0516a36f83440c664bd61ce5d8397b445971ce4

Observation e5d92926-ae7b-4961-aa24-17b5e79640b3 · outbound

This paper cites Debug like a Hu- man: A Large Language Model Debugger via Ver- ifying Runtime Execution Step-by-step,.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Debug like a Hu- man: A Large Language Model Debugger via Ver- ifying Runtime Execution Step-by-step,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:44.032991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.368839Z digest=sha256:5459448b82e25780eb86fb61b1cb57433426ae15d160d6bb6718b436f14310aa

Observation 5bbc7345-4b66-4432-9907-3920a25ce4e7 · outbound

This paper cites Leveraging Print Debugging to Improve Code Gen- eration in Large Language Models,.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Leveraging Print Debugging to Improve Code Gen- eration in Large Language Models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:44.015816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.373749Z digest=sha256:8779c5db838c6f44ead3e570fc4505a199ec549756fb09a1257d5a1c88ba4676

Observation e5ea0a8e-e10e-46fa-b1b4-1e235c1b6bec · outbound

This paper cites RGD: Multi-LLM Based Agent Debugger via Refinement and Generation Guidance.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency RGD: Multi-LLM Based Agent Debugger via Refinement and Generation Guidance

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.378374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.378374Z digest=sha256:94783a1e6342941ba6c53d7b50b92f99279bdbc6b15fd77c6a149e49ec99c80b

Observation 1d3b70fd-2043-4a3b-b87b-43f975e65c9c · outbound

This paper cites From Code to Correctness: Closing the Last Mile of Code Gen- eration with Hierarchical Debugging,.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency From Code to Correctness: Closing the Last Mile of Code Gen- eration with Hierarchical Debugging,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:43.999252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.383365Z digest=sha256:6e6ab670abc3889ef9bf6c861e0ca694edfce31ee3b6b77cfb3cd827c22d7c57

Observation 48863728-1fc6-4104-b23d-878133734bb2 · outbound

This paper cites SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.392179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.392179Z digest=sha256:c00cf616d010aa8a08515c1dce778ef87f27db16ffe56592aee0632f1c7fe7c3

Observation 04b168da-97fc-41e5-b35e-d44deba62306 · outbound

This paper cites ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.397129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.397129Z digest=sha256:2592e8d9db04229d057a6d6ca6cee18553809f858e75cd3172b372a8584ff1e6

Observation ce530737-14de-4e1a-9d60-7924b50b05db · outbound

This paper cites Teaching Large Language Models to Self-Debug,.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Teaching Large Language Models to Self-Debug,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:43.985492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.401659Z digest=sha256:dd1c5af6a87da01a7c891f974bd0569afb133f956e53546ce5363ba3f85281f3

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.406271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.406271Z digest=sha256:3b0f2b7d461901e1e6587a729618d5e2e8e8e9efd529984f61a0ecf86540694b

Observation dc747d94-56fa-43eb-bcd3-e24efdf0b5cc · outbound

This paper cites Is Your Code Generated by ChatGPT Really Cor- rect? Rigorous Evaluation of Large Language Mod- els for Code Generation.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Is Your Code Generated by ChatGPT Really Cor- rect? Rigorous Evaluation of Large Language Mod- els for Code Generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:43.969505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.411102Z digest=sha256:a8eadd991433bf0fdeb43196792c5424a0284e1e9cafbe31c85a9450472a13ce

Observation b74e3a94-0f41-4441-a020-d4bfe7c69838 · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.415696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.415696Z digest=sha256:1f9ed4b46267abbd7ed17ab6cc55ae37d1b7ec8ed39d91a7ac96a78dca80774f

Observation 4c15df1f-98a2-470d-b476-8660ae97f09a · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways,.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency PaLM: Scaling Language Modeling with Pathways,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:43.953782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.420769Z digest=sha256:674aad8cc281d0ad9e7bea042867b1ae41aa97061db4f6fe65766c61ac783e61

Observation 87f81ae2-f6df-43d1-bf1b-ee63a9359b47 · outbound

This paper cites Qwen Technical Report,.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Qwen Technical Report,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T01:04:43.937052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T01:04:43.424965Z digest=sha256:3282ed0002f3b7fc11f42c8e53175803b7d4b31cce30e53ca12a5bb5fdb1d6c2

Observation a5cb156e-f217-4a1c-a89a-9c90034eb762 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.428961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.428961Z digest=sha256:4c6e63007a5f5e55d94fc5f321b2d6706017158d7b5347b9b3c9d3b1cc1e42fe

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.433479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.433479Z digest=sha256:0aaa357cbf0bb438a56ed7f7408eea806a20b077cffc58032e20b2e59e394f74

Observation 37e5dbd0-3c1b-4c40-8042-9fb50580b369 · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.437476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.437476Z digest=sha256:b8a6d51487ced35d4c3dcc95c583502b1756abe86656f02d67c3f5a7aff842c3

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.441349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.441349Z digest=sha256:5482401f8eae57ff6df665997733ca2b247e451acb5a08965e1946e9e64f3fec

Observation 39909963-21db-4165-8d7f-221b5d1efbf5 · outbound

This paper cites Magicoder: Empowering Code Generation with OSS-Instruct.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Magicoder: Empowering Code Generation with OSS-Instruct

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.445511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.445511Z digest=sha256:1dc1c857ed9f8d9f92367b362071501e8f28ab970d98a80216629649d31e7f1d

Observation 3399c2de-76b1-4782-a35a-ff29a92b24bb · outbound

This paper cites OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.449698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.449698Z digest=sha256:71a5ffd5b2bcf969e3849a6a988ae6c807202aa6df4e20a4d1107d94cccb0586

Observation 15b5655c-2c1c-4b8a-9e5e-d27cce6b770b · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency StarCoder 2 and The Stack v2: The Next Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.454572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.454572Z digest=sha256:bcd6c2269bfeb0a149bb51215f2e40c15213bf2ef373678b38b0c97688b6db5a

Observation dee68f70-8c12-4d26-a275-425b2f933eac · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T01:04:43.459105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:04:43.459105Z digest=sha256:5cced282b2a44ea24e97a7181eb6a5c8544ff533ae1ea0fd615c684356ed5d15

Pith citing papers

Observation 0793178c-9d13-43b2-9189-2a21611f57ec · inbound

Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI cites this paper.

Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency

Reference 176

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:48.502829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:48.502829Z digest=sha256:78b9e813cb7726dd8050bd16173c83912b9929546ce885f1d353a4083a3f5bd4

Observation 0f8e2ab5-3c98-4d47-9e65-b73712e18143 · inbound

Position Paper: Programming Language Techniques for Bridging LLM Code Generation Semantic Gaps cites this paper.

Position Paper: Programming Language Techniques for Bridging LLM Code Generation Semantic Gaps Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:07:08.007222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:07:08.007222Z digest=sha256:9fbbbc6b8e4885c285dd9b783a3c5f4a10899075dab050872b81ee4c6c43310f

Observation 349621a9-5105-4119-ad89-854ca063cbbb · inbound

WildCode Revisited: A Comprehensive Empirical Study on the Security of LLM-Generated Code cites this paper.

WildCode Revisited: A Comprehensive Empirical Study on the Security of LLM-Generated Code Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T18:41:17.920359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:41:17.920359Z digest=sha256:863ce522ffbed063a484190207901ae1969c149d92ecc0c71bcbd305567205fb

Observation 8ac82032-2948-4021-94d7-9ffae6761838 · inbound

Cascaded Code Editing: Large-Small Model Collaboration for Effective and Efficient Code Editing cites this paper.

Cascaded Code Editing: Large-Small Model Collaboration for Effective and Efficient Code Editing Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:56:05.616333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T02:35:55.075730Z digest=sha256:7ce4e91b341ecbf8e56f40137baf0a3182897447e5146dcfdfec7fe75e7cc4bd

Observation 6d8f5d4b-db41-4aa0-9144-2998e43f7c0a · inbound

The Infinite Mutation Engine? Measuring Polymorphism in LLM-Generated Offensive Code cites this paper.

The Infinite Mutation Engine? Measuring Polymorphism in LLM-Generated Offensive Code Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:56:30.177624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T15:35:02.934397Z digest=sha256:a3f13129a2d190bd71a4d014fffb0d70d4ef9078d1fb69c379d6430a76ffd5f6

Observation 914550fb-d923-462a-b161-56bb1db492a3 · inbound

The Infinite Mutation Engine? Measuring Polymorphism in LLM-Generated Offensive Code cites this paper.

The Infinite Mutation Engine? Measuring Polymorphism in LLM-Generated Offensive Code Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:39.553334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T18:39:02.963388Z digest=sha256:8f07bb881abeb9fe5d9dcde819acf53b76ca406363f2bb4ea7d852684c72107c

Observation de17b9d3-59f7-4d1a-aee2-20ecf94db210 · inbound

How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval cites this paper.

How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval Enhancing LLM Code Generation: A Systematic Evaluation of Multi-Agent Collaboration and Runtime Debugging for Improved Accuracy, Reliability, and Latency

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.299701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T21:14:01.070601Z digest=sha256:b1a725883cb8cf17bc6a8566aa4b0c4349278eda27696bae5e2db20cc45c71d6