Pith. sign in

Paper Citation Record · LEDGER

A Survey on Evaluating Large Language Models in Code Generation Tasks

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2408.16498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.16498 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:39:48.434745Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 58a63251-dcd7-4788-9ef4-8bb4a215dfd2 · inbound

A Study of LLMs' Preferences for Libraries and Programming Languages cites this paper.

A Study of LLMs' Preferences for Libraries and Programming Languages A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:55:12.174917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T22:53:16.951417Z digest=sha256:2d466cd77cde682e29fa1fa8030c1f861ff61ecdbb39d3b28ecd621bffd818c7

Observation bd44c6f3-ecca-4d4e-ac0e-f6ae2a08d05f · inbound

CIgrate: Automating CI Service Migration with Large Language Models cites this paper.

CIgrate: Automating CI Service Migration with Large Language Models A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:48.434745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:48.434745Z digest=sha256:b2408d8e9eda60c0d059752dadd9554ef1fd88376b6dc0600e5514786c161e0c

Observation 29e1aff2-75d2-47e1-9d3b-26e46f87d92d · inbound

Rethinking Technology Stack Selection with AI Coding Proficiency cites this paper.

Rethinking Technology Stack Selection with AI Coding Proficiency A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T17:07:43.607070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:07:43.607070Z digest=sha256:52583c2bea60e9d26264d949184bcc74fdca81979df8214ccecb20af205ab0bb

Observation 70e86167-3ef9-40d7-ba54-f8481c8985b8 · inbound

Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries cites this paper.

Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:34:23.698211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T22:31:20.358758Z digest=sha256:b538a4a72b9b359c352272f39617671de579673e4bc11726478889d1e611f9ba

Observation 6d148974-2d34-4eb6-a86d-94d87adc80a9 · inbound

Sustainable Code Generation Using Large Language Models: A Systematic Literature Review cites this paper.

Sustainable Code Generation Using Large Language Models: A Systematic Literature Review A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:50:16.475540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T18:49:01.097179Z digest=sha256:5487d35f1a0178619003259247e513c30490a210c45294befa5a42bcbbf3bdd0

Observation 1307cc67-2121-4df9-98f5-2c49b7f36862 · inbound

Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios cites this paper.

Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:35:42.457281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:35:03.746555Z digest=sha256:21176222a055de76e34414f804d5ca58d4c5785a053ad7e61b0e8c9bbb83961c

Observation d979dbcb-b38f-4977-b0d5-73b14d51a5d4 · inbound

Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios cites this paper.

Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T16:42:13.327116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:42:13.327116Z digest=sha256:0cd1a2e9e31dba473f4ec9b17d1e9c09201bcc2a75715a11a58250befd1fef4d

Observation 0f744cc3-4b5c-4133-8937-b6b1320ac142 · inbound

When LLMs Lag Behind: Knowledge Conflicts from Evolving APIs in Code Generation cites this paper.

When LLMs Lag Behind: Knowledge Conflicts from Evolving APIs in Code Generation A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:01.515692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:48:11.705307Z digest=sha256:2305b6c8d6dfbd0722daa728884d1e99187152bda7a7116b22d581ca26371a1f

Observation dff3068f-f048-4f86-845d-7308681cc812 · inbound

Evaluation of LLM-Based Software Engineering Tools: Practices, Challenges, and Future Directions cites this paper.

Evaluation of LLM-Based Software Engineering Tools: Practices, Challenges, and Future Directions A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:21:52.985054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T02:52:55.120013Z digest=sha256:5e579e6dfe381a8298bd82c5dc8651e59dfda8588a0168483d71ae43f09f98bb

Observation c11e77b1-a70f-4c26-ac52-016eac86f864 · inbound

Leveraging LLMs for Multi-File DSL Code Generation: An Industrial Case Study cites this paper.

Leveraging LLMs for Multi-File DSL Code Generation: An Industrial Case Study A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:26:13.447868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T02:49:14.533263Z digest=sha256:0bf4e2b7d30930ec2dce631c7891058ffd0577223da0a3b7c30c875507b0ae54

Observation 179bd47b-4d3c-4d7a-b1a3-307365621512 · inbound

An Empirical Security Evaluation of LLM-Generated Cryptographic Rust Code cites this paper.

An Empirical Security Evaluation of LLM-Generated Cryptographic Rust Code A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:56:26.782309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T13:26:17.100399Z digest=sha256:3ced3eef1d03168c2fe72b7b500e9ebd06177f5e43ebbd9f609542e9134db0a4

Observation 1fc82384-6a5f-4bc2-af0c-a9d46c474e89 · inbound

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code cites this paper.

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:11.078592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T17:37:51.790000Z digest=sha256:dc786e5dc2a625ef7b7df3d0b9f2d47e02a60976f61cc27376918a460734f6b0

Observation 02537f36-3b2a-42e7-a1e8-21cbf4040725 · inbound

Contextualized Code Pretraining for Code Generation cites this paper.

Contextualized Code Pretraining for Code Generation A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:38:10.829807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T09:36:15.468902Z digest=sha256:440e44bdc3f37c851346cc1ab5875ccc51be241046e09a823c1d163e76a1a55d

Observation d29147dc-ee30-48b2-8434-2cf3b772f042 · inbound

Inferring Code Correctness from Specification cites this paper.

Inferring Code Correctness from Specification A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:33:31.547450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T06:33:36.835860Z digest=sha256:c4eb61b41de398feeb2cfd81f46fb031b8bc26ed3fabacba91c69aeca87a7bd9

Observation a2c76825-01b9-4849-a316-7590d42e7ff8 · inbound

LLM-Based Code Documentation Generation and Multi-Judge Evaluation cites this paper.

LLM-Based Code Documentation Generation and Multi-Judge Evaluation A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:55:45.140465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:38:55.804639Z digest=sha256:b6e187ce6ed5626255b273611d4c21a0f69cd76177f27d585c2730ecb946e5d5

Observation 1539c3cc-ffb3-4493-958c-66ee305803a5 · inbound

BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling cites this paper.

BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:29.732057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T17:41:23.922715Z digest=sha256:488bd313d5f0fabb688aa24a22b6cad8338ee8bce24bfbc3fe22e6db79c955b2

Observation ade76ee1-a86e-4feb-a2c4-7a3e1d4c0fe2 · inbound

Effectiveness of LLM-based Software Diversity for Reliability Improvement -- an Empirical Study cites this paper.

Effectiveness of LLM-based Software Diversity for Reliability Improvement -- an Empirical Study A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T04:21:23.867336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:21:23.867336Z digest=sha256:5f78e9eb38607f5f26470dfffc0ef63c22266c9468857224f4579a7306b716e3

Observation 76f4e29d-dd2f-4adb-862a-58c0aaad8dff · inbound

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs cites this paper.

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-10T13:57:06.865446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-10T13:49:17.343893Z digest=sha256:3428f92f62da95e2c84d13dc8e4569ee37144f0744ef067e6184cb9c2754cbff

Observation 9e1daf9c-3cdc-4972-9fc0-d32fc56393d8 · inbound

Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality cites this paper.

Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:01:33.319622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:01:33.319622Z digest=sha256:33e2b5e29abce9db59f797f8fe3dc6e7adb5d682c574f8a76610f107a6714aa9