Pith. sign in

Paper Citation Record · LEDGER

Evaluating LLM-Generated Code: A Benchmark and Developer Study

As of 14 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2605.09059.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.09059 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:55:49.377456Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact12
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c81cd13-b026-4d1d-85f3-527c77d0b89d · outbound

This paper cites Multi-lingual Evaluation of Code Generation Models.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Multi-lingual Evaluation of Code Generation Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.801549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:d286a86c1206b87f8a5f73a02df8ec85accf3cfbbce0c99cfd709dbcf60e2b07

Observation d14f189c-a05e-45b4-b15f-b5f1eb49c966 · outbound

This paper cites Program Synthesis with Large Language Models.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Program Synthesis with Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:35:46.804017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:737fed6f87103c4c49eae5dcbf74a7a8c6cf67a16a68cf4e9a78780d54c35aa1

Observation dc07e2ef-d11f-4c9c-9190-4e39167c3f3b · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:35:46.806529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:c14a2688cb1f5b3083c32a0466a39b37da535c0f06d58f60712d91c65190dddc

Observation 75aebb0f-9b15-4152-a351-77a32e5f0155 · outbound

This paper cites 2025.The Temperature Parameter | DeepSeek API Docs.

Evaluating LLM-Generated Code: A Benchmark and Developer Study 2025.The Temperature Parameter | DeepSeek API Docs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T12:23:48.901394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:308ce5bd7f0a5e129cbc720bcdc30fda9bca40c21c63ccaf76c6ccb76dc25e6c

Observation 2d450392-3896-42c8-ad21-3a69b35ea363 · outbound

This paper cites Evaluating Large Language Models in Class-Level Code Generation.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Evaluating Large Language Models in Class-Level Code Generation

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:05:06.703990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:bf5d937d25ed0c1d59b0176791b3642feeedf52dbb75fefef35404384d7fec96

Observation 9c21b4fd-54fb-404a-9f69-7c29d61068cf · outbound

This paper cites Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.809139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:41a34f2bff6f67153f795ac591552c45fa9931ba37db4a82097aa3a9bdfb4e41

Observation 89f2bcf6-def9-438e-902a-6baa08435917 · outbound

This paper cites Mapping Language to Code in Programmatic Context.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Mapping Language to Code in Programmatic Context

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:35:46.796290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:ed9ee6bb95efd26b2a9ed8db81cfba2aeabe1f84115a9e57e2eb56c0c3aff91a

Observation 2b046d03-c903-4df3-bb30-a480c0342ff2 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Evaluating LLM-Generated Code: A Benchmark and Developer Study SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T13:35:46.793435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:3ed1cb6a233faf383c12527853dd72874b02cae89e059014c4217e9870efaa10

Observation cacf0642-5b7b-4c18-9d8d-cc6fc6eec3a1 · outbound

This paper cites Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:35:46.788186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:9891e6b5bf9fa8d9bb49bac6ac870c61382164cce72d81d1c672b4d74e4d7a01

Observation fe160874-8e8e-4a3a-b6a2-43b1ad12a4b1 · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

Evaluating LLM-Generated Code: A Benchmark and Developer Study RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:35:46.790813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:a0f5ed3761df83aa1503ed4e7d53f179337de9a2685cbec30d0c46c1aec31382

Observation eb790eb8-54f5-4a2f-a241-1ca8858310ef · outbound

This paper cites an unresolved cited work.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Unresolved cited work

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:05:06.697378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:e89772d0c6271a161bcc6130f0d4ff88bba7c694473c388297705468dc8f4f87

Observation a8410229-2d97-43c4-8988-e9103464f083 · outbound

This paper cites an unresolved cited work.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Unresolved cited work

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:05:06.700633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:eaefcae4ae1e3f9bcd3afa6274e6eee39187bb8123d7caf48242ae2066c1da50

Observation 0025d0a0-b272-4444-86a4-1a5cb2296a60 · outbound

This paper cites 2025.Homepage | SonarQube Cloud | Sonar Documentation.

Evaluating LLM-Generated Code: A Benchmark and Developer Study 2025.Homepage | SonarQube Cloud | Sonar Documentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T12:23:48.903665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:d37f2d92cd924ebbcc365a74df3869db6aa0071939ba14afab78e2280dd22461

Observation db7bf431-6895-415d-bcb9-3e6dba8af913 · outbound

This paper cites 2025.Software qualities | SonarQube Cloud Documentation.

Evaluating LLM-Generated Code: A Benchmark and Developer Study 2025.Software qualities | SonarQube Cloud Documentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T12:23:48.905912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:6473f74a500df26d27635e5c3f517f5283c661b94d82bb5f5725d879049d75da

Observation 1e5a5a2e-0bce-44f4-a776-8dd1f92558a8 · outbound

This paper cites 2026.Evaluating LLM-Generated Code: Benchmarking on complex assignment.

Evaluating LLM-Generated Code: A Benchmark and Developer Study 2026.Evaluating LLM-Generated Code: Benchmarking on complex assignment

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T12:23:48.899190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:2ef750078e8aab16af227335c9a8339ff44aeafa470aacb559026e3a79881fbd

Observation d95a9970-f9e1-47e9-ae78-eaf4f8a18da4 · outbound

This paper cites 2026.Evaluating LLM-Generated Code: Developer Study.

Evaluating LLM-Generated Code: A Benchmark and Developer Study 2026.Evaluating LLM-Generated Code: Developer Study

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T12:23:48.896553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:72a403937fcceb7e46726100df57bd4afdb9358291945c19c2b5fe0f140d8d5e

Observation 14f43bf3-7144-4d2a-819d-2c0d495bce81 · outbound

This paper cites 2025.Building The Tree of Life from Scratch.

Evaluating LLM-Generated Code: A Benchmark and Developer Study 2025.Building The Tree of Life from Scratch

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T12:23:48.908516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:dc80c33306bf9a322d52cc2ad8831d7845351d7fa48d336050ec2cbccd12a217

Observation 3b30561b-8c93-4f2e-ad70-9a1500715ace · outbound

This paper cites URLhttps://doi.org/10.1145/3728963.

Evaluating LLM-Generated Code: A Benchmark and Developer Study URLhttps://doi.org/10.1145/3728963

Reference 18

Resolution
verified exact
doi, observed 2026-06-30T23:05:06.694698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:cdd1851264c63145a698e500478ccc9bb6df9534c2e8ecdf4dffc24826d4f9a4

Observation b50b43ce-c120-4f84-82d3-e54e4de422ff · outbound

This paper cites an unresolved cited work.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-06-30T23:05:06.706074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:16da56a05b67efa7d714dc9e9471602fb67420f8c68ee0215516e7473ed863de

Observation dfc8c545-533c-4b3f-ba39-f3f366af027d · outbound

This paper cites Available: https://doi.org/10.1145/3597503.3623316.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Available: https://doi.org/10.1145/3597503.3623316

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:05:06.711359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:2b1a6f88e8181a9c7ac70d6a7e8c30b305787c212c3c271c3ed2e2fa39338528

Observation 638025d0-ebd4-4cb3-8444-8adcc56e8c97 · outbound

This paper cites Repocoder: Repository-level code completion through iterative retrieval and generation.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Repocoder: Repository-level code completion through iterative retrieval and generation

Reference 21

Resolution
verified exact
doi, observed 2026-06-30T23:05:06.708072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:c5ace472083c41d66acb37cace187887ba5d520e4ca1d5f9eec9425e7e46b58c

Observation 8fce7454-f870-4b31-a362-2c52a31330b7 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T13:35:46.798946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:0beaec309ae89bad8f7afabcdb1b2bc57b872fefaf86e747709047341089da36

Observation 1a43b0bd-2abc-460e-a5f7-7e7be53f3465 · outbound

This paper cites CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X.

Evaluating LLM-Generated Code: A Benchmark and Developer Study CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.785333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:ed5b1fa08042a08ecd4726f4b80cca9022d3e45ee9ccdb138e4922a573cce59d

Pith citing papers

No inbound Pith citation observations are available.