Pith. sign in

Paper Citation Record · LEDGER

Multi-lingual Evaluation of Code Generation Models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2210.14868.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2210.14868 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:53:25.956850Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

28
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 92af7c64-b7d9-4e51-a2ac-3fbc1a18825c · inbound

StarCoder: may the source be with you! cites this paper.

StarCoder: may the source be with you! Multi-lingual Evaluation of Code Generation Models

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:33:00.359835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T23:32:59.517389Z digest=sha256:eec62680bc63c7feda4ebc575dedc7f7f44c605999ade2df6d3893bf58b73aba

Observation 1420d9b2-81a0-433c-be42-3a062685c8a6 · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Multi-lingual Evaluation of Code Generation Models

Reference 242

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:34:42.863830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:309331193b347214f3f8136749422d4adc27bd6784aba34179154ad7179c333a

Observation 0196dbba-4ba1-401f-ab37-e2d1050ea114 · inbound

A Survey on Large Language Models for Code Generation cites this paper.

A Survey on Large Language Models for Code Generation Multi-lingual Evaluation of Code Generation Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:18:06.667539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T20:18:06.304134Z digest=sha256:452e7bf9599ed223b68f68eff49d92eaa09b57711bbbbf0113389b6149a48c79

Observation 127d747b-b722-4093-a323-5fdfad651925 · inbound

ProSec: Fortifying Code LLMs with Proactive Security Alignment cites this paper.

ProSec: Fortifying Code LLMs with Proactive Security Alignment Multi-lingual Evaluation of Code Generation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T17:10:55.406265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:10:55.406265Z digest=sha256:1b32a031b83848791b486c710dbbe5c777c52a12b16bacaf80b87cad92abb6a3

Observation bf8a3c85-3d47-400c-b678-b1483b7fc7b7 · inbound

A Preliminary Study of Multilingual Code Language Models for Code Generation Task Using Translated Benchmarks cites this paper.

A Preliminary Study of Multilingual Code Language Models for Code Generation Task Using Translated Benchmarks Multi-lingual Evaluation of Code Generation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:18:33.026719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:18:33.026719Z digest=sha256:8cd1efc12ec475337cf536f3d62b98072c18e6ecf0ff3061670e97e596b8da0a

Observation 910dc0c6-c405-4c96-ac07-85fd4e810a6b · inbound

The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model cites this paper.

The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model Multi-lingual Evaluation of Code Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:03:47.980691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:03:47.980691Z digest=sha256:8e1c909ee6246b948a9f3df1ff2eb77721534c544793488ea357d64c13b7d205

Observation 033749d6-46d4-4338-83ba-565e6f3dc325 · inbound

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation cites this paper.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Multi-lingual Evaluation of Code Generation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.683347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.683347Z digest=sha256:dcd735d295d2da4499e8d35c0be6d38d889971b4738a0ef6bee61e44a9e0c978

Observation 14e8420b-741f-4028-b57b-70af239ee457 · inbound

A Study of LLMs' Preferences for Libraries and Programming Languages cites this paper.

A Study of LLMs' Preferences for Libraries and Programming Languages Multi-lingual Evaluation of Code Generation Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:55:12.191333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T22:53:16.951417Z digest=sha256:0e762c82b65a91b9066222d92ec6cd90d406eb338c7b1f6e90f01503fee1f36a

Observation 1a3a6e20-84a9-4fa8-8932-036bd34b9600 · inbound

Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving cites this paper.

Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving Multi-lingual Evaluation of Code Generation Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:48:50.496587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T06:48:50.416649Z digest=sha256:981133eec45b53358c09996aa8e83f9c27ace6e8f941c7fcae10966a7448e115

Observation e96b423d-d89a-46a9-a65c-b0025a2d457e · inbound

OpenClassGen: A Large-Scale Corpus of Real-World Python Classes for LLM Research cites this paper.

OpenClassGen: A Large-Scale Corpus of Real-World Python Classes for LLM Research Multi-lingual Evaluation of Code Generation Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:19.883923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T19:20:54.644575Z digest=sha256:7ebc8ff0c3173d5c00d6b9eb07f2b9da009f5abf64c4462be4918f4dd7e5ff0f

Observation 1fc24f99-0845-40e4-91ff-529b89c40048 · inbound

ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies cites this paper.

ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies Multi-lingual Evaluation of Code Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:25.956850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:53:25.956850Z digest=sha256:d0a7d3095f646e1161479c8ed061554c4c00f1edce979ce0eddef7354a2adf4d

Observation b7d697aa-76ea-4ea8-b57e-b7dbe5ee8fff · inbound

OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution cites this paper.

OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution Multi-lingual Evaluation of Code Generation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:27:58.276972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:27:58.276972Z digest=sha256:6bfca927cbc14e25bb52005b748742e01b7c8b8bd8d9911ba97b0795633ae5d8

Observation e2d3ac28-05d7-4368-b368-5a434d8bbbcc · inbound

QiMeng-TensorOp: Automatically Generating High-Performance Tensor Operators with Hardware Primitives cites this paper.

QiMeng-TensorOp: Automatically Generating High-Performance Tensor Operators with Hardware Primitives Multi-lingual Evaluation of Code Generation Models

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:44.521745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:21:44.521745Z digest=sha256:760609049217d1de9a14b55b60d1d692673007e0cc806df41dabf91658c3f071

Observation 8bbdaf7d-2b61-4f8b-a052-81f5602257b0 · inbound

MigrationBench: Repository-Level Code Migration Benchmark from Java 8 cites this paper.

MigrationBench: Repository-Level Code Migration Benchmark from Java 8 Multi-lingual Evaluation of Code Generation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:07.671800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:07.671800Z digest=sha256:e7a2d16f4b139cb94f21c5c0837bd402e0249c71d8a57d75b512a7aa892eb83c

Observation 3bde5860-92bd-4aa0-a342-d77b95197d2e · inbound

Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences cites this paper.

Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences Multi-lingual Evaluation of Code Generation Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:10:29.962269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:10:29.962269Z digest=sha256:f252a2f3221fa6f273c8818ebf64cdd02b66946614a3167883995b3ae6536dcc

Observation 25c7c523-b62a-400f-96bd-04ed7f28a192 · inbound

Mutation-Guided Unit Test Generation with a Large Language Model cites this paper.

Mutation-Guided Unit Test Generation with a Large Language Model Multi-lingual Evaluation of Code Generation Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:07:15.188725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T11:05:15.898419Z digest=sha256:442387bb8163368e0e63e086e2d85e2bc8e7d7f034e7c7839360ba8373fefe23

Observation 94857e75-89c2-405e-b321-1c7072033fd0 · inbound

In-Context Learning as an Effective Estimator of Functional Correctness of LLM-Generated Code cites this paper.

In-Context Learning as an Effective Estimator of Functional Correctness of LLM-Generated Code Multi-lingual Evaluation of Code Generation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:34:34.587748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:34:34.587748Z digest=sha256:fac05506d2265cd45b3bd5c4081284af1dd8d85cdf8ee23165c2cceb03fd03d0

Observation bb257d84-d21b-4eaa-ae39-01df4ebd3ef2 · inbound

SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation cites this paper.

SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation Multi-lingual Evaluation of Code Generation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:35.892082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:35.892082Z digest=sha256:115e6881d8c9507d1e1e3ca49a15f7b86681278e026ee20ad070d3b6e859aabf

Observation 6276f757-6ee0-49cb-b9f8-bfd14b2591fd · inbound

Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference cites this paper.

Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference Multi-lingual Evaluation of Code Generation Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:17:50.325921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T16:17:50.215935Z digest=sha256:f6d8c9354642ea9984c24ead6cbe05df14452c56de8343b35fce18dcdc29907f

Observation 7d567052-2cde-46b6-82b1-7b690d8db3da · inbound

A Taxonomy of Programming Languages for Code Generation cites this paper.

A Taxonomy of Programming Languages for Code Generation Multi-lingual Evaluation of Code Generation Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:56:14.178719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T02:17:18.127674Z digest=sha256:6123af0a5e6d6cf93db332c81760b8ac12588913150033da5fd82321faa37183

Observation 07e6d07e-56a6-4c4a-ac5c-3cb24345ea5c · inbound

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review cites this paper.

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review Multi-lingual Evaluation of Code Generation Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.893735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T19:52:49.324500Z digest=sha256:2754e606472ea3834670c932b833033f2df13fab36245249e1377b74216016bf

Observation 22c6f3ed-bb50-4015-8f0d-926367a311e8 · inbound

Assessing the Impact of Requirement Ambiguity on LLM-based Function-Level Code Generation cites this paper.

Assessing the Impact of Requirement Ambiguity on LLM-based Function-Level Code Generation Multi-lingual Evaluation of Code Generation Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:41:23.479400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T21:15:51.570343Z digest=sha256:02079f9ef8a13a2c54004d0317f67618c052c4fa2276d6f94ff32c7fefa869af

Observation 9a93e213-ca79-45b0-bd6c-6345c1df53dc · inbound

RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices cites this paper.

RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices Multi-lingual Evaluation of Code Generation Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:36:15.364569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T11:24:07.839469Z digest=sha256:34e43c4775a8ce58afc10c527cc64ec92efeeb62ae9dc7821d55aeb9782b3e2d

Observation bcf03380-7c0c-4a78-b7c9-382997f99e5f · inbound

Learning from Execution: Self-Evolving Memory for Private-Library Code Generation cites this paper.

Learning from Execution: Self-Evolving Memory for Private-Library Code Generation Multi-lingual Evaluation of Code Generation Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:38.575910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T03:40:05.380630Z digest=sha256:c0b655281ecc4c8e4c20db115cd41a47510757544d129f3c851af880f592b1e7

Observation c1974c29-71d2-451c-9384-04b3c0e37381 · inbound

Learning from Execution: Self-Evolving Memory for Private-Library Code Generation cites this paper.

Learning from Execution: Self-Evolving Memory for Private-Library Code Generation Multi-lingual Evaluation of Code Generation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T18:16:50.322933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:16:50.322933Z digest=sha256:cf3579af6854fe7cc20f80031dabfc65b9e1133398a2e88c626bd94cb625ec07

Observation f82767ea-48ab-4be6-9bb1-77e1ef813282 · inbound

Evaluating LLM-Generated Code: A Benchmark and Developer Study cites this paper.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Multi-lingual Evaluation of Code Generation Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:37:02.619902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T02:27:11.573347Z digest=sha256:2ef20a4875f5cf9135abbf2665677637ab765a961b8676dd4b4c6eb8a52de6e8

Observation 3c81cd13-b026-4d1d-85f3-527c77d0b89d · inbound

Evaluating LLM-Generated Code: A Benchmark and Developer Study cites this paper.

Evaluating LLM-Generated Code: A Benchmark and Developer Study Multi-lingual Evaluation of Code Generation Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.801549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T22:55:49.377456Z digest=sha256:49dbae3375fa7d7e58e661a0ae8660b11265fc499998d674a15b975bd0ee635b

Observation c7bbb054-735f-449d-a92e-78d6b9d90019 · inbound

CodegenBench: Can LLMs Write Efficient Code Across Architectures? cites this paper.

CodegenBench: Can LLMs Write Efficient Code Across Architectures? Multi-lingual Evaluation of Code Generation Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:06:23.764000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T13:38:48.308683Z digest=sha256:89d1d16a514ee03fec9ba707c26ac21872d9336e9de68f7896b185644b0f356d

Observation 5dcc1e99-c490-4ea4-9d2f-5a3e4f3a3954 · inbound

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code cites this paper.

SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code Multi-lingual Evaluation of Code Generation Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T09:56:51.758679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T05:21:34.772199Z digest=sha256:d960706d22664532904f26356b70e53b6b77fb4195c991d8b9663a2d1d0e344b

Observation 415bb5b3-f6e9-4076-acad-8b24e9d09b86 · inbound

Beyond Pass Rate: A Multilingual, Execution-Grounded Evaluation of Open Code LLMs cites this paper.

Beyond Pass Rate: A Multilingual, Execution-Grounded Evaluation of Open Code LLMs Multi-lingual Evaluation of Code Generation Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:17:29.743533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T18:19:26.108492Z digest=sha256:e95f790de2c16fc0465a3a81809355efc867ce9b6c44f7ed601058b28c02bb0b

Observation be4ed27b-54d6-4ab4-9587-436f81870666 · inbound

Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages cites this paper.

Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages Multi-lingual Evaluation of Code Generation Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.014285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T17:39:27.723253Z digest=sha256:1c43fad37b0f9714b06ed0345a06cabc94a209320e9c4853754bda6756f9c228

Observation 0095572b-c0f4-4ae0-9f99-a69a76488c7f · inbound

The Best Programming Language for Tokenmaxxing: An Investigation of Coding Agent Behavior Across Programming Languages cites this paper.

The Best Programming Language for Tokenmaxxing: An Investigation of Coding Agent Behavior Across Programming Languages Multi-lingual Evaluation of Code Generation Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T04:41:31.803525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:41:31.803525Z digest=sha256:f5a4b5f5c13028fab14d5f6445674d9a95bfba77c19c61b0cba9d5577ef6cf97

Observation d604f600-da62-41c1-b6ab-1a998ce71052 · inbound

Efficient Grammar-Constrained Decoding via Parser Stack Classification cites this paper.

Efficient Grammar-Constrained Decoding via Parser Stack Classification Multi-lingual Evaluation of Code Generation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:06:44.611094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:06:44.611094Z digest=sha256:842beb5599ae03e00c393447ef1c88c1a69d357437c04e81fbc5874c8cadea4c