Pith. sign in

Paper Citation Record · LEDGER

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

As of 21 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 13 inbound Pith citation observations for arXiv:2412.21199.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.21199 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:37.864289Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:06:00.821737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T14:33:31.671093Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d876dc4f-c33c-4ac7-accd-52835c08fa0a · outbound

This paper cites online" 'onlinestring :=.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.661344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.661344Z digest=sha256:539dac7b3918e8cf0f677a887447993ccb0748e8cf775fc35badc8b2d5a2d66d

Observation 0d88c956-7f5e-47f4-9ad4-f0a8c7f1f6ec · outbound

This paper cites write newline.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.666987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.666987Z digest=sha256:2f9926dcc29079f0240647ab82244929a371fb65abea94df59aa971bfe16a49f

Observation 080a5760-a8d3-474a-b9cc-8212ebbedc5f · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.472932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.673044Z digest=sha256:fe08427e01a8b7715a379ef70d751337dfd4f2b9d061a77e4103e2854655f5c7

Observation f8053017-7e59-4e98-8250-3d4326e88a78 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.678098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.678098Z digest=sha256:47caa8862676262cba965b3d8dd70eb07899066f700b5f5d4111f7cce377f4c3

Observation 033749d6-46d4-4338-83ba-565e6f3dc325 · outbound

This paper cites Multi-lingual Evaluation of Code Generation Models.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Multi-lingual Evaluation of Code Generation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.683347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.683347Z digest=sha256:dcd735d295d2da4499e8d35c0be6d38d889971b4738a0ef6bee61e44a9e0c978

Observation 081d3982-20cd-4ed4-b84d-6b6827ac32a0 · outbound

This paper cites Program Synthesis with Large Language Models.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Program Synthesis with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.688411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.688411Z digest=sha256:73cbc9965c38e60aa35eb652b2daef6db495437295a981a784a2a363c74f0337

Observation c6c353f6-ca3b-4472-8e11-8c009960e6c4 · outbound

This paper cites A parallel corpus of Python functions and documentation strings for automated code documentation and code generation.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation A parallel corpus of Python functions and documentation strings for automated code documentation and code generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.692932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.692932Z digest=sha256:7b143cd647e5ec35dcac3b5f05bed8220815b82412736ca3f44426c21ffbecd2

Observation 7a4491f1-207e-4155-8203-e5941a53d6c1 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.698545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.698545Z digest=sha256:801d146992c9fa6f6027eb3b033e311244569ec46c5073a3b40b4bc7e8eb5b32

Observation 2c5be030-5340-4c47-aee8-273f73006960 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.704255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.704255Z digest=sha256:f94aca732d94cbff40c1b2783729ce4ee9f3836c507c32169fe99d4730bebb6d

Observation e24e4516-ad1c-4218-9688-cfc5990073b4 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.445866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.709869Z digest=sha256:acc6ce04c27e4f570730986fc3943ecd0b2dd8a1cc7ec330c000081c042e76bb

Observation 7e684a6e-11d6-4abd-96f8-514f5772a171 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.715051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.715051Z digest=sha256:46d83d98ad30594bf34c29b36730f827adf6c3a652f25cdc285f785e9ec18dc6

Observation b26248b7-fc10-45f4-97c7-8c4562ad2f08 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.431581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.719215Z digest=sha256:81c0c37a02023629ab3383ae18ec5c72ed7dfaa3bf87df84dcac8409deb66202

Observation cf1cb287-f0a1-4222-bcd3-33b5898fb3d1 · outbound

This paper cites CoDesc: A Large Code-Description Parallel Dataset.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation CoDesc: A Large Code-Description Parallel Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.723893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.723893Z digest=sha256:4903a4d1c365b44863ecef14bfe4363016b4b9ce241ecb4e207f6953f0cef5d9

Observation 8ce31397-d5bc-4d4f-a982-7e2798b26cba · outbound

This paper cites OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.728230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.728230Z digest=sha256:816eb59fc93c5186dde1bf1c8e1c5b446ea3ce4ce4cd2789f71275bd767dcdc0

Observation 977957ed-a753-4bde-8f22-ffae09f7d481 · outbound

This paper cites Qwen2.5-Coder Technical Report.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Qwen2.5-Coder Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.732820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.732820Z digest=sha256:cd5fefc32b0c4e0e5cc792658e832e820861c6f7522bbdbb34b16cb28d6871a5

Observation d6d5afda-6d22-456c-b57f-bf05192c3346 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.739566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.739566Z digest=sha256:1b024f9777ccd20190b47587b5ded8e093cfe9f73c3e5eddc5d3bdd599f81992

Observation 4546af62-288e-4803-a79f-764d74d0ae2a · outbound

This paper cites Impact of Code Language Models on Automated Program Repair.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Impact of Code Language Models on Automated Program Repair

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.745220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.745220Z digest=sha256:23c5b1f4799b89fdf7dc59f99183add5758faf27cd747e0b859b45d1afd89cc0

Observation e313ef28-8728-4a80-9525-5fa9c17d9f06 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.749484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.749484Z digest=sha256:d9d20678e4f63047ebb0244056ba2d3af8d7faadc704aa4faccb34f64928e169

Observation 019ef2cc-36e5-42db-8d4b-fc9d237e9f18 · outbound

This paper cites InferFix: End-to-End Program Repair with LLMs.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation InferFix: End-to-End Program Repair with LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.753760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.753760Z digest=sha256:0f5fe4a724fa5f858858cd2cf154973c7f008ea5403de06b597ae9d0712b11b2

Observation d898fb23-5cf4-43b7-b2ef-a5be47bca329 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.409903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.758747Z digest=sha256:796de4477f5b00c4590352ca4e58878c017b824fb1f3b8c42d6d6153ed4644e0

Observation d659fbe8-440b-44d8-808f-9e09ed6ac76b · outbound

This paper cites StarCoder: may the source be with you!.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation StarCoder: may the source be with you!

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.762895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.762895Z digest=sha256:c8510a2a83e27faba7127c7abe3cff44cabf9020162d6ca92b7fd3327e0b1147

Observation 8426c663-bc41-40d2-a51c-8d9638c55351 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.394955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.767199Z digest=sha256:5386bee354a9e974f8a5924454bc661476dff239b4e399151673801aa53cd86d

Observation 6b7744f3-bc8a-4b86-8175-da019dd24d4b · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.771380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.771380Z digest=sha256:74f3328f94e98309614992cf80ba944d9d30ec4f788cd5d2bd1b74c2d5d3a24b

Observation 020fc0de-60e9-4288-9030-e7c11452b449 · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.775570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.775570Z digest=sha256:ef473c6aec37294eb830259aff472b0f81fa4d28b97d8366d59da690e288f0f9

Observation 9546eaaf-9d9c-4d2b-a67e-48fa0fdf7093 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.378797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.780175Z digest=sha256:d8335cfb995de104b59ddddaa922c46b70ca13f779b4ed431f778ecc84e936ef

Observation 854ad207-6b18-4fbf-a254-ddb76d19f573 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.784703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.784703Z digest=sha256:243a0fb7803cb3968408b133a55fedc5291e5313f390eefc84a47c19c61f316d

Observation c4169652-c38d-4b82-988e-aa3ec49d70b1 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.789006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.789006Z digest=sha256:26081aee9455032fa3b797357a5551d6ef861e283df1977b8200a3e54638e857

Observation 1748b129-b0bb-4a33-8323-a46721ac04b1 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.353939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.795950Z digest=sha256:808ec8c673d41f6bcacc097d0519b7a99e3480d20b96306ebafc404c0ce88bda

Observation 51a1e376-df4a-47c1-b7e8-86bd9ef400f6 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.329912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.799893Z digest=sha256:31215b264176d5b2a1650846da22e905b4f5812021c0913c9fd0bfc9dee1f72a

Observation 6fea44e9-578c-491a-8da9-d674ba2718c3 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.804772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.804772Z digest=sha256:aa1029406e65cf271f41fd54fe938319961fd24a47a20f025825b2142494101e

Observation d97a2495-8882-4308-99fc-a8ab6bf6a230 · outbound

This paper cites ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.808644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.808644Z digest=sha256:dcf3e7fae4cff350f98b272f34d8a51e0f46122bfa74d000b3f6e27f6bc54f51

Observation f25c5153-1218-4f6d-a3f0-5ed636b533a0 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Code Llama: Open Foundation Models for Code

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.812756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.812756Z digest=sha256:045f512ce7b6b6c06d07b1187b9092e9af77c08c158cf7e0262e4a65ec3a224e

Observation 66d931f9-994f-445d-85ad-cbe24d7a93f6 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.816642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.816642Z digest=sha256:0ae6455ce87013fc2bca68bacd794455afd90e219d7e83c84b68180ff5716eb5

Observation d639f24a-63f0-48b2-8fca-fa2fc2f95657 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.820346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.820346Z digest=sha256:479dd3eef731b48ef37bcf9eafea00c3fa2bd3c7040d4474219b847b97e14b75

Observation 4492f096-a198-4925-9dad-596d10ecb4d6 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.824782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.824782Z digest=sha256:ded5ceae5e5a9bb5a75befa0c96507133e9ad67ba4b02c569b5265761c5be0d2

Observation 0250df47-431e-4736-9b2c-6263f50e7bdb · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:38.267828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-10T23:21:37.828851Z digest=sha256:0beee7946dee10e5170ecb0c3db01fd1fc13a05c1d3338ebf4d483eb671d6f9e

Observation 66e6bfa1-eb77-495b-b6a6-c583ae4c3343 · outbound

This paper cites Practical Program Repair in the Era of Large Pre-trained Language Models.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Practical Program Repair in the Era of Large Pre-trained Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.832818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.832818Z digest=sha256:f50890b5145af1c5e8d1a2de9e28f68dd4323b4ce1cf945cff4f956b27d9f8f0

Observation 46bf6f97-003b-4d22-95f6-16f2c6ce59d3 · outbound

This paper cites an unresolved cited work.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.836832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.836832Z digest=sha256:af57f6ad30612161c1fb9a8ccb638806325c58cf9d83e1a0377dfd5f8ef81660

Observation 27cb009c-685a-42c3-8b61-3237d64ddafe · outbound

This paper cites Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.840885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.840885Z digest=sha256:f05ee501e3195fe899e24f1c2a6f81eb789d24dd86fa43e95e1848cbe3311459

Observation c8abdda8-005d-44d1-a10c-884343b3e9e5 · outbound

This paper cites CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.845013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.845013Z digest=sha256:d0de5a567b55f2c25e459060c99a929617c5af8ae8ce5217626a8daa5d6fd42f

Observation b15370ef-e38b-4f6e-962c-f751d784b109 · outbound

This paper cites OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.849093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.849093Z digest=sha256:e8db154f3ff39dbc9267601e412d1d19f7adfac5a8900ded71010ee4d07b37cc

Observation 5b8c92bb-49a5-4c0b-af41-af0d143da344 · outbound

This paper cites XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.853024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.853024Z digest=sha256:607ed4c3f09f458358e43a1abc89ed9e4604a3551d8213b88e51b523ccbbbc69

Observation fdb3ab70-438b-4c3a-81d7-be9d5ef433e4 · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.857588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.857588Z digest=sha256:312ead2274c247e9fd188a358cce833bf907cb8d0e2e2dd68d3d3380a15a4f42

Observation 2bdb5c8e-c31b-4ff2-9365-c2e259feb06c · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:37.864289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:37.864289Z digest=sha256:2b7602c4d562f5a5bbd2587e86f863f27a1c507da0cb5d0b27c5fb0a1b2530b9

Pith citing papers

Observation c0188b4d-f77f-400f-b809-c542f336ef2e · inbound

HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems cites this paper.

HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:23.589246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:23.589246Z digest=sha256:d16f0cbfc1e28340937e586576d4f8da6df3a7f37bd0622cbc3a5265d63d7830

Observation e75e17fe-3ce5-4b1f-9c40-1b8d65f3dfce · inbound

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics cites this paper.

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:00.821737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:00.821737Z digest=sha256:bc6450fefb150434e74cde4b34a0f5ed74015d4ca97b808c925041fa7707e37f

Observation a955b216-27c0-4ebc-9550-e3066daeeb7d · inbound

AutoGEEval: A Multimodal and Automated Framework for Geospatial Code Generation on GEE with Large Language Models cites this paper.

AutoGEEval: A Multimodal and Automated Framework for Geospatial Code Generation on GEE with Large Language Models HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:30:40.665486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:30:40.665486Z digest=sha256:c1d091583a7a3e2fe1a972f06eddd2ee1fdc9d05e7b9fc5281e3025fdca90f3c

Observation c1dc0812-c573-420c-a324-23f2aca5ef02 · inbound

Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Models cites this paper.

Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Models HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:20.105337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:20.105337Z digest=sha256:0477eb7ff3d2e4b89a589797c95368fae2f9b76119169ce855f0b2633f67b54f

Observation 87e3b6f0-bf66-4c79-a1bc-9e4cc155d542 · inbound

Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure cites this paper.

Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:57.050750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:59:57.050750Z digest=sha256:cabcaf5dfb385b64a0cfd0c4867ff12f2579006a9752d00b35e1c168db98a51b

Observation 3b213d7f-3344-4e7e-9e2a-b143a710c997 · inbound

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks cites this paper.

Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T23:43:15.114501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:43:15.114501Z digest=sha256:7c8c27659e80121ec299093c8e2af860e4ce153ad916db972f7d469f724b5aa0

Observation ce9a34e2-15bc-46cd-893c-47c92f4f07ed · inbound

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review cites this paper.

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.775143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T19:52:49.324500Z digest=sha256:3336dd59539ae0394cfcdcd930ded2c6364567b2dca80b1e8de71a99bccbbeb3

Observation d81c616b-3c83-48c7-8893-5aa8fa3fb860 · inbound

Agentic Frameworks for Reasoning Tasks: An Empirical Study cites this paper.

Agentic Frameworks for Reasoning Tasks: An Empirical Study HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:08:26.182245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T08:24:51.573913Z digest=sha256:b8c0ce089858a8e884d93f67976e782ab9303c45b75553d4532d7c278e08e79d

Observation ae7d3cc9-89ac-4a6d-bc98-253c89ec33fc · inbound

Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions cites this paper.

Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:29.331881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T05:49:57.957009Z digest=sha256:a372a0bd043052d343c64de0120aaa8bdbdc9e5bac3b5155a10083cb94ab3c52

Observation 4524c26e-f802-4aa7-888b-6e408c19af54 · inbound

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution cites this paper.

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:57.029369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T02:28:07.557119Z digest=sha256:6347a26f47268f7fea58b1573ba03d383513b60ec91719bbcbf4cb50725e39ae

Observation 60f2ccbc-e987-460d-a2fc-12697d43e592 · inbound

A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback cites this paper.

A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:23:10.609159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T09:22:06.285118Z digest=sha256:022121ca4e938976a9d0c09d7796f8ee9d6153a234a092802002be2f7641fe16

Observation 7d49d5d0-ece7-4be7-bb7e-e28556ea6543 · inbound

CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models cites this paper.

CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:33:31.672803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T06:25:56.404246Z digest=sha256:8b7059d139c40486ef6d3ddd4e5ae4dda340236f8db30f230a78ab483c3974a6

Observation 9e3c6695-b408-4083-a1cc-6c9803c17d38 · inbound

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy cites this paper.

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-12T04:43:45.592808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:43:45.592808Z digest=sha256:98c87804a27c35c1fbd25b04fcdd5cf5141e9713d7ea12ed9adf9dae5e9f3949