Pith. sign in

Paper Citation Record · LEDGER

StackEval: Benchmarking LLMs in Coding Assistance

As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2412.05288.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05288 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:40:26.135061Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T12:16:40.086291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T12:17:52.011169Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved20
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65ea882b-0010-4620-a691-58eb41b566fb · outbound

This paper cites Large Enough — mistral.ai.

StackEval: Benchmarking LLMs in Coding Assistance Large Enough — mistral.ai

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:27.021130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.853782Z digest=sha256:7a0dbbd3169c7d795b7a5ce8afe3ee9bc4847feaf14ac6e2583e7fdbe4a39510

Observation eb39ff62-e240-4c65-bcc3-db7c736b27f1 · outbound

This paper cites Mistral NeMo — mistral.ai.

StackEval: Benchmarking LLMs in Coding Assistance Mistral NeMo — mistral.ai

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:27.002993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.859644Z digest=sha256:c02599ba54ed04ed00e87096c262e92d3574f5163ece002ad2ec10f7f646ce26

Observation 1a147000-5500-460a-9326-e888991e5421 · outbound

This paper cites Introducing claude 3.5 sonnet.

StackEval: Benchmarking LLMs in Coding Assistance Introducing claude 3.5 sonnet

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.962868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.875607Z digest=sha256:8de9aa38cc814f5159ea280d3560532a7a7c589495f8cbc5a55e52065e79e1ca

Observation fe80a1f8-817c-400b-b3e5-1522a3687d16 · outbound

This paper cites Introducing the Claude 3 family.

StackEval: Benchmarking LLMs in Coding Assistance Introducing the Claude 3 family

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.948216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.883671Z digest=sha256:d68bd3638671d516e720310ef95684b3d6e5aa5d1146c505d967a5d8a84af461

Observation f7e9907b-bf29-48b9-92f4-9d6ca1bc1b7b · outbound

This paper cites Multi-lingual evaluation of code generation models, 2023.

StackEval: Benchmarking LLMs in Coding Assistance Multi-lingual evaluation of code generation models, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.934080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.892073Z digest=sha256:64eefd306a137dcc6a7c7639a028eb7958bb1830282da41054e997e596d61f73

Observation 9955b0ad-c798-40a2-8461-f81f88692459 · outbound

This paper cites Program synthesis with large language models, 2021.

StackEval: Benchmarking LLMs in Coding Assistance Program synthesis with large language models, 2021

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.915004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.897567Z digest=sha256:ea90a03b56db77de927bd057d014cce93c840dfa5540f5b228ad5c8536825966

Observation 3f48137f-d357-4eba-bd06-e9a40f9847ae · outbound

This paper cites an unresolved cited work.

StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:25.905345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:25.905345Z digest=sha256:1054e5e0d9f3224ecbe778e348729dc921cf65e06505fed4b3da3d4ed8c0d1c5

Observation d3f416a6-215b-4b7c-b368-d66d5ca48bc3 · outbound

This paper cites Multipl-e: A scalable and extensible approach to benchmarking neural code generation, 2022.

StackEval: Benchmarking LLMs in Coding Assistance Multipl-e: A scalable and extensible approach to benchmarking neural code generation, 2022

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:25.910185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:25.910185Z digest=sha256:312ebea381dfb55ab08b51a1097000d509f67c61d94cbfe5c2b43180bff7574c

Observation 27331acb-5ca4-45c2-be37-545a66d1c496 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

StackEval: Benchmarking LLMs in Coding Assistance Evaluating Large Language Models Trained on Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:25.915910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:25.915910Z digest=sha256:8d7e24bba997e7e7c9702bbec46cd3a2052a08c3f540134fd69250114cdc2493

Observation 98b31dc7-6dff-425e-9575-9f36703bc96d · outbound

This paper cites Meta large language model compiler: Foundation models of compiler optimization, 2024.

StackEval: Benchmarking LLMs in Coding Assistance Meta large language model compiler: Foundation models of compiler optimization, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:25.927162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:25.927162Z digest=sha256:a3402c5cfaf15732ce3df20599d5ab8232b2e2343c40f5345400b99161da9e7b

Observation 6da4e34c-5f15-483e-a404-b0731ef07ae9 · outbound

This paper cites Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

StackEval: Benchmarking LLMs in Coding Assistance Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.845959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.934298Z digest=sha256:c94a52e02daedf47679d65a6459dd69ee118a883186d65b9fb015cbb848dcde5

Observation 239b3626-51c5-4bbc-bd90-7d9d37bfe464 · outbound

This paper cites The llama 3 herd of models, 2024.

StackEval: Benchmarking LLMs in Coding Assistance The llama 3 herd of models, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.822923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.939477Z digest=sha256:ea35c94d2860325166cfbbb1137a68ea7b66322ccb6807482e577c3003fa557f

Observation 326cd81b-c638-4de3-982d-455e01fccca1 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

StackEval: Benchmarking LLMs in Coding Assistance Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:25.945175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:25.945175Z digest=sha256:b751560bbde6fa9fae9fa85777e36e537223f680166838d41b1bc10b9e908924

Observation 491d141e-d9e8-409a-b7f1-e6a6d6bd7fa4 · outbound

This paper cites Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024.

StackEval: Benchmarking LLMs in Coding Assistance Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.805360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.956859Z digest=sha256:8718859f3960c1617ad95308a4aab2acb9e4b5cce8e2701f9dbc4ba99f23f7be

Observation db24aa86-8dec-47c9-97de-7b693e0bd739 · outbound

This paper cites Gemma 2: Improving open language models at a practical size, 2024.

StackEval: Benchmarking LLMs in Coding Assistance Gemma 2: Improving open language models at a practical size, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.787116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.962940Z digest=sha256:3ff72d6cd43389643ea667e6d4583f9135bcc8e8a5458f60b1d88f23e79b7222

Observation a779f551-ce0a-488f-89d7-802581ee57c3 · outbound

This paper cites an unresolved cited work.

StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:40:26.766887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.967866Z digest=sha256:d70d7114674a97cce8baac96981edf78532396faaedf98db75fc06bd6cb31ede

Observation 11c20530-8ee6-46ef-8a76-8b897cac8c57 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

StackEval: Benchmarking LLMs in Coding Assistance SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:25.972607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:25.972607Z digest=sha256:80dde18d52c0a197f8bdfc6e3f1dad82792d0e09b6b2be4b971b50951b96d576

Observation 13050429-d2e8-4a4d-b2be-0576d511c097 · outbound

This paper cites Gravity theories with local energy-momentum exchange: a closer look at Rastall-like gravity.

StackEval: Benchmarking LLMs in Coding Assistance Gravity theories with local energy-momentum exchange: a closer look at Rastall-like gravity

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T15:40:26.197442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.977597Z digest=sha256:7ef9cd28d4cf1cd01946a0f1101e2882ed8cc2319b027595e17888b797fc9678

Observation 39595990-f61a-48b7-bd17-631d8b460e40 · outbound

This paper cites Hashimoto.

StackEval: Benchmarking LLMs in Coding Assistance Hashimoto

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:25.983311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:25.983311Z digest=sha256:c20aed97e173c6ad9b3bccb8b039ba2319a6630eb129752597d216620a47c7b9

Observation 34100273-9e7b-415b-a52f-5ca3ea3a9665 · outbound

This paper cites Wildbench: Benchmarking language models with challenging tasks from real users in the wild, 2024.

StackEval: Benchmarking LLMs in Coding Assistance Wildbench: Benchmarking language models with challenging tasks from real users in the wild, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.722743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.990456Z digest=sha256:08ed34e81947d809e731e01c51e52b94d23ece3e830e729596c5344b5346e800

Observation 84334a51-918d-4467-99dc-2f21a3097627 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

StackEval: Benchmarking LLMs in Coding Assistance Rouge: A package for automatic evaluation of summaries

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:25.996802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:25.996802Z digest=sha256:e670641fa5a75ccd830b4d0b2c13e65978bc8e38ebfec58c7019355c54f25535

Observation ba35ff5f-d32e-49f0-aa0b-933a916571b7 · outbound

This paper cites llama-3_1-nemotron-70b-instruct | NVIDIA NIM — build.nvidia.com.

StackEval: Benchmarking LLMs in Coding Assistance llama-3_1-nemotron-70b-instruct | NVIDIA NIM — build.nvidia.com

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.695011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.004907Z digest=sha256:72cf54287c553b53d1f473d6e1b05218d84af499920f84bff358597c00330ec6

Observation a3894810-abc8-4ba5-a310-edc7c67c0e97 · outbound

This paper cites New models and developer products announced at devday.

StackEval: Benchmarking LLMs in Coding Assistance New models and developer products announced at devday

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.678352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.010059Z digest=sha256:dd839aae48908ca9cc3a6676b8e3a8eee649578bb034cc784dc54d270bd8455c

Observation f781adb2-fb6d-4d70-9618-fe232a51aea0 · outbound

This paper cites Introducing openai o1-preview.

StackEval: Benchmarking LLMs in Coding Assistance Introducing openai o1-preview

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.659211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.017756Z digest=sha256:9612c066ce015bc36f6a0bd4c82f45e41a593ed102c9ad07737ff1b8da94781f

Observation 2ccaf991-8a35-4cf0-9476-ebe78c2a8f53 · outbound

This paper cites Gpt-4 technical report, 2024.

StackEval: Benchmarking LLMs in Coding Assistance Gpt-4 technical report, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.635690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.024936Z digest=sha256:59f1a7a157284151d2d8cfd0dcf2bf6b759a44e541dbb11e2ba62090ab0b1bcb

Observation 0d23d370-f193-49c8-8914-f59a7e7ee38a · outbound

This paper cites Stack overflow developer survey 2023, 2023.

StackEval: Benchmarking LLMs in Coding Assistance Stack overflow developer survey 2023, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.611296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.030417Z digest=sha256:88d1848cc8ed3c9a3ad3dd68f3dc274702b72c4b69dc435c7954b2a6c0f1f738

Observation a22ab879-8454-4b48-b730-fb885af3f4c7 · outbound

This paper cites Bowman, and Shi Feng.

StackEval: Benchmarking LLMs in Coding Assistance Bowman, and Shi Feng

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.596482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.035627Z digest=sha256:223eaa9ff971be1adcc786d991a8b36638850ae176f6363c42d293ad8d868237

Observation 8121e2be-55ef-4fca-9239-8d1caea92fab · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

StackEval: Benchmarking LLMs in Coding Assistance Bleu: a method for automatic evaluation of machine translation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.580488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.040480Z digest=sha256:3391f5b21f0d0615439025f9dc1de3f44820b596e8dc2fd4fdff286c7fc26265

Observation e08fd9d7-3279-47ed-97f0-7a3ae546050e · outbound

This paper cites Can foundation models label data like humans? Hugging Face Blog, 2023.

StackEval: Benchmarking LLMs in Coding Assistance Can foundation models label data like humans? Hugging Face Blog, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.564127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.044744Z digest=sha256:ac1eb5e1c127277f3388a2c33a89785904bae7105c737dd438c45b0236f75fe8

Observation f0a70ca9-41ec-4b19-80bf-3a37129838b0 · outbound

This paper cites Codebleu: a method for automatic evaluation of code synthesis, 2020.

StackEval: Benchmarking LLMs in Coding Assistance Codebleu: a method for automatic evaluation of code synthesis, 2020

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:26.049132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:26.049132Z digest=sha256:539cb4899e8fa2d9745de7f7beebbc4e23fdf5a9e2def124c492a4310482652f

Observation 15a40489-797b-4e54-802a-abc3cc7091a7 · outbound

This paper cites Code llama: Open foundation models for code, 2024.

StackEval: Benchmarking LLMs in Coding Assistance Code llama: Open foundation models for code, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:26.053117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:26.053117Z digest=sha256:b131b6ac27100f04b7dc7821f920434f8fa8d9df509acc0f12a63af9d7fa04f3

Observation 570a897c-4658-4730-a51e-d66550d74b03 · outbound

This paper cites Learning performance-improving code edits, 2024.

StackEval: Benchmarking LLMs in Coding Assistance Learning performance-improving code edits, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:26.057338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:26.057338Z digest=sha256:87812b2c18879cb3de0f72ac61bbd9306ee5bb79de9b85bca843f2b5ce5d14e3

Observation cd168baa-b7fd-4c41-9c10-f9f7262830ec · outbound

This paper cites Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024.

StackEval: Benchmarking LLMs in Coding Assistance Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.513131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.062816Z digest=sha256:e59389ab10bcf537bc91091f7410ce1f1a6b16cdd93402876a57f38a90e6c238

Observation f81df30c-5256-4881-bd7c-5931704dfb26 · outbound

This paper cites Large language models are incon- sistent and biased evaluators, 2024.

StackEval: Benchmarking LLMs in Coding Assistance Large language models are incon- sistent and biased evaluators, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.495780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.066823Z digest=sha256:708b86d4779f7b408a35ae48f312e708b28171209140d584eaaf6a4a6f40fb0a

Observation 51221124-dbaf-4b49-b581-fc16bf209623 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

StackEval: Benchmarking LLMs in Coding Assistance Qwen2.5: A party of foundation models, September 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:26.071690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:26.071690Z digest=sha256:b515dd46e0e5525ac24f757d0593e7bfc8c3640217acd948ae155bc1dc75f5ce

Observation cee711e6-e6fd-4bd1-ad52-1cae2075a15b · outbound

This paper cites Large language models are not fair evaluators, 2023.

StackEval: Benchmarking LLMs in Coding Assistance Large language models are not fair evaluators, 2023

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:26.075618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:26.075618Z digest=sha256:ac886be3bf6397dfcd1dfe042c6241f50a8767344f708bc3e1046b26d2979d4a

Observation 92befad6-f4ef-492e-9024-98f8a240e561 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models, 2023.

StackEval: Benchmarking LLMs in Coding Assistance Chain-of-thought prompting elicits reasoning in large language models, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:26.079768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:26.079768Z digest=sha256:25b044abcb85323da944c4e84c0772943cffe68f95d828fda8f5ea69df84f665

Observation 77c2d629-eeb8-4fe8-9ca6-0763a77477b3 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

StackEval: Benchmarking LLMs in Coding Assistance WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.440778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.085086Z digest=sha256:2c00deb839cd9d0b91d76b543b6fcba462cecada20389e0d4162d628d671d023

Observation a1d42979-98cb-4d31-a0d5-17d9b4888b15 · outbound

This paper cites an unresolved cited work.

StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:40:26.411116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.089419Z digest=sha256:74156ea733340a7acbed64ac0cf3c1b31d7fb37ddd9d6cd12fbf2e6b887d2cd5

Observation 16d73fa3-9b9b-4519-a766-6012279b1f9f · outbound

This paper cites Thinking before speaking: A role-playing model with mindset, 2024.

StackEval: Benchmarking LLMs in Coding Assistance Thinking before speaking: A role-playing model with mindset, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.392289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.094463Z digest=sha256:5951d09573fcc1cef06b1a209fc2ab12e08ed38c88fb840771286932bff794fd

Observation 5e3b4a28-d257-4bb5-ba59-0c1d21dbc2e0 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

StackEval: Benchmarking LLMs in Coding Assistance Xing, Hao Zhang, Joseph E

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:26.098783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:26.098783Z digest=sha256:7f80a28d1efd6244bbeb2c0c0a5d549d943e067fad8b42529c1b9570f704ef29

Observation 1c9a51e0-32ed-4afe-97e2-6b7a6c360bb1 · outbound

This paper cites Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x, 2023.

StackEval: Benchmarking LLMs in Coding Assistance Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x, 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.350514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.104062Z digest=sha256:8cf95f93ce1733cf2d13bf3c92dc788fa5efb7037a300902256590428d334d0f

Observation 23f00ec4-0965-41c1-aa23-b3baf7879777 · outbound

This paper cites an unresolved cited work.

StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:40:26.331682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.109316Z digest=sha256:399ed32ca251114af114f0841a9a86516904ef29bfe476cea7a128b5ac78d698

Observation d909a09e-781d-47c6-8879-446dbee3b482 · outbound

This paper cites an unresolved cited work.

StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:40:26.316220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.117338Z digest=sha256:1baeb881ed9e2adf16337c7f7230fb8b408e3626347ab75e3efa7c0eee327d79

Observation 5b865bf3-50b7-4e3e-a9f5-06ea51dfbac5 · outbound

This paper cites an unresolved cited work.

StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:40:26.295482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.127631Z digest=sha256:d0bafe07499384b93b9ee80ff8a9de2307b224f0446cf165d7ba4921597ace0a

Observation e1129173-6bcf-44be-b3c3-76a9cb51f359 · outbound

This paper cites questionAnalysis.

StackEval: Benchmarking LLMs in Coding Assistance questionAnalysis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:40:26.275429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:26.135061Z digest=sha256:4fe1bd3d7a854fb6f6fb4f0a3c77b7cfab9fc519434cc0c640aaea52932185c2

Observation 6718e648-5e30-405d-b1b8-58c346afb64c · outbound

This paper cites an unresolved cited work.

StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-12T15:40:26.981076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:40:25.867499Z digest=sha256:8793ed3d6d011b6e7ece7999af41566d308f15661584d2304cf0f3520e5be58a

Pith citing papers

Observation fc622f00-0f46-499d-bdff-6e8086ba1788 · inbound

RubberDuckBench: A Benchmark for AI Coding Assistants cites this paper.

RubberDuckBench: A Benchmark for AI Coding Assistants StackEval: Benchmarking LLMs in Coding Assistance

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:17:52.012845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T12:16:40.086291Z digest=sha256:26e571ff672f39133ddd836b3ec6d8c4a85384ed210be29348cfe8c71d37a479