Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:40:26.135061Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2412.05288.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:40:26.135061Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-16T12:16:40.086291Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-16T12:17:52.011169Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 65ea882b-0010-4620-a691-58eb41b566fb · outbound
StackEval: Benchmarking LLMs in Coding Assistance Large Enough — mistral.ai
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eb39ff62-e240-4c65-bcc3-db7c736b27f1 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Mistral NeMo — mistral.ai
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1a147000-5500-460a-9326-e888991e5421 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Introducing claude 3.5 sonnet
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fe80a1f8-817c-400b-b3e5-1522a3687d16 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Introducing the Claude 3 family
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f7e9907b-bf29-48b9-92f4-9d6ca1bc1b7b · outbound
StackEval: Benchmarking LLMs in Coding Assistance Multi-lingual evaluation of code generation models, 2023
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9955b0ad-c798-40a2-8461-f81f88692459 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Program synthesis with large language models, 2021
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3f48137f-d357-4eba-bd06-e9a40f9847ae · outbound
StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3f416a6-215b-4b7c-b368-d66d5ca48bc3 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Multipl-e: A scalable and extensible approach to benchmarking neural code generation, 2022
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27331acb-5ca4-45c2-be37-545a66d1c496 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Evaluating Large Language Models Trained on Code
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b31dc7-6dff-425e-9575-9f36703bc96d · outbound
StackEval: Benchmarking LLMs in Coding Assistance Meta large language model compiler: Foundation models of compiler optimization, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6da4e34c-5f15-483e-a404-b0731ef07ae9 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 239b3626-51c5-4bbc-bd90-7d9d37bfe464 · outbound
StackEval: Benchmarking LLMs in Coding Assistance The llama 3 herd of models, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 326cd81b-c638-4de3-982d-455e01fccca1 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491d141e-d9e8-409a-b7f1-e6a6d6bd7fa4 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation db24aa86-8dec-47c9-97de-7b693e0bd739 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Gemma 2: Improving open language models at a practical size, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a779f551-ce0a-488f-89d7-802581ee57c3 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 11c20530-8ee6-46ef-8a76-8b897cac8c57 · outbound
StackEval: Benchmarking LLMs in Coding Assistance SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13050429-d2e8-4a4d-b2be-0576d511c097 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Gravity theories with local energy-momentum exchange: a closer look at Rastall-like gravity
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 39595990-f61a-48b7-bd17-631d8b460e40 · outbound
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34100273-9e7b-415b-a52f-5ca3ea3a9665 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Wildbench: Benchmarking language models with challenging tasks from real users in the wild, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 84334a51-918d-4467-99dc-2f21a3097627 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Rouge: A package for automatic evaluation of summaries
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba35ff5f-d32e-49f0-aa0b-933a916571b7 · outbound
StackEval: Benchmarking LLMs in Coding Assistance llama-3_1-nemotron-70b-instruct | NVIDIA NIM — build.nvidia.com
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a3894810-abc8-4ba5-a310-edc7c67c0e97 · outbound
StackEval: Benchmarking LLMs in Coding Assistance New models and developer products announced at devday
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f781adb2-fb6d-4d70-9618-fe232a51aea0 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Introducing openai o1-preview
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2ccaf991-8a35-4cf0-9476-ebe78c2a8f53 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Gpt-4 technical report, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0d23d370-f193-49c8-8914-f59a7e7ee38a · outbound
StackEval: Benchmarking LLMs in Coding Assistance Stack overflow developer survey 2023, 2023
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a22ab879-8454-4b48-b730-fb885af3f4c7 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Bowman, and Shi Feng
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8121e2be-55ef-4fca-9239-8d1caea92fab · outbound
StackEval: Benchmarking LLMs in Coding Assistance Bleu: a method for automatic evaluation of machine translation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e08fd9d7-3279-47ed-97f0-7a3ae546050e · outbound
StackEval: Benchmarking LLMs in Coding Assistance Can foundation models label data like humans? Hugging Face Blog, 2023
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f0a70ca9-41ec-4b19-80bf-3a37129838b0 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Codebleu: a method for automatic evaluation of code synthesis, 2020
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a40489-797b-4e54-802a-abc3cc7091a7 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Code llama: Open foundation models for code, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 570a897c-4658-4730-a51e-d66550d74b03 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Learning performance-improving code edits, 2024
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd168baa-b7fd-4c41-9c10-f9f7262830ec · outbound
StackEval: Benchmarking LLMs in Coding Assistance Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f81df30c-5256-4881-bd7c-5931704dfb26 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Large language models are incon- sistent and biased evaluators, 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 51221124-dbaf-4b49-b581-fc16bf209623 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Qwen2.5: A party of foundation models, September 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee711e6-e6fd-4bd1-ad52-1cae2075a15b · outbound
StackEval: Benchmarking LLMs in Coding Assistance Large language models are not fair evaluators, 2023
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92befad6-f4ef-492e-9024-98f8a240e561 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Chain-of-thought prompting elicits reasoning in large language models, 2023
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c2d629-eeb8-4fe8-9ca6-0763a77477b3 · outbound
StackEval: Benchmarking LLMs in Coding Assistance WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a1d42979-98cb-4d31-a0d5-17d9b4888b15 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 16d73fa3-9b9b-4519-a766-6012279b1f9f · outbound
StackEval: Benchmarking LLMs in Coding Assistance Thinking before speaking: A role-playing model with mindset, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5e3b4a28-d257-4bb5-ba59-0c1d21dbc2e0 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Xing, Hao Zhang, Joseph E
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c9a51e0-32ed-4afe-97e2-6b7a6c360bb1 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x, 2023
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 23f00ec4-0965-41c1-aa23-b3baf7879777 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d909a09e-781d-47c6-8879-446dbee3b482 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5b865bf3-50b7-4e3e-a9f5-06ea51dfbac5 · outbound
StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e1129173-6bcf-44be-b3c3-76a9cb51f359 · outbound
StackEval: Benchmarking LLMs in Coding Assistance questionAnalysis
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6718e648-5e30-405d-b1b8-58c346afb64c · outbound
StackEval: Benchmarking LLMs in Coding Assistance Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fc622f00-0f46-499d-bdff-6e8086ba1788 · inbound
RubberDuckBench: A Benchmark for AI Coding Assistants StackEval: Benchmarking LLMs in Coding Assistance
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.