Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:20:50.633822Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2506.17369.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:20:50.633822Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f2422f55-53cc-465c-9c21-58868e402185 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Gpt-4o system card, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5d26ee0-9de2-4d96-9f17-8d93436d4c10 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation The llama 3 herd of models, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2be190f-bc40-473b-8477-142654602955 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Qwen technical report, 2023
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 87ef3a44-bad2-404e-b7e9-4e0bd24eea1e · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Code llama: Open foundation models for code, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6d7f402e-b0f7-4372-a1f4-c2fe0ff0987a · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Evaluating Large Language Models Trained on Code
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 777103b3-d166-4ac5-9662-2640fa7e437a · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Unit Test Case Generation with Transformers and Focal Context
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fffcaa7-719d-40d7-9cf6-717d3502eada · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Codexglue: A machine learning benchmark dataset for code understanding and generation, 2021
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2781ec7e-b651-4fb6-8330-79fba66fd3e9 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Reasoning Runtime Behavior of a Pro- gram with LLM: How Far Are We?
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a5088587-b051-4b46-94aa-e6167b478d26 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Program Synthesis with Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9412ad1b-52d5-45be-8d94-b581a6c06dbf · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation, 2023
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74019c44-5f57-417a-8553-926958b0089e · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Testeval: Benchmarking large language models for test case generation, 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 39813682-0044-4eeb-9f1c-ee69e70f3b9c · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation CRUXEval: A benchmark for code reasoning, understanding and execution
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4b5b0d50-b789-452e-8210-2b45c46561c8 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Lyu, and Shing-Chi Cheung
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a6cd81f4-6176-49d2-b34a-69d1c12127f4 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Introduction to prompting, 2025
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 88d9b5f6-d6f7-4505-a202-3bce5976a16b · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Prompt design and engineering: Introduction and advanced methods, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 34e56316-2aa4-4e29-99f3-21efc74708e6 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Benchmarking knowledge boundary for large language models: A different perspective on model evaluation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9c17d1a2-27f6-46cd-b47b-5d7ecc03cffd · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Quantifying language models’ sensitivity to spu- rious features in prompt design or: How i learned to start worrying about prompt formatting
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 00aca115-afa3-4182-9ac0-d1065dc7805d · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation State of what art? a call for multi-prompt LLM evaluation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e318fa2c-ed33-4d2b-83c0-7df9e4df91a0 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation On the worst prompt performance of large language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ba74e0c4-95f7-4f3a-99cc-ec28288e9bd2 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation LMentry: A language model benchmark of elementary language tasks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 348cc597-9d1d-4f20-b5e4-4d2fa13b2dd0 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 155db46a-46af-4811-a26a-37eef804d556 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 62e23409-1589-4e05-a51d-6b2751ea8ee7 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Challenging BIG-bench tasks and whether chain- of-thought can solve them
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 980fc10c-626f-40cd-8fae-4306f9425379 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Codegemma: Open code models based on gemma, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7a6932e3-3d8c-4c2f-9d75-0c9b07339c8f · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Qwen2.5-coder technical report, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e85633ee-c547-42c5-84e0-7423a57e23b7 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31963a75-6e26-4452-99c1-71dc0dee66cd · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Automated program repair in the era of large pre- trained language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3944686a-dad6-4e40-9365-defa4fbb8e3d · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8faff32-1ea7-48fa-92c9-c78bd3fd9e64 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Codereval: A benchmark of pragmatic code generation with generative pre-trained models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf526847-cdb1-48f2-9cd7-64feb4f735dd · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db0b2538-262a-4f45-9504-ca8cb79f0891 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Lessleak-bench: A first investigation of data leakage in llms across 83 software engineering benchmarks, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e766b9-4327-4716-8a00-82177324cd4c · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Coderujb: An executable and unified java benchmark for practical programming scenarios
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b6a90951-c02a-461b-8304-b8d9412123e2 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Promptbench: A unified library for evaluation of large language models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b46c1b59-afc2-40cb-a484-dea90317cd90 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation The butterfly effect of altering prompts: How small changes and jailbreaks affect large language model performance
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 91a317a7-8d2b-4029-a5fe-73e544e312dd · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation ProSA: Assessing and understanding the prompt sensitivity of LLMs
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 09391a27-6337-4f6b-99c4-76fc77367936 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d60dae33-54f0-43bc-8440-daf4a5ee9182 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Fuzz4all: Universal fuzzing with large language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5b8f6f35-8d4d-4387-a6e3-c2d8131958ae · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Large language models are edge-case generators: Crafting unusual programs for fuzzing deep learning libraries
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8258dabb-ac57-4c1e-8dfb-f37273c555bf · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Large language model guided protocol fuzzing
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation afb50b15-a813-4953-8ab4-cee21c04e365 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation From one thousand pages of specification to unveiling hidden bugs: Large language model assisted fuzzing of matter IoT devices
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8e279e49-1aa6-4ff5-b455-c3218742cd26 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Similarity thresholds in retrieval-augmented generation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 65e3681c-01b6-4fc2-ab77-4182c0982050 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Jiang et al
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d87fef2-c12f-4f1e-9435-57a6a3e1a4f6 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 69c80e0a-c26f-4c9d-9512-38374249dfbd · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Gonzalez, Hao Zhang, and Ion Stoica
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0357e228-efcb-43c6-a8a0-958d101eaf0e · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8a15e3f5-9b20-4544-9796-a4d94fca6833 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Statistics (international student edition)
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4da002f-49c1-46fe-aae4-4327ff203974 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation ANOVA: Repeated measures
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ce1ec95b-6df7-46ae-ac17-02eed4dfd964 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Openai o3-mini, 2025
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1a5bdc81-6fb9-4358-915e-bff228e8ec50 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 333950cd-977b-4d32-90f7-91d2bd506a90 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation deepseek-ai/deepseek-r1 - hugging face, 2025
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 69596613-8210-421a-a98d-39a8837da3c5 · outbound
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.