Pith. sign in

Paper Citation Record · LEDGER

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation

As of 23 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2506.17369.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17369 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:20:50.633822Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f2422f55-53cc-465c-9c21-58868e402185 · outbound

This paper cites Gpt-4o system card, 2024.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Gpt-4o system card, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.400321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.400321Z digest=sha256:ccd2f4ca2ee4d0df37e368c433c8a1db736396c938f5e4ae4770782705a1974d

Observation f5d26ee0-9de2-4d96-9f17-8d93436d4c10 · outbound

This paper cites The llama 3 herd of models, 2024.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation The llama 3 herd of models, 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.405378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.405378Z digest=sha256:4328b2637fe86701cddd74457b965c1e47971dfd94a91d0ac833a349504af39a

Observation d2be190f-bc40-473b-8477-142654602955 · outbound

This paper cites Qwen technical report, 2023.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Qwen technical report, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.251251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.410101Z digest=sha256:26146ecd0d2bb1161dfebbb0768d844ffda2185da0cfbf3e6d5ff756b2354dea

Observation 87ef3a44-bad2-404e-b7e9-4e0bd24eea1e · outbound

This paper cites Code llama: Open foundation models for code, 2024.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Code llama: Open foundation models for code, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.238847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.414462Z digest=sha256:94654983b324fdc0ca0d7f3edd77a2f56eae93ad0fc63f010233a210b1dd97be

Observation 6d7f402e-b0f7-4372-a1f4-c2fe0ff0987a · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Evaluating Large Language Models Trained on Code

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.418886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.418886Z digest=sha256:82bfe462868714e65dddd4a612e610695106f990738ea687a410be357cf84faa

Observation 777103b3-d166-4ac5-9662-2640fa7e437a · outbound

This paper cites Unit Test Case Generation with Transformers and Focal Context.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Unit Test Case Generation with Transformers and Focal Context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.423467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.423467Z digest=sha256:2912b3a58729c53481f7705ead669420db114e557c60ee7b22226f9b6191c79e

Observation 0fffcaa7-719d-40d7-9cf6-717d3502eada · outbound

This paper cites Codexglue: A machine learning benchmark dataset for code understanding and generation, 2021.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Codexglue: A machine learning benchmark dataset for code understanding and generation, 2021

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.226651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.429212Z digest=sha256:0812ff1515c0e8700804ccf9df2f5913fcf506dc52240eb90634e7a94cd5d394

Observation 2781ec7e-b651-4fb6-8330-79fba66fd3e9 · outbound

This paper cites Reasoning Runtime Behavior of a Pro- gram with LLM: How Far Are We?.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Reasoning Runtime Behavior of a Pro- gram with LLM: How Far Are We?

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.213948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.433159Z digest=sha256:3506dfda1a91eb80f94a8e116fdfedb2a12e05910823db4395991ca3468b8cf1

Observation a5088587-b051-4b46-94aa-e6167b478d26 · outbound

This paper cites Program Synthesis with Large Language Models.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Program Synthesis with Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.437120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.437120Z digest=sha256:a4060897f794edb8b85a9d006356440fc1b5ec7ded99107f88f8088a3b74cbe7

Observation 9412ad1b-52d5-45be-8d94-b581a6c06dbf · outbound

This paper cites Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation, 2023.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.441548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.441548Z digest=sha256:961197fe28bb5f3f4cb3f9faa79f3888c573bdae7f484cfde533dbecfce509bd

Observation 74019c44-5f57-417a-8553-926958b0089e · outbound

This paper cites Testeval: Benchmarking large language models for test case generation, 2025.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Testeval: Benchmarking large language models for test case generation, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.191443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.446004Z digest=sha256:1add1f19f63688332787142c77eb6e7af3c80a1e6f9ffb3876fe49247585135e

Observation 39813682-0044-4eeb-9f1c-ee69e70f3b9c · outbound

This paper cites CRUXEval: A benchmark for code reasoning, understanding and execution.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation CRUXEval: A benchmark for code reasoning, understanding and execution

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.178984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.450317Z digest=sha256:c430530a32c22aff2b50771cb16bd4ae49725a2cca21ca211f44e58efc26a622

Observation 4b5b0d50-b789-452e-8210-2b45c46561c8 · outbound

This paper cites Lyu, and Shing-Chi Cheung.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Lyu, and Shing-Chi Cheung

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.167502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.454525Z digest=sha256:c05f101982d535e0946220d6074e12ffdb78da847dc56524525f332915453511

Observation a6cd81f4-6176-49d2-b34a-69d1c12127f4 · outbound

This paper cites Introduction to prompting, 2025.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Introduction to prompting, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.156187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.458806Z digest=sha256:5159172b2613f44bb499258355188fd8437cc1a7429c038655208d823a5033c2

Observation 88d9b5f6-d6f7-4505-a202-3bce5976a16b · outbound

This paper cites Prompt design and engineering: Introduction and advanced methods, 2024.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Prompt design and engineering: Introduction and advanced methods, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.145563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.462927Z digest=sha256:0753fa4133290eed8179a205301539fffa3dcd6b3bf81c96f86603e105d9fb5b

Observation 34e56316-2aa4-4e29-99f3-21efc74708e6 · outbound

This paper cites Benchmarking knowledge boundary for large language models: A different perspective on model evaluation.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Benchmarking knowledge boundary for large language models: A different perspective on model evaluation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.134180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.467575Z digest=sha256:0cbe6ea9a24ac57e9488a639649f44197774e94e3f6c8071bbf6134765ac9890

Observation 9c17d1a2-27f6-46cd-b47b-5d7ecc03cffd · outbound

This paper cites Quantifying language models’ sensitivity to spu- rious features in prompt design or: How i learned to start worrying about prompt formatting.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Quantifying language models’ sensitivity to spu- rious features in prompt design or: How i learned to start worrying about prompt formatting

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.122683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.472303Z digest=sha256:5f33e27b0937bfe89794e48d6389edcb0b76355af1157f0d4034fd81d06ee640

Observation 00aca115-afa3-4182-9ac0-d1065dc7805d · outbound

This paper cites State of what art? a call for multi-prompt LLM evaluation.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation State of what art? a call for multi-prompt LLM evaluation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.109881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.476573Z digest=sha256:ac28b98a3c76399197d81a2fcbe520a24017787afcb0d1eb605a02034f17be98

Observation e318fa2c-ed33-4d2b-83c0-7df9e4df91a0 · outbound

This paper cites On the worst prompt performance of large language models.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation On the worst prompt performance of large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.097797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.481004Z digest=sha256:5624ae16d705eb5632fd1a282f96bc316089fad3b6d3393f0a0d865ff7b53f9f

Observation ba74e0c4-95f7-4f3a-99cc-ec28288e9bd2 · outbound

This paper cites LMentry: A language model benchmark of elementary language tasks.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation LMentry: A language model benchmark of elementary language tasks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.084693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.485788Z digest=sha256:35ac8fe5bb4fd02ea8d3e2dff4359bbeda1bde220a08994699befa10d05ade28

Observation 348cc597-9d1d-4f20-b5e4-4d2fa13b2dd0 · outbound

This paper cites Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.071363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.489929Z digest=sha256:a8cf254c990ee4e43c699840a43fee4f6582d3ced2250d972ae7b7a5846a4dfc

Observation 155db46a-46af-4811-a26a-37eef804d556 · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.058533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.494046Z digest=sha256:51738ad38367da77d8434f6d28a68e475d35acaa01a6b984e51a3288b6fea6ae

Observation 62e23409-1589-4e05-a51d-6b2751ea8ee7 · outbound

This paper cites Challenging BIG-bench tasks and whether chain- of-thought can solve them.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Challenging BIG-bench tasks and whether chain- of-thought can solve them

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.046674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.498681Z digest=sha256:1e3679f908c7ca9cba86c92cb03038358c1c49e407e68ecc23bad58cb8b4623e

Observation 980fc10c-626f-40cd-8fae-4306f9425379 · outbound

This paper cites Codegemma: Open code models based on gemma, 2024.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Codegemma: Open code models based on gemma, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.034430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.502517Z digest=sha256:124e18326e2b394f836131de14db31200bccb448b000f4a0b1a76b84c62b5046

Observation 7a6932e3-3d8c-4c2f-9d75-0c9b07339c8f · outbound

This paper cites Qwen2.5-coder technical report, 2024.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Qwen2.5-coder technical report, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.022783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.506159Z digest=sha256:2102d7b02eea5123d8c767cafb5241b77d841ba0fbae6b7093d7d759a5ab840b

Observation e85633ee-c547-42c5-84e0-7423a57e23b7 · outbound

This paper cites an unresolved cited work.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.509963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.509963Z digest=sha256:922d339f774e15d2f858d1b70f982e17dac7c75dd050703c565fb9b597423356

Observation 31963a75-6e26-4452-99c1-71dc0dee66cd · outbound

This paper cites Automated program repair in the era of large pre- trained language models.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Automated program repair in the era of large pre- trained language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:51.004396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.513833Z digest=sha256:a20d04917f189ab1b84eb69bde7916d61380143480ed1d0075d41a76aba05c0d

Observation 3944686a-dad6-4e40-9365-defa4fbb8e3d · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.518057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.518057Z digest=sha256:63970f1f8b64692d0a73bd21cf1fbfe8e4972f403ee9b50e45a378d9822fa9d7

Observation e8faff32-1ea7-48fa-92c9-c78bd3fd9e64 · outbound

This paper cites Codereval: A benchmark of pragmatic code generation with generative pre-trained models.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Codereval: A benchmark of pragmatic code generation with generative pre-trained models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.522004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.522004Z digest=sha256:92342d3f743cf70c35dfecca0bc6fe2c3763dfa91e662c9e813433d3fa0f3b12

Observation bf526847-cdb1-48f2-9cd7-64feb4f735dd · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.525905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.525905Z digest=sha256:774a83ad55874e206eaeb78576061c46384011d87df6f74b9ff6f5142c26be3f

Observation db0b2538-262a-4f45-9504-ca8cb79f0891 · outbound

This paper cites Lessleak-bench: A first investigation of data leakage in llms across 83 software engineering benchmarks, 2025.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Lessleak-bench: A first investigation of data leakage in llms across 83 software engineering benchmarks, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.529717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.529717Z digest=sha256:da4ebceb9d6289629789a20d0391f56bb6b4e0fa55055e3d332620bf9ce024af

Observation 94e766b9-4327-4716-8a00-82177324cd4c · outbound

This paper cites Coderujb: An executable and unified java benchmark for practical programming scenarios.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Coderujb: An executable and unified java benchmark for practical programming scenarios

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.972269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.534820Z digest=sha256:fda9e89395cdbb4b9644d62e125b0f573836e9849ee5bebe0525da703008d4b3

Observation b6a90951-c02a-461b-8304-b8d9412123e2 · outbound

This paper cites Promptbench: A unified library for evaluation of large language models.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Promptbench: A unified library for evaluation of large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.958985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.548767Z digest=sha256:2d01636b642a70dd0a34cfd05c513dc8472855f1c3cba0b64b9a62e0643e1481

Observation b46c1b59-afc2-40cb-a484-dea90317cd90 · outbound

This paper cites The butterfly effect of altering prompts: How small changes and jailbreaks affect large language model performance.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation The butterfly effect of altering prompts: How small changes and jailbreaks affect large language model performance

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.945664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.553438Z digest=sha256:5b700a026dba25b76e95f851be780c0fe2e9ef5a0266c2f09752c46c357ac398

Observation 91a317a7-8d2b-4029-a5fe-73e544e312dd · outbound

This paper cites ProSA: Assessing and understanding the prompt sensitivity of LLMs.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation ProSA: Assessing and understanding the prompt sensitivity of LLMs

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.923816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.563127Z digest=sha256:45f79bbaf20d73f7455749f331eb5d55ee8d0958831a98f94ec6a3f83046a9e9

Observation 09391a27-6337-4f6b-99c4-76fc77367936 · outbound

This paper cites Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.912029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.567660Z digest=sha256:9876ceed4f139341c4814b84fe8840b1bb54e29a93fb95dc32dfa4e82224f302

Observation d60dae33-54f0-43bc-8440-daf4a5ee9182 · outbound

This paper cites Fuzz4all: Universal fuzzing with large language models.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Fuzz4all: Universal fuzzing with large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.898962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.572063Z digest=sha256:98881fd78a42ff85d222a23a29095448773489ed4f6c6f16dc73f91e789d7f36

Observation 5b8f6f35-8d4d-4387-a6e3-c2d8131958ae · outbound

This paper cites Large language models are edge-case generators: Crafting unusual programs for fuzzing deep learning libraries.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Large language models are edge-case generators: Crafting unusual programs for fuzzing deep learning libraries

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.887905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.576815Z digest=sha256:bb2abe1e1bab0f32bc3e25b46b7497a7fc4aa38597482b2c99a6e51f81d1917e

Observation 8258dabb-ac57-4c1e-8dfb-f37273c555bf · outbound

This paper cites Large language model guided protocol fuzzing.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Large language model guided protocol fuzzing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.876745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.581413Z digest=sha256:ac7b728f29e449c6ec6cc79509082477e0c03c6746e4369dce0aa81cd6978269

Observation afb50b15-a813-4953-8ab4-cee21c04e365 · outbound

This paper cites From one thousand pages of specification to unveiling hidden bugs: Large language model assisted fuzzing of matter IoT devices.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation From one thousand pages of specification to unveiling hidden bugs: Large language model assisted fuzzing of matter IoT devices

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.864665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.585790Z digest=sha256:56b8919dd694d9462c0c0bbbe161f0d3168dbe48065c125e9e1df891a8da423f

Observation 8e279e49-1aa6-4ff5-b455-c3218742cd26 · outbound

This paper cites Similarity thresholds in retrieval-augmented generation.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Similarity thresholds in retrieval-augmented generation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.851991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.590332Z digest=sha256:9218756d4b98c073243df21fce22b80779bb5b187523c3c4efa6ad69f92f24b0

Observation 65e3681c-01b6-4fc2-ab77-4182c0982050 · outbound

This paper cites Jiang et al.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Jiang et al

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.594879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.594879Z digest=sha256:b5e3a9f2a3cbb06620bc42ba98f7ffc6ce1be89d92bdd99671255bfc893fdbb2

Observation 4d87fef2-c12f-4f1e-9435-57a6a3e1a4f6 · outbound

This paper cites an unresolved cited work.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:20:50.830919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.599706Z digest=sha256:3c37e8524fd3c1e8c5f806014d2ed17700351f06be60a9d0b87d68c35f20be4e

Observation 69c80e0a-c26f-4c9d-9512-38374249dfbd · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Gonzalez, Hao Zhang, and Ion Stoica

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.604632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.604632Z digest=sha256:6fe5d100ab79e98293b679e4b487ee591b8b59ae2590af0028f612b26c828e66

Observation 0357e228-efcb-43c6-a8a0-958d101eaf0e · outbound

This paper cites an unresolved cited work.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:20:50.809727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.609444Z digest=sha256:fb095a78e3282a08031ed962850f09fcb47c62827fd9cb6ec0a1d6c874b79b5f

Observation 8a15e3f5-9b20-4544-9796-a4d94fca6833 · outbound

This paper cites Statistics (international student edition).

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Statistics (international student edition)

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.614189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.614189Z digest=sha256:f69c3e3ddc2e6c311ab9d66ccd3a102d20d3789e2ac1057e673fbfeff860a64d

Observation b4da002f-49c1-46fe-aae4-4327ff203974 · outbound

This paper cites ANOVA: Repeated measures.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation ANOVA: Repeated measures

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.789356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.618670Z digest=sha256:24ea4c8f3aa595a32ba9220e7d935da655e69a97f82ed4fbadc5e1d42ee2724a

Observation ce1ec95b-6df7-46ae-ac17-02eed4dfd964 · outbound

This paper cites Openai o3-mini, 2025.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Openai o3-mini, 2025

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.776736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.623327Z digest=sha256:5c4b61fbe26a9c4a3538dd4a340b3033929c388ebc8679bb4d3187754f245834

Observation 1a5bdc81-6fb9-4358-915e-bff228e8ec50 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.763314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.628881Z digest=sha256:fc30e45953eebb54596ce177c7267a587c668fcc601e77adae3a177186be9534

Observation 333950cd-977b-4d32-90f7-91d2bd506a90 · outbound

This paper cites deepseek-ai/deepseek-r1 - hugging face, 2025.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation deepseek-ai/deepseek-r1 - hugging face, 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:20:50.748707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T19:20:50.633822Z digest=sha256:208cb12e28ad113656a28d92aa732f1e4d4db007351a2385ca5590a1db9d7591

Observation 69596613-8210-421a-a98d-39a8837da3c5 · outbound

This paper cites an unresolved cited work.

Re-Evaluating Code LLM Benchmarks Under Semantic Mutation Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:20:50.557986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:20:50.557986Z digest=sha256:ede58a68a622c42b4633490f4b6b3d722a7cb0853a40be02406735dd622f7c50

Pith citing papers

No inbound Pith citation observations are available.