Pith. sign in

Paper Citation Record · LEDGER

OSS-Bench: Benchmark Generator for Coding LLMs

As of 19 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.12331.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12331 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:41:09.087338Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 421c77fd-dd77-4d41-8c94-8255de3985fa · outbound

This paper cites Github copilot.https://copilot.github.com, 2021.

OSS-Bench: Benchmark Generator for Coding LLMs Github copilot.https://copilot.github.com, 2021

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.765896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.882885Z digest=sha256:8cf34dc7e453c8d2344437694de6dc3ab23561aaf5d0991b085bad5062960945

Observation 9420a2ab-6a59-4c0a-ad9c-199c6e3e0c0d · outbound

This paper cites Cursor: The ai-powered code editor.https://cursor.so, 2023.

OSS-Bench: Benchmark Generator for Coding LLMs Cursor: The ai-powered code editor.https://cursor.so, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.755646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.887126Z digest=sha256:699f4480a9315e246c9e7b471f9d1d3b48e5341237c8925beeed45bad5de7966

Observation 37a4dd52-72cb-4904-9535-2acb97562cc9 · outbound

This paper cites an unresolved cited work.

OSS-Bench: Benchmark Generator for Coding LLMs Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.890792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.890792Z digest=sha256:e652a2b32807a436f872f18bb4949292ba1a788060c1dd4f7ef7f199dee0b713

Observation ffd97942-0d28-4e3f-8692-fa5f96fc972c · outbound

This paper cites Codelmsec benchmark: Systematically evaluating and finding security vulnerabilities in black-box code language models.

OSS-Bench: Benchmark Generator for Coding LLMs Codelmsec benchmark: Systematically evaluating and finding security vulnerabilities in black-box code language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.738087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.894506Z digest=sha256:d381d388aeac7bea906e9e89c1df2565fb8a5e78272adb8dd8d27a4de6f30135

Observation f4362cce-2d41-45cc-9720-d8e55ab24d87 · outbound

This paper cites DynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code Generation.

OSS-Bench: Benchmark Generator for Coding LLMs DynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.898224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.898224Z digest=sha256:3e69b705f5a015a4da2fdd38c82589c99a257bab0ed3a956add52933aac994fa

Observation 92931cb0-3c53-4536-bab2-27d63c5c177c · outbound

This paper cites Humanevo: An evolution-aware benchmark for more realistic evaluation of repository-level code generation.

OSS-Bench: Benchmark Generator for Coding LLMs Humanevo: An evolution-aware benchmark for more realistic evaluation of repository-level code generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.728038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.902726Z digest=sha256:0c52d9ea4c803187a4256999b7507c2207f7c30f0d75725891f9e1f3305625fd

Observation be0cf7b6-15bc-4966-9283-e7b95304764b · outbound

This paper cites Codeif-bench: Evaluating instruction-following capabilities of large language models in interactive code generation.arXiv preprint arXiv:2503.22688, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Codeif-bench: Evaluating instruction-following capabilities of large language models in interactive code generation.arXiv preprint arXiv:2503.22688, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.906493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.906493Z digest=sha256:87f5df045d255e11a39dc004e74b52567294e3c3c147e0f0cf12cf50a62bd63b

Observation 19ebd076-1d70-427d-912f-77193fece6a3 · outbound

This paper cites ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code.

OSS-Bench: Benchmark Generator for Coding LLMs ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.909862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.909862Z digest=sha256:723ff996bfb6f40e90cd5c5fab44f0d584fd56a4bf687d4f3bd947927876adaf

Observation 0fcf29d9-48f7-44a6-bf52-a8bd0d4fd83b · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

OSS-Bench: Benchmark Generator for Coding LLMs SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.913882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.913882Z digest=sha256:d3d8895add91afac0f342afb5bd2e54bfa1e42ecf4f1fe6645afae0c4a35d58f

Observation be14a1d2-d442-401c-a987-89953bee9903 · outbound

This paper cites Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica.

OSS-Bench: Benchmark Generator for Coding LLMs Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.718194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.917414Z digest=sha256:fbb1d609aace80c86a44dd87da743751b3d47b4f39d019435cb2bd9cd012e33b

Observation 1f926510-652a-45e2-81e3-6297ad342f41 · outbound

This paper cites an unresolved cited work.

OSS-Bench: Benchmark Generator for Coding LLMs Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:41:09.708228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.920715Z digest=sha256:49afa9caa76734824d34d817d168fcb97b9de42c052e97d89a208f3d7d5b043e

Observation 5aabb163-a1e4-4c76-8171-de64910dd20d · outbound

This paper cites Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving.

OSS-Bench: Benchmark Generator for Coding LLMs Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.924377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.924377Z digest=sha256:ee0bd1036a954fff5f79fd38ad9b0facb6ac4a4029fd870c637b4939d38d483b

Observation 4f899914-f0ea-4ffe-be68-ef5008893aec · outbound

This paper cites SecRepoBench: Benchmarking LLMs for secure code generation in real-world repositories, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs SecRepoBench: Benchmarking LLMs for secure code generation in real-world repositories, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.698118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.928133Z digest=sha256:a32508b16e0bde604cda516b6578426c89f591650b19c0f4c7e8ec2d2a3099bf

Observation 175ab0dc-7966-411f-b6c0-7ddeb40f37a1 · outbound

This paper cites CWEval: Outcome-driven evaluation on functionality and security of LLM code generation, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs CWEval: Outcome-driven evaluation on functionality and security of LLM code generation, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.687983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.931450Z digest=sha256:635d21195c32e073700ed402549e77d95709561d267edaa438421f6f356d1f10

Observation d267b0db-c1a3-4ef8-b916-cecb179222a6 · outbound

This paper cites Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models.

OSS-Bench: Benchmark Generator for Coding LLMs Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.934788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.934788Z digest=sha256:b392067a813c9902da4cf3da55336a0be28a7e07fcfcdf549c9bbd31511dd8fa

Observation 8097f06e-e716-4baa-a9f5-f247dd6ddd16 · outbound

This paper cites Codearena: Inspecting and improving code quality metrics using minecraft.

OSS-Bench: Benchmark Generator for Coding LLMs Codearena: Inspecting and improving code quality metrics using minecraft

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.677609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.938516Z digest=sha256:54e8cb4efdc1a37e1056884db1a18d8390f60c5d717977d15575f4deb02efeca

Observation 838d148b-e976-4a6a-8eae-01547b8c5a40 · outbound

This paper cites CodeElo: Benchmarking competition-level code generation of LLMs with human-comparable elo ratings, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs CodeElo: Benchmarking competition-level code generation of LLMs with human-comparable elo ratings, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.667110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.942026Z digest=sha256:ede0b3e2033c2e9dcacffb9d76cef6cd58050e992932d77afb1a6f9e4e579982

Observation f7e97651-5964-4935-a92f-b6670019008f · outbound

This paper cites ComplexCodeEval: A benchmark for evaluating large code models on more complex code.

OSS-Bench: Benchmark Generator for Coding LLMs ComplexCodeEval: A benchmark for evaluating large code models on more complex code

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.656575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.945410Z digest=sha256:47d049e57f275bfea449c3dcc6b2a70eb5233e554f10b06c7664db3fee93637c

Observation d54edff5-10a6-4393-a1e0-24f00abd278f · outbound

This paper cites PythonSaga: Redefining the benchmark for code generating LLMs, 2024.

OSS-Bench: Benchmark Generator for Coding LLMs PythonSaga: Redefining the benchmark for code generating LLMs, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.646182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.948865Z digest=sha256:75601937052fe0619aca0da3788e1bc11cba211d52ce12f437d23ff6674ba63f

Observation b56dff38-2269-4311-9d09-2d1988dba53b · outbound

This paper cites Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark.

OSS-Bench: Benchmark Generator for Coding LLMs Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.952533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.952533Z digest=sha256:a98bffc7258af1f87f2e04102ddf60eb167e228b014e39c3a6a44f4ac46ca15c

Observation 5d8a4b47-36f1-4ff6-8f67-cc13055f3cd5 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

OSS-Bench: Benchmark Generator for Coding LLMs Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.635334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.956417Z digest=sha256:56a10e9625ee159a2d610e323a721b93b48219c68d83163605892b431843d580

Observation f99279ee-a255-4536-9367-6778d381a976 · outbound

This paper cites ClassEval: A manually-crafted benchmark for evaluating LLMs on class-level code generation, 2023.

OSS-Bench: Benchmark Generator for Coding LLMs ClassEval: A manually-crafted benchmark for evaluating LLMs on class-level code generation, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.625092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.959822Z digest=sha256:1df4c8ccb9a76b6a35be374931b0982688004834335911cf2342d0076262b3fc

Observation 2cbeb47d-f80b-4c4f-86fb-534d76cc5f87 · outbound

This paper cites HumanEval-XL: A multilingual code generation benchmark for cross-lingual natural language generalization, 2024.

OSS-Bench: Benchmark Generator for Coding LLMs HumanEval-XL: A multilingual code generation benchmark for cross-lingual natural language generalization, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.613243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.963333Z digest=sha256:1256b78a5cc27989a728352f57b1a433d41e63fc73f1ff48ad0e10512b8d2e95

Observation 996c0b14-7950-4375-9405-0055fdad5d9e · outbound

This paper cites Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x.

OSS-Bench: Benchmark Generator for Coding LLMs Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.966684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.966684Z digest=sha256:46db9646d5d2be499c15877f6c49a5392b159e7d39afc8fde40c6e1b6b348374

Observation 25f43123-634e-4ac1-998e-991732e8da59 · outbound

This paper cites {AddressSanitizer}: A fast address sanity checker.

OSS-Bench: Benchmark Generator for Coding LLMs {AddressSanitizer}: A fast address sanity checker

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.970147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.970147Z digest=sha256:4b93a34c62c44514dae54a236e3bed89398b2b6f84cd33a36cad346c49caafe3

Observation 74e819d9-d3f8-47ba-9843-c946e1a8f32b · outbound

This paper cites php-src: The php interpreter.https://github.com/php/php-src, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs php-src: The php interpreter.https://github.com/php/php-src, 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.596794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.973421Z digest=sha256:a3530286d194a4e597ea8c9f16005067536ab7f18e053749498fbef9c468e8b1

Observation b09990d6-ebf5-434e-9d9b-0c5d93593900 · outbound

This paper cites Richard Hipp.

OSS-Bench: Benchmark Generator for Coding LLMs Richard Hipp

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.586385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.976672Z digest=sha256:4eb8ee91e74e464c1adf5c327012c076a5c79c8e90e04d4e62e91b99fb3391a9

Observation 6e377465-2058-4c79-ac59-b801173648b5 · outbound

This paper cites https://testing.googleblog.com/2020/08/ code-coverage-best-practices.html, 2020.

OSS-Bench: Benchmark Generator for Coding LLMs https://testing.googleblog.com/2020/08/ code-coverage-best-practices.html, 2020

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.574840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.979969Z digest=sha256:92541a180809a826644a6263148a83fc9cfdb8fc59df7298721c81f4b35d2bbf

Observation 88dc70a5-32bc-440f-abeb-a04a7fac21bc · outbound

This paper cites libclang: C interface to the clang library.

OSS-Bench: Benchmark Generator for Coding LLMs libclang: C interface to the clang library

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.564370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.983338Z digest=sha256:1aebb12bbe338da054b87406a94931c7e1c072e7269c54ea1daf24f324e2d6b2

Observation d1450324-a4f6-45df-a943-c1b8d6b4b472 · outbound

This paper cites an unresolved cited work.

OSS-Bench: Benchmark Generator for Coding LLMs Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.987369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.987369Z digest=sha256:d572b9208a601ad75d46ac4cdd79ceb8222e46f812b50866c16c1a0b5966f72e

Observation 1978e6ba-1cb1-4ecc-bc52-25c65abf6ff3 · outbound

This paper cites difflib — helpers for computing deltas between objects.

OSS-Bench: Benchmark Generator for Coding LLMs difflib — helpers for computing deltas between objects

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.548044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.990624Z digest=sha256:3e56399854f7878fbff842eaef0677c86c8f46501b874840d58700de9c11b5e0

Observation cd2d49f0-5dc8-4fae-98e3-42c38be67156 · outbound

This paper cites gpt-o1 model.https://platform.openai.com/docs/models/o1, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs gpt-o1 model.https://platform.openai.com/docs/models/o1, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.538208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.993964Z digest=sha256:afb84d758da9db697a3aaca6fc732781d9b0ebf560666cdc4a85a30dae04883b

Observation ab3f1842-84bd-444f-b18d-d960ecb00322 · outbound

This paper cites o3-mini model.https://platform.openai.com/docs/models/o3-mini, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs o3-mini model.https://platform.openai.com/docs/models/o3-mini, 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.528623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:08.997232Z digest=sha256:7095befafb12ad18e3a3daea9a936c28995abbfee514470abc3a91b31a5dcab8

Observation c5d584d7-b454-482d-bd96-1006e253fffb · outbound

This paper cites Claude 3.7 sonnet.https://www.anthropic.com/claude/sonnet, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Claude 3.7 sonnet.https://www.anthropic.com/claude/sonnet, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.518845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.000387Z digest=sha256:0c4a16894d1e3d7eeef16686188d43853c91606fdaf57765959434002851ff59

Observation 5ac39b2b-67e2-42b1-a827-d6712bfb99da · outbound

This paper cites Claude 3.5 haiku.https://www.anthropic.com/claude/haiku, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Claude 3.5 haiku.https://www.anthropic.com/claude/haiku, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.509216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.003876Z digest=sha256:eeba262ea77f63204b30161463ee01591fddb683aacc098bc24745a6a037fc27

Observation ecb39fce-e3dc-4376-b46c-2dc6b95967f7 · outbound

This paper cites Gemini 2.5 flash.https://developers.generativelanguage.google/, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Gemini 2.5 flash.https://developers.generativelanguage.google/, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.499364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.007301Z digest=sha256:d69a76e9b345204f3b8b458171781be1917cbd73bccedd31f02ed7d9c1498a8d

Observation 614a4c67-f17b-462e-afc7-18d9722d4bb9 · outbound

This paper cites Llama 3.3 70b instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Llama 3.3 70b instruct (fp16)

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.489480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.010502Z digest=sha256:8eb053b84a0fb4823dd1aaee094729ba13167491348d0fa9952bfef338cc2791

Observation 3e19ca2f-c7eb-4839-9edb-39d90a5268ed · outbound

This paper cites Codellama 70b instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Codellama 70b instruct (fp16)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.479853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.013703Z digest=sha256:18e280622450cc2e73ad834f9eb7c92d6fd1068284378d4730eea3a977c67c48

Observation d97379d9-4aec-40a6-af08-2b98a80f12bc · outbound

This paper cites Qwen 2.5 coder 32b instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Qwen 2.5 coder 32b instruct (fp16)

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.470476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.016914Z digest=sha256:8be1d60c517e8690bc72f088b0a342066ec7644bad0fec98a38b2f1cf385d1d7

Observation 16c0acef-3449-42ba-8439-7306efa9a235 · outbound

This paper cites Qwen 3.0 30b-a3b fp16.https://ollama.com/library/qwen3:30b-a3b-fp16, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Qwen 3.0 30b-a3b fp16.https://ollama.com/library/qwen3:30b-a3b-fp16, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.460304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.020235Z digest=sha256:69aaaa9dca4bdaf3af24ec7b7a72b5bb4b67aa9da9be5c427a26d6f096f6dbe4

Observation 36fdb130-8cbe-47e1-ad88-2d20377819a9 · outbound

This paper cites Qwen 3 8b fp16.https://ollama.com/library/qwen3:8b-fp16, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Qwen 3 8b fp16.https://ollama.com/library/qwen3:8b-fp16, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.450714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.023486Z digest=sha256:97286da36b67fafed44d981578037b4f60747e90ff765d1de48d7ebdb459b467

Observation ed415a7d-841a-4574-9f59-df739406e666 · outbound

This paper cites Gemma 3 27b-it fp16.https://ollama.com/library/gemma3:27b-it-fp16, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Gemma 3 27b-it fp16.https://ollama.com/library/gemma3:27b-it-fp16, 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.440770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.026695Z digest=sha256:e62fbc30a9ebfe56d7324ff5828ccd384627199e4dcc4be8d823f3c627be3f6b

Observation f14f87fd-c5ed-4fa4-9171-ec76af6b46ed · outbound

This paper cites Qwen 2.5 coder 14b instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Qwen 2.5 coder 14b instruct (fp16)

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.430956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.029928Z digest=sha256:b8476b154c679bec8a759f512d5defa84885641c7ad5319ba8103a09983bbfaf

Observation 2e09b457-c428-4822-8ef3-01b67b7c49c2 · outbound

This paper cites Deepseek coder v2 16b lite instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Deepseek coder v2 16b lite instruct (fp16)

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.421381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.033335Z digest=sha256:094c2242f3618cc5258b6be2518fbed8dc2a0958f667dccef11203601ced5a47

Observation 2651a8ef-fb93-47ae-a2dc-546876a5fc10 · outbound

This paper cites Starcoder2-15b-instruct-v0.1 (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Starcoder2-15b-instruct-v0.1 (fp16)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.411806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.036683Z digest=sha256:4940f6426bb00ace8e37437f3b940f64e2523c847b980eeffc706496fd4e10a6

Observation 065a418e-8f61-4bad-9e28-8c86c4b37fce · outbound

This paper cites Phi-4 14b fp16.https://ollama.com/library/phi4:14b-fp16, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Phi-4 14b fp16.https://ollama.com/library/phi4:14b-fp16, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.402729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.039967Z digest=sha256:0f1bb336c3faf4997fb014967a29cbcf6022487f6534b174e11def47acbe53ac

Observation e8a8fc19-035d-41bc-9fb8-27aa8ccfdf02 · outbound

This paper cites Mistral 7b instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Mistral 7b instruct (fp16)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.392942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.043072Z digest=sha256:77530c9c5f4e2d709356872a0485f1c1a760c8e065d9c0724feb11cf9ec56094

Observation 56c62f9f-203a-46fc-bb32-8424f361b066 · outbound

This paper cites Codegemma 7b instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Codegemma 7b instruct (fp16)

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.382442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.045955Z digest=sha256:3947574990b6977ba0b53f724f8e538eb02d547855d03a5a0d77dba6a3f7fc20

Observation 51c00e5f-a141-4348-b0c7-1541b8d6090b · outbound

This paper cites Openai: Advances in safe and beneficial ai.https://openai.com, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Openai: Advances in safe and beneficial ai.https://openai.com, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.372657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.048955Z digest=sha256:f1511c27f3dcbe76dab974850f667229abe2b713b1e4482389654bf3e0c3c1c2

Observation 1074e2ff-727f-496b-800a-30dcd2e6b1b3 · outbound

This paper cites Anthropic: Building reliable, steerable ai systems.

OSS-Bench: Benchmark Generator for Coding LLMs Anthropic: Building reliable, steerable ai systems

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.362447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.052241Z digest=sha256:f20e455a11d7e3544f9d59284a74f14619fdd83741c8a1367f959697e16ace69

Observation faa5be50-d86b-462e-97bc-7f4e4e873a70 · outbound

This paper cites Google: Organizing the world’s information.https://www.google.com, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Google: Organizing the world’s information.https://www.google.com, 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.352265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.055397Z digest=sha256:46a96eb91ad6907785f497babc8dacda4ea9784be8eaad982f9db47a665dfbe1

Observation b029a42d-0535-4919-b5d9-939351530fc5 · outbound

This paper cites Deepseek: Developer of high-performance open-source llms.

OSS-Bench: Benchmark Generator for Coding LLMs Deepseek: Developer of high-performance open-source llms

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.342614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.058832Z digest=sha256:2ecae03e5a1b6a6b08d01e69a90a71c45c68cff6da3ce1abc3ac5b73031a1009

Observation c12238fa-e33f-4420-ac49-bcf5605bd33e · outbound

This paper cites Alibaba group: Global trade and technology.https://www.alibaba.com, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Alibaba group: Global trade and technology.https://www.alibaba.com, 2025

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.332122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.061970Z digest=sha256:d1652e87084cbef372107b52ea24eb6ec62492ef755fc5f1fbd59cbee8a6ce01

Observation c9f9aa06-2519-4215-acfd-e96bd8cb4b54 · outbound

This paper cites Meta: Bringing the metaverse and social technology together.

OSS-Bench: Benchmark Generator for Coding LLMs Meta: Bringing the metaverse and social technology together

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.321033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.065031Z digest=sha256:151315f5f14bbf2d36295a35388c23b2e26a6ab8241d4dd15fd671111f0ef076

Observation cded8149-9ca2-4f3e-80e8-5a24d16e8f85 · outbound

This paper cites Microsoft: Empowering every person and organization.

OSS-Bench: Benchmark Generator for Coding LLMs Microsoft: Empowering every person and organization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.310513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.068195Z digest=sha256:d00f444bb038911970efb473e032e36a468a8079f4c64fbff1a066bfe85942d4

Observation 1ab6e7d1-79b9-4dc6-904f-a9092340bc36 · outbound

This paper cites Bigcode: Open and responsible development of code llms.

OSS-Bench: Benchmark Generator for Coding LLMs Bigcode: Open and responsible development of code llms

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.299952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.071519Z digest=sha256:60c431d3b9a52b2b705738603c103b819084a2551041ef014b022355bd109421

Observation 2c5f0684-a7fa-4f47-ab91-988aacd0d958 · outbound

This paper cites Mistral ai: Frontier ai in your hands.https://mistral.ai, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Mistral ai: Frontier ai in your hands.https://mistral.ai, 2025

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.289957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.074554Z digest=sha256:a6074056cce7cfe1848365e0150256b6fedb3d12738f27d05c700d65d8a42b8d

Observation 77e66e68-6bcb-4ffd-a89b-8398e909a334 · outbound

This paper cites Ollama: Get up and running with large language models locally.

OSS-Bench: Benchmark Generator for Coding LLMs Ollama: Get up and running with large language models locally

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.279340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.077778Z digest=sha256:855386d13df290d376200542eb5eab6f673ae04597899289b3f985147e2447bd

Observation 6376ec4b-63b3-4a79-bb6d-4e07616eb100 · outbound

This paper cites SPoC: Search-based Pseudocode to Code.

OSS-Bench: Benchmark Generator for Coding LLMs SPoC: Search-based Pseudocode to Code

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:09.080649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:09.080649Z digest=sha256:62c7814588bed830141481b7f5142f38252dadee7f7e57113ea62d7be323c9b4

Observation c79ca41e-c891-4d6a-b74f-17d8593de675 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

OSS-Bench: Benchmark Generator for Coding LLMs Evaluating Large Language Models Trained on Code

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:09.084092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:09.084092Z digest=sha256:059e408add49683be7a965cda02476705c48e6d38be3eaa91e82b97927e8736a

Observation ca45235d-09f5-492e-9102-3bbbf33dedfd · outbound

This paper cites Fuzzing the PHP Interpreter via Dataflow Fusion.

OSS-Bench: Benchmark Generator for Coding LLMs Fuzzing the PHP Interpreter via Dataflow Fusion

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T20:41:09.129077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:41:09.087338Z digest=sha256:dbb15fbc6b435aa541e0427702fb5eb502e458a7abbfa6382231f1b256a03e57

Pith citing papers

No inbound Pith citation observations are available.