Pith. sign in

Paper Citation Record · LEDGER

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

As of 7 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 11 inbound Pith citation observations for arXiv:2507.10535.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10535 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:32:34.292050Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:21:21.808023Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T12:53:50.345403Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy31
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0c9225aa-51d4-4752-85d4-1bcb44500da7 · outbound

This paper cites Phi-4 Technical Report.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:29.662597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:29.662597Z digest=sha256:f23c8daa31f59782522c176516eebba2d1fafb473b75a7918406a22a99492000

Observation 9c8434c2-b472-434c-b1f1-b5318c68cdd3 · outbound

This paper cites GPT-4 Technical Report.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:29.722527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:29.722527Z digest=sha256:d3444c9029f2ffd455d45117832771b67cdd59b900d228a774b20fa99d340b5e

Observation e1e60d71-f17a-48b4-a3ee-b24feed19bc5 · outbound

This paper cites Automated unit test improvement using large language models at meta.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Automated unit test improvement using large language models at meta

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:42.690068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:29.910326Z digest=sha256:8cb56740dd470edab301012167f245fb83f3643a408c179efd94afd37c24e1a4

Observation 62df6330-63b9-4093-95d4-df8bb968ef8c · outbound

This paper cites Claude 3.7.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Claude 3.7

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:42.581478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:30.002094Z digest=sha256:478688e07df60a3320e4c7dbc225d069b1b3ffb7a7898395d739b79348fc1782

Observation 9341a860-3a15-4d0d-b84c-f39b31c1ff6a · outbound

This paper cites Claude 4.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Claude 4

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:42.378283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:30.072314Z digest=sha256:03d418a65a17172f97a03c6844bd943fac0cd95f1f3e6d80e35fd54dae29f508

Observation 3bd8cac9-39a9-4a64-95fb-fc03c8fa2508 · outbound

This paper cites Program Synthesis with Large Language Models.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Program Synthesis with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.162463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.162463Z digest=sha256:feb82513f7c966654d4be36bc8fa2f1c4c1bc1924298d80d43a05d3ced7226e8

Observation 0e4ec585-5f9d-41e7-979f-cb0bb4d58296 · outbound

This paper cites Codet: Code generation with generated tests.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Codet: Code generation with generated tests

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.256357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.256357Z digest=sha256:a87934fc1a640aac628876a2812b13ff812bbaacc066983e26904cf09fc7e938

Observation 45ae6b93-0aed-4333-8e5a-2843f38df120 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.358202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.358202Z digest=sha256:a8b8e396bede0e346bc68dc57c0c9a6cc272e913fc9b194afb9bfa02e19ae3df

Observation d409d1fb-c907-4c60-987b-00016a2eb356 · outbound

This paper cites Teaching large language models to self-debug.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Teaching large language models to self-debug

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:42.179788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:30.467485Z digest=sha256:e2ed0df0640c6151970039603d41065bcd158892c525b8e1dfff91ec68b02379

Observation 8f725fbe-d2e9-48e0-9fb4-00a9eb87b637 · outbound

This paper cites Rm-r1: Reward modeling as reasoning.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Rm-r1: Reward modeling as reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.562154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.562154Z digest=sha256:edff1a525da6984f547be84116fd6092430eb5585ea3d29d82db55ed3b1ec5bd

Observation 6712258d-8282-40f6-ad1f-31d0bc269b16 · outbound

This paper cites AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.634578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.634578Z digest=sha256:ecfccd463ad600311c62dc948fd3e284cc16bc2bf95dbdc5374d337f9b8038c6

Observation 695a6cdc-7378-42b7-bcb5-426c16d294c5 · outbound

This paper cites an unresolved cited work.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:32:42.080154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:30.697316Z digest=sha256:d83fe899152f29ea7f553f794c5eae5c014ed43b6ec1781d074a9359d1ae7f0d

Observation 410a27dd-5851-4e7c-8697-13ed5db91dc1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.805802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.805802Z digest=sha256:1abf1d1596e6184838b68837a7c47771fe1e022861eaa94e8f3907d10837afb1

Observation 3b843fee-cf32-4efb-88a2-681611e71858 · outbound

This paper cites PentestGPT: An LLM-empowered Automatic Penetration Testing Tool.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks PentestGPT: An LLM-empowered Automatic Penetration Testing Tool

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.900133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.900133Z digest=sha256:bd4d90322d302e55e8d79e8eee52e2d920345c2340e52c1b6acbcfa9040787a0

Observation a506cdaf-03d9-4bbf-8546-4787acd178c7 · outbound

This paper cites CodeMonkeys: Scaling Test-Time Compute for Software Engineering.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CodeMonkeys: Scaling Test-Time Compute for Software Engineering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.995056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.995056Z digest=sha256:20883b37e4f33b7046de409b9055a2e70e9941801ee6c9c1f4f558ca73543f7b

Observation e5c1bca7-a3d5-4dbe-8616-345224c88e81 · outbound

This paper cites Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:31.064570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:31.064570Z digest=sha256:18d97ca686a409a7eb4f02019b38b3dd3f68a0f3a3fe7792293aa058fec7a6a6

Observation 95dd83da-1adb-44f4-ad80-4d83218f093d · outbound

This paper cites Gonzalez, and Ion Stoica.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Gonzalez, and Ion Stoica

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:41.850357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:31.155235Z digest=sha256:59e03854eb8073d5b7aefceaa22c857c1aac97ac53bf166e01d46f639b36dbf9

Observation d941a680-1c4d-4da4-924e-50dd858da513 · outbound

This paper cites The Llama 3 Herd of Models.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:31.253381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:31.253381Z digest=sha256:60d266ab887eb9a9d0b7497b8df5b09658c6aa63b90efeff8cd3c51deceaaf4e

Observation 316e3584-dd13-49d2-bfa7-fb34d882d5d8 · outbound

This paper cites A Survey on LLM-as-a-Judge.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks A Survey on LLM-as-a-Judge

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:31.353188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:31.353188Z digest=sha256:f4bd852ca9519f9fbaa0f85332dde7aabf83db757f2b026d2cd3034bc5ff31a1

Observation 6c19e092-2e61-4550-8117-b1a7ec4a4f62 · outbound

This paper cites From Code to Courtroom: LLMs as the New Software Judges.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks From Code to Courtroom: LLMs as the New Software Judges

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:31.444375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:31.444375Z digest=sha256:952d0708e64a7bec78c650ae24becb8cf27bbef95955416d32f0214da42046e3

Observation 3647b23d-a7eb-49a6-befb-0dbbc17e4742 · outbound

This paper cites An empirical study on fine-tuning large language models of code for automated program repair.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks An empirical study on fine-tuning large language models of code for automated program repair

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:41.666376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:31.535547Z digest=sha256:798a3b0a88b96f6207e97e937a47c0700a50050403acd5949dbdd499afd6821d

Observation eaa332d2-4a91-4484-ba4b-004a39388cfc · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:41.298959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:31.724734Z digest=sha256:acdc3fda88aaf7cad69a81ed6f463bae63f7315c12a4a67ceba64a91481dfb2d

Observation 578fa9a7-ae6c-47e5-849d-3318038fe9c7 · outbound

This paper cites Self-planning code generation with large language models.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Self-planning code generation with large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:41.116337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:31.808480Z digest=sha256:377beebfcf0c5156144b0804d5ad766e9fd31f604c87773f1aaa663dfcd9937b

Observation ef81ede8-6741-4d9a-be3c-9706e88a8122 · outbound

This paper cites Critiquellm: Towards an informative critique generation model for evaluation of large language model generation.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Critiquellm: Towards an informative critique generation model for evaluation of large language model generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:40.948898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:31.884269Z digest=sha256:6790b8e15dd9680c3b345118a6ae30bac051340da73c1a9932c7414bdff942ce

Observation a3d91791-856c-4d30-bea3-b492ecddfe3b · outbound

This paper cites Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:40.717471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:31.947036Z digest=sha256:2a401d6938ce6a36f543c2273090925cbee47d6eaab226c9a409a7c59b861fb0

Observation b74c1185-6ad2-44c5-af6c-4095df969310 · outbound

This paper cites Overfitting in semantics-based automated program repair.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Overfitting in semantics-based automated program repair

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:40.627958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:31.986536Z digest=sha256:da80093354454591e2513dd33ad5c9056c6271a4a141645efe58d41c1c4338d2

Observation 2ddd0d21-3a93-4d42-b139-cdbb69f0486d · outbound

This paper cites Generative judge for evaluating alignment.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Generative judge for evaluating alignment

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:40.377848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:32.050286Z digest=sha256:4e61f908d4a4f3020c43d0cfc8458964c0fb81778b917c8f67d6ecc5b6efcd1c

Observation 73900c1d-eef2-4ca9-9e17-3bbe1bb4b175 · outbound

This paper cites Competition-level code generation with alphacode.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Competition-level code generation with alphacode

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:40.177026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:32.125628Z digest=sha256:3d857b3f6b3c8a28221b3354077cfd91d589c82f4caab71f6478a81424e6dd3d

Observation e923a080-abd6-415f-a6fd-116562346624 · outbound

This paper cites Llms for relational reasoning: How far are we? In Proceedings of the 1st International Workshop on Large Language Models for Code, pages 119–126, 2024.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Llms for relational reasoning: How far are we? In Proceedings of the 1st International Workshop on Large Language Models for Code, pages 119–126, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:39.982049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:32.215870Z digest=sha256:ec3b5555d0bf5320239bb23493c937d310f27c18e39138909d5b0c8f68f58dea

Observation 82359399-10d2-4abb-8e8b-46f46f96a630 · outbound

This paper cites RM-bench: Benchmarking reward models of language models with subtlety and style.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks RM-bench: Benchmarking reward models of language models with subtlety and style

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:39.801044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:32.300774Z digest=sha256:621a8a44d94bdd18cac030813baad1b9786c672fb7c6fdf6d26a73cc2928c9fc

Observation 631f1674-bce1-43d5-83b6-4f0f63460f6d · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Deepcoder: A fully open-source 14b coder at o3-mini level

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:39.578462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:32.397991Z digest=sha256:29842e4621ed6af74409363e1944c35842a5cfb83869afb9f262740ae9f82dfe

Observation 6868fb9f-f436-4f7d-959a-d7d92e1ba5c1 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks LLM Critics Help Catch LLM Bugs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:32.445876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:32.445876Z digest=sha256:b58d5a3d07f903e042cc46f3acd6cbb80771ca462e27f71a2154c253928c9b66

Observation 829352c5-1163-4e92-8643-563012c9b409 · outbound

This paper cites Swt-bench: Testing and validating real- world bug-fixes with code agents.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Swt-bench: Testing and validating real- world bug-fixes with code agents

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:39.338250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:32.554314Z digest=sha256:525535d416f4755835ffc5c1f0dc607a6e2d24f80dbed67722ab61f08f6dc022

Observation 472ed59f-d8a6-4962-bb05-5172541db030 · outbound

This paper cites Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:39.158021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:32.619413Z digest=sha256:fbcf9d5af3ad7634605ef6a0a55193d678d0f1006997138228ab17852302150f

Observation 5ebb1ec5-b98a-4307-8344-71d3e751d1ce · outbound

This paper cites M-prometheus: A suite of open multilingual llm judges.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks M-prometheus: A suite of open multilingual llm judges

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:32.748374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:32.748374Z digest=sha256:ea384cf52f086df39ce1f02a9be4721925e4273aebeb79035501a3ee469fe82d

Observation 37fc0660-7375-480d-81b2-a5f9e0a455bc · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:32.851096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:32.851096Z digest=sha256:e534c0d15c0603de92006359d6889449806ad044b08605e22ceaebda5e6bbbcb

Observation 021bad87-7e6c-4c79-9237-9370d53118ce · outbound

This paper cites Skywork critic model series.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Skywork critic model series

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:38.843015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:32.948957Z digest=sha256:2c00dab29feb941678604074e0f6ce4c5df0493c2e9448221965295b4743ad08

Observation 0501d0ec-bd1a-40fc-9468-b84909a2b173 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.016659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.016659Z digest=sha256:657a7337895c86a07c765a40491a313eb33aad055e08168fa4b3a28d79ed8fb7

Observation 8511f9d0-f646-4aae-9251-be4967b4aa68 · outbound

This paper cites Judgebench: A benchmark for evaluating LLM-based judges.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Judgebench: A benchmark for evaluating LLM-based judges

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:38.624158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:33.074682Z digest=sha256:f4841fa1db8dce2ac14993f23d4c070eed2adfbb6c0f2c06d4e104acc914cbf7

Observation c7f9eda7-12be-4143-a706-61e625c8f76d · outbound

This paper cites Code repair with LLMs gives an exploration-exploitation tradeoff.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Code repair with LLMs gives an exploration-exploitation tradeoff

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:38.439815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:33.171459Z digest=sha256:e67e0416f62a517138ff5d4ffb4715f977af63223cd32c65174ce20c1a3be343

Observation c5763663-0db8-4bcc-9788-2daf9c7148fb · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.228386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.228386Z digest=sha256:c253658b46ab6af0379decc06ed2304a2ef7e6702172eedf7ce7d48eb1d2999e

Observation 140f308e-9e28-4b3a-b437-16554802428b · outbound

This paper cites Can llms replace human evaluators? an empirical study of llm-as-a-judge in software engineering.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Can llms replace human evaluators? an empirical study of llm-as-a-judge in software engineering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:38.150125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:33.302701Z digest=sha256:e9532fe4a8fd3736209f3d6a062d57b99c7f4a75a577c2d05bcf079c892651ee

Observation ac0be8b9-04c1-4ea2-a14f-64a21bb03f77 · outbound

This paper cites Self-Taught Evaluators.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Self-Taught Evaluators

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.368009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.368009Z digest=sha256:d033131b5e06a33d1f772fd69bd01519836a4c05dc7ec13f033f81c45c098c64

Observation 96dfe862-aef5-47d9-8556-c1c967cc724d · outbound

This paper cites PandaLM: An automatic evaluation benchmark for LLM instruction tuning optimization.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks PandaLM: An automatic evaluation benchmark for LLM instruction tuning optimization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:37.938874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:33.423579Z digest=sha256:d32aadaf7a8fc9368dda196677620b3e5178c1ddfd50f8d239203f65292c693e

Observation d02dbe15-a1e3-44b9-8637-b54352e5b1d8 · outbound

This paper cites Chi, Tatsunori Hashimoto, O.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Chi, Tatsunori Hashimoto, O

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:37.753167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:33.480092Z digest=sha256:d270c4c57966d39343e05d1eddafef252b4140ba36df333e7e35c51e42dbf150

Observation 88c200d2-e87a-4020-99f4-0866db8a631e · outbound

This paper cites Weyssow, Aton Kamanda, Xin Zhou, and H.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Weyssow, Aton Kamanda, Xin Zhou, and H

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:37.467984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:33.542361Z digest=sha256:25cd9754bf7a5df2a84dae5ffe272f53c56f9fae0b39794f6f5cff2a86cbce60

Observation 315537e2-45d4-4d6b-8b39-5b45057ab3ba · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks BloombergGPT: A Large Language Model for Finance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.592955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.592955Z digest=sha256:3d747ea1cd9bab81d528104d16c69431b64aa8fe7e9361c727b2d4a01abf8f28

Observation e1b57fa7-d8c3-4db2-9623-1166c2b32fa1 · outbound

This paper cites Qwen3 Technical Report.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Qwen3 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.649376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.649376Z digest=sha256:bf5314df95384a3c8275cc195794b03813187745ad0085bae785e477732c389c

Observation c8630a18-a370-41d3-b810-fce5a54e322b · outbound

This paper cites Qwen2.5 Technical Report.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.704634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.704634Z digest=sha256:b3771c28eb33da801f04f5d123773a8dec8235dce669537251c01e638b74ea5f

Observation 3578691e-6982-4de8-8a26-b88ee4ce7fbf · outbound

This paper cites CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:32:34.645100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:33.789468Z digest=sha256:f5f88dcab5eef3c53f014759b2819f4461966b209bc2582a32b654286ed07c78

Observation 03322681-7dee-498c-a3c3-3894d5dd2a95 · outbound

This paper cites Fingpt: Open-source financial large language models.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Fingpt: Open-source financial large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.841936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.841936Z digest=sha256:9d62465e681b1f95fb61f68e0b5323358ef6c26ee4ea2f1fe400ed509abfb10a

Observation c3182861-bd31-4aff-8784-a6e4d46520d5 · outbound

This paper cites Demystifying long chain-of- thought reasoning in llms, 2025.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Demystifying long chain-of- thought reasoning in llms, 2025

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:37.208972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:33.881950Z digest=sha256:8dba0d466ec3d45812694bbcf958e622132d91e9e158135d0e87a7b1d98dabb8

Observation de824f34-39f5-4063-9f92-02c42a138841 · outbound

This paper cites ACECODER: Acing Coder RL via Automated Test-Case Synthesis.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks ACECODER: Acing Coder RL via Automated Test-Case Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.941396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.941396Z digest=sha256:21961b37392276d082e6e73d287697d9a106f2e8d809a2ff8c4b0d15844bfba7

Observation 14f188bb-f344-4d63-b28a-ffffff9a306f · outbound

This paper cites Codecriticbench: A holistic code critique benchmark for large language models, 2025.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Codecriticbench: A holistic code critique benchmark for large language models, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:36.936289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:33.991009Z digest=sha256:5aef32201d6bfc0ba3bf2b2c076db6a71ca550a6d6284bc8f990d84bffe88101

Observation e11e4bc7-896e-4b54-b34e-7ea4f720d65c · outbound

This paper cites an unresolved cited work.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:32:36.639897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:34.024035Z digest=sha256:fd922d83f62496e16cea70b1426a7853fcebb53dc2505d488dc46612817ebe10

Observation 9f89e6cd-5739-484b-be82-2f5872ce34bf · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:34.062676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:34.062676Z digest=sha256:3a0560c8eb5b857f89c502ed3493d72b25595a50ded2c39924b434a12c2ec73e

Observation 7fd513e5-0d7e-4684-89bc-9a3e21ce592c · outbound

This paper cites RMB: Compre- hensively benchmarking reward models in LLM alignment.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks RMB: Compre- hensively benchmarking reward models in LLM alignment

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:36.342972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:34.116865Z digest=sha256:668049c42b8deefaf4833a16ab58c1c7b2e57afc7547fbe14d0dbd351af05f07

Observation 748fb1ef-1c7e-4a3a-a8da-615210c80fcb · outbound

This paper cites Leveraging large language model for automatic patch correctness assessment.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Leveraging large language model for automatic patch correctness assessment

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:36.057972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:34.165260Z digest=sha256:497ae9112f08b14344599091ed262b9c9264cf31c2e2dbdf615692a6f9233c7b

Observation 31548f35-18a2-4c74-b9fa-d3c5a84bfa18 · outbound

This paper cites Evaluating judges as evaluators: The JETTS benchmark of LLM-as-judges as test-time scaling evaluators.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Evaluating judges as evaluators: The JETTS benchmark of LLM-as-judges as test-time scaling evaluators

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:35.747660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:34.238243Z digest=sha256:79969d7038c76f740a28751c4a2803c04d4435c74486d57bfa667092117b75a6

Observation a1aa5af6-0d44-43b3-ba26-2d6a0bcf856e · outbound

This paper cites JudgeLM: Fine-tuned large language models are scalable judges.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks JudgeLM: Fine-tuned large language models are scalable judges

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:35.431422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:34.292050Z digest=sha256:292092fb6ef5dd24a34ffb323e1363e23fac955a8877e6b3749c875f4859062f

Observation c34bd653-6094-41c6-abab-36d6b78352cc · outbound

This paper cites an unresolved cited work.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Unresolved cited work

Reference 1174

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:32:41.495630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:32:31.633118Z digest=sha256:55a0e2f45aea69f641e55f6a348ae79ad6f17b3d226cc92b85cf4f6b80dfb8a7

Pith citing papers

Observation 081fa9fd-8790-4a89-a31e-c299e3a060b9 · inbound

Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference cites this paper.

Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T20:21:21.808023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:21:21.808023Z digest=sha256:edcf2aa7bb9b2e996e91f72670cfe3280829e2178c795d788f12fac7a3c34f9f

Observation 0b3b896f-db59-46f6-85d3-2005ee87ec9d · inbound

SciML Agents: Write the Solver, Not the Solution cites this paper.

SciML Agents: Write the Solver, Not the Solution CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:01:43.843517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:57:51.444493Z digest=sha256:18a7cf92a4f91596f99628f1f674c92026f70d34ffb9cfabe43ac0f686cd0956

Observation 9c5510e2-213a-4a3a-98d2-e90d1936e3c6 · inbound

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering cites this paper.

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:32:00.313227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:29:03.994957Z digest=sha256:8a368591547569469bcdde116bdb8e9fb2791f73fca7c6dd3c4a2be202c605a7

Observation f0897a24-f187-4a50-ba93-1a1b3bbfa912 · inbound

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding cites this paper.

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:56:28.093044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T08:39:55.256518Z digest=sha256:3b6a81cfee11356c852227203fd3a81ce9c523cc7a4aae88a06cf12221a36dd5

Observation 5095cda1-e4e3-45de-acf0-126397d21c30 · inbound

ReMedi: Reasoner for Medical Clinical Prediction cites this paper.

ReMedi: Reasoner for Medical Clinical Prediction CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:01:05.917444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T14:20:29.672994Z digest=sha256:ff53cf816dd8ab45664623d36166f35e5717cc6b93fa38d906d913d3189d259a

Observation 4ab9dd30-515b-4947-a9b1-41a95316c4bc · inbound

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning cites this paper.

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:48.632378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:07:07.548384Z digest=sha256:0df9385c6fa69ac4cfcbd1566c57af83b08847571ad6ac87a2365b834134d236

Observation 5e398c30-299a-4459-83a9-72ed8fe09a01 · inbound

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning cites this paper.

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:27:24.609554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:25:16.650306Z digest=sha256:dfddbcf66ff01a9e9606d461129222a2403593b2453e2ce4a04f175b0ec636c7

Observation 53877acc-d601-4396-a639-446a0da43499 · inbound

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning cites this paper.

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T02:30:32.418305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:30:32.418305Z digest=sha256:c5b1f72b59d2fd964f021a7ea75105a81aab36d206bbaed1c1675cd0a045ebc8

Observation 7e8a0566-9712-4552-86ea-45b3bb2a9231 · inbound

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle cites this paper.

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:37:35.746653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T18:34:39.997353Z digest=sha256:038e994c0dfd075088fcdfa4a1327045e78506573fe4cf9fc3feabb4be7932ea

Observation c47fe0ad-dcbc-4e60-ba75-5ed87aa6e5fe · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-07-07T12:53:50.347229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-07T12:47:29.552283Z digest=sha256:bd49b3796932d3394c1543d4e644e875db22ebb3683c8200c2a659f39370ef0b

Observation b7a9f655-9f30-4eb0-99d8-c68ab6d5b659 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:e95eb2626d632659163da9ecf42aee1ef4a24f5ee6f9b0bb372d12ae5b9c94da