Pith. sign in

Paper Citation Record · LEDGER

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

As of 18 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 11 inbound Pith citation observations for arXiv:2507.10535.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10535 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:32:34.292050Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:21:21.808023Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T12:53:50.345403Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy31
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0c9225aa-51d4-4752-85d4-1bcb44500da7 · outbound

This paper cites Phi-4 Technical Report.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:29.662597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:29.662597Z digest=sha256:3607616d53da8f1ee026af20026dfaf42d6b5ad412840abeea6b69e660ca9d7a

Observation 9c8434c2-b472-434c-b1f1-b5318c68cdd3 · outbound

This paper cites GPT-4 Technical Report.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:29.722527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:29.722527Z digest=sha256:5d99487094d813af59479e0602d468854f25e32cca92aab94f6dd01dd3e8b56e

Observation e1e60d71-f17a-48b4-a3ee-b24feed19bc5 · outbound

This paper cites Automated unit test improvement using large language models at meta.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Automated unit test improvement using large language models at meta

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:42.690068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:29.910326Z digest=sha256:af545d5f205ca8f393f880194eaaea650facf711ffef9ad9264689e0471f3f2b

Observation 62df6330-63b9-4093-95d4-df8bb968ef8c · outbound

This paper cites Claude 3.7.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Claude 3.7

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:42.581478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:30.002094Z digest=sha256:4f6e412d2c3ad016f54639108b131f7d2a3d806f077363f6727ccd3ffe490ad1

Observation 9341a860-3a15-4d0d-b84c-f39b31c1ff6a · outbound

This paper cites Claude 4.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Claude 4

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:42.378283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:30.072314Z digest=sha256:9a32620f75d8a0453a44673fdeaf1f8cf6c08429ee62f39de1b7ed30a976378a

Observation 3bd8cac9-39a9-4a64-95fb-fc03c8fa2508 · outbound

This paper cites Program Synthesis with Large Language Models.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Program Synthesis with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.162463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.162463Z digest=sha256:90b6c60825b31e17d9e5160d49ff0ad204f32a5e972118d4e44ea8d3f88cec4b

Observation 0e4ec585-5f9d-41e7-979f-cb0bb4d58296 · outbound

This paper cites Codet: Code generation with generated tests.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Codet: Code generation with generated tests

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.256357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.256357Z digest=sha256:5d959393e5d8536bc70eaf98a606d525e7c7af45de8755e056523c50946fa83c

Observation 45ae6b93-0aed-4333-8e5a-2843f38df120 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Evaluating Large Language Models Trained on Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.358202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.358202Z digest=sha256:0dcc771ea029aa07d9886eecc98152bf6bea0b6d382f7184ca3c759d21aba4cf

Observation d409d1fb-c907-4c60-987b-00016a2eb356 · outbound

This paper cites Teaching large language models to self-debug.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Teaching large language models to self-debug

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:42.179788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:30.467485Z digest=sha256:2e8f6d63da643e58d247cd6ac984094d19c71a194df3486b313fe25ec4fb3998

Observation 8f725fbe-d2e9-48e0-9fb4-00a9eb87b637 · outbound

This paper cites Rm-r1: Reward modeling as reasoning.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Rm-r1: Reward modeling as reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.562154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.562154Z digest=sha256:544260b51189abfa9ad48c49c052bb81eebec630c60a9f182a398cca0b1dc096

Observation 6712258d-8282-40f6-ad1f-31d0bc269b16 · outbound

This paper cites AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.634578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.634578Z digest=sha256:470cb5c8c927ee1522a203d7cea959622d129d41753a9f8ef4c49e18bff42126

Observation 695a6cdc-7378-42b7-bcb5-426c16d294c5 · outbound

This paper cites an unresolved cited work.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:32:42.080154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:30.697316Z digest=sha256:1a5306d14d3668afbfebde13038eba5b8e8321029d4c1868302d840906bac4f6

Observation 410a27dd-5851-4e7c-8697-13ed5db91dc1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.805802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.805802Z digest=sha256:2dc0313c01dc2b3953b2225c201bec4194901618a04a192f3732898626a79125

Observation 3b843fee-cf32-4efb-88a2-681611e71858 · outbound

This paper cites PentestGPT: An LLM-empowered Automatic Penetration Testing Tool.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks PentestGPT: An LLM-empowered Automatic Penetration Testing Tool

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.900133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.900133Z digest=sha256:871602116a14e7b039770c0e9a3612c49464b3796ca6e7be0989176088940d5d

Observation a506cdaf-03d9-4bbf-8546-4787acd178c7 · outbound

This paper cites CodeMonkeys: Scaling Test-Time Compute for Software Engineering.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CodeMonkeys: Scaling Test-Time Compute for Software Engineering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.995056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.995056Z digest=sha256:2a187941e76f524877cbf03de18600e4d347bdccec309b4db0e9e2fee5d9899f

Observation e5c1bca7-a3d5-4dbe-8616-345224c88e81 · outbound

This paper cites Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:31.064570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:31.064570Z digest=sha256:da2e7fe96c3ea99f6d06f2232d7aff05b86f4608202fb6eb45f624b41734088a

Observation 95dd83da-1adb-44f4-ad80-4d83218f093d · outbound

This paper cites Gonzalez, and Ion Stoica.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Gonzalez, and Ion Stoica

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:41.850357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:31.155235Z digest=sha256:5cdbf1a67c15db53243ceaa725970f243e4e226f84c1244e7f64d048c4b520ad

Observation d941a680-1c4d-4da4-924e-50dd858da513 · outbound

This paper cites The Llama 3 Herd of Models.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:31.253381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:31.253381Z digest=sha256:6fe8d9a09b3024ff79e3e1473ae9d263028014b5511ec7933c0981d3ec375ec2

Observation 316e3584-dd13-49d2-bfa7-fb34d882d5d8 · outbound

This paper cites A Survey on LLM-as-a-Judge.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks A Survey on LLM-as-a-Judge

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:31.353188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:31.353188Z digest=sha256:f8a7b87bd5882c5ae28b04b3c45c92f5dd260bf1716875a05e16bac8f9c5cf68

Observation 6c19e092-2e61-4550-8117-b1a7ec4a4f62 · outbound

This paper cites From Code to Courtroom: LLMs as the New Software Judges.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks From Code to Courtroom: LLMs as the New Software Judges

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:31.444375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:31.444375Z digest=sha256:4919d64b09e985e34fa206793cf0dd3b95129f08cb3c98c57201eab100aa3262

Observation 3647b23d-a7eb-49a6-befb-0dbbc17e4742 · outbound

This paper cites An empirical study on fine-tuning large language models of code for automated program repair.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks An empirical study on fine-tuning large language models of code for automated program repair

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:41.666376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:31.535547Z digest=sha256:c22b14a129e05ffb4e72757f2e0345d97aaade68b7d1ec58fd8910dbee1a4ea8

Observation eaa332d2-4a91-4484-ba4b-004a39388cfc · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:41.298959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:31.724734Z digest=sha256:b0f1eb99afe16f4d9a7a6b3b8af25cd8e67678c81f5f13a23e3922654a53123b

Observation 578fa9a7-ae6c-47e5-849d-3318038fe9c7 · outbound

This paper cites Self-planning code generation with large language models.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Self-planning code generation with large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:41.116337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:31.808480Z digest=sha256:8cfa2057bcb9fd695880f981aecb9bf9a88c84979bb51a2791147b408554f823

Observation ef81ede8-6741-4d9a-be3c-9706e88a8122 · outbound

This paper cites Critiquellm: Towards an informative critique generation model for evaluation of large language model generation.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Critiquellm: Towards an informative critique generation model for evaluation of large language model generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:40.948898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:31.884269Z digest=sha256:957f21761bd2b9ef3852853ab74d7df19ae4f3ed8fc0ff982f9864d438e95ad2

Observation a3d91791-856c-4d30-bea3-b492ecddfe3b · outbound

This paper cites Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:40.717471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:31.947036Z digest=sha256:4e80c39988ba69c517d2d71afa24695058cc73b316a4fdf19d2c1e870a84b50b

Observation b74c1185-6ad2-44c5-af6c-4095df969310 · outbound

This paper cites Overfitting in semantics-based automated program repair.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Overfitting in semantics-based automated program repair

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:40.627958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:31.986536Z digest=sha256:0590f280c9a1b0ba9f5f815d52a9e7588edf647ecf98bad87ebfeb0f313d9e77

Observation 2ddd0d21-3a93-4d42-b139-cdbb69f0486d · outbound

This paper cites Generative judge for evaluating alignment.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Generative judge for evaluating alignment

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:40.377848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:32.050286Z digest=sha256:f2d575c186f83cf9b0c97e6de7a058485fdca6fc39c78d30adb3067f5a4c8c65

Observation 73900c1d-eef2-4ca9-9e17-3bbe1bb4b175 · outbound

This paper cites Competition-level code generation with alphacode.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Competition-level code generation with alphacode

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:40.177026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:32.125628Z digest=sha256:f85f31df61af39d8c158d7d180826b85037c60f2d95b0a1fcf1580a830d8b600

Observation e923a080-abd6-415f-a6fd-116562346624 · outbound

This paper cites Llms for relational reasoning: How far are we? In Proceedings of the 1st International Workshop on Large Language Models for Code, pages 119–126, 2024.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Llms for relational reasoning: How far are we? In Proceedings of the 1st International Workshop on Large Language Models for Code, pages 119–126, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:39.982049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:32.215870Z digest=sha256:6c7e7b2895c40570d873661f2e12f83e2be86ff524734a1974f073c5d1968859

Observation 82359399-10d2-4abb-8e8b-46f46f96a630 · outbound

This paper cites RM-bench: Benchmarking reward models of language models with subtlety and style.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks RM-bench: Benchmarking reward models of language models with subtlety and style

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:39.801044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:32.300774Z digest=sha256:128afc3aaa1c1022df597dd42a7e156c1cfab60288611a0a7654b09f493a3e84

Observation 631f1674-bce1-43d5-83b6-4f0f63460f6d · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Deepcoder: A fully open-source 14b coder at o3-mini level

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:39.578462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:32.397991Z digest=sha256:6034379b8783dca17a64a33fed6872f40756909c6b6f975dfa51b8aaa427dd71

Observation 6868fb9f-f436-4f7d-959a-d7d92e1ba5c1 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks LLM Critics Help Catch LLM Bugs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:32.445876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:32.445876Z digest=sha256:8da9b85222c85cbe056ad314b724045525886a4cd5552afd9e57bcfb5e14a6ab

Observation 829352c5-1163-4e92-8643-563012c9b409 · outbound

This paper cites Swt-bench: Testing and validating real- world bug-fixes with code agents.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Swt-bench: Testing and validating real- world bug-fixes with code agents

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:39.338250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:32.554314Z digest=sha256:12fb6594767ceed55e02e566c9bf945dc48672c0211aba8294622dcaa6f636d6

Observation 472ed59f-d8a6-4962-bb05-5172541db030 · outbound

This paper cites Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:39.158021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:32.619413Z digest=sha256:f722ebf1ef9156bc65bd34b7136dfdb0338be04a06e519b5c036b4dab5d9bb52

Observation 5ebb1ec5-b98a-4307-8344-71d3e751d1ce · outbound

This paper cites M-prometheus: A suite of open multilingual llm judges.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks M-prometheus: A suite of open multilingual llm judges

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:32.748374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:32.748374Z digest=sha256:a9fafba39865a519915f5bc9ecf61eedd35d8ab4745fd58dc6fd23f53b88e893

Observation 37fc0660-7375-480d-81b2-a5f9e0a455bc · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:32.851096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:32.851096Z digest=sha256:8c821bfd775fab32eb8297298397fe86ab5a07d97161818a57c6038a228f6949

Observation 021bad87-7e6c-4c79-9237-9370d53118ce · outbound

This paper cites Skywork critic model series.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Skywork critic model series

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:38.843015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:32.948957Z digest=sha256:7a1bc98240d1901f4a115a43121ad2b028a772b3cf613e35ec266c903feb9fac

Observation 0501d0ec-bd1a-40fc-9468-b84909a2b173 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.016659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.016659Z digest=sha256:a72d1af7bc9b41d5c5da58aa560f7382cfde4b83dae4ed3775fc44267f09205c

Observation 8511f9d0-f646-4aae-9251-be4967b4aa68 · outbound

This paper cites Judgebench: A benchmark for evaluating LLM-based judges.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Judgebench: A benchmark for evaluating LLM-based judges

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:38.624158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:33.074682Z digest=sha256:aed1947fb34af3445c88950e74d8a211b9ad791f8b447df1a4700ccf3fc86a81

Observation c7f9eda7-12be-4143-a706-61e625c8f76d · outbound

This paper cites Code repair with LLMs gives an exploration-exploitation tradeoff.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Code repair with LLMs gives an exploration-exploitation tradeoff

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:38.439815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:33.171459Z digest=sha256:04a385b953933cc3fd07855ac3cd845c1f7054fce4258629659b6cac46c4cf42

Observation c5763663-0db8-4bcc-9788-2daf9c7148fb · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.228386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.228386Z digest=sha256:2dff7c957061234218cb8dce6a9e2c7f305bb160c99cc81df23ce142ab19dd16

Observation 140f308e-9e28-4b3a-b437-16554802428b · outbound

This paper cites Can llms replace human evaluators? an empirical study of llm-as-a-judge in software engineering.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Can llms replace human evaluators? an empirical study of llm-as-a-judge in software engineering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:38.150125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:33.302701Z digest=sha256:c2a456da5ee404b47c8cc99d0870d476f5786bc028dce9c14725e39db63d593c

Observation ac0be8b9-04c1-4ea2-a14f-64a21bb03f77 · outbound

This paper cites Self-Taught Evaluators.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Self-Taught Evaluators

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.368009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.368009Z digest=sha256:b5a5007132c041864d70642e0556d32356df57ed005c53c22b52fadb40a0ab7f

Observation 96dfe862-aef5-47d9-8556-c1c967cc724d · outbound

This paper cites PandaLM: An automatic evaluation benchmark for LLM instruction tuning optimization.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks PandaLM: An automatic evaluation benchmark for LLM instruction tuning optimization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:37.938874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:33.423579Z digest=sha256:cb9ae3b1149dd0f01514e940b5a1d7da474a62191d11a7d3da0554dd3c5abf6b

Observation d02dbe15-a1e3-44b9-8637-b54352e5b1d8 · outbound

This paper cites Chi, Tatsunori Hashimoto, O.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Chi, Tatsunori Hashimoto, O

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:37.753167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:33.480092Z digest=sha256:47ef074097fbb61644c76713947a3fb5509fd515fc28a94ca56c6ef99fb1da3d

Observation 88c200d2-e87a-4020-99f4-0866db8a631e · outbound

This paper cites Weyssow, Aton Kamanda, Xin Zhou, and H.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Weyssow, Aton Kamanda, Xin Zhou, and H

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:37.467984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:33.542361Z digest=sha256:f2e50f0b4cee552dc1aec86da5ec5be0c0f591a48f0964acfe84e9f19bde4e0d

Observation 315537e2-45d4-4d6b-8b39-5b45057ab3ba · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks BloombergGPT: A Large Language Model for Finance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.592955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.592955Z digest=sha256:bf380bb21da246559e12f4222b5b01c128fbea7d1c2770d93b134cddb5f09648

Observation e1b57fa7-d8c3-4db2-9623-1166c2b32fa1 · outbound

This paper cites Qwen3 Technical Report.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Qwen3 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.649376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.649376Z digest=sha256:0258d72aa0c37c9763e4212a5bf974c284f5102692b00ca237bb2ae537ed917e

Observation c8630a18-a370-41d3-b810-fce5a54e322b · outbound

This paper cites Qwen2.5 Technical Report.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.704634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.704634Z digest=sha256:957211b758523c2a9c9fdfa906b5455c66519eb32ff6ff7eea8148961a440a7d

Observation 3578691e-6982-4de8-8a26-b88ee4ce7fbf · outbound

This paper cites CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:32:34.645100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:33.789468Z digest=sha256:a5b26dfb0cfa2cec45b8e99a794693a76e9dd2664f89644a0a34f35e57fe3da6

Observation 03322681-7dee-498c-a3c3-3894d5dd2a95 · outbound

This paper cites Fingpt: Open-source financial large language models.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Fingpt: Open-source financial large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.841936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.841936Z digest=sha256:2933774b315e0409f83b1707905c8d9057609180670db67447c34bbe198ab77e

Observation c3182861-bd31-4aff-8784-a6e4d46520d5 · outbound

This paper cites Demystifying long chain-of- thought reasoning in llms, 2025.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Demystifying long chain-of- thought reasoning in llms, 2025

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:37.208972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:33.881950Z digest=sha256:e70724f8792fb72b730ed703b61a9a24d75581cf688f6d6ec3cba73c361fead5

Observation de824f34-39f5-4063-9f92-02c42a138841 · outbound

This paper cites ACECODER: Acing Coder RL via Automated Test-Case Synthesis.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks ACECODER: Acing Coder RL via Automated Test-Case Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.941396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.941396Z digest=sha256:a27b0afd521d5bb858bfc56ac86cf2e54ef7dfb8240e073ec11c420929e8d9f0

Observation 14f188bb-f344-4d63-b28a-ffffff9a306f · outbound

This paper cites Codecriticbench: A holistic code critique benchmark for large language models, 2025.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Codecriticbench: A holistic code critique benchmark for large language models, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:36.936289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:33.991009Z digest=sha256:88235ce2b9146bb9d537d0b7ecc6aefc54e71caf0ed3d94d1f6ef627ecad8016

Observation e11e4bc7-896e-4b54-b34e-7ea4f720d65c · outbound

This paper cites an unresolved cited work.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:32:36.639897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:34.024035Z digest=sha256:0943d9e084a7e31e1d522b60b14fe5270ab0ba5347fc010cf5bd6e7b71938ce8

Observation 9f89e6cd-5739-484b-be82-2f5872ce34bf · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:34.062676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:34.062676Z digest=sha256:3b06ce2eb0ab65fadfc15db0afaa72385a85fe4ae23e8ab5f03e28214ab8642e

Observation 7fd513e5-0d7e-4684-89bc-9a3e21ce592c · outbound

This paper cites RMB: Compre- hensively benchmarking reward models in LLM alignment.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks RMB: Compre- hensively benchmarking reward models in LLM alignment

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:36.342972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:34.116865Z digest=sha256:cd96221eac3e5579030dd0d82ef37b74a55d7fdd11e2ad87dafb449af6956a22

Observation 748fb1ef-1c7e-4a3a-a8da-615210c80fcb · outbound

This paper cites Leveraging large language model for automatic patch correctness assessment.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Leveraging large language model for automatic patch correctness assessment

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:36.057972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:34.165260Z digest=sha256:a1c7c8cf9e0634f94094f11391101084005e363d6e51696463ac6ac0eb371b0e

Observation 31548f35-18a2-4c74-b9fa-d3c5a84bfa18 · outbound

This paper cites Evaluating judges as evaluators: The JETTS benchmark of LLM-as-judges as test-time scaling evaluators.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Evaluating judges as evaluators: The JETTS benchmark of LLM-as-judges as test-time scaling evaluators

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:35.747660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:34.238243Z digest=sha256:5e8d0722a7dcfc8ab951b6711d139ce413f64da8f7111be9d534c19e9acdda2f

Observation a1aa5af6-0d44-43b3-ba26-2d6a0bcf856e · outbound

This paper cites JudgeLM: Fine-tuned large language models are scalable judges.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks JudgeLM: Fine-tuned large language models are scalable judges

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:32:35.431422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:34.292050Z digest=sha256:5abdd1bb329aa0d92bfc6ebc4227c7a274d8cd71e7366670c615ec67fea3eb38

Observation c34bd653-6094-41c6-abab-36d6b78352cc · outbound

This paper cites an unresolved cited work.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Unresolved cited work

Reference 1174

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:32:41.495630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T17:32:31.633118Z digest=sha256:58bebb67a39dd0d484ce940e8d799c39b3da3b146086b7436406fb7f465da1be

Pith citing papers

Observation 081fa9fd-8790-4a89-a31e-c299e3a060b9 · inbound

Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference cites this paper.

Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T20:21:21.808023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:21:21.808023Z digest=sha256:e98b013ca5754e456452367e1039f3a164a3869a3625d12dd892dfd6ba901ec1

Observation 0b3b896f-db59-46f6-85d3-2005ee87ec9d · inbound

SciML Agents: Write the Solver, Not the Solution cites this paper.

SciML Agents: Write the Solver, Not the Solution CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:01:43.843517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T17:57:51.444493Z digest=sha256:539b771d0b70b86dc27f27e4b6e79f53a856f886df01f0d1eded92d3bf22bbe0

Observation 9c5510e2-213a-4a3a-98d2-e90d1936e3c6 · inbound

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering cites this paper.

Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:32:00.313227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T07:29:03.994957Z digest=sha256:9a7c93b7d08c4c5f234b45cb16e5da5bc5a495ccea2628b136541ca5668a3f5f

Observation f0897a24-f187-4a50-ba93-1a1b3bbfa912 · inbound

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding cites this paper.

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:56:28.093044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T08:39:55.256518Z digest=sha256:34b812373dabc96a4c97c070356bc3799433cd95f1db2acfd49cda9fe04c1e51

Observation 5095cda1-e4e3-45de-acf0-126397d21c30 · inbound

ReMedi: Reasoner for Medical Clinical Prediction cites this paper.

ReMedi: Reasoner for Medical Clinical Prediction CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:01:05.917444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-09T14:20:29.672994Z digest=sha256:743af0934aac7f48c7edad8c644906b086a194dfc8e315ebc0dbaf87b6e04635

Observation 4ab9dd30-515b-4947-a9b1-41a95316c4bc · inbound

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning cites this paper.

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:15:48.632378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T20:07:07.548384Z digest=sha256:531f42827e96c0f3f7f6da603868fd4d18cc17d80b1d61b09319e5afe6df20ee

Observation 5e398c30-299a-4459-83a9-72ed8fe09a01 · inbound

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning cites this paper.

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:27:24.609554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T06:25:16.650306Z digest=sha256:be31a6a03310845393469f28a4ac5b259befabfee7e2d89d08090ebe5431480c

Observation 53877acc-d601-4396-a639-446a0da43499 · inbound

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning cites this paper.

OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T02:30:32.418305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:30:32.418305Z digest=sha256:cf2f9fface8faf9ffff229d239e8e032722b7fc01b17402be99699e8afe1e7fd

Observation 7e8a0566-9712-4552-86ea-45b3bb2a9231 · inbound

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle cites this paper.

SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:37:35.746653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T18:34:39.997353Z digest=sha256:cb858ecf45bcd1419be9dbbe643e2122f396d784f362d32b7a339770dd6d2cab

Observation c47fe0ad-dcbc-4e60-ba75-5ed87aa6e5fe · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-07-07T12:53:50.347229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-07T12:47:29.552283Z digest=sha256:7012e7a281a866b40e0cfd41ce8caee37330fa415d69512ee269e42134d54beb

Observation b7a9f655-9f30-4eb0-99d8-c68ab6d5b659 · inbound

LLM-as-a-Verifier: A General-Purpose Verification Framework cites this paper.

LLM-as-a-Verifier: A General-Purpose Verification Framework CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-11T07:02:51.850836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:02:51.850836Z digest=sha256:b17082428dabd054f36af87efacdc0a58272bc2fb522fa472b25282af2bf06a5