Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:32:34.292050Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 11 inbound Pith citation observations for arXiv:2507.10535.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:32:34.292050Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T20:21:21.808023Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-07T12:53:50.345403Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0c9225aa-51d4-4752-85d4-1bcb44500da7 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Phi-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c8434c2-b472-434c-b1f1-b5318c68cdd3 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1e60d71-f17a-48b4-a3ee-b24feed19bc5 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Automated unit test improvement using large language models at meta
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62df6330-63b9-4093-95d4-df8bb968ef8c · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Claude 3.7
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9341a860-3a15-4d0d-b84c-f39b31c1ff6a · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Claude 4
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3bd8cac9-39a9-4a64-95fb-fc03c8fa2508 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Program Synthesis with Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e4ec585-5f9d-41e7-979f-cb0bb4d58296 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Codet: Code generation with generated tests
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ae6b93-0aed-4333-8e5a-2843f38df120 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Evaluating Large Language Models Trained on Code
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d409d1fb-c907-4c60-987b-00016a2eb356 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Teaching large language models to self-debug
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f725fbe-d2e9-48e0-9fb4-00a9eb87b637 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Rm-r1: Reward modeling as reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6712258d-8282-40f6-ad1f-31d0bc269b16 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 695a6cdc-7378-42b7-bcb5-426c16d294c5 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 410a27dd-5851-4e7c-8697-13ed5db91dc1 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b843fee-cf32-4efb-88a2-681611e71858 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a506cdaf-03d9-4bbf-8546-4787acd178c7 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CodeMonkeys: Scaling Test-Time Compute for Software Engineering
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5c1bca7-a3d5-4dbe-8616-345224c88e81 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95dd83da-1adb-44f4-ad80-4d83218f093d · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Gonzalez, and Ion Stoica
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d941a680-1c4d-4da4-924e-50dd858da513 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 316e3584-dd13-49d2-bfa7-fb34d882d5d8 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks A Survey on LLM-as-a-Judge
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c19e092-2e61-4550-8117-b1a7ec4a4f62 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks From Code to Courtroom: LLMs as the New Software Judges
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3647b23d-a7eb-49a6-befb-0dbbc17e4742 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks An empirical study on fine-tuning large language models of code for automated program repair
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eaa332d2-4a91-4484-ba4b-004a39388cfc · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Livecodebench: Holistic and contamination free evaluation of large language models for code
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 578fa9a7-ae6c-47e5-849d-3318038fe9c7 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Self-planning code generation with large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef81ede8-6741-4d9a-be3c-9706e88a8122 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Critiquellm: Towards an informative critique generation model for evaluation of large language model generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3d91791-856c-4d30-bea3-b492ecddfe3b · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b74c1185-6ad2-44c5-af6c-4095df969310 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Overfitting in semantics-based automated program repair
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ddd0d21-3a93-4d42-b139-cdbb69f0486d · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Generative judge for evaluating alignment
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73900c1d-eef2-4ca9-9e17-3bbe1bb4b175 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Competition-level code generation with alphacode
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e923a080-abd6-415f-a6fd-116562346624 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Llms for relational reasoning: How far are we? In Proceedings of the 1st International Workshop on Large Language Models for Code, pages 119–126, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82359399-10d2-4abb-8e8b-46f46f96a630 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks RM-bench: Benchmarking reward models of language models with subtlety and style
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 631f1674-bce1-43d5-83b6-4f0f63460f6d · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Deepcoder: A fully open-source 14b coder at o3-mini level
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6868fb9f-f436-4f7d-959a-d7d92e1ba5c1 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks LLM Critics Help Catch LLM Bugs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 829352c5-1163-4e92-8643-563012c9b409 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Swt-bench: Testing and validating real- world bug-fixes with code agents
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 472ed59f-d8a6-4962-bb05-5172541db030 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ebb1ec5-b98a-4307-8344-71d3e751d1ce · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks M-prometheus: A suite of open multilingual llm judges
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37fc0660-7375-480d-81b2-a5f9e0a455bc · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 021bad87-7e6c-4c79-9237-9370d53118ce · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Skywork critic model series
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0501d0ec-bd1a-40fc-9468-b84909a2b173 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8511f9d0-f646-4aae-9251-be4967b4aa68 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Judgebench: A benchmark for evaluating LLM-based judges
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7f9eda7-12be-4143-a706-61e625c8f76d · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Code repair with LLMs gives an exploration-exploitation tradeoff
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5763663-0db8-4bcc-9788-2daf9c7148fb · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 140f308e-9e28-4b3a-b437-16554802428b · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Can llms replace human evaluators? an empirical study of llm-as-a-judge in software engineering
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac0be8b9-04c1-4ea2-a14f-64a21bb03f77 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Self-Taught Evaluators
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96dfe862-aef5-47d9-8556-c1c967cc724d · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks PandaLM: An automatic evaluation benchmark for LLM instruction tuning optimization
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d02dbe15-a1e3-44b9-8637-b54352e5b1d8 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Chi, Tatsunori Hashimoto, O
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88c200d2-e87a-4020-99f4-0866db8a631e · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Weyssow, Aton Kamanda, Xin Zhou, and H
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 315537e2-45d4-4d6b-8b39-5b45057ab3ba · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks BloombergGPT: A Large Language Model for Finance
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1b57fa7-d8c3-4db2-9623-1166c2b32fa1 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Qwen3 Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8630a18-a370-41d3-b810-fce5a54e322b · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Qwen2.5 Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3578691e-6982-4de8-8a26-b88ee4ce7fbf · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03322681-7dee-498c-a3c3-3894d5dd2a95 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Fingpt: Open-source financial large language models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3182861-bd31-4aff-8784-a6e4d46520d5 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Demystifying long chain-of- thought reasoning in llms, 2025
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de824f34-39f5-4063-9f92-02c42a138841 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks ACECODER: Acing Coder RL via Automated Test-Case Synthesis
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14f188bb-f344-4d63-b28a-ffffff9a306f · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Codecriticbench: A holistic code critique benchmark for large language models, 2025
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e11e4bc7-896e-4b54-b34e-7ea4f720d65c · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f89e6cd-5739-484b-be82-2f5872ce34bf · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fd513e5-0d7e-4684-89bc-9a3e21ce592c · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks RMB: Compre- hensively benchmarking reward models in LLM alignment
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 748fb1ef-1c7e-4a3a-a8da-615210c80fcb · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Leveraging large language model for automatic patch correctness assessment
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31548f35-18a2-4c74-b9fa-d3c5a84bfa18 · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Evaluating judges as evaluators: The JETTS benchmark of LLM-as-judges as test-time scaling evaluators
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1aa5af6-0d44-43b3-ba26-2d6a0bcf856e · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks JudgeLM: Fine-tuned large language models are scalable judges
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c34bd653-6094-41c6-abab-36d6b78352cc · outbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Unresolved cited work
Reference 1174
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 081fa9fd-8790-4a89-a31e-c299e3a060b9 · inbound
Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b3b896f-db59-46f6-85d3-2005ee87ec9d · inbound
SciML Agents: Write the Solver, Not the Solution CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c5510e2-213a-4a3a-98d2-e90d1936e3c6 · inbound
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0897a24-f187-4a50-ba93-1a1b3bbfa912 · inbound
LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5095cda1-e4e3-45de-acf0-126397d21c30 · inbound
ReMedi: Reasoner for Medical Clinical Prediction CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ab9dd30-515b-4947-a9b1-41a95316c4bc · inbound
OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e398c30-299a-4459-83a9-72ed8fe09a01 · inbound
OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53877acc-d601-4396-a639-446a0da43499 · inbound
OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e8a0566-9712-4552-86ea-45b3bb2a9231 · inbound
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c47fe0ad-dcbc-4e60-ba75-5ed87aa6e5fe · inbound
LLM-as-a-Verifier: A General-Purpose Verification Framework CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7a9f655-9f30-4eb0-99d8-c68ab6d5b659 · inbound
LLM-as-a-Verifier: A General-Purpose Verification Framework CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.