REVIEW 3 cited by
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advancements in large language models (LLMs) have showcased impressive code generation capabilities, primarily evaluated through language-to-code benchmarks. However, these benchmarks may not fully capture a model's code understanding abilities. We introduce CodeJudge-Eval (CJ-Eval), a novel benchmark designed to assess LLMs' code understanding abilities from the perspective of code judging rather than code generation. CJ-Eval challenges models to determine the correctness of provided code solutions, encompassing various error types and compilation issues. By leveraging a diverse set of problems and a fine-grained judging system, CJ-Eval addresses the limitations of traditional benchmarks, including the potential memorization of solutions. Evaluation of 12 well-known LLMs on CJ-Eval reveals that even state-of-the-art models struggle, highlighting the benchmark's ability to probe deeper into models' code understanding abilities. Our codes and benchmark are available at \url{https://github.com/CodeLLM-Research/CodeJudge-Eval}.
Forward citations
Cited by 3 Pith papers
-
ReCatcher: Towards LLMs Regression Testing for Code Generation
ReCatcher systematically measures regressions in LLM code generation across correctness, static quality, and performance, and its evaluation shows fine-tuning, merging, and new releases each introduce specific regressions.
-
Is Your Automated Software Engineer Trustworthy?
BouncerBench evaluates whether LLM coding agents can abstain from acting on underspecified tickets and incorrect patches, and shows current models rarely abstain correctly.
-
Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge
Fine-tuned EP-Shield detects and purifies identifier-substitution attacks on code models, restoring correct predictions on roughly 90 to 99 percent of adversarial examples across three tasks.
Discussion (0). Continue with ORCID to comment.