REVIEW 3 cited by
Measuring Determinism in Large Language Models for Software Code Review
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) promise to streamline software code reviews, but their ability to produce consistent assessments remains an open question. In this study, we tested four leading LLMs -- GPT-4o mini, GPT-4o, Claude 3.5 Sonnet, and LLaMA 3.2 90B Vision -- on 70 Java commits from both private and public repositories. By setting each model's temperature to zero, clearing context, and repeating the exact same prompts five times, we measured how consistently each model generated code-review assessments. Our results reveal that even with temperature minimized, LLM responses varied to different degrees. These findings highlight a consideration about the inherently limited consistency (test-retest reliability) of LLMs -- even when the temperature is set to zero -- and the need for caution when using LLM-generated code reviews to make real-world decisions.
Forward citations
Cited by 3 Pith papers
-
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
DF3DV-1K supplies 1,048 scenes with clean and cluttered image pairs plus a challenging 41-scene subset to benchmark and improve distractor-free radiance field methods.
-
Linear-LLM-SCM: Benchmarking LLMs for Coefficient Elicitation in Linear-Gaussian Causal Models
LLMs asked to fill in linear-Gaussian causal equations give inaccurate, unstable, and perturbation-sensitive coefficients; the open-source Linear-LLM-SCM benchmark measures this, with Gemini 2.5 Flash leading on scale...
-
Automated Bug Frame Retrieval from Gameplay Videos Using Vision-Language Models
A keyframe-plus-GPT-4o pipeline retrieves the single most representative frame for a reported gameplay bug, with F1@1 of 0.79 and Accuracy@1 of 0.89 on industrial bug-report videos.
Discussion (0). Continue with ORCID to comment.