Pith. sign in

REVIEW 32 cited by

Non-Determinism of "Deterministic" LLM Settings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.04667 v5 pith:STT55N72 submitted 2024-08-06 cs.CL cs.AIcs.LGcs.SE

classification cs.CLcs.AIcs.LGcs.SE
keywords acrossdeterministicnon-determinismrunssettingsaccuracyagreementdata
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

LLM (large language model) practitioners commonly notice that outputs can vary for the same inputs under settings expected to be deterministic. Yet the questions of how pervasive this is, and with what impact on results, have not to our knowledge been systematically investigated. We investigate non-determinism in five LLMs configured to be deterministic when applied to eight common tasks in across 10 runs, in both zero-shot and few-shot settings. We see accuracy variations up to 15% across naturally occurring runs with a gap of best possible performance to worst possible performance up to 70%. In fact, none of the LLMs consistently delivers repeatable accuracy across all tasks, much less identical output strings. Sharing preliminary results with insiders has revealed that non-determinism perhaps essential to the efficient use of compute resources via co-mingled data in input buffers so this issue is not going away anytime soon. To better quantify our observations, we introduce metrics focused on quantifying determinism, TARr@N for the total agreement rate at N runs over raw output, and TARa@N for total agreement rate of parsed-out answers. Our code and data are publicly available at https://github.com/breckbaldwin/llm-stability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 32 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Cross-model value distances from single draws are inflated by response determinism and confounded by the deployment client; a repeated counterbalanced protocol plus flip/magnitude decomposition separates them.

  2. Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam

    econ.GN 2026-07 conditional novelty 6.5 of 10

    Inserting an irrelevant passage into graduate microeconomics problems lowers LLM final-answer accuracy by 12.3 percentage points, corrupts the reasoning, preserves response form, and makes models rate the corrupted ta...

  3. Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit

    cs.CR 2026-04 unverdicted novelty 6.5 of 10

    Security practitioners use LLMs independently for low-risk productivity tasks while showing interest in enterprise platforms, but reliability, verification needs, and security risks limit broader autonomy.

  4. Governing Agentic AI in FinTech

    cs.CY 2026-08 conditional novelty 6.0 of 10

    Financial institutions can lose the ability to explain or reproduce agentic AI decisions even when the system is capable and seemingly stable; the paper names this a Verifiability Gap and dissects its mechanisms.

  5. Detecting Behavioral Changes in Python Refactoring Implementations with Foundation Models

    cs.SE 2026-08 conditional novelty 6.0 of 10

    A diff-reading foundation-model oracle detects behavior-changing Python refactorings with 91% recall and surfaced 13 Rope bugs, 12 accepted by maintainers.

  6. KGCache: Amortized Subgraph Retrieval for KG Reasoning with LLMs

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Caching one-hop knowledge-graph neighborhoods with LRU or LFU prevents repeated graph queries in KGQA systems, giving up to 1.91x faster graph retrieval but only about 1.06x end-to-end speedup.

  7. A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Gemini 2.5 Flash matches human rank correlation within 0.07 on 5 of 8 dimensions and within-1 agreement on 6 of 8 when judging raw stereo full-duplex agent conversations, with rank order transferring across later Gemi...

  8. Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Spider 2.0-AIFunc is a 465-instance benchmark for evaluating text-to-SQL systems on queries that incorporate Snowflake Cortex AI functions, with evaluations of ten models showing proprietary models reach 67-70% accuracy.

  9. Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement

    cs.CL 2026-06 conditional novelty 6.0 of 10

    A prompt that forces LLMs to separate facts, inferences, and emotions reduces repeated-answer variability (+0.016 to +0.021 SI on a ~0.95 baseline) and, under injected state persistence, cuts decision-flip rate by 82%...

  10. An Empirical Study of Foundation Models for Variability-Induced Compilation Errors in Configurable C Code

    cs.SE 2026-01 conditional novelty 6.0 of 10

    On a synthetic benchmark of configurable C snippets, GPT-OSS-20B detected affected configurations with 84.7% precision and 52.1% recall and repaired 72.4% of faulty snippets; a smaller real-world study suggests practi...

  11. CAPE: Context-Aware Personality Evaluation Framework for Large Language Models

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Conversational history changes LLM personality-test answers: it increases answer consistency through in-context learning but shifts OCEAN scores, especially for GPT-3.5/4, while smaller models rely heavily on prior in...

  12. Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation

    cs.CL 2025-08 conditional novelty 6.0 of 10

    FineCE trains a language model to emit fine-grained, continuous confidence scores during generation, outperforming existing coarse confidence estimators on six benchmarks.

  13. Efficiency of turbulence

    physics.flu-dyn 2025-08 unverdicted novelty 6.0 of 10

    The efficiency of turbulence, the fraction of input energy stored in the flow, appears bounded and may saturate in a power-law manner across several turbulent flows.

  14. LLMs on Trial: Evaluating Judicial Fairness for Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new 177,100-case benchmark shows that 16 LLMs systematically vary criminal sentences based on extra-legal demographic and procedural details, revealing pervasive judicial unfairness.

  15. Linguistic Features Extracted by GPT-4 Improve Alzheimer's Disease Detection based on Spontaneous Speech

    cs.CL 2024-12 conditional novelty 6.0 of 10

    GPT-4 ratings of five dementia-related language symptoms, added to 40 standard linguistic features, improve automatic Alzheimer's detection from spontaneous speech transcripts, reaching AUROC 0.931 on ADReSS.

  16. Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens

    cs.CV 2024-11 conditional novelty 6.0 of 10

    AbilityLens unifies 11 public benchmarks into six perception abilities with accuracy and stability metrics, and reveals ability conflicts during MLLM training linked to data mixing and model size.

  17. Alignment Plausibility: A New Standard for Assuring AI in Healthcare

    cs.AI 2026-07 conditional novelty 5.5 of 10

    Alignment plausibility—evidence that an AI system's values, training, and oversight cohere with safe positive health outcomes—should be the regulatory analogue of biological plausibility for LLMs in healthcare.

  18. PROBE: Benchmarking Code Generation in Large Language Models

    cs.SE 2026-07 conditional novelty 5.0 of 10

    A multi-language evaluation framework measuring correctness, solution proximity, and code quality finds current LLMs pass at most ~0.70 per language and worsen sharply with problem difficulty.

  19. Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers

    cs.IR 2026-07 conditional novelty 5.0 of 10

    For AI brand answers, query language is the biggest source of answer noise (26.5%) while brand identity is tiny (1.5%), so repeating the same prompt is the least efficient way to spend a query budget.

  20. Stop Testing Attacks, Start Diagnosing Defenses: The Four-Checkpoint Framework Reveals Where LLM Safety Breaks

    cs.CR 2026-02 conditional novelty 5.0 of 10

    A graded-leakage measure raises reported LLM jailbreak success from 22.6% to 52.7%, with output-stage and intent-level defenses emerging as the weak checkpoints.

  21. From Risk Perception to Behavior Large Language Models-Based Simulation of Pandemic Prevention Behaviors

    cs.SI 2026-01 reject novelty 5.0 of 10

    LLM-based simulations of pandemic prevention behaviors show moderate distributional agreement with Beijing survey data, but the validation uses a lenient threshold and selective reporting.

  22. The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Small open-source LLMs give inconsistent answers on 20% to 50% of multiple-choice questions even at low temperature, while 50B-80B models are consistent on over 90% of questions.

  23. AllSummedUp: un framework open-source pour comparer les metriques d'evaluation de resume

    cs.CL 2025-08 conditional novelty 5.0 of 10

    On SummEval, LLM-based summary evaluators are expensive and unstable, and several published correlations did not reproduce when using open-weight models.

  24. Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation

    cs.CL 2025-05 conditional novelty 5.0 of 10

    With temperature set to zero and ten repeated runs, all tested LLMs kept semantic consistency above 96% on clinical note generation, while Llama 70B and Mistral Small had the best combined consistency and correctness.

  25. Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems

    cs.AI 2025-05 conditional novelty 5.0 of 10

    The paper shows that evaluating prompts by the semantic similarity of repeated LLM outputs, and refining prompts toward higher similarity, improves task success in general-purpose multi-agent systems.

  26. The Accuracy, Robustness, and Readability of LLM-Generated Sustainability-Related Word Definitions

    cs.CL 2025-02 conditional novelty 5.0 of 10

    LLM-generated definitions of IPCC climate terms achieve moderate semantic similarity (0.57 to 0.59) to official definitions and lower readability than the originals.

  27. A statistically consistent measure of semantic uncertainty using Language Models

    cs.CL 2025-02 reject novelty 5.0 of 10

    Semantic spectral entropy is proposed as a consistent measure of LLM output uncertainty, but a key proof step is invalid.

  28. Relative Bias: A Comparative Framework for Quantifying Bias in LLMs

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A model is 'relatively biased' when its responses deviate from the consensus of a baseline LLM set, and this deviation can be scored by embedding distances or LLM judges plus equivalence tests.

  29. A Structured Literature Review on Traditional Approaches in Current Natural Language Processing

    cs.CL 2025-05 accept novelty 4.0 of 10

    A structured literature review of 2023 ACM papers finds that traditional, non-neural NLP techniques are still used in classification, information extraction, relation extraction, text simplification, and text summariz...

  30. Enhancing Annotated Bibliography Generation with LLM Ensembles

    cs.CL 2024-12 reject novelty 4.0 of 10

    An LLM ensemble with a judge and a summarizer improves readability and conciseness of generated annotated bibliographies, but the evaluation is too thin to support the stated quality gains.

  31. Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge

    cs.CL 2024-12 conditional novelty 4.0 of 10

    LLM judges are often internally inconsistent across seed variations, with McDonald's omega reliability scores mostly below acceptable thresholds.

  32. On the Surprising Efficacy of LLMs for Penetration-Testing

    cs.CR 2025-07 conditional novelty 3.0 of 10

    A critical review arguing that LLMs are surprisingly effective for penetration testing because the task is largely pattern-matching, while noting serious reliability, safety, and cost barriers to autonomous use.

Pith tools