Pith. sign in

REVIEW 17 cited by

Is Temperature the Creativity Parameter of Large Language Models?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.00492 v1 pith:V44WXGM5 submitted 2024-05-01 cs.CL cs.AI

Is Temperature the Creativity Parameter of Large Language Models?

classification cs.CL cs.AI
keywords creativitytemperatureparameteroutputsclaimcohesioncorrelatedgeneration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large language models (LLMs) are applied to all sorts of creative tasks, and their outputs vary from beautiful, to peculiar, to pastiche, into plain plagiarism. The temperature parameter of an LLM regulates the amount of randomness, leading to more diverse outputs; therefore, it is often claimed to be the creativity parameter. Here, we investigate this claim using a narrative generation task with a predetermined fixed context, model and prompt. Specifically, we present an empirical analysis of the LLM output for different temperature values using four necessary conditions for creativity in narrative generation: novelty, typicality, cohesion, and coherence. We find that temperature is weakly correlated with novelty, and unsurprisingly, moderately correlated with incoherence, but there is no relationship with either cohesion or typicality. However, the influence of temperature on creativity is far more nuanced and weak than suggested by the "creativity parameter" claim; overall results suggest that the LLM generates slightly more novel outputs as temperatures get higher. Finally, we discuss ideas to allow more controlled LLM creativity, rather than relying on chance via changing the temperature parameter.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference

    cs.DC 2026-06 unverdicted novelty 7.0

    A new fault-injection framework enables a systematic empirical study that produces 17 takeaways on error propagation in LLM inference and four software-only mitigation directions.

  2. More Is Not More: What Matters for Diversity in LLM Opinions?

    cs.CL 2026-05 conditional novelty 7.0

    Diversity in LLM opinions comes mostly from the first persona sentence and from combining different interaction architectures, not from richer personas, temperature, or diversity instructions.

  3. Retrieval-Augmented Large Language Models for Evidence-Informed Guidance on Cannabidiol Use in Older Adults

    cs.IR 2026-01 unverdicted novelty 7.0

    Retrieval-augmented LLMs produce more cautious and guideline-aligned recommendations on cannabidiol for older adults than standalone models, demonstrated via automated evaluation on 64 diverse scenarios.

  4. Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

    cs.AI 2026-07 conditional novelty 6.0

    A sequential greedy-then-high-temperature decoding method with a learned refusal-prefix gate preserves 91-99% of greedy refusal responses at high temperatures.

  5. AGC-Bench: Measuring Artificial General Creativity

    cs.CL 2026-07 unverdicted novelty 6.0

    AGC-Bench standardizes measurement of LLM creativity across domains, recovers a dominant 'c' factor explaining 81.5% variance separable from reasoning, and shows humans still lead on matched tasks.

  6. AGC-Bench: Measuring Artificial General Creativity

    cs.CL 2026-07 unverdicted novelty 6.0

    AGC-Bench introduces a multi-domain creativity benchmark for LLMs, recovers a general 'c' factor explaining 81.5% of variance, and finds humans still outperform top models on matched tasks.

  7. Automated Creativity Evaluation of Language Models Across Open-Ended Tasks

    cs.CL 2026-06 unverdicted novelty 6.0

    Authors propose a new framework for automated LLM creativity evaluation that separates measurement from the task, using semantic entropy and multi-agent judges, validated on problem-solving, research ideation, and cre...

  8. "I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise

    cs.CL 2026-06 unverdicted novelty 6.0

    Decan (D_Ca_n = C × a_n) measures text diversity as progressive conditional surprise from base LM log-probabilities, scoring 0.846 OCA on McDiv benchmark and detecting monotonic diversity drop across base→SFT→DPO→RLVR stages.

  9. Expressing Social Emotions: Misalignment Between LLMs and Human Cultural Emotion Norms

    cs.CL 2026-04 unverdicted novelty 6.0

    Frontier LLMs over-express engaging emotions relative to disengaging ones and generate deterministic responses that fail to match the cultural and individual diversity observed in human social emotion expression.

  10. Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning

    cs.CL 2026-04 unverdicted novelty 6.0

    A hybrid fine-tuning objective using KL divergence for token calibration and Kahneman-Tversky optimization for semantic binding enables LLMs to produce outputs that match desired attribute distributions across repeate...

  11. Assessing the Business Process Modeling Competences of Large Language Models

    cs.SE 2026-01 conditional novelty 6.0

    Open-source LLMs can produce BPMN process models that rival human experts on syntax and readability, but they lag on semantic accuracy and frequently generate invalid BPMN-XML.

  12. A Study of LLMs' Preferences for Libraries and Programming Languages

    cs.SE 2025-03 unverdicted novelty 6.0

    Empirical study of eight LLMs finds overuse of popular libraries like NumPy in up to 45% of unnecessary cases and strong default preference for Python even when suboptimal.

  13. Emotion Profiling in LLM-Based Literary Translation: Systematic Shifts Across MT and Post-Editing

    cs.CL 2026-06 unverdicted novelty 5.0

    LLM translations introduce model-specific statistically significant emotional fingerprints that limit preservation of author voice, with post-editing providing partial alignment to human norms.

  14. IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs

    cs.CL 2026-05 unverdicted novelty 5.0

    IDEAFix is an evaluation framework that varies task attributes and defixation prompts in LLM idea generation, showing task formulation affects performance while simple prompts boost originality but homogenization persists.

  15. Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior

    cs.CL 2026-04 unverdicted novelty 5.0

    CDS is a new synthetic corpus of LLM-generated texts on vaccines, disinformation, gender gaps, and STEM stereotypes, linked to persona attributes to enable bias and alignment audits.

  16. DORA Explorer: Improving the Exploration Ability of LLMs Without Training

    cs.CL 2026-04 unverdicted novelty 5.0

    DORA Explorer boosts LLM agent exploration without training by ranking diverse actions using log-probabilities and a tunable parameter, yielding UCB-competitive results on multi-armed bandits and gains on text adventu...

  17. Mapping and Comparing Climate Equity Policy Practices Using RAG LLM-Based Semantic Analysis and Recommendation Systems

    cs.CY 2026-01 conditional novelty 5.0

    A RAG-LLM pipeline extracts transportation and energy policy items from U.S. climate equity plans and recommends cities with similar policy practices, but extraction is not validated against human coding.