REVIEW 17 cited by
Is Temperature the Creativity Parameter of Large Language Models?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Is Temperature the Creativity Parameter of Large Language Models?
read the original abstract
Large language models (LLMs) are applied to all sorts of creative tasks, and their outputs vary from beautiful, to peculiar, to pastiche, into plain plagiarism. The temperature parameter of an LLM regulates the amount of randomness, leading to more diverse outputs; therefore, it is often claimed to be the creativity parameter. Here, we investigate this claim using a narrative generation task with a predetermined fixed context, model and prompt. Specifically, we present an empirical analysis of the LLM output for different temperature values using four necessary conditions for creativity in narrative generation: novelty, typicality, cohesion, and coherence. We find that temperature is weakly correlated with novelty, and unsurprisingly, moderately correlated with incoherence, but there is no relationship with either cohesion or typicality. However, the influence of temperature on creativity is far more nuanced and weak than suggested by the "creativity parameter" claim; overall results suggest that the LLM generates slightly more novel outputs as temperatures get higher. Finally, we discuss ideas to allow more controlled LLM creativity, rather than relying on chance via changing the temperature parameter.
Forward citations
Cited by 17 Pith papers
-
Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference
A new fault-injection framework enables a systematic empirical study that produces 17 takeaways on error propagation in LLM inference and four software-only mitigation directions.
-
More Is Not More: What Matters for Diversity in LLM Opinions?
Diversity in LLM opinions comes mostly from the first persona sentence and from combining different interaction architectures, not from richer personas, temperature, or diversity instructions.
-
Retrieval-Augmented Large Language Models for Evidence-Informed Guidance on Cannabidiol Use in Older Adults
Retrieval-augmented LLMs produce more cautious and guideline-aligned recommendations on cannabidiol for older adults than standalone models, demonstrated via automated evaluation on 64 diverse scenarios.
-
Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling
A sequential greedy-then-high-temperature decoding method with a learned refusal-prefix gate preserves 91-99% of greedy refusal responses at high temperatures.
-
AGC-Bench: Measuring Artificial General Creativity
AGC-Bench standardizes measurement of LLM creativity across domains, recovers a dominant 'c' factor explaining 81.5% variance separable from reasoning, and shows humans still lead on matched tasks.
-
AGC-Bench: Measuring Artificial General Creativity
AGC-Bench introduces a multi-domain creativity benchmark for LLMs, recovers a general 'c' factor explaining 81.5% of variance, and finds humans still outperform top models on matched tasks.
-
Automated Creativity Evaluation of Language Models Across Open-Ended Tasks
Authors propose a new framework for automated LLM creativity evaluation that separates measurement from the task, using semantic entropy and multi-agent judges, validated on problem-solving, research ideation, and cre...
-
"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise
Decan (D_Ca_n = C × a_n) measures text diversity as progressive conditional surprise from base LM log-probabilities, scoring 0.846 OCA on McDiv benchmark and detecting monotonic diversity drop across base→SFT→DPO→RLVR stages.
-
Expressing Social Emotions: Misalignment Between LLMs and Human Cultural Emotion Norms
Frontier LLMs over-express engaging emotions relative to disengaging ones and generate deterministic responses that fail to match the cultural and individual diversity observed in human social emotion expression.
-
Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning
A hybrid fine-tuning objective using KL divergence for token calibration and Kahneman-Tversky optimization for semantic binding enables LLMs to produce outputs that match desired attribute distributions across repeate...
-
Assessing the Business Process Modeling Competences of Large Language Models
Open-source LLMs can produce BPMN process models that rival human experts on syntax and readability, but they lag on semantic accuracy and frequently generate invalid BPMN-XML.
-
A Study of LLMs' Preferences for Libraries and Programming Languages
Empirical study of eight LLMs finds overuse of popular libraries like NumPy in up to 45% of unnecessary cases and strong default preference for Python even when suboptimal.
-
Emotion Profiling in LLM-Based Literary Translation: Systematic Shifts Across MT and Post-Editing
LLM translations introduce model-specific statistically significant emotional fingerprints that limit preservation of author voice, with post-editing providing partial alignment to human norms.
-
IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs
IDEAFix is an evaluation framework that varies task attributes and defixation prompts in LLM idea generation, showing task formulation affects performance while simple prompts boost originality but homogenization persists.
-
Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior
CDS is a new synthetic corpus of LLM-generated texts on vaccines, disinformation, gender gaps, and STEM stereotypes, linked to persona attributes to enable bias and alignment audits.
-
DORA Explorer: Improving the Exploration Ability of LLMs Without Training
DORA Explorer boosts LLM agent exploration without training by ranking diverse actions using log-probabilities and a tunable parameter, yielding UCB-competitive results on multi-armed bandits and gains on text adventu...
-
Mapping and Comparing Climate Equity Policy Practices Using RAG LLM-Based Semantic Analysis and Recommendation Systems
A RAG-LLM pipeline extracts transportation and energy policy items from U.S. climate equity plans and recommends cities with similar policy practices, but extraction is not validated against human coding.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.