With temperature set to zero and ten repeated runs, all tested LLMs kept semantic consistency above 96% on clinical note generation, while Llama 70B and Mistral Small had the best combined consistency and correctness.
Title resolution pending
1 Pith paper cite this work, alongside 4 external citations. Polarity classification is still indexing.
1
Pith paper citing it
4
external citations · OpenAlex
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation
With temperature set to zero and ten repeated runs, all tested LLMs kept semantic consistency above 96% on clinical note generation, while Llama 70B and Mistral Small had the best combined consistency and correctness.