With temperature set to zero and ten repeated runs, all tested LLMs kept semantic consistency above 96% on clinical note generation, while Llama 70B and Mistral Small had the best combined consistency and correctness.
Rahmani, and Youlin Li
1 Pith paper cite this work, alongside 41 external citations. Polarity classification is still indexing.
1
Pith paper citing it
41
external citations · OpenAlex
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation
With temperature set to zero and ten repeated runs, all tested LLMs kept semantic consistency above 96% on clinical note generation, while Llama 70B and Mistral Small had the best combined consistency and correctness.