Higher generator-evaluator self-consistency across 10 frontier LLMs correlates with increased vulnerability to physician-validated mistakes in clinical settings.
An empirical study on large language models in accuracy and robustness under chinese industrial scenarios
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
years
2026 2roles
background 1polarities
background 1representative citing papers
A systematic survey of 93 studies that maps the bidirectional relationship between metamorphic testing and LLMs, proposing a taxonomy for MT applied to LLMs and LLMs applied to MT.
citing papers explorer
-
The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes
Higher generator-evaluator self-consistency across 10 frontier LLMs correlates with increased vulnerability to physician-validated mistakes in clinical settings.
-
Bidirectional Empowerment of Metamorphic Testing and Large Language Models: A Systematic Survey
A systematic survey of 93 studies that maps the bidirectional relationship between metamorphic testing and LLMs, proposing a taxonomy for MT applied to LLMs and LLMs applied to MT.