Across 224 problems and twelve models, LLM-generated implementations show strongly correlated failures, so majority-vote N-version ensembles realize only about 0.43–0.44 of the reliability gain expected under independence.
hub
System structure for software fault tolerance,
1 Pith paper cite this work, alongside 1,527 external citations. Polarity classification is still indexing.
1
Pith paper citing it
1,527
external citations · external index
hub tools
fields
cs.SE 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Systematic Methodology for Evaluating Failure Independence in LLM-Generated Code
Across 224 problems and twelve models, LLM-generated implementations show strongly correlated failures, so majority-vote N-version ensembles realize only about 0.43–0.44 of the reliability gain expected under independence.