On a small synthetic four-language dataset, GPT-4.0 detects code smells with higher precision than DeepSeek-V3, while both models miss most annotated smells and the cost comparison is unreliable.
Are sonarqube rules inducing bugs? In 2020 IEEE 27th international conference on software analysis, evolution and reengineering (SANER), pages 501–511
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3
On a small synthetic four-language dataset, GPT-4.0 detects code smells with higher precision than DeepSeek-V3, while both models miss most annotated smells and the cost comparison is unreliable.