CWEval is a new benchmark that simultaneously checks functional correctness and security of AI-generated code with dynamic test oracles, exposing a large correct-but-insecure gap in current LLMs.
Safecoder: A machine- learning-based encoding system to embed safety identification informa- tion into qr codes,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation
CWEval is a new benchmark that simultaneously checks functional correctness and security of AI-generated code with dynamic test oracles, exposing a large correct-but-insecure gap in current LLMs.