REVIEW 12 cited by
Robustness, Security, Privacy, Explainability, Efficiency, and Usability of Large Language Models for Code
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models for code (LLM4Code), which demonstrate strong performance (e.g., high accuracy) in processing source code, have significantly transformed software engineering. Many studies separately investigate the non-functional properties of LM4Code, but there is no systematic review of how these properties are evaluated and enhanced. This paper fills this gap by thoroughly examining 146 relevant studies, thereby presenting the first systematic literature review to identify seven important properties beyond accuracy, including robustness, security, privacy, explainability, efficiency, and usability. We discuss the current state-of-the-art methods and trends, identify gaps in existing research, and present promising directions for future study.
Forward citations
Cited by 12 Pith papers
-
Insecure Coding Preferences in Long-Term Memory: Security Risks for LLM-based Code Generation
Injected insecure coding preferences in LLM long-term memory raise vulnerability rates by 2.7-50.3 pp and suppress warnings; memory-level filtering restores safe behavior in the tested set.
-
Unified Communication Compression Beyond Global Error Bounds for Distributed Nonconvex Optimization
A unified compression algorithm for distributed nonconvex optimization achieves O(1/sqrt(T)) convergence for locally-bounded compressors, matching centralized 1-bit methods, with an improved O(1/T^{2/3}) rate after on...
-
Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?
Standard gender-bias detectors for text-to-image models deviate substantially from human-annotated bias, and a face-filtering plus CLIP pipeline measures bias more accurately.
-
Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation
An MLLM-driven agentic pipeline generates PCB component symbols and footprints from datasheets with reported 86%/80% accuracy and builds a 1,000-component library.
-
"My productivity is boosted, but ..." Demystifying Users' Perception on AI Coding Assistants
Using 1,085 marketplace extensions and reviews of 32 assistants, this study maps developer satisfaction and criticism into a taxonomy and five design implications.
-
Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
Quantized code LLMs appear more robust than full-precision ones in a majority of tested adversarial and noise scenarios, but the proposed Relative Robustness Score is misspecified.
-
Deconstructing Obfuscation: A four-dimensional framework for evaluating Large Language Models assembly code deobfuscation capabilities
Commercial LLMs deobfuscate simple OLLVM-obfuscated assembly well for some techniques but universally fail when three obfuscations are combined.
-
CodeImprove: Program Adaptation for Deep Code Models
Adapting program inputs with a layerwise validity score and genetic search improves deep code model accuracy by up to 8.78% without retraining.
-
An Empirical Study of Vulnerable Package Dependencies in LLM Repositories
In 52 open-source LLM projects, 75.8% of those with dependency configs use at least one vulnerable package, and half of supply chain vulnerabilities stay undisclosed for over 56 months.
-
A Functional Software Reference Architecture for LLM-Integrated Systems
A preliminary four-layer functional reference architecture for LLM-integrated systems, illustrated on three open-source projects.
-
Towards Advancing Code Generation with Large Language Models: A Research Roadmap
A roadmap paper that organizes LLM code generation into a six-layer architecture and a four-phase human-in-the-loop workflow, and lists open challenges and recommendations.
-
Position Paper: Programming Language Techniques for Bridging LLM Code Generation Semantic Gaps
A position paper arguing that PL techniques, especially formal verification and structure-aware representations, should be deeply integrated into LLM code generation.
Discussion (0). Continue with ORCID to comment.