ARIADNE combines blackboard architecture with MCTS to coordinate strategy, code, test, evaluation, and repair stages, yielding higher Pass@1 scores than prior LLM baselines on APPS, CodeContests, and related benchmarks.
Un- veiling inefficiencies in llm-generated code: Toward a comprehensive taxonomy
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
SpecDB generates a 23,779-line Rust database via LLM subagents that matches PostgreSQL and MySQL tpmC on TPC-C while using roughly 3% of their code size.
A review of 114 studies creates taxonomies for code and data quality issues, formalizes 18 propagation mechanisms from training data defects to LLM-generated code defects, and synthesizes detection and mitigation techniques.
Empirical study of Stack Overflow logging posts identifies 11 topics where containerized environments show the highest difficulty via unanswered rates and resolution time.
Develops a conceptual distinction between human-cognitive and artificial-stochastic error architectures in code generation, drawing on Dennett, Rescher, and Floridi to explore implications for AI-human collaboration.
citing papers explorer
-
ARIADNE: Agentic Reward-Informed Adaptive Decision Exploration via Blackboard-Driven MCTS for Competitive Program Generation
ARIADNE combines blackboard architecture with MCTS to coordinate strategy, code, test, evaluation, and repair stages, yielding higher Pass@1 scores than prior LLM baselines on APPS, CodeContests, and related benchmarks.
-
SpecDB: LLM-Generated Customized Databases via Feature-Oriented Decomposition
SpecDB generates a 23,779-line Rust database via LLM subagents that matches PostgreSQL and MySQL tpmC on TPC-C while using roughly 3% of their code size.
-
Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code
A review of 114 studies creates taxonomies for code and data quality issues, formalizes 18 propagation mechanisms from training data defects to LLM-generated code defects, and synthesizes detection and mitigation techniques.
-
An Empirical Study on Logging Evolution On Stack Overflow: Trends, Topics, and Challenges
Empirical study of Stack Overflow logging posts identifies 11 topics where containerized environments show the highest difficulty via unanswered rates and resolution time.
-
Architectures of Error: A Philosophical Inquiry into AI and Human Code Generation
Develops a conceptual distinction between human-cognitive and artificial-stochastic error architectures in code generation, drawing on Dennett, Rescher, and Floridi to explore implications for AI-human collaboration.