PEEM is a multi-criteria LLM-based evaluator for prompts and responses that aligns with standard accuracy while enabling zero-shot prompt optimization via feedback.
Which prompting technique should I use? An empirical investigation of prompting tech- niques for software engineering tasks
6 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 6representative citing papers
A study of seven LLMs finds that realistic prompt variations such as one-character misspellings trigger library hallucinations in up to 26% of cases, fabricated names in up to 99%, and time-based prompts in up to 85%, and introduces LibHalluBench for evaluation.
A 400-entry benchmark and protocol shows tool-augmented agents reach 89.5% compilation but only 60.5% consensus faithfulness, with a 29-point gap; elaboration feedback improves validity most but increases unfaithful compiles.
An AI-native TDD framework operationalizes classical TDD principles as prompt-level and workflow-level governance mechanisms in a layered multi-agent architecture to improve stability and reproducibility of LLM code generation.
Gemini 3 Flash achieved the highest accuracy on PSM I-style questions among three tested LLMs, with low intra-model variability and systematic error patterns by question format and topic.
GPT-5 with source-citation prompting achieves 89.1% accuracy on 993 PSM questions, outperforming zero-shot and chain-of-thought while errors cluster in multi-select and interpretive topics.
citing papers explorer
-
PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses
PEEM is a multi-criteria LLM-based evaluator for prompts and responses that aligns with standard accuracy while enabling zero-shot prompt optimization via feedback.
-
Library Hallucinations in LLM-Generated Code: A Risk Analysis Grounded in Developer Queries
A study of seven LLMs finds that realistic prompt variations such as one-character misspellings trigger library hallucinations in up to 26% of cases, fabricated names in up to 99%, and time-based prompts in up to 85%, and introduces LibHalluBench for evaluation.
-
Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization
A 400-entry benchmark and protocol shows tool-augmented agents reach 89.5% compilation but only 60.5% consensus faithfulness, with a 29-point gap; elaboration feedback improves validity most but increases unfaithful compiles.
-
TDD Governance for Multi-Agent Code Generation via Prompt Engineering
An AI-native TDD framework operationalizes classical TDD principles as prompt-level and workflow-level governance mechanisms in a layered multi-agent architecture to improve stability and reproducibility of LLM code generation.
-
Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns
Gemini 3 Flash achieved the highest accuracy on PSM I-style questions among three tested LLMs, with low intra-model variability and systematic error patterns by question format and topic.
-
Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study
GPT-5 with source-citation prompting achieves 89.1% accuracy on 993 PSM questions, outperforming zero-shot and chain-of-thought while errors cluster in multi-select and interpretive topics.