REVIEW 7 cited by
Prompt Engineering a Prompt Engineer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Prompt engineering is a challenging yet crucial task for optimizing the performance of large language models on customized tasks. It requires complex reasoning to examine the model's errors, hypothesize what is missing or misleading in the current prompt, and communicate the task with clarity. While recent works indicate that large language models can be meta-prompted to perform automatic prompt engineering, we argue that their potential is limited due to insufficient guidance for complex reasoning in the meta-prompt. We fill this gap by infusing into the meta-prompt three key components: detailed descriptions, context specification, and a step-by-step reasoning template. The resulting method, named PE2, exhibits remarkable versatility across diverse language tasks. It finds prompts that outperform "let's think step by step" by 6.3% on MultiArith and 3.1% on GSM8K, and outperforms competitive baselines on counterfactual tasks by 6.9%. Further, we show that PE2 can make targeted and highly specific prompt edits, rectify erroneous prompts, and induce multi-step plans for complex tasks.
Forward citations
Cited by 7 Pith papers
-
ISTQB Certifications Under the Lens: Their Contributions to the Software-Testing Profession; and AI-assisted Synthesis of Practitioners' Endorsements and Criticisms
ISTQB certifications deliver career and communication benefits yet remain contested for theoretical bias and weak practical skill assessment, per AI-synthesized practitioner sources and expert validation.
-
Towards Efficient and Robust Linguistic Emotion Diagnosis for Mental Health via Multi-Agent Instruction Refinement
A multi-agent prompt-rewriting loop is claimed to improve LLM emotion diagnosis accuracy, but its evaluation appears to optimize on the test set and lacks replication details.
-
Retrieval augmented generation based dynamic prompting for few-shot biomedical named entity recognition using large language models
Retrieval-based selection of in-context examples improves few-shot biomedical named entity recognition F1 over random selection, with TF-IDF and SBERT outperforming ColBERT and DPR.
-
Prompt Smart, Pay Less: Cost-Aware APO for Real-World Applications
APE-OPRO, a hybrid of APE and OPRO, achieves similar weighted F1 to OPRO at roughly 18% lower API cost on a 2,500-product commercial classification benchmark.
-
SI-Agent: An Agentic Framework for Feedback-Driven Generation and Tuning of Human-Readable System Instructions for Large Language Models
The paper proposes a multi-agent loop (instructor, follower, feedback) to auto-generate human-readable system prompts, claiming good benchmark performance and readability, but the supporting experiments are not reprod...
-
Small Language Models in the Real World: Insights from Industrial Text Classification
For 1B to 3B models, prompting alone is near random, while training a small classification head on frozen weights is the most accurate and VRAM-efficient path, with data volume and pretraining domain as the main bottlenecks.
-
Evolutionary Computation and Large Language Models: A Survey of Methods, Synergies, and Applications
A survey that maps bidirectional synergies between evolutionary computation and large language models and proposes a taxonomy plus research gaps.
Discussion (0). Sign in to comment.