REVIEW 21 cited by
Unleashing the potential of prompt engineering for large language models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This comprehensive review delves into the pivotal role of prompt engineering in unleashing the capabilities of Large Language Models (LLMs). The development of Artificial Intelligence (AI), from its inception in the 1950s to the emergence of advanced neural networks and deep learning architectures, has made a breakthrough in LLMs, with models such as GPT-4o and Claude-3, and in Vision-Language Models (VLMs), with models such as CLIP and ALIGN. Prompt engineering is the process of structuring inputs, which has emerged as a crucial technique to maximize the utility and accuracy of these models. This paper explores both foundational and advanced methodologies of prompt engineering, including techniques such as self-consistency, chain-of-thought, and generated knowledge, which significantly enhance model performance. Additionally, it examines the prompt method of VLMs through innovative approaches such as Context Optimization (CoOp), Conditional Context Optimization (CoCoOp), and Multimodal Prompt Learning (MaPLe). Critical to this discussion is the aspect of AI security, particularly adversarial attacks that exploit vulnerabilities in prompt engineering. Strategies to mitigate these risks and enhance model robustness are thoroughly reviewed. The evaluation of prompt methods is also addressed through both subjective and objective metrics, ensuring a robust analysis of their efficacy. This review also reflects the essential role of prompt engineering in advancing AI capabilities, providing a structured framework for future research and application.
Forward citations
Cited by 21 Pith papers
-
Prompt engineering using order-of-addition experiments: An application to generating two-level fractional factorial designs
Order-of-addition designs and logistic pairwise-ordering models measure and optimize prompt-element order, lifting LLM success on 16-run fractional factorial design tasks from low teens or mid-thirties to near 100%.
-
LLM-based Question-Answer Framework for Sensor-driven HVAC System Interaction
JARVIS, an LLM-based HVAC question-answering framework with an Expert-LLM, a parameterized SQL builder, and bottom-up planning, outperforms a text-to-SQL baseline and its own ablations on a small expert-curated dataset.
-
Sword and Shield: Uses and Strategies of LLMs in Navigating Disinformation
In a 25-participant Werewolf-style game, all roles used an LLM chatbot strategically, as a sword for disinformation and a shield against it.
-
From Parameters to Prompts: Understanding and Mitigating the Factuality Gap between Fine-Tuned LLMs
Fine-tuning LLMs on known versus unknown facts creates a factuality gap that in-context prompting can largely erase, according to experiments and a knowledge-graph model.
-
Privacy-preserving Prompt Personalization in Federated Learning for Multimodal Large Language Models
SecFPP combines hierarchical prompt decomposition with secret-sharing-based adaptive clustering to protect user prompts in federated learning while preserving personalization accuracy.
-
Transforming Remanufacturing Automation with Large Language Models: A Forward-Looking Analysis with Case Studies
The authors propose ReManGPT, a conceptual orchestration framework for applying LLMs to remanufacturing, and illustrate it with case studies in disassembly planning, repair guidance, and robotic execution.
-
Using LLMs to create analytical datasets: A case study of reconstructing the historical memory of Colombia
A GPT-based pipeline converted 235,000 scanned newspaper articles into a dataset of 78,685 violent events in Colombia, enabling descriptive analysis and a coca eradication regression that found no significant relationship.
-
Extracting Research Instruments from Educational Literature Using LLMs
A multi-step LLM pipeline extracts research instruments and their attributes from education literature, reporting moderate F1 scores but no released code or baseline statistics.
-
Using LLMs to Detect Growth in Computational Thinking in Introductory Physics
An LLM scored students' written computational thinking responses in an introductory physics course with human-level agreement on well-defined practices and reproduced pre-post growth trends at scale.
-
Prompts Blend Requirements and Solutions: From Intent to Implementation
Prompts in AI-assisted development can be decomposed into functionality/quality, general solutions, and specific solutions — the 'Prompt Triangle' framework.
-
Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial
A survey and tutorial that organizes LLM-enabled wireless network optimization into formulation, solution, and verification stages, with case studies drawn from the authors' own prior papers.
-
Investigation of the Inter-Rater Reliability between Large Language Models and Human Raters in Qualitative Analysis
GPT-4o and GPT-4.5-preview reached moderate to substantial Cohen's kappa with human coders for three of four student thinking themes after per-theme hyperparameter and prompt tuning.
-
CaTE Data Curation for Trustworthy AI
A synthesis of data curation practices for trustworthy AI, framed around an actionable definition of trustworthiness and a decision tree.
-
Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models
A zero-configuration prompt optimization system built on DSPy and MIPROv2 that reports competitive performance on five NLP tasks, with results measured on synthetic validation data.
-
An Integrated Framework of Prompt Engineering and Multidimensional Knowledge Graphs for Legal Dispute Analysis
A prompt-plus-knowledge-graph framework for legal dispute analysis reports improved LLM sensitivity and citation accuracy on a 100-pair test set, but with limited statistical support.
-
An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
A broad benchmark of six open-weights LLMs shows prompt design and chunking affect summarization quality more than model size alone.
-
Enhancing COBOL Code Explanations: A Multi-Agents Approach Using Large Language Models
A two-agent LLM pipeline with hierarchical merging explains COBOL functions, files, and projects, outperforming zero-shot baselines on several text-quality metrics.
-
Evaluating and Improving Large Language Models for Competitive Program Generation
DeepSeek-R1 solves only 5 of 80 recent ICPC/CCPC competitive programming problems with a basic prompt, and 46 of 80 after a taxonomy-guided repair and regeneration pipeline.
-
AI-Facilitated Analysis of Abstracts and Conclusions: Flagging Unsubstantiated Claims and Ambiguous Pronouns
Structured prompts can steer LLMs to flag certain unsupported claims and ambiguous pronouns, but performance varies sharply by model, context, and the syntactic role of the target, and the single test case was also th...
-
Designing Effective LLM-Assisted Interfaces for Curriculum Development
A clickable interface with predefined commands improves teachers' perceived usability and reduces workload compared with a ChatGPT-style chat interface for course outline creation.
-
FAA Framework: A Large Language Model-Based Approach for Credit Card Fraud Investigations
The FAA framework automates credit card fraud investigations with GPT-4o agents and reports 98-99% fraud-detection F1, though the evaluation is weakened by self-referential LLM scoring.
Discussion (0). Sign in to comment.