Pith. sign in

REVIEW 9 cited by

Prompt Design and Engineering: Introduction and Advanced Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.14423 v4 pith:SY7P6F6X submitted 2024-01-24 cs.SE cs.LG

classification cs.SEcs.LG
keywords promptadvanceddesignengineeringagentsbecomebehindbuilding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Prompt design and engineering has rapidly become essential for maximizing the potential of large language models. In this paper, we introduce core concepts, advanced techniques like Chain-of-Thought and Reflection, and the principles behind building LLM-based agents. Finally, we provide a survey of tools for prompt engineers.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large language models replicate and predict human cooperation across experiments in game theory

    cs.AI 2025-11 conditional novelty 6.0 of 10

    Llama-3.1-8B with a multi-step reasoning-and-filter prompt reproduces human cooperation rates across 121 dyadic games (MSD=0.031, r=0.89), outperforming Nash-equilibrium predictions (MSD=0.096, r=0.78).

  2. Limited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Generation in Low-Data Regimes

    cs.LG 2025-09 conditional novelty 6.0 of 10

    ReFine combines rule-guided prompting and dual-granularity filtering to improve LLM-based tabular data generation when only 30 to 90 labeled rows exist, achieving top average rank over baselines.

  3. Ratas framework: A comprehensive genai-based approach to rubric-based marking of real-world textual exams

    cs.CL 2025-05 conditional novelty 6.0 of 10

    RATAS decomposes rubrics into simplified rules, scores each rule with GPT-4o, and cascades scores to grade long textual exam answers with reported near-human accuracy.

  4. LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents

    cs.AI 2025-09 reject novelty 5.0 of 10

    On CUAD legal contracts, a prompt-engineered QWEN-2 pipeline with chunking and two answer-selection heuristics reportedly outperforms the fine-tuned DeBERTa-large baseline by about 9%, reaching claimed state-of-the-ar...

  5. Decoding ML Decision: An Agentic Reasoning Framework for Large-Scale Ranking System

    cs.AI 2026-02 conditional novelty 4.0 of 10

    GEARS outperforms prompting baselines at selecting ranking policies from GAS-generated candidate sets, using tool-based filtering and feature-stability checks.

  6. AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito

    cs.AI 2026-01 conditional novelty 4.0 of 10

    An AI agent combining GraphRAG, static Fortran analysis, and LLM code generation is reported to translate legacy Fortran finite-difference code into Devito, with Grade-A results claimed on roughly three-quarters of 13...

  7. An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A broad benchmark of six open-weights LLMs shows prompt design and chunking affect summarization quality more than model size alone.

  8. Exploring Prompt Patterns in AI-Assisted Code Generation: Towards Faster and More Effective Developer-AI Collaboration

    cs.SE 2025-06 reject novelty 3.0 of 10

    Using keyword matching on DevGPT conversations, the authors rank seven prompt patterns by a hand-weighted effectiveness score and claim 'Context and Instruction' and 'Recipe' reduce iterations.

  9. A Short Survey on Formalising Software Requirements using Large Language Models

    cs.SE 2025-06 unverdicted novelty 1.0 of 10

    A survey summarizing 35 papers on using LLMs to formalize software requirements, but it contains no new experimental results and its classification tables have errors.

Pith tools