Pith. sign in

REVIEW 12 cited by

TEMPERA: Test-Time Prompting via Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.11890 v1 pith:BOC52NC6 submitted 2022-11-21 cs.CL cs.AI

classification cs.CLcs.AI
keywords promptdesignlearningmethodstemperaachievescomparedediting
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Careful prompt design is critical to the use of large language models in zero-shot or few-shot learning. As a consequence, there is a growing interest in automated methods to design optimal prompts. In this work, we propose Test-time Prompt Editing using Reinforcement learning (TEMPERA). In contrast to prior prompt generation methods, TEMPERA can efficiently leverage prior knowledge, is adaptive to different queries and provides an interpretable prompt for every query. To achieve this, we design a novel action space that allows flexible editing of the initial prompts covering a wide set of commonly-used components like instructions, few-shot exemplars, and verbalizers. The proposed method achieves significant gains compared with recent SoTA approaches like prompt tuning, AutoPrompt, and RLPrompt, across a variety of tasks including sentiment analysis, topic classification, natural language inference, and reading comprehension. Our method achieves 5.33x on average improvement in sample efficiency when compared to the traditional fine-tuning methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A small LLM trained with GRPO and LLM-judge rewards rewrites simple prompts into more effective ones, improving question-answering and arithmetic accuracy over base prompts while giving mixed, often negligible gains o...

  2. TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards

    cs.CL 2025-07 conditional novelty 6.0 of 10

    TRPrompt trains an 8B prompt model directly on natural-language textual rewards and reports the highest accuracies on GSMHard and MATH among the compared methods.

  3. DEMONSTRATE: Zero-shot Language to Robotic Control via Multi-task Demonstration Learning

    cs.RO 2025-07 conditional novelty 6.0 of 10

    DEMONSTRATE learns a zero-shot mapping from natural-language embeddings to MPC cost parameters from demonstrations, achieving tabletop manipulation success rates comparable to prior LLM-based pipelines.

  4. Grammar-Guided Evolutionary Search for Discrete Prompt Optimisation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A grammar-guided evolutionary search that composes prompt edits outperformed PromptWizard, OPRO, and RL-Prompt on small LLMs across four domain-specific tasks.

  5. Stabilizing Black-Box Prompt Optimization with Textual Regularization and Signal Aggregation

    cs.LG 2025-07 conditional novelty 6.0 of 10

    TRAS adds success-based textual regularization and Monte Carlo signal aggregation to black-box prompt optimization, improving accuracy and reducing instruction loss when moving prompts across models.

  6. Model Performance-Guided Evaluation Data Selection for Effective Prompt Optimization

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An evaluation-set selection method that adds real-time model feedback to semantic sampling improves the accuracy and stability of three prompt optimization methods on two datasets.

  7. Online Prompt Selection for Program Synthesis

    cs.AI 2025-01 conditional novelty 6.0 of 10

    An online multi-armed bandit that selects among symbolic solvers and LLM-prompt combinations for program synthesis solves 37.2% more queries than the best single solver and reaches 96% of the virtual best solver's per...

  8. In-Context Learning as Implicit Policy Gradient

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Score-conditioned in-context learning is structurally analogous to a REINFORCE policy-gradient update in a simplified hidden-state model, with attention performing score-weighted aggregation and a bounded KL shift.

  9. LLM-AutoDiff: Auto-Differentiate Any LLM Workflow

    cs.CL 2025-01 conditional novelty 5.0 of 10

    An automatic differentiation-style framework optimizes prompts across multi-component and cyclic LLM workflows, outperforming single-node textual-gradient baselines on several small benchmarks.

  10. TAPO: Task-Referenced Adaptation for Prompt Optimization

    cs.CL 2025-01 conditional novelty 5.0 of 10

    TAPO is a prompt optimization framework that selects task-specific evaluation metrics and evolves prompts, reporting small and partly conflicting gains over four baselines on six reasoning datasets.

  11. Evolutionary Computation and Large Language Models: A Survey of Methods, Synergies, and Applications

    cs.NE 2025-05 conditional novelty 4.0 of 10

    A survey that maps bidirectional synergies between evolutionary computation and large language models and proposes a taxonomy plus research gaps.

  12. The Evolution of Natural Language Processing: How Prompt Optimization and Language Models are Shaping the Future

    cs.CL 2025-06 reject novelty 3.0 of 10

    A review that categorizes 45 prompt optimization strategies into 11 classes and surveys their use across NLP tasks, models, and datasets, but with inconsistent counts and overlapping categories.

Pith tools